A novel view synthesis method for 3D scenes based on matching light
Optimizing the 3D Gaussian ellipsoid by matching ray strategies, the problem of poor rendering effect of 3DGS in sparse viewing angles is solved, and more accurate rendering effect and stability is achieved, avoiding geometric distortion and missing details.
Patent Information
- Application Number
- CN202411717551.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing 3DGS rendering effect is poor in sparse perspectives, and is prone to problems such as overfitting and geometric artifacts.
Through a strategy based on matching light, the correspondence between light and pixels between different viewing angles is captured, the position and shape of the 3D Gaussian ellipsoid is optimized, and the matching prior information is introduced as a constraint to guide the model to avoid geometric conflicts during rendering depth and improve rendering effect.
From a sparse perspective, the rendered geometric structure more accurately reflects the real structure of the scene, avoids geometric distortion and missing details, and improves rendering quality and stability.
Smart Images

Figure CN119478173B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of view synthesis technology, and in particular relates to a novel view synthesis method for a three-dimensional scene based on matching light. Background Art
[0002] In the field of 3D computer vision, generating novel views from a small number of perspectives (Few-shot Novel View Synthesis) has always been a challenging research topic. With the increasing demand for applications such as virtual reality (VR), augmented reality (AR), and autonomous driving, how to reconstruct 3D scenes from sparse 2D images and generate unseen perspectives has become a key technical problem that needs to be solved urgently. In recent years, Neural Radiance Field (NeRF), as an emerging 3D representation method, has achieved remarkable achievements in the field of novel view image generation (NVS) with its excellent rendering effect. However, NeRF usually relies on a large number of training images, and its optimization and rendering speed are slow, which brings a large computational burden in practical applications, especially in scenarios that require fast response such as autonomous driving and robotics.
[0003] To address this issue, 3D Gaussian Splatting (3DGS) is a new 3D representation method that replaces volume rendering in NeRF with efficient, high-resolution, and fast Gaussian primitives, achieving better performance. 3DGS uses Gaussian points initialized from sparse structured light (SfM) point clouds to explicitly represent 3D scenes, and optimizes its Gaussian parameters through photometric loss, ultimately achieving high-quality, real-time rendering. However, with a small number of view inputs, especially when the scene geometry information is insufficient, the performance of 3DGS will still degrade significantly, and the rendering results are prone to problems such as overfitting and geometric artifacts. In other words, the current rendering effect of 3DGS under sparse view angles is poor. Summary of the Invention
[0004] The embodiments of the present application provide a novel view synthesis method for a three-dimensional scene based on matched light, which can solve the problem of poor rendering effect of 3DGS under sparse viewing angles.
[0005] The present invention provides a novel view synthesis method for a three-dimensional scene based on matched light, including:
[0006] Generate multiple 3D Gaussian ellipsoids based on multiple perspective images of the target scene;
[0007] Randomly select multiple target perspective images from multiple perspective images;
[0008] According to the multiple target view images, multiple target 3D Gaussian ellipsoids are screened out from the multiple 3D Gaussian ellipsoids;
[0009] Based on the matching light, a plurality of Gaussian matching pairs are screened out from the plurality of target 3D Gaussian ellipsoids; the Gaussian matching pairs include two target 3D Gaussian ellipsoids from the plurality of target 3D Gaussian ellipsoids;
[0010] Optimize the target 3D Gaussian ellipsoid in all Gaussian matching pairs;
[0011] The optimized multiple target 3D Gaussian ellipsoids are fused to obtain a novel view of the target scene.
[0012] Optionally, based on the multiple target view images, multiple target 3D Gaussian ellipsoids are screened out from the multiple 3D Gaussian ellipsoids, including:
[0013] For each target view image, the L1 loss and SSIM loss are used to calculate the photometric loss L between the target view image and the rendered image. p ; The rendered image is the rendering result of the process of generating a 3D Gaussian ellipsoid using the target perspective image;
[0014] The luminosity loss L p The 3D Gaussian ellipsoid corresponding to the target perspective image that is less than the luminosity loss threshold is used as the target 3D Gaussian ellipsoid.
[0015] Optionally, L1 loss and SSIM loss are used to calculate the photometric loss between the target view image and the rendered image, including:
[0016] By formula Calculate the photometric loss L between the target view image and the rendered image p ;
[0017] Among them, λ is the weight factor, The target perspective image I and the rendered image The L1 loss between The target perspective image I and the rendered image The SSIM loss between .
[0018] Optionally, based on the matching rays, multiple Gaussian matching pairs are screened from multiple target 3D Gaussian ellipsoids, including:
[0019] For the i-th target 3D Gaussian ellipsoid G among multiple target 3D Gaussian ellipsoids i and the j-th target 3D Gaussian ellipsoid G j , calculate the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G jGaussian position loss; i = 1, ..., Q, j = 1, ..., Q, i ≠ j, Q is the number of target 3D Gaussian ellipsoids;
[0020] If the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j The Gaussian position loss of the i-th target 3D Gaussian ellipsoid G is less than the Gaussian position loss threshold. i and the j-th target 3D Gaussian ellipsoid G j as Gaussian matched pairs.
[0021] Optionally, calculate the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss, including:
[0022] The i-th target 3D Gaussian ellipsoid G is obtained by the projection error calculation formula i To the jth target 3D Gaussian ellipsoid G j Projection error and the j-th target 3D Gaussian ellipsoid G j To the i-th target 3D Gaussian ellipsoid G i Projection error
[0023] The projection error and projection error The average value of the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss L g_primitives ;
[0024] The projection error calculation formula is:
[0025]
[0026] Among them, p i and p j They are target perspective images I i and target perspective image I j The pixel coordinates of the 2D pixels corresponding to the same 3D point, {p i ,p j} is the target perspective image I i and target perspective image I j A pair of matching rays {r i ,r j}Corresponding pixel coordinates, ray r i Collect target perspective image I for the camera i When the three-dimensional point points to the camera focus, the ray rj Collect target perspective image I for the camera j The ray from the three-dimensional point to the camera focus, p i→j (μ' i ) is p i Projected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, p j→i (μ' j ) is p j Projected to the target perspective image I i The corresponding 2D projection coordinates of the 2D image plane, μ' i =o i +z i d i , μ' j =o j +z j d j , o i Represents the i-th target 3D Gaussian ellipsoid G i The origin of the coordinate system on the corresponding two-dimensional image plane, z i Represents the i-th target 3D Gaussian ellipsoid G i The depth of the corresponding observation point in three-dimensional space, d i Represents the i-th target 3D Gaussian ellipsoid G i The corresponding unit direction vector on the two-dimensional image plane, o j Represents the j-th target 3D Gaussian ellipsoid G j The origin of the coordinate system on the corresponding two-dimensional image plane, z j Represents the j-th target 3D Gaussian ellipsoid G j The depth of the corresponding observation point in three-dimensional space, d j Represents the j-th target 3D Gaussian ellipsoid G j The corresponding unit direction vector on the two-dimensional image plane, the target perspective image I i is the i-th target 3D Gaussian ellipsoid G i The corresponding target perspective image, target perspective image I j is the jth target 3D Gaussian ellipsoid G j The corresponding target perspective image.
[0027] Optionally, optimize the target 3D Gaussian ellipsoids of all Gaussian matching pairs, including:
[0028] Calculate the geometric loss L for each Gaussian matching pair separately render_g ;
[0029] Based on the calculated geometric loss L render_g The average value of all Gaussian position losses L g_primitivesThe average value and all luminosity losses L p Calculate the final error;
[0030] Iteratively update the parameters of the target 3D Gaussian ellipsoid based on the final error, and return to calculate the geometric loss L for each Gaussian matching pair separately. render_g The target 3D Gaussian ellipsoid in each Gaussian matching pair when the iteration termination condition is met is used as the optimized target 3D Gaussian ellipsoid.
[0031] Optionally, calculate the geometric loss L for each Gaussian matching pair render_g ,include:
[0032] The target 3D Gaussian ellipsoid G' of the Gaussian matching pair i and target 3D Gaussian ellipsoid G′ j Projection to the 2D image plane;
[0033] Get the depth value D in the 2D image plane depth_i (p i ) and depth value D depth_j (p j );
[0034] The depth value D depth_i (p i ) is transformed into 3D space and the position v in 3D space is obtained i , and the depth value D depth_j (p j ) is transformed into 3D space and the position v in 3D space is obtained j ;
[0035] The target 3D Gaussian ellipsoid G' is obtained by calculating the geometric error formula i To the target 3D Gaussian ellipsoid G' j Geometric error and target 3D Gaussian ellipsoid G' j To the target 3D Gaussian ellipsoid G' i Geometric error
[0036] The geometric error and geometric errors The average value of the geometric loss L of the Gaussian matching pair render_g .
[0037] Optionally, the geometric error is calculated as:
[0038]
[0039] Among them, p i→j (v' i ) is v iProjected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, R i is the rotation matrix of the observation camera under the target perspective i, D i (p i ) is D depth_i (p i ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective i, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle i, t i is the translation vector of the observation camera under the target perspective i, p j→i (v' j ) is v j Projected to the target perspective image I i The corresponding two-dimensional projection coordinates of the 2D image plane, R j is the rotation matrix of the observation camera under the target perspective j, D j (p j ) is D depth_j (p j ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective j, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle j, t j is the translation vector of the observation camera under the target perspective j.
[0040] Optionally, based on the calculated total geometric loss L render_g The average value of all Gaussian position losses L g_primitives The average value and all luminosity losses L p Calculate the final error, including:
[0041] Through the formula Loss = L' p +βL' g_primitives +δL' render_g Calculate the final error;
[0042] Among them, Loss is the final error, L' p is the calculated total luminosity loss L p The sum of β and δ are weight factors, L' g_primitives is the calculated Gaussian position loss L g_primitives The average value, L' render_g is the calculated geometric loss L render_g The average value of .
[0043] The above solution of the present application has the following beneficial effects:
[0044] In an embodiment of the present application, by adopting a strategy of matching light, the correspondence between light and pixels between different perspectives is captured, which helps to optimize the model to maintain a consistent 3D structure under multiple perspectives. In this way, these prior matched light rays are introduced as constraints into the optimization process, thereby guiding the model to avoid geometric conflicts caused by multi-perspective inconsistencies or scale problems when rendering depth. Through the guidance of matching priors, the rendered geometric structure can more accurately reflect the real structure of the scene, thereby avoiding the common geometric distortion and detail loss problems in the case of a small number of perspectives, and improving the rendering effect of 3DGS under sparse perspectives.
[0045] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0047] Figure 1 This is a flowchart of a novel view synthesis method for a three-dimensional scene provided in one embodiment of the present application. DETAILED DESCRIPTION
[0048] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0049] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0050] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0051] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0052] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0053] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0054] To address the current problem of poor rendering performance of 3DGS under sparse view angles, an embodiment of the present application provides a novel view synthesis method for three-dimensional scenes based on matched light. By adopting a matching light strategy, the correspondence between light and pixels between different view angles is captured, helping the optimization model maintain a consistent 3D structure under multiple view angles. In this way, these prior matched light rays are introduced as constraints into the optimization process, thereby guiding the model to avoid geometric conflicts caused by multi-view inconsistencies or scale issues when rendering depth. Through the guidance of matching priors, the rendered geometric structure can more accurately reflect the real structure of the scene, thereby avoiding the common geometric distortion and detail loss problems in the case of a small number of view angles, and improving the rendering effect of 3DGS under sparse view angles.
[0055] The novel view synthesis method of a three-dimensional scene based on matching light provided by the present application is exemplarily described below with reference to specific embodiments.
[0056] like Figure 1 As shown, the novel view synthesis method of a three-dimensional scene based on matching light provided in an embodiment of the present application includes the following steps:
[0057] Step 11: Generate multiple 3D Gaussian ellipsoids based on multiple perspective images of the target scene.
[0058] In some embodiments of the present application, a camera may be used to capture multiple perspective images of a target scene (i.e., image information from multiple perspectives), and then these perspective images, combined with internal and external parameters of the camera, may be used to generate multiple 3D Gaussian ellipsoids.
[0059] Specifically, a semi-positive covariance matrix Σ can be defined based on the input camera view information. The covariance matrix is decomposed into two parts: a scaling matrix S and a rotation matrix R. This allows for more efficient adjustment and optimization of the shape and orientation of the Gaussian primitives (i.e., 3D Gaussian ellipsoids) during the optimization process, so that the geometric structure of each Gaussian primitive maintains reasonable physical constraints in three-dimensional space, thereby ensuring the accuracy and stability of the scene representation: Σ = RSS T R T Then, according to the defined covariance matrix Σ, combined with the Gaussian center vector information and its position x, explicitly parameterizing the 3D Gaussian ellipsoid:
[0060] For a 3D Gaussian ellipsoid, it is necessary to project the 3D Gaussian onto the 2D image plane via the Jacobian matrix of the view transformation matrix W and the affine approximation of the projection transformation J: ∑' = JW∑W T J T , for point-based rendering, the color C of a pixel can be calculated by mixing N ordered Gaussians that overlap with the pixel:
[0061]
[0062] Similarly, the corresponding depth image D is rendered depth :
[0063]
[0064] In the above formula, N is the number of generated 3D Gaussian ellipsoids, c i is the color value of the i-th Gaussian distribution (i.e., 3D Gaussian ellipsoid), α' i is the opacity of the i-th Gaussian distribution, ranging from [0,1], which is used to measure the contribution of the distribution to the final pixel color. is the cumulative transmittance, which represents the remaining part that is not completely covered when the i-th Gaussian is covered and blocked by the previously generated Gaussian j, d i is the depth value of the i-th observed object, which represents the distance from the observation point to the observed object.
[0065] Step 12: randomly select multiple target perspective images from the multiple perspective images.
[0066] In some embodiments of the present application, a limited number of M training perspectives can be selected by setting a random seed. (i.e., M target perspective images are selected) to simulate the imaging process under sparse perspective.
[0067]
[0068] Among them, RandomSubset refers to the process of randomly selecting a part of data from a larger data set as a subset, and V is the Vth perspective image of the target scene.
[0069] Step 13: Screening out multiple target 3D Gaussian ellipsoids from the multiple 3D Gaussian ellipsoids according to the multiple target view angle images.
[0070] In some embodiments of the present application, the photometric loss L between the target view image and the rendered image (the rendered image is the rendering result in the process of generating a 3D Gaussian ellipsoid using the target view image) can be calculated using L1 loss and SSIM loss for each target view image. p ; Then the luminosity loss L p The 3D Gaussian ellipsoid corresponding to the target perspective image that is less than the luminosity loss threshold is used as the target 3D Gaussian ellipsoid.
[0071] In some embodiments of the present application, the formula Calculate the photometric loss L between the target view image and the rendered image p .
[0072] Among them, λ is the weight factor, which can be set to 0.15. The target perspective image I and the rendered image The L1 loss between The target perspective image I and the rendered image The SSIM loss between .
[0073] In order to allow 3DGS to learn a consistent 3D scene structure under sparse view input without SFM (SFM is a method for recovering structure from motion, mainly used for 3D reconstruction) point initialization, after selecting the better Gaussian primitives under sparse view, it is necessary to complete the matching of prior information. The specific steps are as follows:
[0074] (1) Matching is completed through prior information. Matching ray pairs represent the 2D pixel positions corresponding to the same 3D point (referring to a point in the target scene) in different views (i.e., perspective images). These matching rays can be used as key multi-view constraints. Given an image I i and I j A pair of matching rays {r i ,r j}, and its corresponding pixel coordinates are {p i ,p j}, assuming that the surface points intersecting each ray have been calculated to be P i and P j , then we can get: P i =P j ;
[0075] (2) Using the camera's intrinsic parameter matrix {K i ,K j} and the external parameter matrix {[R i ,t i ],[R j ,t j ]}, project the three-dimensional surface point onto another two-dimensional image plane to obtain the corresponding projected pixel coordinates p i→j (P i ):
[0076]
[0077] Get pixel coordinates p j =p i→j , similarly, we have p i =p j→i .
[0078] Among them, r i Refers to a given image I i The corresponding light, r j Refers to a given image I j The corresponding light, K i is the internal parameter matrix of the observation camera under the target perspective i, K j is the intrinsic parameter matrix of the observation camera under the target perspective j, R i is the rotation matrix of the observation camera under the target perspective i, t i is the translation vector of the observation camera under the target perspective i, R j is the rotation matrix of the observation camera under the target perspective j, t j is the translation vector of the observation camera under the target perspective j, p i→j is the viewing angle i (which can be understood as the image I captured by the camera at viewing angle i) i ) to the perspective j (which can be understood as the image I taken by the camera at the perspective j j )’s two-dimensional pixel plane projection pixel coordinates, p j→i is the viewing angle j (which can be understood as the image I captured by the camera at the viewing angle j) j ) to the viewing angle i (which can be understood as the image I taken by the camera at the viewing angle i) i )'s two-dimensional pixel plane projection pixel coordinates.
[0079] In the matching prior, the positions of matching rays precisely indicate regions that are visible in at least two views. These multi-view visible regions are crucial for 3D scene reconstruction because they provide valuable information about the scene's geometric structure. When overlapping regions exist between different views, the model can leverage these regions to enforce consistency constraints, effectively optimizing and aligning the 3D structure.
[0080] In step 14, based on the matching rays, a plurality of Gaussian matching pairs are selected from the plurality of target 3D Gaussian ellipsoids.
[0081] The Gaussian matching pair includes two target 3D Gaussian ellipsoids from a plurality of target 3D Gaussian ellipsoids.
[0082] In some embodiments of the present application, the i-th target 3D Gaussian ellipsoid G among the multiple target 3D Gaussian ellipsoids may be i and the j-th target 3D Gaussian ellipsoid G j , calculate the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss; if the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j The Gaussian position loss of the i-th target 3D Gaussian ellipsoid G is less than the Gaussian position loss threshold. i and the j-th target 3D Gaussian ellipsoid G j As Gaussian matching pairs, where i = 1, ..., Q, j = 1, ..., Q, i ≠ j, and Q is the number of target 3D Gaussian ellipsoids.
[0083] Specifically, calculate the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j The calculation method of the Gaussian position loss can be: the i-th target 3D Gaussian ellipsoid G is calculated by the projection error calculation formula i To the jth target 3D Gaussian ellipsoid G j Projection error and the j-th target 3D Gaussian ellipsoid G j To the i-th target 3D Gaussian ellipsoid G i Projection error Then the projection error and projection error The average value of the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss L g_primitives .
[0084] The above projection error calculation formula is:
[0085]
[0086] Among them, p i and p j They are target perspective images I i and target perspective image I j The pixel coordinates of the 2D pixels corresponding to the same 3D point, {p i ,p j} is the target perspective image I i and target perspective image I j A pair of matching rays {r i ,r j}Corresponding pixel coordinates, ray r i Collect target perspective image I for the camera i When the three-dimensional point points to the camera focus, the ray r j Collect target perspective image I for the camera j The ray from the three-dimensional point to the camera focus, p i→j (μ' i ) is p i Projected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, p j→i (μ' j ) is p j Projected to the target perspective image I i The corresponding 2D projection coordinates of the 2D image plane, μ' i =o i +z i d i , μ' j =o j +z j d j , o i Represents the i-th target 3D Gaussian ellipsoid G i The origin of the coordinate system on the corresponding two-dimensional image plane, z i Represents the i-th target 3D Gaussian ellipsoid G i The depth of the observation point in the corresponding three-dimensional space (the value of the z axis), that is, the distance from the observation point to the projection plane, d i Represents the i-th target 3D Gaussian ellipsoid G i The corresponding unit direction vector on the two-dimensional image plane defines the direction of the projection transformation, o j Represents the j-th target 3D Gaussian ellipsoid G j The origin of the coordinate system on the corresponding two-dimensional image plane, z j Represents the j-th target 3D Gaussian ellipsoid G jThe depth of the observation point in the corresponding three-dimensional space (the value of the z axis), that is, the distance from the observation point to the projection plane, d j Represents the j-th target 3D Gaussian ellipsoid G j The corresponding unit direction vector on the two-dimensional image plane defines the direction of the projection transformation, the target perspective image I i is the i-th target 3D Gaussian ellipsoid G i The corresponding target perspective image, target perspective image I j is the jth target 3D Gaussian ellipsoid G j The corresponding target perspective image.
[0087] The projection error is and The derivation process is explained.
[0088] In addition to the common unstructured Gaussian primitives used to recover the visible background area in a single view, the novel view synthesis method for a 3D scene in the embodiments of this application also introduces ray-based Gaussian primitives (i.e., Gaussian matching pairs). These primitives are bound to matching rays and optimized using the geometric relationship between the rays. The specific implementation scheme is as follows:
[0089] (1) Initialize with ray-based Gaussian primitives and bind them to matching rays to complete Gaussian initialization and densification. Assume there are N pairs of matching rays For the kth pair of matching rays, initialize N pairs of ray-based Gaussian primitives For the kth Gaussian matching pair, the position of the ray-based Gaussian basis element μ' is defined as: μ' = o + zd, and the average amplitude of the view space position gradient is used to identify the underreconstructed area, where o represents the origin of the coordinate system on the corresponding two-dimensional image plane, z represents the depth of the observation point in the three-dimensional space (the value of the z-axis), that is, the distance from the observation point to the projection plane, and d represents the unit direction vector on the corresponding two-dimensional image plane, which defines the direction of the projection transformation.
[0090] (2) Optimize the position of the original Gaussian: For image I i and I j A pair of matching rays {r i ,r j}, we get a pair of bound Gaussian basis elements {G i ,G j}, their positions in 3D space are μ' i =o i +z i d i and μ' j =o j +z j d j. Thus we get the two-dimensional projection coordinates from image i to image j: p i→j (μ' i ), the two-dimensional projection coordinates from j to i: p j→i (μ' j ). Therefore, the projection error of this pair of Gaussian basis elements can be obtained as:
[0091]
[0092] Step 15: Optimize the target 3D Gaussian ellipsoids in all Gaussian matching pairs.
[0093] In some embodiments of the present application, the specific implementation method of optimizing the target 3D Gaussian ellipsoid in all Gaussian matching pairs is:
[0094] Step 15.1, calculate the geometric loss L for each Gaussian matching pair separately render_g .
[0095] It should be noted that the geometric loss L of each Gaussian matching pair render_g The calculation method is the same, where the geometric loss L for any Gaussian matching pair is render_g Specifically, the geometric loss L of the Gaussian matching pair is render_g The calculation process is:
[0096] The target 3D Gaussian ellipsoid G' of the Gaussian matching pair i and target 3D Gaussian ellipsoid G′ j Projection to the 2D image plane;
[0097] Get the depth value D in the 2D image plane depth_i (p i ) and depth value D depth_j (p j );D depth_i (p i ) is the target 3D Gaussian ellipsoid G' i The corresponding depth value in the 2D image plane, D depth_j (p j ) is the target 3D Gaussian ellipsoid G′ j The depth value corresponding to the 2D image plane can be calculated according to the depth image D depth (i.e. the depth value of the 3D Gaussian ellipsoid);
[0098] The depth value D depth_i (p i ) is transformed into 3D space and the position v in 3D space is obtained i , and the depth value D depth_j (p j) is transformed into 3D space and the position v in 3D space is obtained j ;
[0099] The target 3D Gaussian ellipsoid G' is obtained by calculating the geometric error formula i To the target 3D Gaussian ellipsoid G' j Geometric error and target 3D Gaussian ellipsoid G' j To the target 3D Gaussian ellipsoid G' i Geometric error
[0100] The geometric error and geometric errors The average value of the geometric loss L of the Gaussian matching pair render_g .
[0101] The geometric error calculation formula is:
[0102]
[0103] In the above formula, p i→j (v' i ) is v i Projected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, R i is the target view angle i (i.e., the camera captures the target view angle image I i The rotation matrix of the observation camera under the viewing angle (time), D i (p i ) is D depth_i (p i ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective i, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle i, t i is the translation vector of the observation camera under the target perspective i, p j→i (v′ j ) is v j Projected to the target perspective image I i The corresponding two-dimensional projection coordinates of the 2D image plane, R j is the target view angle j (i.e., the camera captures the target view angle image I j The rotation matrix of the observation camera under the viewing angle (time), D j (p j ) is D depth_j (p j ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective j, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle j, t j is the translation vector of the observation camera under the target perspective j.
[0104] Step 15.2, based on the calculated geometric loss L render_g The average value of all Gaussian position losses L g_primitives The average value and all luminosity losses L p Calculate the final error. Specifically, the formula Loss = L' p +βL' g_primitives +δL' render_g Calculate the final error.
[0105] Among them, Loss is the final error, L' p is the calculated total luminosity loss L p The sum of β and δ are weight factors, L' g_primitives is the calculated Gaussian position loss L g_primitives The average value, L' render_g is the calculated geometric loss L render_g The average value of .
[0106] Step 15.3, based on the final error, adjust the parameters of the target 3D Gaussian ellipsoid (such as learning rate Ir, spatial position o of the 3D Gaussian ellipsoid, scale d i (indicates the size or diffusion range of Gaussian), color C and opacity α (controls the degree of light occlusion), etc.) are iteratively updated and returned to calculate the geometric loss L for each Gaussian matching pair separately. render_g The step (i.e., returning to step 15.1) is performed, and the target 3D Gaussian ellipsoid in each Gaussian matching pair when the iteration termination condition is met is used as the optimized target 3D Gaussian ellipsoid.
[0107] In some embodiments of the present application, the above-mentioned iteration termination condition may be that the final error is less than an error threshold, or the number of iterations reaches a preset number.
[0108] During training, β is set to 1.0 to ensure that the various losses are balanced during the optimization process. To prevent the model from falling into a suboptimal solution at the beginning of training, δ is first set to 0 and then gradually increased to 0.4 after 1000 iterations to gradually introduce more constraints. In the first 1000 iterations, to ensure that the Gaussian primitives can converge to the optimal position, a caching strategy is adopted, that is, the Gaussian primitive position loss (L) is saved at each iteration. g_primitives ) is the smallest position. In addition, since there may be unmatched light pairs in the matching prior, we further filter out those positions where the loss L g_primitivesGaussian primitives with a value greater than a certain threshold η. During the optimization process, ray-based Gaussian primitives will not be pruned to ensure the consistency of the model structure.
[0109] For example, the number of iterations can be set to 3000. At the beginning of training, the learning rate of the learnable distance factor z is set to 0.1, and gradually reduced to 1.5×10 -6 , in order to fine-tune the position of the Gaussian primitives.
[0110] It is worth mentioning that by introducing these error terms, all aspects of the scene can be optimized more comprehensively. Photometric loss is mainly used to ensure the consistency of color and brightness between the generated image and the real image, ensuring high quality of the rendering effect. Gaussian position loss helps optimize the spatial position of the Gaussian basis units, ensuring their reasonable distribution in three-dimensional space, thereby improving geometric consistency. Rendering geometry loss further constrains the geometric structure of the rendering result to keep it consistent under changes in perspective, especially under sparse perspective conditions, to prevent rendering problems caused by overfitting or geometric distortion. Combining these three losses can better balance geometric structure and texture details, and improve the stability and accuracy of few-view novel view synthesis.
[0111] In step 16, the optimized multiple target 3D Gaussian ellipsoids are fused to obtain a novel view of the target scene.
[0112] In summary, this application addresses the key challenges in generating novel views from a small number of perspectives and proposes an effective solution that can maintain the consistency of scene structure and improve rendering quality under sparse view conditions. This solution is mainly reflected in the following aspects:
[0113] 1. Introduction and optimization of matching priors:
[0114] By introducing matching prior information and leveraging the correspondence between light and pixels across viewpoints, the geometric structure of the 3D scene is optimized. This matching prior helps ensure that the model maintains a consistent 3D structure under sparse viewpoints, avoiding geometric artifacts or missing details caused by sparse viewpoints and inconsistent multi-viewpoints. The core value of this strategy lies in its ability to effectively avoid geometric distortion issues under a single viewpoint or a small number of viewpoints, particularly scale and geometric inconsistencies that can occur in depth estimation and multi-view fusion.
[0115] 2. Mixed representation of Gaussian primitives and light constraints:
[0116] To address the high interdependence between the shape and position of Gaussian primitives, a hybrid Gaussian representation is designed. In addition to traditional unstructured Gaussian primitives, ray-based Gaussian primitives are introduced. These primitives can be bound to matching rays and optimized along the ray direction. This method precisely constrains the position of Gaussian primitives to the real 3D geometric surface through ray matching and optimization, thereby improving the geometric accuracy and consistency of the rendered results.
[0117] 3. Double optimization strategy:
[0118] A dual optimization strategy is employed, simultaneously optimizing the position and shape of Gaussian primitives. This approach not only improves the detail and quality of rendered images, but also ensures that the spatial distribution of Gaussian primitives is consistent with the scene structure. This avoids overfitting and texture blurring, particularly in conditions with a small number of viewpoints. This dual optimization strategy ensures geometric stability in sparse viewpoint conditions and significantly improves rendering performance.
[0119] 4. Seamless connection from matching priors to deep optimization:
[0120] By leveraging pre-trained matching models (e.g., geometric constraints on matching light pairs), the limitations of traditional methods that rely on SFM points or random initialization are avoided, enabling effective scene reconstruction without a high-quality initial point cloud. This allows the framework to operate stably in data-scarce or complex scenes, with better adaptability.
[0121] The core of the invention lies in the introduction of an innovative strategy of matching prior information and optimizing Gaussian primitives with light constraints, which solves the common problems of geometric inconsistency, detail loss and texture blur under a small number of viewing angles. At the same time, through a dual optimization strategy, it ensures a significant improvement in the accuracy and quality of the rendering results.
[0122] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A novel view synthesis method for a three-dimensional scene based on matching light, characterized in that: include: Generate multiple 3D Gaussian ellipsoids based on multiple perspective images of the target scene; Randomly selecting a plurality of target perspective images from the plurality of perspective images; screening out a plurality of target 3D Gaussian ellipsoids from the plurality of 3D Gaussian ellipsoids according to the plurality of target perspective images; Based on the matching light, a plurality of Gaussian matching pairs are screened out from the plurality of target 3D Gaussian ellipsoids; the Gaussian matching pairs include two target 3D Gaussian ellipsoids from the plurality of target 3D Gaussian ellipsoids; Optimize the target 3D Gaussian ellipsoid in all Gaussian matching pairs; fusing the optimized multiple target 3D Gaussian ellipsoids to obtain a novel view of the target scene; The step of selecting a plurality of target 3D Gaussian ellipsoids from the plurality of 3D Gaussian ellipsoids according to the plurality of target perspective images includes: For each target perspective image, the photometric loss L between the target perspective image and the rendered image is calculated using the L1 loss and the SSIM loss. p ; The rendered image is a rendering result of a process of generating a 3D Gaussian ellipsoid using the target perspective image; The luminosity loss L p The 3D Gaussian ellipsoid corresponding to the target perspective image that is less than the light loss threshold is used as the target 3D Gaussian ellipsoid; The step of selecting a plurality of Gaussian matching pairs from the plurality of target 3D Gaussian ellipsoids based on the matching light rays includes: For the i-th target 3D Gaussian ellipsoid G among the multiple target 3D Gaussian ellipsoids i and the j-th target 3D Gaussian ellipsoid G j , calculate the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss; i = 1, ..., Q, j = 1, ..., Q, i ≠ j, Q is the number of target 3D Gaussian ellipsoids; If the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j The Gaussian position loss of the i-th target 3D Gaussian ellipsoid G is less than the Gaussian position loss threshold. i and the j-th target 3D Gaussian ellipsoid G j as Gaussian matched pairs; The optimization of the target 3D Gaussian ellipsoid in all Gaussian matching pairs includes: Calculate the geometric loss L for each Gaussian matching pair separately render_g ; Based on the calculated geometric loss L render_g The average value of all Gaussian position losses L g_primitives The average value and all luminosity losses L p Calculate the final error; Iteratively update the parameters of the target 3D Gaussian ellipsoid based on the final error, and return to execute the geometric loss L calculated for each Gaussian matching pair render_g The target 3D Gaussian ellipsoid in each Gaussian matching pair when the iteration termination condition is met is used as the optimized target 3D Gaussian ellipsoid.
2. The novel view synthesis method for a 3D scene according to claim 1, characterized in that: The calculating the photometric loss between the target perspective image and the rendered image using the L1 loss and the SSIM loss includes: By formula Calculate the photometric loss L between the target perspective image and the rendered image p ; Among them, λ is the weight factor, The target perspective image I and the rendered image The L1 loss between The target perspective image I and the rendered image The SSIM loss between .
3. The novel view synthesis method for a 3D scene according to claim 1, characterized in that: The calculation of the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss, including: The i-th target 3D Gaussian ellipsoid G is obtained by the projection error calculation formula i To the jth target 3D Gaussian ellipsoid G j Projection error and the j-th target 3D Gaussian ellipsoid G j To the i-th target 3D Gaussian ellipsoid G i Projection error The projection error and projection error The average value of the i-th target 3D Gaussian ellipsoid G i and the j-th target 3D Gaussian ellipsoid G j Gaussian position loss L g_primitives ; The projection error calculation formula is: Among them, p i and p j They are target perspective images I i and target perspective image I j The pixel coordinates of the 2D pixels corresponding to the same 3D point, {p i ,p j } is the target perspective image I i and target perspective image I j A pair of matching rays {r i ,r j }Corresponding pixel coordinates, ray r i Collect target perspective image I for the camera i When the three-dimensional point points to the ray of the camera focus, the ray r j Collect target perspective image I for the camera j When the three-dimensional point points to the ray of the camera focus, p i→j (μ' i ) is p i Projected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, p j→i (μ′ j ) is p j Projected to the target perspective image I i The corresponding 2D projection coordinates of the 2D image plane, μ' i =o i +z i d i , μ' j =o j +z j d j , o i Represents the i-th target 3D Gaussian ellipsoid G i The origin of the coordinate system on the corresponding two-dimensional image plane, z i Represents the i-th target 3D Gaussian ellipsoid G i The depth of the observation point in the corresponding three-dimensional space, d i Represents the i-th target 3D Gaussian ellipsoid G i The corresponding unit direction vector on the two-dimensional image plane, o j Represents the j-th target 3D Gaussian ellipsoid G j The origin of the coordinate system on the corresponding two-dimensional image plane, z j Represents the j-th target 3D Gaussian ellipsoid G j The depth of the corresponding observation point in three-dimensional space, d j Represents the j-th target 3D Gaussian ellipsoid G j The corresponding unit direction vector on the two-dimensional image plane, the target perspective image I i is the i-th target 3D Gaussian ellipsoid G i The corresponding target perspective image, target perspective image I j is the jth target 3D Gaussian ellipsoid G j The corresponding target perspective image.
4. The novel view synthesis method for a 3D scene according to claim 3, characterized in that: The geometric loss L is calculated for each Gaussian matching pair render_g ,include: The target 3D Gaussian ellipsoid G' of the Gaussian matching pair i and target 3D Gaussian ellipsoid G' j Projection to the 2D image plane; Get the depth value D in the 2D image plane depth_i (p i ) and depth value D depth_j (p j ); The depth value D depth_i (p i ) is transformed into 3D space and the position v in 3D space is obtained i , and the depth value D depth_j (p j ) is transformed into 3D space and the position v in 3D space is obtained j ; The target 3D Gaussian ellipsoid G' is obtained by calculating the geometric error formula i To the target 3D Gaussian ellipsoid G' j Geometric error and target 3D Gaussian ellipsoid G' j To the target 3D Gaussian ellipsoid G' i Geometric error The geometric error and geometric errors The average value of the geometric loss L of the Gaussian matching pair render_g .
5. The novel view synthesis method for a 3D scene according to claim 4, characterized in that: The geometric error calculation formula is: Among them, p i→j (v' i ) is v i Projected to the target perspective image I j The corresponding two-dimensional projection coordinates of the 2D image plane, R i is the rotation matrix of the observation camera under the target perspective i, D i (p i ) is D depth_i (p i ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective i, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle i, t i is the translation vector of the observation camera under the target perspective i, p j→i (v′ j ) is v j Projected to the target perspective image I i The corresponding two-dimensional projection coordinates of the 2D image plane, R j is the rotation matrix of the observation camera under the target perspective j, D j (p j ) is D depth_j (p j ), is the inverse matrix of the intrinsic parameter matrix of the observation camera under the target perspective j, is the homogeneous coordinate of the pixel coordinate on the two-dimensional image plane corresponding to the target view angle j, t j is the translation vector of the observation camera under the target perspective j.
6. The novel view synthesis method for a 3D scene according to claim 1, characterized in that: All geometric losses L calculated based on render_g The average value of all Gaussian position losses L g_primitives The average value and all luminosity losses L p Calculate the final error, including: Through the formula Loss = L' p +βL' g_primitives +δL' render_g Calculate the final error; Among them, Loss is the final error, L' p is the calculated total luminosity loss L p The sum of β and δ are weight factors, L' g_primitives is the calculated Gaussian position loss L g_primitives The average value, L' render_g is the calculated geometric loss L render_g The average value of .
Citation Information
Patent Citations
Rapid high-precision dense reconstruction method and system based on 3D Gaussian rasterization
CN118314271A
Cutter three-dimensional scene reconstruction method and system based on 3DGS
CN118918237A