Method and apparatus for training three-dimensional gaussian splatting model, three-dimensional reconstruction method and apparatus, and storage medium and program product

WO2026200674A1PCT designated stage Publication Date: 2026-10-01MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084483
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-19
Publication Date
2026-10-01

Smart Images

  • Figure CN2026084483_01102026_PF_FP_ABST
    Figure CN2026084483_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for training a three-dimensional Gaussian splatting model, a three-dimensional reconstruction method and apparatus, and a storage medium and a program product. The method for training a three-dimensional Gaussian splatting model comprises: obtaining a reference image pair of a target scene; inputting a first camera parameter corresponding to a first reference image into a three-dimensional Gaussian splatting model corresponding to the target scene, so as to generate a first depth map corresponding to the first camera parameter, and inputting a second camera parameter corresponding to a second reference image into the three-dimensional Gaussian splatting model, so as to generate a second depth map corresponding to the second camera parameter; on the basis of difference information between the first depth map and the second depth map, determining a depth consistency loss value of the three-dimensional Gaussian splatting model; and on the basis of the depth consistency loss value, updating parameters of the three-dimensional Gaussian splatting model. The present disclosure enables effective elimination of floaters in a training process of a three-dimensional Gaussian splatting model, thereby achieving more stable and efficient effects of 3D reconstruction and new perspective generation.
Need to check novelty before this filing date? Find Prior Art

Description

Training methods, 3D reconstruction methods, devices, storage media, and program products for 3D Gaussian point rendering models.

[0001] This application claims priority to Chinese Patent Application No. 202510352468.5, filed on March 24, 2025, entitled “Training Method, 3D Reconstruction Method, Apparatus, Storage Medium and Program Product for 3D Gaussian Point Drawing Model”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of computer technology, and in particular to a training method for a three-dimensional Gaussian point drawing model, a three-dimensional reconstruction method, a training device for a three-dimensional Gaussian point drawing model, a three-dimensional reconstruction device, a non-volatile computer-readable storage medium, and a computer program product. Background Technology

[0003] With the continuous development of computer vision and 3D reconstruction technologies, novel perspective generation techniques have been widely applied in fields such as film and television special effects, virtual reality, and augmented reality. Among them, 3DGS (3Dimensions Gaussian Splatting), as an efficient novel perspective generation technique, has attracted much attention due to its efficiency and flexibility in handling sparse input data. By using 3D Gaussian points to represent spatial scenes, 3DGS can infer images of unobserved perspectives from multi-frame RGB (Red, Green, Blue) image datasets, providing an effective solution for 3D reconstruction and novel perspective generation.

[0004] However, existing 3DGS technology still faces some challenges in practical applications. One of the most prominent problems is the floater problem, where incorrect Gaussian points in space occlude correct image content, resulting in blurred color blocks in the rendered result, severely affecting image quality and visual effects. This problem not only affects the accuracy of 3D reconstruction but also limits the application of 3DGS technology in scenarios with high image quality requirements. Summary of the Invention

[0005] In view of this, this disclosure provides a training technique for a three-dimensional Gaussian point rendering model.

[0006] According to one aspect of this disclosure, a method for training a three-dimensional Gaussian point rendering model is provided, comprising:

[0007] Obtain a pair of reference images of the target scene, wherein the pair of reference images includes a first reference image and a second reference image;

[0008] The first camera parameters corresponding to the first reference image are input into the three-dimensional Gaussian point rendering model corresponding to the target scene, and a first depth map corresponding to the first camera parameters is generated through the three-dimensional Gaussian point rendering model. The second camera parameters corresponding to the second reference image are input into the three-dimensional Gaussian point rendering model, and a second depth map corresponding to the second camera parameters is generated through the three-dimensional Gaussian point rendering model.

[0009] Based on the difference information between the first depth map and the second depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined;

[0010] The parameters of the 3D Gaussian point rendering model are updated based on the depth consistency loss value.

[0011] In one possible implementation, determining the depth consistency loss value of the 3D Gaussian point rendering model based on the difference information between the first depth map and the second depth map includes:

[0012] The first depth map is projected onto the coordinate system of the second reference image to obtain the first projected depth map corresponding to the first depth map, and the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the first depth difference information between the first projected depth map and the second depth map.

[0013] And / or,

[0014] The second depth map is projected onto the coordinate system of the first reference image to obtain the second projected depth map corresponding to the second depth map. Based on the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0015] In one possible implementation, during the projection process, for any pixel location, the minimum value among all depth values ​​projected to that pixel location is taken as the depth value of that pixel location in the projection depth map.

[0016] In one possible implementation,

[0017] The step of determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the first depth difference information between the first projection depth map and the second depth map includes: determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the first depth difference information between the effective projection area in the first projection depth map and the corresponding area in the second depth map;

[0018] And / or,

[0019] The step of determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the second depth difference information between the second projection depth map and the first depth map includes: determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the second depth difference information between the effective projection area in the second projection depth map and the corresponding area in the first depth map.

[0020] In one possible implementation, the method further includes:

[0021] Determine the color loss value of the three-dimensional Gaussian point rendering model;

[0022] The parameters of the 3D Gaussian point rendering model are updated based on the color loss value.

[0023] In one possible implementation, determining the color loss value of the 3D Gaussian point rendering model includes:

[0024] The three-dimensional Gaussian point rendering model is used to render a first image corresponding to the first camera parameters based on the first opacity parameter; the three-dimensional Gaussian point rendering model is used to render a first improved rendering image corresponding to the first camera parameters based on the first opacity parameter and the second opacity parameter; the color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the first rendering image and the first reference image, and the difference information between the first improved rendering image and the first reference image.

[0025] And / or,

[0026] The second rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter. The second improved rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter and the second opacity parameter. The color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0027] In one possible implementation, the reference image pair satisfies at least two of the following conditions:

[0028] The number of key point matches between the first reference image and the second reference image is greater than or equal to a preset number;

[0029] The difference in rotation angle between the first reference image and the second reference image is within a preset angle range;

[0030] The ratio of the z-axis component of the translation vector between the first reference image and the second reference image to the magnitude of the translation vector is less than or equal to a preset ratio.

[0031] According to another aspect of this disclosure, a three-dimensional reconstruction method is provided, comprising:

[0032] Obtain camera parameters from the target's perspective;

[0033] The camera parameters of the target viewpoint are input into a pre-trained 3D Gaussian point rendering model, and the rendered image corresponding to the target viewpoint is output through the 3D Gaussian point rendering model. The 3D Gaussian point rendering model is trained using the training method of the 3D Gaussian point rendering model.

[0034] According to another aspect of this disclosure, a training apparatus for a three-dimensional Gaussian point rendering model is provided, comprising:

[0035] The acquisition module is used to acquire a pair of reference images of the target scene, wherein the pair of reference images includes a first reference image and a second reference image;

[0036] The first generation module is used to input the first camera parameters corresponding to the first reference image into the three-dimensional Gaussian point rendering model corresponding to the target scene, generate a first depth map corresponding to the first camera parameters through the three-dimensional Gaussian point rendering model, and input the second camera parameters corresponding to the second reference image into the three-dimensional Gaussian point rendering model, generate a second depth map corresponding to the second camera parameters through the three-dimensional Gaussian point rendering model.

[0037] The first determining module is used to determine the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the difference information between the first depth map and the second depth map;

[0038] An update module is used to update the parameters of the 3D Gaussian point rendering model based on the depth consistency loss value.

[0039] In one possible implementation, the first determining module is used to:

[0040] The first depth map is projected onto the coordinate system of the second reference image to obtain the first projected depth map corresponding to the first depth map, and the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the first depth difference information between the first projected depth map and the second depth map.

[0041] And / or,

[0042] The second depth map is projected onto the coordinate system of the first reference image to obtain the second projected depth map corresponding to the second depth map. Based on the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0043] In one possible implementation, during the projection process, for any pixel location, the minimum value among all depth values ​​projected to that pixel location is taken as the depth value of that pixel location in the projection depth map.

[0044] In one possible implementation, the first determining module is used to:

[0045] Based on the first depth difference information between the effective projection area in the first projection depth map and the corresponding area in the second depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0046] And / or,

[0047] The depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the second depth difference information between the effective projection area in the second projection depth map and the corresponding area in the first depth map.

[0048] In one possible implementation, the device further includes:

[0049] The second determining module is used to determine the color loss value of the three-dimensional Gaussian point rendering model;

[0050] An update module is used to update the parameters of the three-dimensional Gaussian point rendering model based on the color loss value.

[0051] In one possible implementation, the second determining module is used to:

[0052] The three-dimensional Gaussian point rendering model is used to render a first image corresponding to the first camera parameters based on the first opacity parameter; the three-dimensional Gaussian point rendering model is used to render a first improved rendering image corresponding to the first camera parameters based on the first opacity parameter and the second opacity parameter; the color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the first rendering image and the first reference image, and the difference information between the first improved rendering image and the first reference image.

[0053] And / or,

[0054] The second rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter. The second improved rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter and the second opacity parameter. The color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0055] In one possible implementation, the reference image pair satisfies at least two of the following conditions:

[0056] The number of key point matches between the first reference image and the second reference image is greater than or equal to a preset number;

[0057] The difference in rotation angle between the first reference image and the second reference image is within a preset angle range;

[0058] The ratio of the z-axis component of the translation vector between the first reference image and the second reference image to the magnitude of the translation vector is less than or equal to a preset ratio.

[0059] According to another aspect of this disclosure, a three-dimensional reconstruction apparatus is provided, comprising:

[0060] The acquisition module is used to acquire camera parameters from the target's viewpoint;

[0061] The second generation module is used to input the camera parameters of the target viewpoint into a pre-trained three-dimensional Gaussian point rendering model, and output the rendered image corresponding to the target viewpoint through the three-dimensional Gaussian point rendering model. The three-dimensional Gaussian point rendering model is trained using a training device for the three-dimensional Gaussian point rendering model.

[0062] According to another aspect of this disclosure, a training apparatus for a three-dimensional Gaussian point drawing model is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described training method for a three-dimensional Gaussian point drawing model.

[0063] According to another aspect of this disclosure, a three-dimensional reconstruction apparatus is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the three-dimensional reconstruction method described above.

[0064] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.

[0065] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0066] In this embodiment of the disclosure, by obtaining a pair of reference images of a target scene, wherein the pair of reference images includes a first reference image and a second reference image, the first camera parameters corresponding to the first reference image are input into a 3D Gaussian point rendering model corresponding to the target scene, and a first depth map corresponding to the first camera parameters is generated by the 3D Gaussian point rendering model. The second camera parameters corresponding to the second reference image are input into the 3D Gaussian point rendering model, and a second depth map corresponding to the second camera parameters is generated by the 3D Gaussian point rendering model. Based on the difference information between the first depth map and the second depth map, the depth consistency loss value of the 3D Gaussian point rendering model is determined, and the parameters of the 3D Gaussian point rendering model are updated according to the depth consistency loss value. This effectively eliminates floating points during the training process of the 3D Gaussian point rendering model, achieves more stable and efficient 3D reconstruction and new perspective generation effects, and improves the quality of new perspective rendered images.

[0067] This disclosure employs a dual-view depth constraint strategy, eliminating the need to rely on monocular depth estimation to obtain supervisory data. Therefore, this disclosure avoids the scale ambiguity and estimation error problems associated with monocular depth estimation, thus better preserving geometric details in the image.

[0068] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0069] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0070] Figure 1a illustrates a schematic diagram of the detail loss problem in 3DGS models in related technologies.

[0071] Figure 1b shows the result of using the GOF point splitting strategy after applying the multi-resolution point rendering (Mip-Splatting) method, which originally did not have the floating point problem.

[0072] Figure 1c shows the result of the original 3DGS halving the threshold of its point splitting strategy.

[0073] Figure 2 shows a flowchart of the training method for the three-dimensional Gaussian point rendering model provided in the embodiments of this disclosure.

[0074] Figure 3a shows a schematic diagram of the first depth map in the training method of the three-dimensional Gaussian point rendering model provided in the embodiments of this disclosure.

[0075] Figure 3b shows a schematic diagram of the second depth map in the training method of the three-dimensional Gaussian point rendering model provided in the embodiments of this disclosure.

[0076] Figure 3c shows a schematic diagram of the second projection depth map in the training method of the three-dimensional Gaussian point rendering model provided in the embodiments of this disclosure.

[0077] Figure 3d shows a schematic diagram of the first projection depth map in the training method of the three-dimensional Gaussian point rendering model provided in the embodiments of this disclosure.

[0078] Figure 4 shows a block diagram of a training device for a three-dimensional Gaussian point rendering model provided in an embodiment of this disclosure.

[0079] Figure 5 is a block diagram of a training device or a 3D reconstruction device 1900 for drawing a 3D Gaussian point model according to an exemplary embodiment. Detailed Implementation

[0080] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0081] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.

[0082] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.

[0083] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0084] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0085] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0086] The standard 3DGS (3Dimensions Gaussian Splatting) model uses a learnable set of 3D Gaussian ellipsoids to represent the scene, denoted as . For a single Gaussian sphere G k , including parameter {α k ,μ k ,q k ,s k ,c k}. Wherein, α k ∈[0,1] represents the Gaussian sphere G k The opacity at the center, G represents the Gaussian sphere. k The coordinates of the center, quaternion Used to represent Gaussian sphere G k rotation matrix Used to represent Gaussian sphere G k diagonal scaling matrix Used to represent Gaussian sphere G k The color. Standard 3DGS typically calculates c using spherical harmonics. k To simplify the problem, we will directly use C. k To perform color parameterization. Gaussian sphere G k At a point in space The opacity at this location is:

[0087] in, For Gaussian sphere G k The covariance matrix is ​​used to control the spatial distribution of the ellipsoid.

[0088] For a ray r(t) = o + td (t ≥ 0), 3DGS renders the color of this ray using a raster method, where o represents the origin of the ray and d represents the direction vector of the ray. Gaussian sphere G k The depth along the light ray r(t) can be approximated by the Gaussian sphere G. k center coordinates μ k Projection on the direction d of the ray Gaussian ball G k center coordinates μk The projection point on the ray r(t) is Let K r A Gaussian sphere intersects with ray r(t), according to... The rendering equation is as follows: (Sorted in ascending order)

[0089] in, and These represent 3D Gaussian sphere scenes. The color and depth rendered on the ray r(t). In this paper, Both refer to the result after rendering.

[0090] A frame of image I contains multiple rays, and its corresponding ray set is denoted as O = {r}. r∈I C gt This represents the image pixel value corresponding to each ray of light. (Based on a 3DGS model) The rendered image is an RGB (Red, Green, Blue) image. and depth map The loss function of the original 3DGS can be expressed as:

[0091] in, L1 loss, L D-SSIM (·) represents the SSIM (Structural Similarity Index Measure) loss.

[0092] 3DGS models are widely used in novel perspective generation tasks due to their ability to efficiently represent spatial structures in a differentiable manner. However, 3DGS models in related technologies suffer from two types of defects: loss of detail and color block artifacts. Figure 1a illustrates the detail loss problem in 3DGS models in related technologies. As shown in Figure 1a, the first type of defect manifests as the inability to reconstruct high-frequency details such as grass, which we can call the blurry problem. This stems from insufficient Gaussian distribution caused by the failure of the 3DGS densification strategy. The second type of defect manifests as occasional color block artifacts in the image, mainly due to the presence of floating semi-transparent Gaussian spheres in space, which we can call the floater problem. The densification strategy proposed using GOF (Gaussian Opacity Fields) can effectively alleviate the first type of problem, but it is ineffective for the second type of problem and may even worsen it. Figure 1b shows the result of using the GOF densification strategy on a multi-resolution mip-splatting method that originally did not have the floater problem. The floater problem can be seen in Figure 1b. Decreasing the point splitting threshold also increases the probability of floating points. Figure 1c shows the result of halving the threshold of the point splitting strategy in the original 3DGS. In summary, whether adopting the point splitting strategy in related techniques or decreasing the point splitting threshold, point splitting becomes more aggressive, thus exacerbating the floating point problem. This also prevents us from achieving better image quality by adding more Gaussian points.

[0093] The floating point problem mainly arises because incorrect Gaussian points in space occlude the correct image content, resulting in blurred color blocks in the rendering result. To solve the floating point problem in 3DGS models, some methods in related technologies (such as 2DGS and Radeon GS) attempt to improve the rendering effect by constraining the rendering normals or increasing the spatial regularization loss, but these methods have failed to fundamentally solve the floating point problem.

[0094] The inventors of this application have discovered that the floating point problem is a fundamental issue in 3DGS differentiable rendering, stemming from the vanishing gradient phenomenon during 3DGS backpropagation. This prevents floating points from being effectively eliminated even in a fully trained model. To illustrate this problem, we use a simplified scenario to explain why related techniques cannot eliminate floating points through training. In this scenario, there are K... f A cluster of floating points composed of Gaussian spheres. Assume there are M light rays passing through this cluster in the dataset. Its color true value is To simplify the problem, we assume that each ray passes through all K... fGiven a Gaussian sphere, and assuming that the scene, excluding floating points, has been fully optimized to match the true values, the rendering equation for light color is:

[0095] We optimize only the opacity α and color c of the Gaussian sphere to minimize the total color error. Theoretically, the globally optimal solution should be Its loss converges to 0. However, our experimental observations reveal that this optimization problem sometimes fails to converge to the global minimum, and this is especially true in K. f Larger or initial This is especially common when the value is high.

[0096] Data analysis revealed that when converging to a local minimum, the alpha blending of the floating point region... Meet the conditions This leads to c f The gradient of approaches zero, thus making and The gradient also vanishes due to the chain rule. We tried various loss functions, including L1 loss, mean squared error (MSE) loss, and cross-entropy loss, but observed this phenomenon in all of them. When the number of floating points K... f When the number of parameters is large, the number of model parameters increases, leading to a significantly higher risk of overfitting. Furthermore, the number of parameters... The sigmoid activation function is used, which is prone to saturation with high initial values, thus slowing down the optimization process. These factors combined make the model more susceptible to getting trapped in local minima.

[0097] This simplified experiment reveals an essential flaw in the 3DGS differentiable rendering optimization equation: color parameters. Over-optimization can easily lead to the model getting trapped in a local minimum. In this case, the opacity parameter can be used to... The geometric optimization was disrupted and failed to converge effectively, resulting in unremovable floating points in the space. Furthermore, both the GOF point splitting strategy and reducing the original point splitting threshold tend to generate more Gaussian spheres, which not only increases the number of floating points but also makes residual color patches even more difficult to eliminate.

[0098] The inventors of this application discovered that the process of constructing a Signed Distance Field (SDF) can be viewed as an aggregation of multi-view ray depths, and the regularization operation of the SDF is essentially a constraint on the consistency of the multi-view depth map. Based on this idea, we experimentally verified that: although the opacity parameter It is still impossible to guarantee complete convergence to 0, but the introduction of a depth loss function can effectively eliminate the Gaussian sphere corresponding to the floating point and significantly reduce its impact on the final image.

[0099] This disclosure provides a training method for a 3D Gaussian point rendering model. By obtaining a pair of reference images of a target scene, including a first reference image and a second reference image, the method inputs first camera parameters corresponding to the first reference image into the 3D Gaussian point rendering model corresponding to the target scene. The 3D Gaussian point rendering model generates a first depth map corresponding to the first camera parameters. Similarly, the method inputs second camera parameters corresponding to the second reference image into the 3D Gaussian point rendering model, generating a second depth map corresponding to the second camera parameters. Based on the difference information between the first and second depth maps, a depth consistency loss value is determined for the 3D Gaussian point rendering model. The parameters of the 3D Gaussian point rendering model are then updated according to the depth consistency loss value. This effectively eliminates floating points during the training process of the 3D Gaussian point rendering model, achieving more stable and efficient 3D reconstruction and new perspective generation effects, and improving the quality of the new perspective rendered image.

[0100] This disclosure employs a dual-view depth constraint strategy, eliminating the need to rely on monocular depth estimation to obtain supervisory data. Therefore, this disclosure avoids the scale ambiguity and estimation error problems associated with monocular depth estimation, thus better preserving geometric details in the image.

[0101] The training method for the three-dimensional Gaussian point rendering model provided in this disclosure will be described in detail below with reference to the accompanying drawings.

[0102] Figure 2 shows a flowchart of the training method for a 3D Gaussian point rendering model provided in an embodiment of this disclosure. In one possible implementation, the execution entity of the training method for the 3D Gaussian point rendering model can be a training device for the 3D Gaussian point rendering model. For example, the training method for the 3D Gaussian point rendering model can be executed by a terminal device, a server, or other electronic devices. The terminal device can be a user equipment (UE), a user terminal, a terminal, or a computing device, etc. In some possible implementations, the training method for the 3D Gaussian point rendering model can be implemented by a processor calling computer-readable instructions stored in memory. As shown in Figure 2, the training method for the 3D Gaussian point rendering model includes steps S21 to S24.

[0103] In step S21, a pair of reference images of the target scene is obtained, wherein the pair of reference images includes a first reference image and a second reference image.

[0104] In step S22, the first camera parameters corresponding to the first reference image are input into the three-dimensional Gaussian point rendering model corresponding to the target scene, and a first depth map corresponding to the first camera parameters is generated through the three-dimensional Gaussian point rendering model. The second camera parameters corresponding to the second reference image are input into the three-dimensional Gaussian point rendering model, and a second depth map corresponding to the second camera parameters is generated through the three-dimensional Gaussian point rendering model.

[0105] In step S23, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the first depth map and the second depth map.

[0106] In step S24, the parameters of the three-dimensional Gaussian point rendering model are updated according to the depth consistency loss value.

[0107] In this embodiment of the disclosure, the target scene can represent a physical space environment that requires 3D reconstruction or the generation of a new perspective. The target scene can consist of a series of 3D objects, surfaces, and backgrounds, and can be captured and reconstructed using multi-view image data. The target scene can be an indoor environment (such as a room, office, studio, etc.) or an outdoor environment (such as a street, building exterior, natural landscape, etc.).

[0108] A reference image for a target scene can represent an image taken from a specific perspective that contains information about the target scene. The reference image can be used to train a 3D Gaussian Splatting (3DGS) model corresponding to the target scene. In some applications, the reference image may also be called a training image, ground truth image, etc., without further elaboration here.

[0109] A reference image pair refers to two images taken from two different perspectives that contain information about the same target scene. The two images in a reference image pair can be called the first reference image and the second reference image. In one example, i and j can be used to represent the frame numbers of the first and second reference images.

[0110] In the embodiments of this disclosure, selecting appropriate reference images is crucial to the training effect of the 3D Gaussian point rendering model, because they need to provide sufficient viewpoint differences and sufficient overlapping areas to ensure that the 3D Gaussian point rendering model can learn the 3D structure and depth information of the target scene.

[0111] In one possible implementation, the reference image pair satisfies at least two of the following conditions: the number of key point matches between the first reference image and the second reference image is greater than or equal to a preset number; the difference in rotation angle between the first reference image and the second reference image is within a preset angle range; and the ratio of the z-axis component of the translation vector between the first reference image and the second reference image to the magnitude of the translation vector is less than or equal to a preset ratio.

[0112] As an example of this implementation, the reference image pair satisfies the following condition: the number of keypoint matches between the first reference image and the second reference image is greater than or equal to a preset number. In one example, the preset number can be 30. In this example, keypoint matching can refer to finding corresponding feature points (such as corner points, edge points, etc.) in the two reference images. These matched points provide the geometric relationship between the two reference images, helping the 3D Gaussian point rendering model understand the 3D structure of the target scene. If the number of matched keypoints is large enough, it indicates that there is sufficient overlap between the two reference images, and the 3D Gaussian point rendering model can more accurately reconstruct the geometric information of the target scene. Therefore, by adopting this example, it is possible to ensure that there is sufficient overlap between the two reference images so that the 3D Gaussian point rendering model can learn the 3D structure of the target scene.

[0113] As an example of this implementation, the reference image pair satisfies the following condition: the difference in rotation angle between the first reference image and the second reference image is within a preset angle range. The difference in rotation angle can refer to the difference in the shooting perspective of the two reference images in the rotation direction. In one example, the preset angle range can be [16°, 60°]. In this example, if the difference in rotation angle between the two reference images is between 16° and 60°, it indicates that the two reference images have a certain difference in perspective, but not excessively large. This moderate difference in perspective helps the 3D Gaussian point rendering model learn the depth information of the target scene. This example ensures that the difference in perspective between the two reference images is moderate, neither lacking sufficient depth information due to an excessively small difference nor lacking sufficient overlap between the reference images due to an excessively large difference.

[0114] As an example of this implementation, the reference image pair satisfies the following condition: the ratio of the z-axis component of the translation vector between the first and second reference images to the magnitude of the translation vector is less than or equal to a preset ratio. Here, the translation vector describes the relative positional change of the two reference images in space. The z-axis component can be related to the depth direction, while the magnitude of the translation vector can represent the total displacement between the two reference images. The ratio in this example can be used to limit the relative motion between the two reference images, avoiding situations where the two reference images in the same reference image pair only move back and forth. For example, the preset ratio can be 0.95. If the ratio of the z-axis component of the translation vector between the first and second reference images to the magnitude of the translation vector is less than or equal to 0.95, it indicates that the motion between the two reference images is not a pure back-and-forth movement. By adopting this example, it is possible to avoid selecting reference image pairs with too small a viewpoint difference, ensuring sufficient viewpoint difference between the reference image pairs, thereby improving the training effect of the 3D Gaussian point rendering model.

[0115] By adopting this implementation method, suitable reference image pairs can be selected so that the 3D Gaussian point rendering model can more effectively learn the 3D structure and depth information of the target scene, thereby improving the training effect and the final rendering quality.

[0116] In this embodiment of the disclosure, camera parameters may include intrinsic parameters (such as focal length and optical center position) and extrinsic parameters (such as camera pose and position). Camera parameters describe the geometric relationship between the camera and the target scene. First camera parameters may represent the camera parameters corresponding to a first reference image, i.e., the camera parameters when acquiring the first reference image. Second camera parameters may represent the camera parameters corresponding to a second reference image, i.e., the camera parameters when acquiring the second reference image.

[0117] In this embodiment, the first camera parameters corresponding to the first reference image can be input into a 3D Gaussian point rendering model corresponding to the target scene, and a first depth map corresponding to the first camera parameters can be generated through the 3D Gaussian point rendering model. The 3D Gaussian point rendering model can calculate the depth value of each pixel based on the input first camera parameters using an internal rendering mechanism (such as ray tracing or rasterization). Furthermore, the 3D Gaussian point rendering model can determine the depth information of each pixel in 3D space based on the parameters of the Gaussian points (such as position, opacity, color, etc.) and the first camera parameters, thereby generating a first depth map corresponding to the first reference image. Similarly, the second camera parameters corresponding to the second reference image can be input into the 3D Gaussian point rendering model, and a second depth map corresponding to the second camera parameters can be generated through the 3D Gaussian point rendering model. The depth map is a two-dimensional array, where each pixel value represents the depth information of that pixel in 3D space. In one example, the first depth map can be... This indicates that the second depth map can be adopted. express.

[0118] In this embodiment of the disclosure, during the training process of the 3D Gaussian point rendering model corresponding to the target scene, the depth consistency loss value can be calculated by comparing the difference between a first depth map and a second depth map. Specifically, these two depth maps correspond to reference images taken from different viewpoints, describing the depth information of the same target scene under different viewpoints. By calculating the difference between these two depth maps (e.g., through pixel-level difference calculation or more complex matching algorithms), the consistency of the depth information reconstructed by the 3D Gaussian point rendering model between different viewpoints can be quantified. If there are significant differences between the depth maps, this may indicate that the depth estimation of the 3D Gaussian point rendering model is inaccurate in some areas, such as the presence of floating points or other geometric distortions. The depth consistency loss value is based on this difference information to evaluate the performance of the 3D Gaussian point rendering model and serves as an optimization objective, guiding the 3D Gaussian point rendering model to adjust its parameters to reduce the depth difference between different viewpoints, thereby improving the accuracy and stability of 3D reconstruction.

[0119] In one possible implementation, determining the depth consistency loss value of the 3D Gaussian point rendering model based on the difference information between the first depth map and the second depth map includes: projecting the first depth map onto the coordinate system of the second reference image to obtain a first projected depth map corresponding to the first depth map, and determining the depth consistency loss value of the 3D Gaussian point rendering model according to the first depth difference information between the first projected depth map and the second depth map; and / or, projecting the second depth map onto the coordinate system of the first reference image to obtain a second projected depth map corresponding to the second depth map, and determining the depth consistency loss value of the 3D Gaussian point rendering model according to the second depth difference information between the second projected depth map and the first depth map.

[0120] In this implementation, a first depth map can be projected onto the coordinate system of a second reference image to obtain a first projected depth map, and a second depth map can be projected onto the coordinate system of the first reference image to obtain a second projected depth map. The projection process involves geometric transformations, including rotation and translation operations, to ensure that the depth values ​​maintain the correct spatial relationships under the new viewpoint. The first projected depth map can represent the depth map obtained by projecting the first depth map onto the coordinate system of the second reference image, and the second projected depth map can represent the depth map obtained by projecting the second depth map onto the coordinate system of the first reference image.

[0121] In this implementation, a first depth difference information can be calculated between the first projected depth map and the second depth map. This first depth difference information represents the depth difference between the first projected depth map and the second depth map. The first depth difference information can be quantized as a pixel-level difference, for example, by calculating the absolute difference (L1 norm) or the squared difference (L2 norm).

[0122] In this implementation, a second depth difference information can be calculated between the second projected depth map and the first depth map. This second depth difference information represents the depth difference between the second projected depth map and the first depth map. The second depth difference information can be quantized as a pixel-level difference, for example, by calculating the absolute difference (L1 norm) or the squared difference (L2 norm).

[0123] In this implementation, the depth consistency loss value can be calculated based on the first depth difference information and / or the second depth difference information.

[0124] As an example of this implementation, the first depth map can be projected onto the coordinate system of the second reference image to obtain a first projected depth map corresponding to the first depth map, and the second depth map can be projected onto the coordinate system of the first reference image to obtain a second projected depth map corresponding to the second depth map. Based on the first depth difference information between the first projected depth map and the second depth map, and the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0125] In 3D space, a floating point is represented as a semi-transparent Gaussian sphere, and its depth is a weighted average of its own depth and the background depth. Therefore, when the depth map is projected onto the coordinate system of another reference image, positions originally affected by the floating point are projected to positions unaffected by it. This process provides the floating point with accurate supervisory information, enabling effective correction.

[0126] Figure 3a shows a schematic diagram of the first depth map in the training method of the 3D Gaussian point rendering model provided in an embodiment of the present disclosure. Figure 3b shows a schematic diagram of the second depth map in the training method of the 3D Gaussian point rendering model provided in an embodiment of the present disclosure. Figure 3c shows a schematic diagram of the second projected depth map in the training method of the 3D Gaussian point rendering model provided in an embodiment of the present disclosure. Figure 3d shows a schematic diagram of the first projected depth map in the training method of the 3D Gaussian point rendering model provided in an embodiment of the present disclosure. The floating point in the upper left corner of the first depth map shown in Figure 3a is projected onto the area marked with a box in the first projected depth map shown in Figure 3d, while the corresponding box area in the second depth map shown in Figure 3b displays the accurate depth value. Therefore, by constraining the depth consistency between the second depth map and the first projected depth map, the floating point can be effectively eliminated.

[0127] In this implementation, the depth consistency loss value serves as the optimization objective, guiding the 3D Gaussian point rendering model to adjust its parameters to reduce depth differences between different viewpoints. This helps improve the accuracy and stability of 3D reconstruction, especially when dealing with floating point problems.

[0128] In one possible implementation, during the projection process, for any pixel location, the minimum value among all depth values ​​projected to that pixel location is taken as the depth value of that pixel location in the projection depth map.

[0129] In this implementation, a depth map of a reference image (e.g., a first depth map) is projected onto the coordinate system of another reference image (e.g., the coordinate system of a second reference image) to obtain a projected depth map (e.g., the first projected depth map). During the projection process, multiple depth values ​​may be projected onto the same pixel location. These depth values ​​may originate from different Gaussian points, and their depth values ​​may differ. In this implementation, to determine the final depth value of the pixel location in the projected depth map, the minimum value among all depth values ​​projected onto that pixel location can be selected. This minimum value represents the nearest point of that pixel location in three-dimensional space.

[0130] Choosing the minimum depth value effectively addresses occlusion issues. In 3D space, closer objects occlude farther objects; therefore, selecting the minimum depth value ensures that the projected depth map reflects the depth information of the foreground object. Furthermore, choosing the minimum depth value improves the consistency of depth maps across different viewpoints, thereby increasing the accuracy of depth consistency loss calculations.

[0131] In one possible implementation, determining the depth consistency loss value of the 3D Gaussian point rendering model based on the first depth difference information between the first projection depth map and the second depth map includes: determining the depth consistency loss value of the 3D Gaussian point rendering model based on the first depth difference information between the effective projection area in the first projection depth map and the corresponding area in the second depth map; and / or, determining the depth consistency loss value of the 3D Gaussian point rendering model based on the second depth difference information between the second projection depth map and the first depth map includes: determining the depth consistency loss value of the 3D Gaussian point rendering model based on the second depth difference information between the effective projection area in the second projection depth map and the corresponding area in the first depth map.

[0132] When projecting one depth map onto the coordinate system of another depth map, the depth values ​​of some areas may be incorrectly projected to other locations, resulting in depth values ​​in these areas being less than the true depth values. These areas can be called shaded areas, invalid projection areas, or ineffective projection areas because they were not properly processed during the projection process, leading to an underestimation of their depth values.

[0133] Because visibility is not considered during projection, the depth value of the shadowed area will be less than the true value. If the depth difference information between the first and second projected depth maps and / or between the second and first projected depth maps is directly used to calculate the depth consistency loss, erroneous supervision signals may be introduced, causing the 3D Gaussian point rendering model to learn incorrect depth information. Therefore, in this implementation, only the depth difference information between the effective projected area in the first projected depth map and the corresponding area in the second depth map (i.e., the first depth difference information), and / or the depth difference information between the effective projected area in the second projected depth map and the corresponding area in the first depth map (i.e., the second depth difference information) can be considered. That is, in this implementation, the depth difference information of non-effectively projected areas can be ignored when calculating the depth consistency loss.

[0134] As an example of this implementation, the depth consistency loss value of the three-dimensional Gaussian point rendering model can be determined based on the first depth difference information between the first projection depth map and the second depth map, and the second depth difference information between the second projection depth map and the first depth map.

[0135] In this implementation, by constraining only the effective projection area, the interference of the shadow part on the depth consistency loss calculation can be avoided, ensuring that the depth consistency loss value can correctly reflect the depth inconsistency problem.

[0136] In one example, the deep consistency loss function can be:

[0137] in, Indicates the effective projection area. This indicates that the first depth map is projected into the coordinate system of the second reference image. This represents the first projected depth map. This indicates that the second depth map is projected onto the coordinate system of the first reference image. This represents the second projected depth map.

[0138] In one possible implementation, the method further includes: determining the color loss value of the three-dimensional Gaussian point rendering model; and updating the parameters of the three-dimensional Gaussian point rendering model based on the color loss value.

[0139] In one example, the loss function corresponding to the 3D Gaussian point rendering model can be: in, This can represent the color loss function corresponding to a 3D Gaussian point rendering model. λ can represent the depth consistency loss function corresponding to the 3D Gaussian point rendering model. consis It can represent the weights of the depth consistency loss function corresponding to the 3D Gaussian point rendering model.

[0140] In this implementation, by introducing color loss values ​​to optimize the parameters of the 3D Gaussian point rendering model, not only can the accuracy of the 3D Gaussian point rendering model in depth estimation be improved, but the color fidelity and visual quality of the rendered image can also be significantly enhanced. This allows the 3D Gaussian point rendering model to better capture and reproduce the color details of objects when dealing with complex scenes, thereby providing more realistic and higher-quality image output in new perspective generation tasks, and enhancing the performance and applicability of the 3D Gaussian point rendering model in practical applications.

[0141] In one possible implementation, determining the color loss value of the 3D Gaussian point rendering model includes: rendering the 3D Gaussian point rendering model based on a first opacity parameter to obtain a first rendered image corresponding to the first camera parameters; rendering the 3D Gaussian point rendering model based on the first opacity parameter and a second opacity parameter to obtain a first improved rendered image corresponding to the first camera parameters; determining the color loss value of the 3D Gaussian point rendering model based on the difference information between the first rendered image and the first reference image, and the difference information between the first improved rendered image and the first reference image; and / or, rendering the 3D Gaussian point rendering model based on the first opacity parameter to obtain a second rendered image corresponding to the second camera parameters, and rendering the 3D Gaussian point rendering model based on the first opacity parameter and the second opacity parameter to obtain a second improved rendered image corresponding to the second camera parameters; determining the color loss value of the 3D Gaussian point rendering model based on the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0142] In this implementation, the first opacity parameter can represent the opacity parameter in the standard 3D Gaussian point rendering model, while the second opacity parameter is a newly introduced auxiliary opacity parameter. Introducing the second opacity parameter helps the 3D Gaussian point rendering model better handle semi-transparent objects.

[0143] In this implementation, the rendering formula based solely on the first opacity parameter can be: The formula for rendering based on the first and second opacity parameters can be: Where, α k This represents the first opacity parameter, α′. k This represents the second opacity parameter.

[0144] In this implementation, the model can be drawn using 3D Gaussian points based solely on the first opacity parameter α. k Rendering is performed to obtain the first rendered image corresponding to the first camera parameters. This image can be drawn using a 3D Gaussian point model based on the first opacity parameter α. k Second opacity parameter α′ k Rendering is performed to obtain the first improved rendered image corresponding to the first camera parameters. The model can be drawn using 3D Gaussian points based solely on the first opacity parameter α. k Rendering is performed to obtain the second rendered image corresponding to the second camera parameters. This image can be generated by drawing a 3D Gaussian point model based on the first opacity parameter α. k Second opacity parameter α′k Rendering is performed to obtain the second improved rendered image corresponding to the second camera parameters.

[0145] As an example of this implementation, the color loss value of the 3D Gaussian point rendering model can be determined based on the difference information between the first rendered image and the first reference image, the difference information between the first improved rendered image and the first reference image, the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0146] In one example, it can be based on Determine the color loss value. Among them, Indicates a reference image. This indicates that the value is based solely on the first opacity parameter α. k The rendered image obtained through rendering. Indicates based on the first opacity parameter α k Second opacity parameter α′ k The improved rendered image obtained through rendering.

[0147] In this implementation, by introducing dual opacity parameters for rendering, the 3D Gaussian point rendering model can more accurately handle the rendering of semi-transparent objects, thereby improving the quality and visual effects of the rendered image. Specifically, the rendering result based on the first opacity parameter can handle opaque objects in the scene, while the improved rendering result combined with the second opacity parameter can better handle the transparency and color blending issues of semi-transparent objects. By comparing the differences between these two rendering results and a reference image to calculate the color loss value, the 3D Gaussian point rendering model can finely adjust the parameters, optimize the rendering effect, and make the final generated image closer to the real scene in terms of color and transparency, thus improving the model's reconstruction and rendering performance in complex scenes.

[0148] In one possible implementation, the step of inputting the first camera parameters corresponding to the first reference image into the 3D Gaussian point rendering model corresponding to the target scene, generating a first depth map corresponding to the first camera parameters through the 3D Gaussian point rendering model, and inputting the second camera parameters corresponding to the second reference image into the 3D Gaussian point rendering model, generating a second depth map corresponding to the second camera parameters through the 3D Gaussian point rendering model, includes: inputting the first camera parameters corresponding to the first reference image into the 3D Gaussian point rendering model corresponding to the target scene, generating a first depth map corresponding to the first camera parameters through rendering based on the first opacity parameter and the second opacity parameter through the 3D Gaussian point rendering model, and inputting the second camera parameters corresponding to the second reference image into the 3D Gaussian point rendering model, generating a second depth map corresponding to the second camera parameters through rendering based on the first opacity parameter and the second opacity parameter through the 3D Gaussian point rendering model.

[0149] Since σ′(x)≤σ(x) always holds true, therefore we always have in, This represents the depth map rendered based solely on the first opacity parameter. This represents the depth map rendered based on the first and second opacity parameters. Since the depth of the floating points is less than the normal value, it is based on the depth map. Calculating the depth consistency loss value can eliminate floating points.

[0150] In another possible implementation, the step of inputting the first camera parameters corresponding to the first reference image into the 3D Gaussian point rendering model corresponding to the target scene, generating a first depth map corresponding to the first camera parameters through the 3D Gaussian point rendering model, and inputting the second camera parameters corresponding to the second reference image into the 3D Gaussian point rendering model, generating a second depth map corresponding to the second camera parameters through the 3D Gaussian point rendering model, includes: inputting the first camera parameters corresponding to the first reference image into the 3D Gaussian point rendering model corresponding to the target scene, generating a first depth map corresponding to the first camera parameters through the 3D Gaussian point rendering model based solely on a first opacity parameter, and inputting the second camera parameters corresponding to the second reference image into the 3D Gaussian point rendering model, generating a second depth map corresponding to the second camera parameters through the 3D Gaussian point rendering model based solely on the first opacity parameter.

[0151] This disclosure also provides a three-dimensional reconstruction method, including: acquiring camera parameters of a target viewpoint; inputting the camera parameters of the target viewpoint into a pre-trained three-dimensional Gaussian point rendering model; and outputting a rendered image corresponding to the target viewpoint through the three-dimensional Gaussian point rendering model, wherein the three-dimensional Gaussian point rendering model is trained using the training method for the three-dimensional Gaussian point rendering model.

[0152] The training method and 3D reconstruction method of the 3D Gaussian point rendering model provided in this disclosure can be applied to the fields of AI (Artificial Intelligence)-CV (Computer Vision)-3D (3D) reconstruction, new perspective generation, 3DGS and other technical fields, and are not limited thereto.

[0153] The training method for the three-dimensional Gaussian point rendering model provided in this embodiment is illustrated below through a specific application scenario.

[0154] In this application scenario, a pair of reference images for the target scene can be obtained. This pair can include a first reference image and a second reference image. The first camera parameters corresponding to the first reference image can be input into a 3D Gaussian point rendering model corresponding to the target scene. The model is then rendered based solely on the first opacity parameter to obtain a first rendered image corresponding to the first camera parameters. Furthermore, the model is rendered based on both the first and second opacity parameters to obtain a first improved rendered image and a first depth map corresponding to the first camera parameters. Similarly, the second camera parameters corresponding to the second reference image can be input into the 3D Gaussian point rendering model. The model is then rendered based solely on the first opacity parameter to obtain a second rendered image corresponding to the second camera parameters. Finally, the model is rendered based on both the first and second opacity parameters to obtain a second improved rendered image and a second depth map corresponding to the second camera parameters.

[0155] The first depth map can be projected onto the coordinate system of the second reference image to obtain the first projected depth map. Similarly, the second depth map can be projected onto the coordinate system of the first reference image to obtain the second projected depth map. Based on the first depth difference information between the first and second projected depth maps, and the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the 3D Gaussian point rendering model can be determined.

[0156] The color loss value of the 3D Gaussian point rendering model can be determined based on the difference information between the first rendered image and the first reference image, the difference information between the first improved rendered image and the first reference image, the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0157] The parameters of the 3D Gaussian point rendering model can be updated based on the depth consistency loss value and the color loss value.

[0158] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0159] In addition, this disclosure also provides a training device for a three-dimensional Gaussian point drawing model, a three-dimensional reconstruction device, a non-volatile computer-readable storage medium, and a computer program product. All of the above can be used to implement any of the training methods or three-dimensional reconstruction methods for a three-dimensional Gaussian point drawing model provided in this disclosure. The corresponding technical solutions and technical effects can be found in the relevant descriptions in the method section, and will not be repeated here.

[0160] Figure 4 shows a block diagram of a training apparatus for a three-dimensional Gaussian point rendering model provided in an embodiment of this disclosure. As shown in Figure 4, the training apparatus for the three-dimensional Gaussian point rendering model includes:

[0161] The acquisition module 41 is used to acquire a pair of reference images of the target scene, wherein the pair of reference images includes a first reference image and a second reference image;

[0162] The first generation module 42 is used to input the first camera parameters corresponding to the first reference image into the three-dimensional Gaussian point drawing model corresponding to the target scene, generate a first depth map corresponding to the first camera parameters through the three-dimensional Gaussian point drawing model, and input the second camera parameters corresponding to the second reference image into the three-dimensional Gaussian point drawing model, generate a second depth map corresponding to the second camera parameters through the three-dimensional Gaussian point drawing model.

[0163] The first determining module 43 is used to determine the depth consistency loss value of the three-dimensional Gaussian point drawing model based on the difference information between the first depth map and the second depth map.

[0164] The update module 44 is used to update the parameters of the three-dimensional Gaussian point rendering model based on the depth consistency loss value.

[0165] In one possible implementation, the first determining module 43 is used to:

[0166] The first depth map is projected onto the coordinate system of the second reference image to obtain the first projected depth map corresponding to the first depth map, and the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the first depth difference information between the first projected depth map and the second depth map.

[0167] And / or,

[0168] The second depth map is projected onto the coordinate system of the first reference image to obtain the second projected depth map corresponding to the second depth map. Based on the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0169] In one possible implementation, during the projection process, for any pixel location, the minimum value among all depth values ​​projected to that pixel location is taken as the depth value of that pixel location in the projection depth map.

[0170] In one possible implementation, the first determining module 43 is used to:

[0171] Based on the first depth difference information between the effective projection area in the first projection depth map and the corresponding area in the second depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

[0172] And / or,

[0173] The depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the second depth difference information between the effective projection area in the second projection depth map and the corresponding area in the first depth map.

[0174] In one possible implementation, the device further includes:

[0175] The second determining module is used to determine the color loss value of the three-dimensional Gaussian point rendering model;

[0176] The update module 44 is used to update the parameters of the three-dimensional Gaussian point drawing model based on the color loss value.

[0177] In one possible implementation, the second determining module is used to:

[0178] The three-dimensional Gaussian point rendering model is used to render a first image corresponding to the first camera parameters based on the first opacity parameter; the three-dimensional Gaussian point rendering model is used to render a first improved rendering image corresponding to the first camera parameters based on the first opacity parameter and the second opacity parameter; the color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the first rendering image and the first reference image, and the difference information between the first improved rendering image and the first reference image.

[0179] And / or,

[0180] The second rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter. The second improved rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter and the second opacity parameter. The color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

[0181] In one possible implementation, the reference image pair satisfies at least two of the following conditions:

[0182] The number of key point matches between the first reference image and the second reference image is greater than or equal to a preset number;

[0183] The difference in rotation angle between the first reference image and the second reference image is within a preset angle range;

[0184] The ratio of the z-axis component of the translation vector between the first reference image and the second reference image to the magnitude of the translation vector is less than or equal to a preset ratio.

[0185] This disclosure also provides a three-dimensional reconstruction apparatus, including:

[0186] The acquisition module is used to acquire camera parameters from the target's viewpoint;

[0187] The second generation module is used to input the camera parameters of the target viewpoint into a pre-trained three-dimensional Gaussian point rendering model, and output the rendered image corresponding to the target viewpoint through the three-dimensional Gaussian point rendering model. The three-dimensional Gaussian point rendering model is trained using a training device for the three-dimensional Gaussian point rendering model.

[0188] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation and technical effects can be referred to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0189] This disclosure also provides a training apparatus for a three-dimensional Gaussian point drawing model, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the above-described training method for a three-dimensional Gaussian point drawing model.

[0190] This disclosure also provides a three-dimensional reconstruction apparatus, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described three-dimensional reconstruction method.

[0191] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.

[0192] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.

[0193] Figure 5 is a block diagram illustrating a training apparatus or 3D reconstruction apparatus 1900 for a 3D Gaussian point rendering model according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server or terminal device. Referring to Figure 5, apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions executable by processing component 1922, such as application programs. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, processing component 1922 is configured to execute instructions to perform the methods described above.

[0194] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Or similar.

[0195] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.

[0196] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0197] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.

[0198] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information of computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this disclosure.

[0199] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0200] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0201] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0203] Computer program products can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0204] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0205] If the technical solution of this disclosure involves personal information, the product applying the technical solution of this disclosure has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this disclosure involves sensitive personal information, the product applying the technical solution of this disclosure has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to indicate that the user has entered the scope of personal information collection and that personal information will be collected. If the user voluntarily enters the collection scope, it is deemed to have consented to the collection of their personal information; or on the personal information processing device, with clear signs / information informing the user of the personal information processing rules, authorization is obtained from the user through pop-up information or by asking the user to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0206] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A training method for a three-dimensional Gaussian point rendering model, characterized in that, include: Obtain a pair of reference images of the target scene, wherein the pair of reference images includes a first reference image and a second reference image; The first camera parameters corresponding to the first reference image are input into the three-dimensional Gaussian point rendering model corresponding to the target scene, and a first depth map corresponding to the first camera parameters is generated through the three-dimensional Gaussian point rendering model. The second camera parameters corresponding to the second reference image are input into the three-dimensional Gaussian point rendering model, and a second depth map corresponding to the second camera parameters is generated through the three-dimensional Gaussian point rendering model. Based on the difference information between the first depth map and the second depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined; The parameters of the 3D Gaussian point rendering model are updated based on the depth consistency loss value.

2. The method according to claim 1, characterized in that, The step of determining the depth consistency loss value of the 3D Gaussian point rendering model based on the difference information between the first depth map and the second depth map includes: The first depth map is projected onto the coordinate system of the second reference image to obtain the first projected depth map corresponding to the first depth map, and the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined based on the first depth difference information between the first projected depth map and the second depth map. And / or, The second depth map is projected onto the coordinate system of the first reference image to obtain the second projected depth map corresponding to the second depth map. Based on the second depth difference information between the second projected depth map and the first depth map, the depth consistency loss value of the three-dimensional Gaussian point rendering model is determined.

3. The method according to claim 2, characterized in that, During the projection process, for any pixel location, the minimum value among all depth values ​​projected to that pixel location is taken as the depth value of that pixel location in the projection depth map.

4. The method according to claim 2, characterized in that, The step of determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the first depth difference information between the first projection depth map and the second depth map includes: determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the first depth difference information between the effective projection area in the first projection depth map and the corresponding area in the second depth map; And / or, The step of determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the second depth difference information between the second projection depth map and the first depth map includes: determining the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the second depth difference information between the effective projection area in the second projection depth map and the corresponding area in the first depth map.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Determine the color loss value of the three-dimensional Gaussian point rendering model; The parameters of the 3D Gaussian point rendering model are updated based on the color loss value.

6. The method according to claim 5, characterized in that, Determining the color loss value of the 3D Gaussian point rendering model includes: The three-dimensional Gaussian point rendering model is used to render a first image corresponding to the first camera parameters based on the first opacity parameter; the three-dimensional Gaussian point rendering model is used to render a first improved rendering image corresponding to the first camera parameters based on the first opacity parameter and the second opacity parameter; the color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the first rendering image and the first reference image, and the difference information between the first improved rendering image and the first reference image. And / or, The second rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter. The second improved rendered image corresponding to the second camera parameters is obtained by rendering the three-dimensional Gaussian point rendering model based on the first opacity parameter and the second opacity parameter. The color loss value of the three-dimensional Gaussian point rendering model is determined based on the difference information between the second rendered image and the second reference image, and the difference information between the second improved rendered image and the second reference image.

7. The method according to any one of claims 1 to 4, characterized in that, The reference image pair satisfies at least two of the following conditions: The number of key point matches between the first reference image and the second reference image is greater than or equal to a preset number; The difference in rotation angle between the first reference image and the second reference image is within a preset angle range; The ratio of the z-axis component of the translation vector between the first reference image and the second reference image to the magnitude of the translation vector is less than or equal to a preset ratio.

8. A three-dimensional reconstruction method, characterized in that, include: Obtain camera parameters from the target's perspective; The camera parameters of the target viewpoint are input into a pre-trained 3D Gaussian point rendering model, and the rendered image corresponding to the target viewpoint is output through the 3D Gaussian point rendering model. The 3D Gaussian point rendering model is trained using the method described in any one of claims 1 to 7.

9. A training device for a three-dimensional Gaussian point rendering model, characterized in that, include: The acquisition module is used to acquire a pair of reference images of the target scene, wherein the pair of reference images includes a first reference image and a second reference image; The first generation module is used to input the first camera parameters corresponding to the first reference image into the three-dimensional Gaussian point rendering model corresponding to the target scene, generate a first depth map corresponding to the first camera parameters through the three-dimensional Gaussian point rendering model, and input the second camera parameters corresponding to the second reference image into the three-dimensional Gaussian point rendering model, generate a second depth map corresponding to the second camera parameters through the three-dimensional Gaussian point rendering model. The first determining module is used to determine the depth consistency loss value of the three-dimensional Gaussian point rendering model based on the difference information between the first depth map and the second depth map; An update module is used to update the parameters of the 3D Gaussian point rendering model based on the depth consistency loss value.

10. A three-dimensional reconstruction device, characterized in that, include: The acquisition module is used to acquire camera parameters from the target's viewpoint; The second generation module is used to input the camera parameters of the target viewpoint into a pre-trained three-dimensional Gaussian point rendering model, and output the rendered image corresponding to the target viewpoint through the three-dimensional Gaussian point rendering model, wherein the three-dimensional Gaussian point rendering model is trained using the device described in claim 9.

11. A training device for a three-dimensional Gaussian point rendering model, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

12. A three-dimensional reconstruction device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 8.

13. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

14. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.