Three-dimensional Gaussian splash optimization method based on depth information
By replacing spherical harmonics with RGB values and using Sigmoid function to limit RGB values, combined with depth guidance operation, the three-dimensional Gaussian splashing method is optimized, and the floating artifacts and incomplete geometric structure problems caused by spherical harmonic coefficient are solved, and the accuracy of three-dimensional reconstruction and the effectiveness of depth guidance are improved.
Patent Information
- Application Number
- CN202510026553.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
AI Technical Summary
In the existing three-dimensional reconstruction technology, spherical harmonic coefficients are prone to floating artifacts and incomplete geometric structures during rendering, affecting the effectiveness of depth guidance and the accuracy of three-dimensional reconstruction.
By replacing the spherical harmonics of each Gaussian point with RGB values and limiting the RGB values to a preset range according to the Sigmoid function, combined with depth guidance operations, the three-dimensional Gaussian splashing method is optimized to improve the accuracy of three-dimensional reconstruction.
Improves the accuracy of depth guidance and the accuracy of three-dimensional reconstruction, reduces the emergence of floating artifacts and incomplete geometric structures, and enhances the interpretability of each Gaussian point in the Gaussian model.
Smart Images

Figure CN120070697A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of computer technology, and in particular, relates to a three-dimensional Gaussian splash optimization method based on depth information. Background Art
[0002] With the rapid development and increasing popularity of artificial intelligence technology, the technology of automatically generating three-dimensional objects using artificial intelligence technology has emerged. Through the three-dimensional object generation technology, a variety of three-dimensional objects can be generated, providing unlimited possibilities for innovation in the fields of games, film and television, and architecture. Therefore, how to achieve three-dimensional reconstruction more accurately is a technical problem that needs to be solved urgently. Summary of the invention
[0003] The purpose of the present invention is to provide a three-dimensional Gaussian splash optimization method based on depth information to solve the technical problems existing in the related art.
[0004] According to a first aspect of an embodiment of the present disclosure, a three-dimensional Gaussian splash optimization method based on depth information is provided, the method comprising: Acquire a multi-view image, and extract a plurality of Gaussian points of the multi-view image, wherein the model corresponding to the plurality of Gaussian points is a Gaussian model; The spherical harmonics of each Gaussian point are replaced with RGB values, and the RGB value of each Gaussian point is limited within a preset range according to a Sigmoid function to obtain an initial model; A depth-guided operation is performed on the initial model to obtain a three-dimensional object.
[0005] Optionally, performing a depth-guided operation on each of the restricted Gaussian points to obtain a three-dimensional object includes: Predicting the depth information of the multi-view images to obtain a predicted depth, and determining a rendering depth of the Gaussian model; A depth-guided operation is performed based on the predicted depth and the rendered depth.
[0006] Optionally, performing a depth guidance operation based on the predicted depth and the rendered depth includes: A Spearman correlation coefficient is determined according to the predicted depth and the rendered depth, and a depth-guided operation is performed based on the Spearman correlation coefficient.
[0007] Optionally, the method further comprises: An adaptive encryption operation is performed on the Gaussian model to obtain a three-dimensional object, wherein the adaptive encryption operation is used to perform densification processing on a target area of the Gaussian model.
[0008] Optionally, performing an adaptive encryption operation on the Gaussian model to obtain the three-dimensional object includes: Determine the target area, where the target area includes an under-reconstructed area and an over-reconstructed area; Encrypt the Gaussian points within the target area based on a first gradient threshold, and encrypt the Gaussian points in other areas except the target area based on a second gradient threshold to obtain the three-dimensional object, where the first gradient threshold is less than the second gradient threshold.
[0009] Optionally, the target area includes a background area. For the Gaussian model, performing an adaptive encryption operation to obtain a three-dimensional object includes: Encrypt the Gaussian points within the background area based on a first gradient threshold, and encrypt the Gaussian points in other areas except the background area based on a second gradient threshold to obtain the three-dimensional object.
[0010] Optionally, extracting a plurality of Gaussian points from the multi-view images includes: Extract a plurality of point clouds in three-dimensional space from the multi-view images, and represent the plurality of point clouds according to the plurality of Gaussian points.
[0011] According to a second aspect of the embodiments of the present disclosure, there is provided a three-dimensional Gaussian splash optimization device based on depth information, including: An acquisition module configured to acquire multi-view images and extract a plurality of Gaussian points of the multi-view images, where the model corresponding to the plurality of Gaussian points is a Gaussian model; A restriction module configured to replace the spherical harmonics of each Gaussian point with RGB values, limit the RGB values of each Gaussian point within a preset range according to the Sigmoid function, and perform a depth guidance operation on the initial model to obtain a three-dimensional object.
[0012] In a third aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any item of the first aspect are implemented.
[0013] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the executable instructions to implement the steps of the method described in any item of the first aspect of the present disclosure.
[0014] In the method and apparatus provided by the exemplary embodiments of the present disclosure, multi-view images are obtained, and a plurality of Gaussian points of the multi-view images are extracted. Among them, the model corresponding to the plurality of Gaussian points is a Gaussian model. On this basis, the spherical harmonics of each Gaussian point are replaced with RGB values, and the RGB values of each Gaussian point are limited within a preset range according to the Sigmoid function, and a depth guidance operation is performed on each of the limited Gaussian points to obtain a three-dimensional object. Since the depth guidance operation is performed after the limiting operation, the accuracy of depth guidance in the embodiments of the present disclosure is relatively high, so that the accuracy of three-dimensional reconstruction can be improved to a certain extent.
[0015] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section.
[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings: Figure 1 is an exemplary diagram showing different degrees of incompleteness for different spherical harmonic orders according to an exemplary embodiment.
[0018] Figure 2 is a flowchart of a three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment.
[0019] Figure 3 is a flowchart of another three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment.
[0020] Figure 4 is a comparison exemplary diagram of adopting and not adopting an adaptive encryption operation in another three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment.
[0021] Figure 5 is a schematic structural diagram of a three-dimensional Gaussian splash optimization device based on depth information according to an exemplary embodiment.
[0022] Figure 6 is an electronic device for implementing a three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment.
[0023] Figure 7 is another electronic device for implementing a three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment. Detailed Implementation Modes
[0024] The following details the specific implementation modes of the present disclosure with reference to the accompanying drawings. It should be understood that the specific implementation modes described herein are only for the purpose of illustrating and explaining the present disclosure, and are not used to limit the present disclosure.
[0025] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0026] It should be understood that the steps recorded in the method implementation modes of the present disclosure can be executed in different orders and / or in parallel. In addition, the method implementation modes may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0027] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the implementation modes of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the users and the authorization of the users should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0032] 3DGS (3D Gaussian Splatting) is a method for 3D scene representation and rendering. It can model the scene by using a set of Gaussian ellipsoids, thus achieving efficient rendering. 3DGS combines the advantages of implicit representation and explicit representation, which can not only provide high-quality rendering effects but also maintain real-time rendering speed.
[0033] However, the spherical harmonics (SH) used by 3DGS to recover the view-dependent effects will have the problem of task ill-posedness to some extent. When the rendering color of individual splats (Gaussian distributions / Gaussian points) exceeds the range, the gradient backpropagation of the spherical harmonics will fail. This will hinder the effectiveness of depth guidance, mainly because these "negative Gaussians" may not be correctly placed according to the depth, as detailed in Figure 1 shown. Based on Figure 1 it can be known that the larger the spherical harmonics coefficient, the more serious the floating artifacts are. That is to say, the larger the spherical harmonics coefficient may lead to an increase in floating artifacts.
[0034] Here, the spherical harmonics coefficient can be a set of orthogonal basis functions defined on the sphere, which provides a method for efficiently representing and processing signals on the sphere. That is, the spherical harmonics coefficient plays an important role in graphics rendering and visual analysis.
[0035] The color characteristics of Gaussian points are mainly expressed by spherical harmonics SH, which mainly takes the viewing angle as the input and generates the output color. Higher-order spherical harmonics allow Gaussian points to produce different colors in different viewing directions. Therefore, higher-order spherical harmonics are more likely to cause floating artifacts and are more likely to produce incomplete geometric structures under limited viewing angles. Although the reconstruction of the viewing directions within the training view can achieve high-quality rendering, the overall reconstructed scene may be affected due to its incomplete geometric structure.
[0036] To solve the above problems, the embodiments of the present disclosure show an optimization method for 3D Gaussian splatting based on depth information. This method can improve the accuracy of 3D reconstruction by replacing the spherical harmonics value with the RGB value.
[0037] Figure 2 is a schematic flowchart of an optimization method for 3D Gaussian splatting based on depth information shown according to an exemplary embodiment. As shown in Figure 2 shown, the optimization method for 3D Gaussian splatting based on depth information can at least include the following steps: In step S110, multi-view images are obtained, and multiple Gaussian points of the multi-view images are extracted. The model corresponding to the multiple Gaussian points is a Gaussian model.
[0038] In the embodiments of the present disclosure, the multi-view images may be a set of images captured from multiple different positions or angles, and the set of images may jointly record different views of the same scene or object. Exemplarily, the multi-view images may be captured by multiple fixed-position cameras, moving cameras, a swarm of drones, or 360-degree panoramic cameras, etc.
[0039] The 3D Gaussian splash (SDGS) may consist of a point cloud. Specifically, in the embodiments of the present disclosure, the three-dimensional structure of the scene may be recovered from the multi-view images based on Structure from Motion (SfM) to generate a sparse point cloud. On this basis, each point in the point cloud is converted into a 3D Gaussian point, that is, color features, opacity, and a covariance matrix are assigned to each point cloud. Each Gaussian point may include attributes such as position, covariance matrix, opacity, or spherical harmonic coefficients, and these attributes are used to determine the "shape" of the point.
[0040] In other words, after the multi-view images are captured, the embodiments of the present disclosure may extract multiple point clouds in the three-dimensional space from the multi-view images. On this basis, multiple Gaussian points are used to represent the multiple point clouds. The multiple Gaussian points of the multi-view images may be three-dimensional Gaussian points, and each Gaussian point may include color features, opacity, and covariance matrix, etc. In the embodiments of the present disclosure, the model corresponding to the multiple Gaussian points may be a Gaussian model.
[0041] Optionally, after the multiple Gaussian points of the multi-view images are extracted, the embodiments of the present disclosure may project each three-dimensional Gaussian point onto the image plane, that is, perform a projection transformation to reorder the Gaussian points according to the distance between each Gaussian point and the image plane. On this basis, a differential rasterization process is performed, and based on this process, the contribution of each Gaussian point to the pixel from front to back can be calculated. Here, the color of each pixel can be generated by the following equation: ; ; where c i refers to the color calculated from the N Gaussian points of the camera ray on the pixel, which may be an RGB value, and c p may be the color information of the rendered image; α i is the opacity of the current Gaussian point, may be the opacity of the previous Gaussian point, and the opacity decreases as the distance from the center point increases. α i may be determined by the covariance and opacity comprehensively. x and μ i may be the coordinates projected into the same coordinate system, and Σ i -1It is the projected 2D covariance matrix. By combining Gaussian splashing with parallel processing of the rasterization process, a real-time and realistic rendering result can be obtained. For example, within about 30,000 times of photometric loss training, the embodiments of the present disclosure can obtain a real-time and realistic rendering result.
[0042] In step S120, the spherical harmonics of each Gaussian point are replaced with RGB values, the RGB values of each Gaussian point are limited within a preset range according to the Sigmoid function, and a depth guidance operation is performed on each Gaussian point after limitation to obtain a three-dimensional object.
[0043] As known from the above introduction, the Gaussian point can include spherical harmonics SH, and the color feature of the Gaussian point can be represented by the spherical harmonics SH. Here, the spherical harmonics SH can take the viewing angle (θ, φ) as the input and generate the output color c. The higher the order of the spherical harmonics SH, the greater the color change of the Gaussian point in different viewing directions, so it is easier to generate floating-point artifacts, or it is very likely to reconstruct an incomplete geometric structure in the case of limited views.
[0044] Although high-quality rendering can be achieved in the viewing directions within the training views, the entire reconstructed scene may be affected by its incomplete geometric structure. Among them, the quantization results of different SH orders can be shown in Table 1.
[0045] In Table 1, PSNR (Peak Signal-to-Noise Ratio) measures the quality of image restoration by comparing the original image and the distorted image. The higher the PSNR value, the better the image quality.
[0046] SSIM (Structural Similarity Index Measure) is an index for measuring the visual similarity of two images. The value of SSIM ranges from -1 to 1, and when the SSIM value is 1, it means that the two images are exactly the same.
[0047] LPIPS ((Learned Perceptual Image Patch Similarity) is an image quality evaluation index based on deep learning. It can evaluate the image quality by comparing the similarity of images in the feature space. The lower the LPIPS value, the better the image quality.
[0048] NIQE (Natural Image Quality Evaluator) is a reference - free image quality assessment method. This method does not require the original image as a reference. Instead, it predicts the image quality by analyzing the natural scene statistical characteristics of the image. The lower the NIQE score, the more natural and better the quality of the image.
[0049] Based on Table 1, it is known that the embodiments of the present disclosure introduce NIQE, which allows reference - free image quality assessment. Generally, artifacts similar to floating - point artifacts that disrupt the picture structure will cause the NIQE value to decrease. As shown in Table 1, the larger the order of the spherical harmonics, the greater the corresponding NIQE may be.
[0050] Table 1
[0051] Based on the above findings, the embodiments of the present disclosure can randomly select views around the scene and check the rendering image quality related to different spherical harmonic orders.
[0052] Although higher spherical harmonic orders can achieve higher rendering quality, using higher spherical harmonic orders on uncovered views often gives a worse structure, resulting in a higher NIQE value, indicating that the rendered image is not natural. Figure 1 An uncovered angle of the "home" scene is shown, which shows different degrees of incompleteness for different spherical harmonic orders, and more floating - point artifacts can be seen in higher - order spherical harmonic models. To enhance the effectiveness of scene geometry reconstruction, the embodiments of the present disclosure can use RGB values to replace spherical harmonics (spherical harmonic functions / spherical harmonic coefficients).
[0053] That is to say, the embodiments of the present disclosure can use RGB values as the color features of each of the said Gaussian points. In summary, after obtaining multiple Gaussian points, the embodiments of the present disclosure can replace the spherical harmonics of each Gaussian point with RGB values, that is, directly use RGB values as the color features of each Gaussian point.
[0054] To enhance the reconstruction effect of the scene geometry structure, the embodiments of the present disclosure can replace spherical harmonics with RGB values, that is, use RGB values to replace the spherical harmonic coefficients of Gaussian points. This is similar to setting the order of spherical harmonics to zero and omitting the coefficients and biases. In other words, the embodiments of the present disclosure can improve the effectiveness of depth guidance by reducing the order of spherical harmonics.
[0055] As an alternative approach, after replacing the spherical harmonics of the Gaussian points with RGB values, the embodiments of the present disclosure can limit the RGB values of each Gaussian point within a preset range according to the Sigmoid function. By using the Sigmoid function to project the RGB values of the Gaussian points into an appropriate range, not only can the effectiveness of depth guidance be improved, but also the interpretability of each Gaussian point in the Gaussian model can be increased, thereby strengthening the geometric and visual significance of each independent Gaussian point.
[0056] The embodiments of the present disclosure can replace the spherical harmonics of each Gaussian point with RGB values and backpropagate the gradient to the spherical harmonics through photometric loss. However, when the Gaussian points are independently rendered, RGB values that are out of range and have unclear geometric meanings may be obtained. For example, some Gaussian points do not contribute to the rendering result in terms of the color elements in RGB.
[0057] Exemplarily, Gaussian points with negative RGB values during rendering are usually restricted to zero and will no longer receive gradients afterwards. These Gaussian points will not contribute to the rendering result in the restricted color, but cannot be pruned because other color elements may still contribute. Additionally, for Gaussian points with rendered RGB values greater than 1, it may cause RGB value overflow during the rendering process. This situation is particularly common in scenes with high exposure values during training.
[0058] To improve the training efficiency and the accuracy of the scene geometric structure, the embodiments of the present disclosure can limit the RGB values of the Gaussian points within a preset range by applying the sigmoid function, where the preset range can be (0, 1), i.e., c r,g,b = σ(f r,g,b ). Here, f can be the color feature value retained in the Gaussian model, and σ(·) can be the sigmoid function, so that the rendering result can be fixed within the specified range to avoid RGB values that are out of range and have unclear geometric meanings.
[0059] In summary, the embodiments of the present disclosure directly use RGB values as the color features of each Gaussian point and use sigmoid projection into an appropriate range. In this way, not only can the effectiveness of depth guidance be enhanced, but also the interpretability of each independent Gaussian point in the Gaussian model can be increased, thereby strengthening the geometric and visual significance of each independent Gaussian point.
[0060] The geometric properties of the Gaussian points can be rendered from the Gaussian model after appropriate processing. Specifically, the embodiments of the present disclosure can generate predicted depth based on a monocular prediction model, and this predicted depth can be used as pseudo ground truth. The depth of each Gaussian point can be the distance between the Gaussian point and the image plane, i.e., the z value of the camera projection.
[0061] It should be noted that in the process of restricting the RGB values of each Gaussian point within a preset range, the embodiments of the present disclosure can simultaneously perform a depth guidance operation, which can be based on a Gaussian model.
[0062] The present disclosure obtains multi-view images and extracts multiple Gaussian points of the multi-view images. Among them, the model corresponding to the multiple Gaussian points is a Gaussian model. On this basis, the spherical harmonics of each Gaussian point are replaced with RGB values, and the RGB values of each Gaussian point are restricted within a preset range according to the Sigmoid function, and a depth guidance operation is performed on each restricted Gaussian point to obtain a three-dimensional object. Since this depth guidance operation is performed after the restriction operation, the accuracy of depth guidance in the embodiments of the present disclosure is relatively high, so that the accuracy of three-dimensional reconstruction can be improved to a certain extent.
[0063] Figure 3 is a flowchart showing another three-dimensional Gaussian splash optimization method based on depth information according to an exemplary embodiment, as Figure 3 shown. The three-dimensional Gaussian splash optimization method based on depth information can at least include the following steps: In step S210, multi-view images are obtained, and multiple Gaussian points of the multi-view images are extracted. The model corresponding to the multiple Gaussian points is a Gaussian model.
[0064] Among them, the specific implementation manner of step S210 has been described in detail in the above embodiments, and will not be elaborated here.
[0065] In step S220, the spherical harmonics of each Gaussian point are replaced with RGB values, and the RGB values of each Gaussian point are restricted within a preset range according to the Sigmoid function, and a depth guidance operation is performed on each restricted Gaussian point to obtain a three-dimensional object.
[0066] As an optional manner, in the process of performing the depth guidance operation, the embodiments of the present disclosure can predict the depth information of the multi-view images to obtain a predicted depth. At the same time, the embodiments of the present disclosure can determine the rendering depth of the Gaussian model. On this basis, the depth guidance operation is performed based on the predicted depth and the rendering depth.
[0067] Among them, the predicted depth can be obtained based on a monocular depth prediction model, that is, by inputting the multi-view images into the monocular depth prediction model, the predicted depth can be obtained. As known from the above introduction, each Gaussian point can include attributes such as position and covariance matrix. Among them, the position can include the x coordinate, y coordinate, and z coordinate. Among them, the z coordinate can be used to represent the depth information of the Gaussian point. The embodiments of the present disclosure can project the position information corresponding to the initial model, that is, project the position information onto the camera position coordinates, and then the projected z information can be used as the rendering depth.
[0068] The specific rendering depth can be calculated based on the following formula: ; where, is the rendering depth, can be the weight ratio of the current Gaussian point, which can be uncertain according to the covariance matrix and opacity of the Gaussian point, can be the ratio of the previous Gaussian point, can be the distance between the current Gaussian point and the camera plane, which can be obtained by projecting the z-axis information in the position information.
[0069] On this basis, based on the predicted depth and the rendering depth, a depth guidance operation can be performed. Specifically, the embodiments of the present disclosure can determine the Spearman correlation coefficient according to the predicted depth and the rendering depth, and perform the depth guidance operation based on the Spearman correlation coefficient. Through this guidance, the Spearman correlation coefficient can be continuously increased until it approaches 1.
[0070] Since the estimation result provided by the monocular depth prediction model is relative depth, there may be a difference from the point cloud structure. Therefore, the embodiments of the present disclosure can use the correlation coefficient between the rendering depth and the predicted depth as a guide.
[0071] Exemplarily, the embodiments of the present disclosure can use the Spearman correlation coefficient to perform the depth guidance operation. The main reason is that its order in the predicted depth map is more reliable than the relative size. Given the ability of the Gaussian point to construct a simple geometric structure, the depth guidance through the Spearman correlation coefficient can remove outliers and floating points that cannot be distinguished only by visual guidance.
[0072] By using the depth guidance operation proposed by the embodiments of the present disclosure, the geometric information of the reconstructed three-dimensional object / three-dimensional model can be made more accurate.
[0073] In step S230, for the Gaussian model, an adaptive encryption operation is performed to obtain a three-dimensional object.
[0074] As an alternative, after obtaining the Gaussian model, the embodiments of the present disclosure can perform an adaptive encryption operation on the Gaussian model to obtain a three-dimensional object. Among them, the adaptive encryption operation is used to densify the target area of the Gaussian model.
[0075] Here, the densification process is mainly used to adaptively adjust the density of the Gaussian points to better capture the geometric details and appearance of the scene. Specifically, during the execution of the adaptive encryption operation, the embodiments of the present disclosure can identify areas with insufficient or excessive reconstruction. On this basis, an encryption operation is performed on this area.
[0076] That is to say, the embodiments of the present disclosure can determine a target area on a three-dimensional object, which can be a defective area or an area to be reconstructed, and can include an under-reconstructed area and an over-reconstructed area. On this basis, the Gaussian points in the target area are encrypted based on a first gradient threshold, and the Gaussian points in other areas except the target area are encrypted based on a second gradient threshold to obtain a three-dimensional object. Wherein, the first gradient threshold can be less than the second gradient threshold.
[0077] As a specific implementation manner, the embodiments of the present disclosure can encrypt the Gaussian points in the background area based on a first gradient threshold, and encrypt the Gaussian points in other areas except the background area based on a second gradient threshold to obtain a three-dimensional object.
[0078] As known from the above introduction, the target area can be an area with insufficient reconstruction or over-reconstruction. The target area can be selected by the average view space gradient of each Gaussian point. Due to incorrect geometric features, there will be a large gradient in the target area. Based on experiments, it is found that in outdoor scenes, the background quality of images usually deteriorates. Gaussian points at closer distances will cover a larger area in the image, and the cumulative view space gradient is also larger than that of more distant areas. At this time, if the same gradient threshold is applied for densification processing, the densification rate of distant areas will be low, resulting in a low appearance completion degree of three-dimensional reconstruction.
[0079] Based on this, the embodiments of the present disclosure introduce an adaptive threshold to allow more splitting and cloning operations on Gaussian points in more distant areas. Specifically, during the training process, the embodiments of the present disclosure can adjust the densification criterion according to the average depth and the view space gradient. For example, the densification criterion is adjusted by multiplying the average depth and the view space gradient by a factor of 1 + nI( >d thres ). Here, the average depth can be the average of the z-axis information after projecting the position information of multiple Gaussian points.
[0080] Wherein, I( ) can be an indicator function, n can be a magnification factor, and d thres can be a threshold that changes according to the depth distribution. Through experiments, it is determined that after allowing more densification of the background, the background quality of the three-dimensional object reconstructed in the embodiments of the present disclosure has been improved. Figure 4 Shows the photometric loss situation with or without using the adaptive encryption strategy. Among them, Figure 4 the left figure of can be an image generated without using the adaptive encryption strategy, and the right figure can be an image generated using the adaptive encryption strategy.
[0081] By comparison, it is known that the pixels in the first region 301 in the left figure are darker than the pixels in the second region 202 in the right figure. The darker pixels indicate lower loss, that is, reconstructing using the adaptive encryption strategy can reduce the loss, thereby making the background rendering result of the obtained three-dimensional object better.
[0082] Exemplarily, embodiments of the present disclosure can set a gradient threshold for adaptive densification, and the gradient threshold d thres = 2d 0.8 − d 0.2 , n = 9, where d 0.2 and d 0.8 respectively refer to the 20th and 80th percentiles of the splat depth.
[0083] Optionally, after obtaining the Gaussian model, embodiments of the present disclosure can perform semantic recognition on the multi-view images in the Gaussian model to determine whether the scene where the multi-view images are located is a preset scene. Here, the preset scene can be an outdoor scene. If it is determined that the scene where the multi-view images are located is a preset scene, then the background region of the Gaussian model is detected. On this basis, the Gaussian points in the background region are encrypted based on the first gradient threshold, and the Gaussian points in other regions except the subject region are encrypted based on the second gradient threshold to obtain the three-dimensional object.
[0084] Optionally, after performing the depth guidance operation on each restricted Gaussian point and obtaining the three-dimensional model, embodiments of the present disclosure can perform semantic recognition on the multi-view images to determine whether the scene where the multi-view images are located is a preset scene. Here, the preset scene can be an outdoor scene. If it is determined that the scene where the multi-view images are located is a preset scene, then the background region of the three-dimensional object is detected. On this basis, the Gaussian points in the background region are encrypted based on the first gradient threshold, and the Gaussian points in other regions except the subject region are encrypted based on the second gradient threshold to obtain the three-dimensional object.
[0085] In addition, embodiments of the present disclosure can obtain the gradient information of the Gaussian points. When it is determined that the gradient information of the Gaussian points is greater than the gradient threshold, embodiments of the present disclosure can obtain the depth information of the Gaussian points. On this basis, it is determined whether the depth information is greater than a preset depth. If it is determined that the depth information is greater than the preset depth, then embodiments of the present disclosure can amplify the gradient of the Gaussian points. For example, multiply the gradient of the Gaussian points by 10. On this basis, densification processing is performed. Conversely, if the depth information is less than the preset depth, the gradient information of the Gaussian points remains unchanged.
[0086] Optionally, when it is determined that the depth information of the Gaussian point is greater than the preset depth, the embodiments of the present disclosure can also obtain the density of the Gaussian points in the background area. If the density is less than the set threshold, the gradient of the Gaussian point can be amplified; otherwise, amplification is not required. This method can also be applied to the above-mentioned first gradient threshold and second gradient threshold, that is, the first gradient threshold can correspond to the density. When it is determined that the depth information is greater than the preset depth and the density is less than the first density, the gradient threshold corresponding to the first density is used as the first gradient threshold. It should be noted that the gradient thresholds in the embodiments of the present disclosure can include not only the first gradient threshold and the second gradient threshold, but also the third gradient threshold, the fourth gradient threshold, etc. The specific number of gradient thresholds included here is not clearly limited and can be selected according to the actual situation. The gradient threshold can be related to the depth information. The ranges of the depth information are different, and the corresponding gradient thresholds may also be different.
[0087] By introducing the adaptive encryption operation, the embodiments of the present disclosure can reconstruct more details in the distant area, and can maintain a high-quality rendering effect both when the training view covers and does not cover the range. In other words, even in the absence of a detailed monocular depth map, the present disclosure can maintain the geometric structure of the object.
[0088] It should be noted that the limitation of the RGB value, the depth guidance operation, and the adaptive encryption operation in the embodiments of the present disclosure can be executed in parallel, and the three can be independent of each other. That is, in the process of replacing the spherical harmonics of the Gaussian model with the RGB value and restricting it within the preset range, the embodiments of the present disclosure can perform depth guidance on the Gaussian model and at the same time perform adaptive encryption operations on the Gaussian points to obtain the final three-dimensional object.
[0089] Optionally, the embodiments of the present disclosure can also perform a depth guidance operation on the restricted Gaussian points after restricting the RGB within the preset range, and can perform an adaptive encryption operation on the Gaussian points (Gaussian model) to obtain a three-dimensional object.
[0090] Optionally, the embodiments of the present disclosure can also perform a depth guidance operation on the restricted Gaussian points after restricting the RGB within the preset range. On this basis, an adaptive encryption operation is performed on the model obtained by the depth guidance operation to obtain a three-dimensional object.
[0091] The specific execution order of the limitation of the RGB value, the depth guidance operation, and the adaptive encryption operation is not clearly limited here and can be selected according to the actual situation.
[0092] Embodiments of the present disclosure can improve the interpretability of the generated three-dimensional object (model) in terms of geometric attributes and appearance by using RGB values to replace spherical harmonics, restricting the RGB values, performing depth guidance, and adaptive encryption operations. Specifically, lifting the constraints on each splashing color feature can provide more accurate guidance for the monocular depth prediction model and enhance the model's ability to reconstruct scene geometry. Experiments show that the embodiments of the present disclosure can effectively reduce the uncertainty of reconstruction problems and allow for a better understanding of the scene structure and individuals, especially in scenes with limited perspectives, while also maintaining high-fidelity rendering in some uncovered views.
[0093] In addition, in an exemplary embodiment of the present disclosure, a three-dimensional Gaussian splash optimization device based on depth information is also provided. Figure 5 A schematic structural diagram of the three-dimensional Gaussian splash optimization device based on depth information is shown, as Figure 5 shown, the three-dimensional Gaussian splash optimization device 300 based on depth information may include: an acquisition module 310, a reconstruction module 320, and a reconstruction module 320.
[0094] The acquisition module 310 is configured to acquire multi-view images and extract multiple Gaussian points of the multi-view images, and the model corresponding to the multiple Gaussian points is a Gaussian model; The reconstruction module 320 is configured to replace the spherical harmonics of each Gaussian point with RGB values, limit the RGB values of each Gaussian point within a preset range according to the Sigmoid function, and perform a depth guidance operation on each of the limited Gaussian points to obtain a three-dimensional object.
[0095] In some embodiments, the reconstruction module 320 may include: a prediction sub-module, configured to predict the depth information of the multi-view images to obtain a predicted depth and determine the rendering depth of the Gaussian model; a guidance sub-module, configured to perform a depth guidance operation based on the predicted depth and the rendering depth.
[0096] In some embodiments, the guidance sub-module is further configured to determine the Spearman correlation coefficient according to the predicted depth and the rendering depth and perform a depth guidance operation based on the Spearman correlation coefficient.
[0097] In some embodiments, the three-dimensional Gaussian splash optimization device 300 based on depth information may further include: an adaptive module, configured to perform an adaptive encryption operation on the Gaussian model to obtain a three-dimensional object, and the adaptive encryption operation is used to perform densification processing on the target area of the Gaussian model.
[0098] In some embodiments, the adaptive module may include: A determination sub-module configured to determine the target region, where the target region includes an under-reconstructed region and an over-reconstructed region; An encryption sub-module configured to encrypt Gaussian points within the target region based on a first gradient threshold and encrypt Gaussian points in other regions except the target region based on a second gradient threshold to obtain the three-dimensional object, where the first gradient threshold is less than the second gradient threshold.
[0099] In some embodiments, the adaptive module is further configured to encrypt Gaussian points within the background region based on a first gradient threshold and encrypt Gaussian points in other regions except the background region based on a second gradient threshold to obtain the three-dimensional object.
[0100] Embodiments of the present disclosure can effectively achieve depth guidance and provide a more accurate geometric structure by simply degrading color features to the zero order and using the sigmoid function for reprojection. In addition, embodiments of the present disclosure can also complete high-fidelity rendering at some uncovered angles and can exceed the rendering effect using monocular depth as guidance. As the integrity of the scene geometric structure improves, embodiments of the present disclosure can introduce other geometric attributes to provide further guidance in a more accurate and effective manner.
[0101] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0102] Figure 6 is a block diagram of an electronic device 600 shown according to an exemplary embodiment. As Figure 6 shown, the electronic device 600 may include: a processor 601, a memory 602. The electronic device 600 may further include one or more of a multimedia component 603, an input / output (I / O) interface 604, and a communication component 605.
[0103] Among them, the processor 601 is used to control the overall operation of the electronic device 600 to complete all or part of the steps in the above-mentioned three-dimensional Gaussian splash optimization method based on depth information. The memory 602 is used to store various types of data to support the operation of the electronic device 600. These data may include, for example, instructions for any application or method operating on the electronic device 600, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The multimedia component 603 may include a screen and an audio component. The screen can be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal can be further stored in the memory 602 or sent through the communication component 605. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 604 provides an interface between the processor 601 and other interface modules, and the above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 605 is used for wired or wireless communication between the electronic device 600 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 605 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.
[0104] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned three-dimensional Gaussian splash optimization method based on depth information.
[0105] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned three-dimensional Gaussian splash optimization method based on depth information are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 602 including program instructions, and the above program instructions may be executed by the processor 601 of the electronic device 600 to complete the above-mentioned three-dimensional Gaussian splash optimization method based on depth information.
[0106] Figure 7 FIG. 7 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. For example, the electronic device 700 may be provided as a server. Referring to Figure 7 , the electronic device 700 includes a processor 722, the number of which may be one or more, and a memory 732 for storing computer programs executable by the processor 722. The computer programs stored in the memory 732 may include one or more modules each corresponding to a set of instructions. In addition, the processor 722 may be configured to execute the computer program to execute the above-mentioned three-dimensional Gaussian splash optimization method based on depth information.
[0107] In addition, the electronic device 700 may further include a power supply component 726 and a communication component 750. The power supply component 726 may be configured to perform power management of the electronic device 700, and the communication component 750 may be configured to implement communication of the electronic device 700, for example, wired or wireless communication. In addition, the electronic device 700 may further include an input / output (I / O) interface 758. The electronic device 700 may operate based on an operating system stored in the memory 732.
[0108] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions, which, when executed by a processor, implement the steps of the above-described three-dimensional Gaussian splash optimization method based on depth information. For example, the non-transitory computer-readable storage medium may be the above-described memory 732 including program instructions, and the above program instructions may be executed by the processor 722 of the electronic device 700 to complete the above-described three-dimensional Gaussian splash optimization method based on depth information.
[0109] In another exemplary embodiment, there is also provided a computer program product, which includes a computer program capable of being executed by a programmable device, and the computer program has a code portion for executing the above-described three-dimensional Gaussian splash optimization method based on depth information when executed by the programmable device.
[0110] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0111] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure does not separately describe various possible combination manners.
[0112] In addition, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
Claims
1. A three-dimensional Gaussian splash optimization method based on depth information, characterized in that: The method comprises: Acquire a multi-view image, and extract a plurality of Gaussian points of the multi-view image, wherein the model corresponding to the plurality of Gaussian points is a Gaussian model; The spherical harmonics of each Gaussian point are replaced with RGB values, and the RGB value of each Gaussian point is limited within a preset range according to a Sigmoid function, and a depth guidance operation is performed on each Gaussian point after limitation to obtain a three-dimensional object.
2. The three-dimensional Gaussian splash optimization method based on depth information according to claim 1, characterized in that: The step of performing a depth-guided operation on each of the restricted Gaussian points to obtain a three-dimensional object comprises: Predicting the depth information of the multi-view images to obtain a predicted depth, and determining a rendering depth of the Gaussian model; The depth-guided operation is performed based on the predicted depth and the rendered depth.
3. The three-dimensional Gaussian splash optimization method based on depth information according to claim 2, characterized in that: The performing the depth guidance operation based on the predicted depth and the rendered depth comprises: A Spearman correlation coefficient is determined according to the predicted depth and the rendered depth, and the depth-guided operation is performed based on the Spearman correlation coefficient.
4. The three-dimensional Gaussian splash optimization method based on depth information according to claim 1, characterized in that: The method further comprises: An adaptive encryption operation is performed on the Gaussian model to obtain the three-dimensional object, wherein the adaptive encryption operation is used to perform densification processing on the target area of the Gaussian model.
5. The three-dimensional Gaussian splash optimization method based on depth information according to claim 4, characterized in that: The step of performing an adaptive encryption operation on the Gaussian model to obtain the three-dimensional object includes: Determine the target area, where the target area includes an under-reconstructed area and an over-reconstructed area; The Gaussian points in the target area are encrypted based on a first gradient threshold, and the Gaussian points in other areas except the target area are encrypted based on a second gradient threshold to obtain the three-dimensional object, wherein the first gradient threshold is less than the second gradient threshold.
6. The three-dimensional Gaussian splash optimization method based on depth information according to claim 5, characterized in that: The target area includes a background area, and the adaptive encryption operation is performed on the Gaussian model to obtain a three-dimensional object, including: The Gaussian points in the background area are encrypted based on a first gradient threshold, and the Gaussian points in other areas except the background area are encrypted based on a second gradient threshold to obtain the three-dimensional object.
7. The three-dimensional Gaussian splash optimization method based on depth information according to any one of claims 1 to 6, characterized in that: The extracting a plurality of Gaussian points of the multi-view image comprises: A plurality of point clouds in a three-dimensional space are extracted from the multi-view images, and the plurality of point clouds are represented according to the plurality of Gaussian points.
8. A three-dimensional Gaussian splash optimization device based on depth information, characterized in that: include: An acquisition module is configured to acquire a multi-view image and extract a plurality of Gaussian points of the multi-view image, wherein the model corresponding to the plurality of Gaussian points is a Gaussian model; The reconstruction module is configured to replace the spherical harmonics of each Gaussian point with an RGB value, and limit the RGB value of each Gaussian point within a preset range according to a Sigmoid function, and perform a depth-guided operation on each Gaussian point after the limit to obtain a three-dimensional object.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Static scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing
CN121414978A