Point cloud reconstruction method and system based on perception and dynamic initialization

Through the 3D Gaussian distribution method with dynamic initialization and perceptual loss optimization, the problem of high time overhead in the iteration process of 3DGS method is solved, and the detail fidelity of reconstruction is improved in complex textures and high-frequency regions.

CN120495542AActive Publication Date: 2025-08-15XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202510990574.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

The existing 3DGS methods are expensive during the iteration process, and the details are prone to blurred or distorted in the reconstruction results of complex textures and high-frequency areas.

Method used

Dynamically initialize the 3D Gaussian distribution and introduce perceptual loss optimization strategy, the color features are processed by spherical harmonic function, the average distance of neighbor points is determined, the initial Gaussian scaling is determined, the quaternion initial rotation angle is initialized, and the initial opacity value is set in reverse sigmoid, and the pixel-level L1 loss, local structure-level SSIM loss and perceptual loss LPIPS are combined for backpropagation optimization.

Benefits of technology

It significantly reduces the time overhead of the iteration process, improves the efficiency and quality of 3D reconstruction, especially retains the clarity of details in complex textures and high-frequency areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495542A_ABST
    Figure CN120495542A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud reconstruction method and system based on perception and dynamic initialization, and relates to the technical field of three-dimensional reconstruction, and the method comprises the steps: extracting point cloud data with position and color information from a multi-view image, converting the point cloud data into a CUDA tensor, and carrying out the processing, thereby achieving the dynamic initialization of 3D Gaussian distribution and parameters; a spherical harmonic function is adopted to process color features, an adjacent point average distance is adopted to determine initial Gaussian scaling, a quaternion is adopted to initialize a rotation angle, and an inverse sigmoid is adopted to set an initial opacity value, so that the initialization of a 3D model is optimized; a rendered image is generated by using a 3D-to-2D sputtering technology, and 3D Gaussian distribution parameters are optimized through pixel-level L1 loss, local structure-level SSIM loss and perceptual loss LPIPS. According to the method, 3D Gaussian distribution is dynamically initialized, and perception loss optimization is introduced, so that the problem of time overhead of a 3DGS method in an iteration process is solved, and the problem that the 3DGS method is easy to have a fuzzy detail or distortion phenomenon in scenes such as complex textures and high-frequency regions is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional reconstruction technology, and in particular to a point cloud reconstruction method and system based on perception and dynamic initialization. Background Art

[0002] 3D reconstruction technology, a core topic in computer vision, plays a vital role in key technological scenarios such as virtual reality, cultural heritage preservation, and autonomous driving. Previous research directions, such as reconstruction methods based on point clouds, voxels, or implicit neural fields (such as NeRF), each have their own advantages and disadvantages. Explicit point cloud representations enable efficient rendering but struggle to model complex surface details. Implicit representations, while offering the advantage of continuity, face inherent bottlenecks such as high computational overhead and poor real-time performance. The emerging 3D Gaussian Splatting (3DGS) technique achieves a significant breakthrough in achieving real-time rendering and geometric accuracy by fusing differentiable radiance fields with explicit point cloud representations. This method first generates a point cloud through random initialization, then represents it using an anisotropic Gaussian distribution function. A differentiable rendering pipeline then establishes a mapping from parameter space to a 2D projection. Finally, a loss is used to calculate the error between the rendered and original images, which is then optimized using backpropagation. Because this process does not involve network modules, it not only improves rendering quality but also significantly reduces training time compared to the previous state-of-the-art NeRF method.

[0003] However, existing 3DGS methods still face two key challenges: First, the random initialization of point clouds leads to low time utilization in the training process. During the training iteration process, multiple cloning and splitting operations are required for correction, which significantly increases the time overhead of the iterative process; second, 3DGS mainly relies on pixel-level photometric loss (L1) and structural similarity loss (SSIM) to optimize geometric and appearance parameters, but its reconstruction results are prone to blurred or distorted details in scenes with complex textures and high-frequency areas.

[0004] Therefore, how to reduce the time overhead of the 3DGS method in the iterative process and solve the problem of blurred or distorted details in scenes with complex textures and high-frequency areas has become an urgent problem to be solved in this field. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a point cloud reconstruction method and system based on perception and dynamic initialization. By dynamically initializing the 3D Gaussian distribution and introducing a perceptual loss optimization strategy, the time overhead of the iterative process is significantly reduced, and the problem of blurred or distorted details in complex textures and high-frequency areas is solved, thereby improving the efficiency and quality of 3D reconstruction.

[0006] The specific plan is as follows:

[0007] On the one hand, point cloud reconstruction methods based on perception and dynamic initialization include:

[0008] S1, obtain the original images to be reconstructed of the same scene shot from different perspectives;

[0009] S2, extracting a 3D sparse point cloud from the original image to be reconstructed by restoring the structure from motion, dynamically initializing the 3D sparse point cloud, and obtaining a 3D Gaussian distribution and initialized 3D Gaussian distribution parameters;

[0010] S3, performs sputtering of 3D to 2D images on the 3D Gaussian distribution to generate a rendered image, compares the rendered image with the original image to be reconstructed, and obtains the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS;

[0011] S4, based on the pixel-level L1 loss, the local structure-level SSIM loss and the perceptual loss LPIPS, uses the back-propagation technology to optimize the initialized 3D Gaussian distribution parameters to obtain the final 3D Gaussian distribution, and uses the final 3D Gaussian distribution as the three-dimensional model of the reconstruction result.

[0012] Furthermore, the S2 specifically includes:

[0013] The original image to be reconstructed is input into the structure-from-motion algorithm to extract and match image features and estimate the camera pose. 3D points are then reconstructed through triangulation. Camera parameters and point positions are optimized through bundled adjustment to generate 3D sparse point cloud data with position and color information.

[0014] The three-dimensional sparse point cloud data with position and color information is converted into a CUDA tensor. The color features in the CUDA tensor are processed by spherical harmonics. The initial 3D Gaussian distribution scale is determined by calculating the average distance between neighboring points of the CUDA tensor. The three-dimensional sparse point cloud data is rotated by quaternion. The initial opacity value is set by inverse sigmoid. The color features, initial Gaussian scale, rotation angle and initial opacity value are encapsulated to obtain the 3D Gaussian distribution and 3D Gaussian distribution parameters.

[0015] Furthermore, the initial 3D Gaussian distribution scaling is determined by calculating the average distance between neighboring points of the CUDA tensor, specifically including:

[0016] Calculate the average distance of the nearest neighbors of the CUDA tensor, then calculate the square root of the average distance and expand the distance by 20% to determine the initial Gaussian scaling, which is calculated as follows:

[0017] ;

[0018] in, represents the initial Gaussian scaling; Representing point clouds The squared distance to the nearest neighbor.

[0019] Furthermore, the rotation angle is initialized by quaternion, including:

[0020] Generate a random normal vector vi for each point cloud data and normalize it. The calculation formula is as follows:

[0021] ;

[0022] Among them, n i Represents the normal vector of each point cloud after normalization;

[0023] Randomly generate a rotation angle , the rotation angle is randomly sampled from the uniform distribution interval [-0.05, 0.05], based on the rotation angle Construct quaternion q i :

[0024] ;

[0025] in, Represents the real part of the quaternion; Represents the imaginary part of the quaternion;

[0026] Through the quaternion q i The 3D point cloud data is rotated and initialized so that the obtained 3D Gaussian distribution has an initial rotation angle.

[0027] Furthermore, the initial opacity value is set by inverse sigmoid, specifically including:

[0028] First, calculate the local density of a single point in the 3D point cloud data :

[0029] ;

[0030] in, represents the neighborhood distance after normalization of the point cloud, Represents the maximum value of the neighborhood distance in all point clouds;

[0031] Perform opacity mapping and Linear mapping to the specified interval, the calculation formula is as follows:

[0032] ;

[0033] in, Indicates opacity; through the inverse sigmoid function Convert opacity to parameter space, the calculation formula is as follows:

[0034] ;

[0035] Get the initial opacity value in parameter space .

[0036] Furthermore, in S3, the 3D Gaussian distribution is sputtered from 3D to 2D images to generate a rendered image. The calculation formula is as follows:

[0037] , ;

[0038] in, is the center position of the 3D Gaussian; is the center position of the 2D Gaussian corresponding to the 3D Gaussian; represents the covariance matrix of a 3D Gaussian; represents the covariance matrix of the 2D Gaussian corresponding to the 3D Gaussian; P is the projection matrix; W is the view transformation matrix; J is the Jacobian matrix of the affine approximation; T represents the transpose.

[0039] Furthermore, in S4, the backpropagation technique is used to optimize the 3D Gaussian distribution parameters based on the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS. Specifically, the following steps are performed:

[0040] The original image to be reconstructed is normalized according to the pre-calculated mean and variance of the ImageNet dataset to eliminate the interference of illumination differences between different acquisition devices and obtain a standardized image. The calculation formula is as follows:

[0041] ;

[0042] in, represents the normalized image, represents the original image to be reconstructed, and Represent the mean and variance of the ImageNet dataset respectively;

[0043] Then the feature extraction of the middle layer is carried out, and the pre-trained convolutional neural network VGG16 with frozen parameters is forward propagated to the conv3_3 layer to obtain a 256-channel feature map , where H and W represent the height and width of the normalized image obtained by normalizing the original image to be reconstructed;

[0044] Perform L2 normalization on the 256-channel feature map, and the calculation formula is as follows:

[0045] ;

[0046] in, represents the L2 norm of the feature map; ε represents a positive number;

[0047] Introducing learnable weight vectors , through back-propagation, the constraint effect of the channel is dynamically strengthened to generate the perceptual loss LPIPS, which is defined as:

[0048] ;

[0049] Among them, N, H and W represent the number of channels, the height and width of the feature map obtained after the original image to be reconstructed is normalized and extracted from the middle layer. represents the learnable weight of the c-th channel; is the original image to be reconstructed I at position The normalized eigenvalue of the cth channel at ; Represents a rendered image; Represents the original image to be reconstructed;

[0050] Optimize the base loss and perceptual loss jointly:

[0051] ;

[0052] in Indicates final loss; Indicates pixel-level loss; represents local structural level loss; is the perceptual loss, and is the weight parameter;

[0053] The 3D Gaussian distribution parameters are optimized based on the final loss to obtain the final 3D Gaussian distribution.

[0054] On the other hand, a point cloud reconstruction system based on perception and dynamic initialization includes:

[0055] The module for acquiring images to be reconstructed obtains the original images to be reconstructed of the same scene shot from different perspectives;

[0056] The point cloud dynamic initialization module extracts a 3D sparse point cloud from the original image to be reconstructed by restoring the structure from motion, dynamically initializes the 3D sparse point cloud, and obtains the 3D Gaussian distribution and the initialized 3D Gaussian distribution parameters;

[0057] The image comparison module performs 3D-to-2D image sputtering on the 3D Gaussian distribution to generate a rendered image. The rendered image is then compared with the original image to be reconstructed to obtain the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS.

[0058] The target reconstruction module uses back-propagation technology to optimize the parameters of the initialized 3D Gaussian distribution based on the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS to obtain the final 3D Gaussian distribution, which is used as the three-dimensional model of the reconstruction result.

[0059] The present invention adopts the above technical solution and has the following beneficial effects:

[0060] (1) The present invention optimizes the initialization of 3D models by using spherical harmonics to process color features, the average distance between neighboring points to determine the initial Gaussian scaling, the quaternion to initialize the rotation angle, and the inverse sigmoid to set the initial opacity value.

[0061] (2) The present invention optimizes the 3D Gaussian distribution parameters by introducing the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS. By using these multi-level loss functions, it can not only accurately capture the subtle differences between images, but also maintain the local structural similarity of the reconstructed image, especially in scenes with complex textures and high-frequency areas. The application of this series of technical means enables the present invention to achieve more refined detail preservation and solves the problem of blurred or distorted details that existed in previous 3D reconstruction technologies.

[0062] (3) The present invention uses 3D to 2D sputtering technology to generate rendered images, and optimizes the 3D Gaussian distribution parameters based on back-propagation technology. By jointly optimizing the basic loss and the perceptual loss, the accuracy of the final 3D reconstruction target is ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a flowchart of a point cloud reconstruction method based on perception and dynamic initialization according to an embodiment of the present invention;

[0064] Figure 2 Detailed flowchart of a point cloud reconstruction algorithm based on perception and dynamic initialization according to an embodiment of the present invention;

[0065] Figure 3 This is a partially enlarged visual comparison diagram of an embodiment of the present invention;

[0066] Figure 4 This is a diagram of a point cloud reconstruction system based on perception and dynamic initialization according to an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The present invention will be described in further detail below with reference to the examples and accompanying drawings, but the embodiments of the present invention are not limited thereto. Figure 1 and Figure 2 As shown, the point cloud reconstruction method based on perception and dynamic initialization of the present invention includes:

[0068] S1, obtain the original images to be reconstructed of the same scene shot from different perspectives.

[0069] S2, extracting a three-dimensional sparse point cloud from the original image to be reconstructed by restoring the structure from motion, dynamically initializing the three-dimensional sparse point cloud, and obtaining a 3D Gaussian distribution and initialized 3D Gaussian distribution parameters.

[0070] Specifically, the S2 includes:

[0071] The original image to be reconstructed is input into the structure-from-motion algorithm to extract and match image features and estimate the camera pose. 3D points are then reconstructed through triangulation. Camera parameters and point positions are optimized through bundled adjustment to generate 3D sparse point cloud data with position and color information.

[0072] The three-dimensional sparse point cloud data with position and color information is converted into a CUDA tensor. The color features in the CUDA tensor are processed by spherical harmonics. The initial 3D Gaussian distribution scale is determined by calculating the average distance between neighboring points of the CUDA tensor. The three-dimensional sparse point cloud data is rotated by quaternion. The initial opacity value is set by inverse sigmoid. The color features, initial Gaussian scale, rotation angle and initial opacity value are encapsulated to obtain the 3D Gaussian distribution and 3D Gaussian distribution parameters.

[0073] The initial 3D Gaussian distribution scaling is determined by calculating the average distance between neighboring points of the CUDA tensor, specifically including:

[0074] Calculate the average distance of the nearest neighbors of the CUDA tensor, then calculate the square root of the average distance and expand the distance by 20% to determine the initial Gaussian scaling, which is calculated as follows:

[0075] ;

[0076] in, represents the initial Gaussian scaling; Representing point clouds The squared distance to the nearest neighbor.

[0077] In this example, the point cloud coordinates and colors are converted to CUDA tensors, and the color features are processed using spherical harmonics. The initial Gaussian scale is then determined by calculating the average distance between neighboring points, the rotation angle is initialized using a quaternion, and the initial opacity value is set using an inverse sigmoid. Finally, the coordinates, color features, scale, rotation, and opacity are encapsulated as trainable parameters to complete the initialization of the Gaussian properties. While expanding the calculation, the maximum scale value is constrained to 5.0 to prevent extreme values from causing numerical instability. This allows for an expanded 3D Gaussian initial coverage and alleviates rendering holes in sparse areas.

[0078] Specifically, the rotation angle is initialized by quaternion, including:

[0079] Generate a random normal vector vi for each point cloud data and normalize it. The calculation formula is as follows:

[0080] ;

[0081] Among them, n i Represents the normal vector of each point cloud after normalization;

[0082] Randomly generate a rotation angle , the rotation angle is randomly sampled from the uniform distribution interval [-0.05, 0.05], based on the rotation angle Construct quaternion q i :

[0083] ;

[0084] in, Represents the real part of the quaternion; Represents the imaginary part of the quaternion;

[0085] Through the quaternion q i The 3D point cloud data is rotated and initialized so that the obtained 3D Gaussian distribution has an initial rotation angle.

[0086] In this embodiment, the rotation angle is limited to ±5.7° to meet the differentiability conditions of Lie algebra. Random small-angle quaternion initialization is used instead of the all-zero initialization of the original method to guide the diversity of optimization directions and accelerate the training process of the model.

[0087] Specifically, the initial opacity value is set by inverse sigmoid, including:

[0088] First, calculate the local density of a single point in the 3D point cloud data :

[0089] ;

[0090] in, represents the neighborhood distance after normalization of the point cloud, Represents the maximum value of the neighborhood distance in all point clouds;

[0091] Perform opacity mapping and Linear mapping to the specified interval, the calculation formula is as follows:

[0092] ;

[0093] in, Indicates opacity; through the inverse sigmoid function Convert opacity to parameter space, the calculation formula is as follows:

[0094] ;

[0095] Get the initial opacity value in parameter space .

[0096] S3 performs 3D to 2D image sputtering on the 3D Gaussian distribution to generate a rendered image, and compares the rendered image with the original image to be reconstructed to obtain the pixel-level L1 loss, local structure-level SSIM loss, and perceptual loss LPIPS.

[0097] Specifically, the 3D Gaussian distribution is sputtered from 3D to 2D images to generate a rendered image. The calculation formula is as follows:

[0098] , ;

[0099] in, is the center position of the 3D Gaussian; is the center position of the 2D Gaussian corresponding to the 3D Gaussian; represents the covariance matrix of a 3D Gaussian; represents the covariance matrix of the 2D Gaussian corresponding to the 3D Gaussian; P is the projection matrix; W is the view transformation matrix; J is the Jacobian matrix of the affine approximation; T represents the transpose.

[0100] S4, based on the pixel-level L1 loss, the local structure-level SSIM loss and the perceptual loss LPIPS, uses the back-propagation technology to optimize the initialized 3D Gaussian distribution parameters to obtain the final 3D Gaussian distribution, and uses the final 3D Gaussian distribution as the three-dimensional model of the reconstruction result.

[0101] Specifically, the back-propagation technology is used to optimize the 3D Gaussian distribution parameters based on the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS, including:

[0102] The original image to be reconstructed is normalized according to the pre-calculated mean and variance of the ImageNet dataset to eliminate the interference of illumination differences between different acquisition devices and obtain a standardized image. The calculation formula is as follows:

[0103] ;

[0104] in, represents the normalized image, represents the original image to be reconstructed, and Represent the mean and variance of the ImageNet dataset respectively;

[0105] Then the feature extraction of the middle layer is carried out, and the pre-trained convolutional neural network VGG16 with frozen parameters is forward propagated to the conv3_3 layer to obtain a 256-channel feature map , where H and W represent the height and width of the normalized image obtained by normalizing the original image to be reconstructed;

[0106] Perform L2 normalization on the 256-channel feature map, and the calculation formula is as follows:

[0107] ;

[0108] in, represents the L2 norm of the feature map; ε represents a positive number;

[0109] Introducing learnable weight vectors , through back-propagation, the constraint effect of the channel is dynamically strengthened to generate the perceptual loss LPIPS, which is defined as:

[0110] ;

[0111] Among them, N, H and W represent the number of channels, the height and width of the feature map obtained after the original image to be reconstructed is normalized and extracted from the middle layer. represents the learnable weight of the c-th channel; is the original image to be reconstructed I at position The normalized eigenvalue of the cth channel at ; Represents a rendered image; Represents the original image to be reconstructed;

[0112] Optimize the base loss and perceptual loss jointly:

[0113] ;

[0114] in Indicates final loss; Indicates pixel-level loss; Represents local structural level loss is the perceptual loss, and is the weight parameter;

[0115] The 3D Gaussian distribution parameters are optimized based on the final loss to obtain the final 3D Gaussian distribution.

[0116] In this embodiment, based on the 3DGS method using pixel-level (L1 loss) and local structure-level (SSIM loss) loss functions for rendering optimization, a perceptual loss that conforms to human perception is introduced. , building a three-level joint constraint of "pixel-structure-perception". The core idea is to establish an alignment relationship between the rendered image and the real image in the high-level visual space by extracting mid-level semantic features from a pre-trained convolutional neural network (VGG16).

[0117] Specifically, the method of the present invention improves the initialization process of the 3DGS method, proposes corresponding initialization schemes for the three parameters of scaling, rotation, and opacity, optimizes the numerical robustness of the 3D Gaussian ellipsoid, and provides a better starting point for subsequent optimization, thereby reducing the overall optimization time; and introduces the learned perceptual image patch similarity (LPIPS) loss based on deep semantics, constructs multiple losses for joint optimization based on the pixel-level loss of the original 3D Gaussian model, and achieves improved reconstruction effect.

[0118] Specifically, this embodiment is carried out on an NVIDIA RTX 3090 GPU server, and the training process is built based on the PyTorch framework. The optimizer uses Adam, and the main parameter settings include: the position learning rate decays exponentially from 0.00016 to 0.0000016 (delay coefficient 0.01), the feature parameter learning rate is 0.0025, the Gaussian initial scaling factor is set to 12 times the logarithmic transformation value of the nearest neighbor distance, the rotation parameter is initialized with a small angle random perturbation in the interval [-0.05rad, 0.05rad], and the feature encoding dimension maintains 3-channel zero initialization. Regularization weight The SSIM coefficient is set to 0.2, and the perceptual loss weight is 0.05. The total number of training iterations is 30,000, with a densification interval of 100 steps and an opacity reset period of 3,000 steps. The random background enhancement strategy is enabled. At the same time, according to the output mode of 3DGS, the corresponding PSNR, SSIM and other specific parameters are output when the number of training iterations reaches 7,000.

[0119] Specifically, the proposed method was comprehensively compared with four representative NeRF-based algorithms: ENeRF and MVSNeRF, as well as 3DGS and its subsequent improved method, MVSGaussian. Multiple comparisons were conducted on two major datasets, Tanks and Temples, with NeRF-LLFF, verifying the proposed method's advantages in both iteration efficiency and reconstruction quality. Detailed information is provided in Tables 1 and 2.

[0120] Table 1 Quantitative index comparison with other advanced methods on the Tanks and Temples dataset;

[0121]

[0122] Table 2 Comparison of quantitative indicators with other advanced methods on the NeRF-LLFF dataset;

[0123]

[0124] As shown in Tables 1 and 2, this method achieves optimal performance on both the Tanks and Temples dataset and the NeRF-LLFF dataset at 30,000 iterations, with PSNR (24.96dB / 27.84dB) and SSIM (0.911 / 0.932) significantly surpassing existing methods, and LPIPS (0.126 / 0.098) dropping to the lowest level, demonstrating its advanced performance in high-frequency detail preservation and perceptual consistency, which proves the effectiveness of the proposed method. It is worth noting that even under lightweight training of 7,000 iterations, this method still achieves better quality with a time consumption close to that of the 30k comparison scheme, verifying the algorithm's ability to converge quickly under a limited number of iterations. Experiments further reveal that the algorithm of the present invention has significant improvements in both training time and visual quality compared to the original 3DGS.

[0125] In addition to quantitative indicators, this embodiment also uses local magnification visualization comparison (such as Figure 3 The results show that the proposed method has the advantage of detail reconstruction: in complex geometric areas (such as dense vegetation areas, textured areas on the train surface, etc.), compared with the edge blur and detail adhesion phenomena of the comparative methods, this method can reconstruct the details more clearly, proving the advanced nature of the proposed method.

[0126] To further verify the effectiveness of the improved module, an ablation experiment was conducted on the Tanks and Temples dataset. The results are shown in Table 3.

[0127] Table 3 Ablation experiments based on the Tanks and Temples dataset;

[0128]

[0129] Experimental results show that the baseline 3DGS model takes approximately 12 minutes and 20 seconds to train on our device after 30,000 iterations, achieving a PSNR of 23.65dB, an SSIM of 0.867, and LPIPS of 0.184. After introducing dynamic initialization, training time is reduced to approximately 11 minutes and 10 seconds (an efficiency improvement of approximately 10%), while maintaining stable PSNR and SSIM (23.67dB / 0.868). This demonstrates that our dynamic initialization method accelerates model convergence and effectively reduces training time without compromising convergence quality. Furthermore, by incorporating a perceptual loss, the model achieves a PSNR of 24.96dB (a 1.31dB improvement over the baseline), an SSIM of 0.911, and an LPIPS of 0.126 (a 31.5% reduction) after 30,000 iterations in 11 minutes and 20 seconds, demonstrating the enhanced effect of our perceptual loss on reconstruction quality. In the 7000-iteration lightweight scenario, dynamic initialization reduced training time from 2 minutes and 30 seconds to 2 minutes and 5 seconds (a 10% reduction). After the joint perceptual loss, the training time increased slightly to 2 minutes and 10 seconds, but the PSNR jumped from 20.13dB to 23.93dB (a 13.6% improvement), and the LPIPS value dropped from 0.319 to 0.195 (a 38.9% decrease). These results demonstrate that the proposed algorithm achieves superior results compared to the original 3DGS, regardless of whether it is 7000 or 30,000 iterations. At 7000 iterations, the reconstruction effect basically meets the viewing requirements of the human eye, and the quality is further enhanced with subsequent iterations.

[0130] like Figure 4 As shown, this embodiment also discloses a point cloud reconstruction system based on perception and dynamic initialization, including:

[0131] The to-be-reconstructed image acquisition module 41 acquires original to-be-reconstructed images of the same scene captured from different perspectives;

[0132] a point cloud dynamic initialization module 42, which extracts a three-dimensional sparse point cloud from the original image to be reconstructed by restoring the structure from motion, and dynamically initializes the three-dimensional sparse point cloud to obtain a 3D Gaussian distribution and initialized 3D Gaussian distribution parameters;

[0133] The image comparison module 43 performs 3D-to-2D image sputtering on the 3D Gaussian distribution to generate a rendered image, compares the rendered image with the original image to be reconstructed, and calculates the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS;

[0134] The target reconstruction module 44 uses back-propagation technology to optimize the parameters of the initialized 3D Gaussian distribution based on the L1 loss at the pixel level, the SSIM loss at the local structure level, and the perceptual loss LPIPS to obtain the final 3D Gaussian distribution, and uses the final 3D Gaussian distribution as the three-dimensional model of the reconstruction result.

[0135] The specific implementation of the point cloud reconstruction system based on perception and dynamic initialization is the same as the point cloud reconstruction method based on perception and dynamic initialization, and will not be repeated in this embodiment.

[0136] Although the present invention has been particularly shown and described in conjunction with preferred embodiments, it will be understood by those skilled in the art that various changes in form and details may be made to the present invention without departing from the spirit and scope of the invention as defined in the appended claims, and all such changes are within the scope of protection of the present invention.

Claims

1. A point cloud reconstruction method based on perception and dynamic initialization, characterized in that: include: S1, obtain the original images to be reconstructed of the same scene shot from different perspectives; S2, extracting a 3D sparse point cloud from the original image to be reconstructed by restoring the structure from motion, dynamically initializing the 3D sparse point cloud, and obtaining a 3D Gaussian distribution and initialized 3D Gaussian distribution parameters; S3, performs sputtering of 3D to 2D images on the 3D Gaussian distribution to generate a rendered image, compares the rendered image with the original image to be reconstructed, and obtains the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS; S4, based on the pixel-level L1 loss, the local structure-level SSIM loss and the perceptual loss LPIPS, uses the back-propagation technology to optimize the initialized 3D Gaussian distribution parameters to obtain the final 3D Gaussian distribution, and uses the final 3D Gaussian distribution as the three-dimensional model of the reconstruction result.

2. The point cloud reconstruction method based on perception and dynamic initialization according to claim 1, characterized in that Said S2 specifically includes: The original image to be reconstructed is input into the structure-from-motion algorithm to extract and match image features and estimate the camera pose. 3D points are then reconstructed through triangulation. Camera parameters and point positions are optimized through bundled adjustment to generate 3D sparse point cloud data with position and color information. The three-dimensional sparse point cloud data with position and color information is converted into a CUDA tensor. The color features in the CUDA tensor are processed by spherical harmonics. The initial 3D Gaussian distribution scale is determined by calculating the average distance between neighboring points of the CUDA tensor. The three-dimensional sparse point cloud data is rotated by quaternion. The initial opacity value is set by inverse sigmoid. The color features, initial Gaussian scale, rotation angle and initial opacity value are encapsulated to obtain the 3D Gaussian distribution and 3D Gaussian distribution parameters.

3. The point cloud reconstruction method based on perception and dynamic initialization according to claim 2, characterized in that: The initial 3D Gaussian distribution scaling is determined by calculating the average distance between neighboring points of the CUDA tensor, specifically including: Calculate the average distance of the nearest neighbors of the CUDA tensor, then calculate the square root of the average distance and expand the distance by 20% to determine the initial Gaussian scaling, which is calculated as follows: ; in, represents the initial Gaussian scaling; Representing point clouds The squared distance to the nearest neighbor.

4. The point cloud reconstruction method based on perception and dynamic initialization according to claim 2, characterized in that Initialize the rotation angle through quaternion, including: Generate a random normal vector vi for each point cloud data and normalize it. The calculation formula is as follows: ; Among them, n i Represents the normal vector of each point cloud after normalization; Randomly generate a rotation angle , the rotation angle is randomly sampled from the uniform distribution interval [-0.05, 0.05], based on the rotation angle Construct quaternion q i : ; in, Represents the real part of the quaternion; Represents the imaginary part of the quaternion; Through the quaternion q i The 3D point cloud data is rotated and initialized so that the obtained 3D Gaussian distribution has an initial rotation angle.

5. The point cloud reconstruction method based on perception and dynamic initialization according to claim 2, characterized in that: The initial opacity value is set by inverse sigmoid, including: First, calculate the local density of a single point in the 3D point cloud data : ; in, represents the neighborhood distance after normalization of the point cloud, Represents the maximum value of the neighborhood distance in all point clouds; Perform opacity mapping and Linear mapping to the specified interval, the calculation formula is as follows: ; in, Indicates opacity; through the inverse sigmoid function Convert opacity to parameter space, the calculation formula is as follows: ; Get the initial opacity value in parameter space .

6. The point cloud reconstruction method based on perception and dynamic initialization according to claim 1, characterized in that In S3, the 3D Gaussian distribution is sputtered from 3D to 2D images to generate a rendered image. The calculation formula is as follows: , ; in, is the center position of the 3D Gaussian; is the center position of the 2D Gaussian corresponding to the 3D Gaussian; represents the covariance matrix of a 3D Gaussian; represents the covariance matrix of the 2D Gaussian corresponding to the 3D Gaussian; P is the projection matrix; W is the view transformation matrix; J is the Jacobian matrix of the affine approximation; T represents the transpose.

7. The point cloud reconstruction method based on perception and dynamic initialization according to claim 1, characterized in that In S4, the backpropagation technology is used to optimize the 3D Gaussian distribution parameters based on the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS, including: The original image to be reconstructed is normalized according to the pre-calculated mean and variance of the ImageNet dataset to eliminate the interference of illumination differences between different acquisition devices and obtain a standardized image. The calculation formula is as follows: ; in, represents the normalized image, represents the original image to be reconstructed, and Represent the mean and variance of the ImageNet dataset respectively; Then the feature extraction of the middle layer is carried out, and the pre-trained convolutional neural network VGG16 with frozen parameters is forward propagated to the conv3_3 layer to obtain a 256-channel feature map , where H and W represent the height and width of the normalized image obtained by normalizing the original image to be reconstructed; Perform L2 normalization on the 256-channel feature map, and the calculation formula is as follows: ; in, represents the L2 norm of the feature map; ε represents a positive number; Introducing learnable weight vectors , through back-propagation, the constraint effect of the channel is dynamically strengthened to generate the perceptual loss LPIPS, which is defined as: ; Among them, N, H and W represent the number of channels, the height and width of the feature map obtained after the original image to be reconstructed is normalized and extracted from the middle layer. represents the learnable weight of the c-th channel; is the original image to be reconstructed I at position The normalized eigenvalue of the cth channel at ; Represents a rendered image; Represents the original image to be reconstructed; Optimize the base loss and perceptual loss jointly: ; in Indicates final loss; Indicates pixel-level loss; represents local structural level loss; is the perceptual loss, and is the weight parameter; The 3D Gaussian distribution parameters are optimized based on the final loss to obtain the final 3D Gaussian distribution.

8. A point cloud reconstruction system based on perception and dynamic initialization, characterized in that: include: The module for acquiring images to be reconstructed obtains the original images to be reconstructed of the same scene shot from different perspectives; The point cloud dynamic initialization module extracts a 3D sparse point cloud from the original image to be reconstructed by restoring the structure from motion, dynamically initializes the 3D sparse point cloud, and obtains the 3D Gaussian distribution and the initialized 3D Gaussian distribution parameters; The image comparison module performs 3D-to-2D image sputtering on the 3D Gaussian distribution to generate a rendered image. The rendered image is then compared with the original image to be reconstructed to obtain the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS. The target reconstruction module uses back-propagation technology to optimize the parameters of the initialized 3D Gaussian distribution based on the pixel-level L1 loss, the local structure-level SSIM loss, and the perceptual loss LPIPS to obtain the final 3D Gaussian distribution, which is used as the three-dimensional model of the reconstruction result.

Citation Information

Patent Citations

  • 3D modeling reconstruction system, method and device based on point cloud information and Gaussian cloud cluster

    CN118196306A

  • Three-dimensional scene arbitrary angle model segmentation method based on 3D Gaussian Splitting

    CN118351277A

  • Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function

    CN118505541A

  • Grid equipment three-dimensional reconstruction method based on Gaussian splashing

    CN119359955A

  • Structural perception three-dimensional scene reconstruction method and device

    CN119888133A

Cited By

  • 3D GS cultural relic digital reconstruction method and system based on block chain

    CN121033287A

  • Three-dimensional object material mapping generation method based on generative prior

    CN121437759A