Point cloud reconstruction method and system based on three-dimensional Gaussian sputtering
By constructing a covariance degradation risk probability map and a gradient convergence dynamic monitoring mechanism, the problem of abnormal gradient propagation in texture-sparse regions in traditional point cloud reconstruction methods is solved, thereby improving the structural stability of 3D point cloud models.
Patent Information
- Application Number
- CN202511282359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-12
AI Technical Summary
Traditional point cloud reconstruction methods are prone to gradient propagation anomalies when dealing with regions with extremely scarce texture information, leading to non-convergence of the optimization process or getting trapped in local minima, resulting in poor structural stability of the 3D point cloud model.
By acquiring multi-view images and reference view images, a covariance degradation risk probability map is constructed, generating multiple Gaussian 3D points to be controlled. 3D point selection and fitting confidence score calculation are performed to construct an initial 3D point cloud model. The covariance update gating mechanism with gradient convergence dynamic monitoring is used for optimization, and the target 3D point cloud model is output.
It effectively avoids gradient propagation anomalies and optimization non-convergence, significantly improving the structural stability of the 3D point cloud model.
Smart Images

Figure CN121120994A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D modeling technology, and in particular to a point cloud reconstruction method and system based on 3D Gaussian sputtering. Background Technology
[0002] In fields such as computer vision, robot perception, and 3D modeling, point cloud reconstruction technology is a core means of recovering the 3D structure of an object from 2D observation data (such as images and depth maps) or sparse sampling points. Its accuracy and robustness directly affect the performance of downstream tasks (such as scene understanding, object recognition, and path planning).
[0003] In point cloud reconstruction based on 3D Gaussian sputtering, in order to accurately characterize the shape, directionality and diffusion characteristics of each Gaussian point in 3D space, it is necessary to characterize the spatial distribution of its probability density through the covariance matrix, which is the key to achieving fine modeling in this method.
[0004] However, in real-world scenarios, if the input image contains regions with extremely scarce texture information, such as areas of high-gloss reflection on metal surfaces, uniform solid-color backgrounds, or locally occluded shadows, traditional point cloud reconstruction methods, when optimizing point cloud parameters through backpropagation of image photometric errors, will suffer from abnormal gradient propagation due to their ill-conditioned covariance matrix. This can cause the optimization process to fail to converge or get trapped in local minima, resulting in poor structural stability of the final 3D point cloud model. Summary of the Invention
[0005] This invention provides a point cloud reconstruction method and system based on three-dimensional Gaussian sputtering, which solves the technical problem that traditional point cloud reconstruction methods cause gradient propagation anomalies, resulting in non-convergence or getting trapped in local minima, leading to poor structural stability of the final three-dimensional point cloud model.
[0006] The first aspect of this invention provides a point cloud reconstruction method based on three-dimensional Gaussian sputtering, comprising:
[0007] Acquire multi-view images and reference view images, and construct a covariance degradation risk probability map based on the multi-view images and the reference view images;
[0008] Based on the covariance degradation risk probability map, multiple Gaussian three-dimensional points to be controlled are generated.
[0009] Based on the three-dimensional Gaussian points to be adjusted, three-dimensional points are screened and fitting confidence scores are calculated. Multiple secondary-adjusted three-dimensional points and the fitting confidence scores corresponding to each secondary-adjusted three-dimensional point are output. An initial three-dimensional point cloud model is constructed based on the multiple secondary-adjusted three-dimensional points.
[0010] Based on the multi-view images and the fitting confidence scores corresponding to each of the secondary-adjusted 3D points, the secondary-adjusted 3D points are optimized to output multiple target Gaussian 3D points.
[0011] The initial three-dimensional point cloud model is updated using multiple target Gaussian three-dimensional points to determine the intermediate three-dimensional point cloud model;
[0012] The intermediate 3D point cloud model is updated using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and the target 3D point cloud model is output.
[0013] Optionally, constructing a covariance degradation risk probability map based on the multi-view image and the reference view image includes:
[0014] The multi-view image is preprocessed to output a target multi-view image, and the target image is converted into a grayscale image.
[0015] The Sobel operator is used to calculate the gradient of the grayscale image, and the first-order reciprocal image in the horizontal and vertical directions is output. Based on the first-order reciprocal image in the horizontal and vertical directions, the sum of squared gradient magnitudes corresponding to the target multi-view image is calculated.
[0016] The average of the squared gradient magnitudes is calculated within a local area of the grayscale image to generate a texture gradient intensity map.
[0017] The texture gradient intensity map is divided to generate multiple equal-sized grids corresponding to the target multi-view image, and the texture response rate of each equal-sized grid is calculated to determine the texture response rate of each equal-sized grid.
[0018] In the reference view image, a target region is selected, and a mutual information weighted block matching algorithm is used to search for the maximum luminous consistency pixel block based on the target region and each of the equal-sized grids, generating multiple matching points corresponding to each of the equal-sized grids;
[0019] Multiple linear triangulations are performed on the multiple matching corresponding points to output multiple three-dimensional points corresponding to each matching corresponding point. Based on the multiple three-dimensional points corresponding to each matching corresponding point, a triangulation consistency vector corresponding to each matching corresponding point is constructed.
[0020] The multiple triangulation consistency vectors are standardized to output the stability evaluation scores of each three-dimensional point corresponding to each matching point. The stability evaluation scores of each three-dimensional point and the texture response rate associated with the equal-sized mesh corresponding to each three-dimensional point are linearly interpolated and combined to calculate the degradation probability score of each three-dimensional point.
[0021] Based on each of the three-dimensional points, a three-dimensional space is constructed, and the three-dimensional space is divided to generate multiple cubic voxel units;
[0022] The degradation probability scores corresponding to multiple three-dimensional points in each cubic voxel are averaged to output the covariance degradation risk level corresponding to each cubic voxel. Based on the covariance degradation risk level corresponding to each cubic voxel, a covariance degradation risk probability map is constructed.
[0023] Optionally, generating multiple Gaussian three-dimensional points to be controlled based on the covariance degradation risk probability map includes:
[0024] In the covariance degradation risk probability map, any cube voxel unit corresponding to a covariance degradation risk level greater than a preset level threshold is taken as a high-risk voxel unit, and each of the high-risk voxel units is divided to output multiple three-dimensional cube meshes.
[0025] Count the number of 3D points in each of the 3D cube meshes;
[0026] The three-dimensional cube mesh corresponding to any number of three-dimensional points less than the preset first point number threshold is taken as a potential degenerate mesh;
[0027] Calculate the cluster density coefficient of the 3D cube mesh corresponding to the number of any 3D points that is greater than or equal to the preset second point number threshold, and compare it with the preset density coefficient threshold.
[0028] Any three-dimensional cubic mesh corresponding to an aggregation density coefficient greater than the preset density coefficient threshold is considered a structurally unstable mesh.
[0029] Perform back projection operations on multiple 3D points in the structurally unstable mesh and multiple 3D points in the potentially degenerate mesh to output the grayscale values of each 3D point in the structurally unstable mesh and the grayscale values of each 3D point in the potentially degenerate mesh from multiple perspectives.
[0030] Based on the gray values of each three-dimensional point in the structurally unstable grid and the gray values of each three-dimensional point in the potentially degenerate grid under multiple views, calculate the sum of squares of the photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squares of the photometric residuals of each three-dimensional point in the potentially degenerate grid.
[0031] The sum of squared photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squared photometric residuals of each three-dimensional point in the potentially degenerate grid are normalized to determine the photometric consistency deviation values of each three-dimensional point in the structurally unstable grid and the photometric consistency deviation values of each three-dimensional point in the potentially degenerate grid, and then compared with preset deviation thresholds respectively.
[0032] Any three-dimensional point corresponding to a photometric consistency deviation value greater than the preset deviation threshold is taken as a three-dimensional reconstruction imbalance point.
[0033] Perform eigenvalue decomposition on the covariance matrix corresponding to each of the three-dimensional reconstruction imbalance points to determine the principal axis eigenvectors and eigenvalues corresponding to each of the three-dimensional reconstruction imbalance points.
[0034] Based on the principal axis direction feature vector and feature value corresponding to each of the three-dimensional reconstruction imbalance points, the three-dimensional reconstruction imbalance points are filtered to output multiple Gaussian three-dimensional points to be adjusted.
[0035] Optionally, the step of screening three-dimensional points and calculating fitting confidence scores based on each of the Gaussian three-dimensional points to be adjusted, and outputting multiple secondary-adjusted three-dimensional points and the fitting confidence scores corresponding to each of the secondary-adjusted three-dimensional points, includes:
[0036] Based on the feature values corresponding to each of the three-dimensional Gaussian points to be adjusted, the feature value ratio of each of the three-dimensional Gaussian points to be adjusted is calculated, and the feature value ratio of each of the three-dimensional Gaussian points to be adjusted is compared with a preset ratio threshold.
[0037] Take any Gaussian 3D point to be adjusted that corresponds to a feature value ratio greater than the preset ratio threshold as a morphological anomaly point, adjust the feature value of each morphological anomaly point, and output the adjusted feature value corresponding to each morphological anomaly point.
[0038] A spherical neighborhood is constructed for each of the aforementioned morphological anomalies, and the spherical neighborhood corresponding to each of the aforementioned morphological anomalies is output.
[0039] The three-dimensional Gaussian points to be controlled in each of the spherical neighborhoods are taken as the nearest neighbors of each of the morphological anomalies. Based on the polar coordinate angle associated with the principal axis direction feature vectors of the nearest neighbors of each of the morphological anomalies, the spherical neighborhoods corresponding to each of the morphological anomalies are divided to determine multiple equally spaced grids corresponding to each of the morphological anomalies.
[0040] The number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each of the morphological anomalies is counted, and the compressed neighborhood space entropy value corresponding to each of the morphological anomalies is calculated based on the number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each of the morphological anomalies.
[0041] Based on the compressed neighborhood space entropy value corresponding to each of the morphological anomalies, the entropy change rate corresponding to each of the morphological anomalies is calculated, and the entropy change rate corresponding to each of the morphological anomalies is compared with a preset change rate threshold.
[0042] Any morphological anomaly point corresponding to an entropy change rate greater than the preset change rate threshold is taken as a three-dimensional point that needs fine-tuning. The adjustment feature value, principal axis direction feature vector, and covariance matrix corresponding to each three-dimensional point that needs fine-tuning are then fine-tuned to determine multiple secondary adjustment three-dimensional points.
[0043] Reproject each of the two-dimensional adjustment points, output the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points, and calculate the area ratio corresponding to each of the two-dimensional adjustment points based on the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points.
[0044] Based on the compressed neighborhood space entropy value corresponding to each of the two-dimensional adjustment points, the change ratio of the neighborhood entropy value corresponding to each of the two-dimensional adjustment points is calculated. Then, a weighted average strategy is used to output the fitting confidence score corresponding to each of the two-dimensional adjustment points based on the area ratio and the change ratio of the neighborhood entropy value.
[0045] Optionally, the optimization of each of the secondary-adjusted 3D points based on the multi-view images and the fitting confidence scores corresponding to each of the secondary-adjusted 3D points, to output multiple target Gaussian 3D points, includes:
[0046] Based on the fitting confidence scores corresponding to each of the aforementioned secondary-adjusted 3D points, multiple moderately confident 3D points are selected.
[0047] A three-dimensional Gaussian ellipsoid is constructed for each of the moderately reliable three-dimensional points, and the Gaussian ellipsoid corresponding to each of the moderately reliable three-dimensional points is output. Based on the multiple views corresponding to the multi-view image, the Gaussian ellipsoid is mapped to output the two-dimensional elliptical region of each of the moderately reliable three-dimensional points in each of the views.
[0048] Based on the contribution weights and opacity values of multiple pixels in the two-dimensional elliptical regions of each medium-confidence 3D point in each viewpoint, a photometric rendering map of each medium-confidence 3D point in each viewpoint is generated.
[0049] Based on the photometric rendering maps of each moderately reliable 3D point at each viewpoint, a residual map of each moderately reliable 3D point at each viewpoint is generated, and standard deviation analysis is performed on the residual maps of each moderately reliable 3D point at each viewpoint to generate the error anomaly target region of each moderately reliable 3D point at each viewpoint.
[0050] Based on the image pixel coordinates associated with the error anomaly target region of each of the medium-confidence 3D points in each of the viewpoints, multiple candidate 3D point positions of each of the medium-confidence 3D points in each of the viewpoints are generated.
[0051] Perform multi-view consistency verification on multiple candidate 3D points of each moderately confident 3D point in each viewpoint, and output the estimated positions of multiple real 3D mutation points of each moderately confident 3D point in each viewpoint.
[0052] Local resampling is performed on the estimated positions of multiple real three-dimensional mutation points of each of the medium-confidence three-dimensional points in each of the aforementioned viewpoints to generate multiple initial Gaussian three-dimensional points;
[0053] Multiple initial Gaussian 3D points are filtered to output multiple target Gaussian 3D points.
[0054] Optionally, the step of updating the intermediate 3D point cloud model using a covariance update gating mechanism based on gradient convergence dynamic monitoring to output the target 3D point cloud model includes:
[0055] The covariance update gating mechanism based on gradient convergence dynamic monitoring is used to determine the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model according to the intermediate 3D point cloud model.
[0056] Based on the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model, the intermediate 3D point cloud model is optimized to output the target 3D point cloud model.
[0057] A second aspect of the present invention provides a point cloud reconstruction system based on three-dimensional Gaussian sputtering, comprising:
[0058] The acquisition module is used to acquire multi-view images and reference view images, and construct a covariance degradation risk probability map based on the multi-view images and the reference view images;
[0059] The generation module is used to generate multiple Gaussian three-dimensional points to be controlled based on the covariance degradation risk probability map.
[0060] The filtering module is used to filter three-dimensional points and calculate fitting confidence scores based on each of the Gaussian three-dimensional points to be adjusted, output multiple secondary-adjusted three-dimensional points and the fitting confidence scores corresponding to each of the secondary-adjusted three-dimensional points, and construct an initial three-dimensional point cloud model based on the multiple secondary-adjusted three-dimensional points.
[0061] The optimization module is used to optimize each of the secondary adjusted 3D points based on the multi-view image and the fitting confidence score corresponding to each of the secondary adjusted 3D points, and output multiple target Gaussian 3D points.
[0062] The first update module is used to update the initial three-dimensional point cloud model using multiple target Gaussian three-dimensional points to determine the intermediate three-dimensional point cloud model.
[0063] The second update module is used to update the intermediate 3D point cloud model using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and output the target 3D point cloud model.
[0064] A computer device provided in a third aspect of the present invention includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the preceding claims.
[0065] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the preceding claims.
[0066] The fifth aspect of the present invention provides a computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein, when the program instructions are executed by a computer, the computer performs the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the preceding claims.
[0067] As can be seen from the above technical solutions, the present invention has the following advantages:
[0068] The above-mentioned technical solution of the present invention provides a point cloud reconstruction method based on three-dimensional Gaussian sputtering. The method acquires multi-view images and a reference view image, and constructs a covariance degradation risk probability map based on the multi-view images and the reference view image. Based on the covariance degradation risk probability map, multiple Gaussian three-dimensional points to be adjusted are generated. Based on each Gaussian three-dimensional point to be adjusted, three-dimensional point selection and fitting confidence score calculation are performed, outputting multiple secondary-adjusted three-dimensional points and the corresponding fitting confidence scores for each secondary-adjusted three-dimensional point. An initial three-dimensional point cloud model is constructed based on the multiple secondary-adjusted three-dimensional points. Based on the multi-view images and the corresponding fitting confidence scores for each secondary-adjusted three-dimensional point, each secondary-adjusted three-dimensional point is optimized, outputting multiple target Gaussian three-dimensional points. The invention employs multiple target Gaussian 3D points to update the initial 3D point cloud model, determining an intermediate 3D point cloud model. A covariance update gating mechanism based on gradient convergence dynamic monitoring is then used to update the intermediate 3D point cloud model, outputting the target 3D point cloud model. Based on this scheme, the invention pre-judges covariance anomaly risks using a covariance degradation risk probability map, reducing the generation of ill-conditioned covariance matrices from the source. By screening 3D points based on each Gaussian 3D point to be controlled, it can specifically screen and correct 3D points prone to gradient anomalies, avoiding the optimization process from getting trapped in local minima. This effectively avoids gradient propagation anomalies, optimization non-convergence, or local minima, significantly improving the structural stability of the 3D point cloud model. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 A flowchart illustrating the steps of a point cloud reconstruction method based on three-dimensional Gaussian sputtering provided in Embodiment 1 of the present invention;
[0071] Figure 2 This is a flowchart illustrating a point cloud reconstruction method based on three-dimensional Gaussian sputtering, provided in Embodiment 1 of the present invention.
[0072] Figure 3 This is a structural block diagram of a point cloud reconstruction system based on three-dimensional Gaussian sputtering, provided in Embodiment 2 of the present invention. Detailed Implementation
[0073] This invention provides a point cloud reconstruction method and system based on three-dimensional Gaussian sputtering, which solves the technical problem that traditional point cloud reconstruction methods cause gradient propagation anomalies, resulting in non-convergence or getting trapped in local minima, leading to poor structural stability of the final three-dimensional point cloud model.
[0074] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0075] Terminology Explanation:
[0076] "Spatial Measurement Based on Point Cloud Reconstruction using 3D Gaussian Sputtering" refers to the process of constructing a 3D point cloud by autonomously acquiring multi-view images using a UAV. Each point is represented as a 3D Gaussian distribution with parameters such as spatial location, covariance structure, color spherical harmonic coefficient, and opacity, thus more realistically expressing the shape and distribution characteristics of points in space. By introducing differentiable rendering technology, this 3D Gaussian point cloud is continuously projected onto a 2D image plane. An elliptical weighted average and transparency fusion are used to generate a realistic rendered image, which is then compared pixel-level with the actual observation image. Based on photometric errors, the parameters of the point cloud are optimized in reverse, gradually improving the accuracy and consistency of the 3D model. Ultimately, the constructed high-precision 3D model not only possesses high fidelity and strong structural continuity but also supports non-contact measurement of spatial location, distance, gaps, and other geometric dimensions from any angle. It is particularly suitable for structural modeling and safety assessment tasks in complex environments such as power operations.
[0077] Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a point cloud reconstruction method based on three-dimensional Gaussian sputtering, as provided in Embodiment 1 of the present invention.
[0078] This invention provides a point cloud reconstruction method based on three-dimensional Gaussian sputtering, comprising:
[0079] Step 101: Obtain multi-view images and reference view images, and construct a covariance degradation risk probability map based on the multi-view images and reference view images.
[0080] It should be noted that multi-view images of the scene to be reconstructed are acquired. Based on the spatial distribution features of local texture gradients in the images and the stability changes of corresponding multi-view matching errors, texture response rate and triangulation consistency parameters are extracted and fused to generate a covariance degradation risk probability map to identify potential regions of covariance degradation. Specifically, to identify regions in the 3D point cloud where covariance parameters may degrade, a covariance robustness identification method based on the spatial distribution features of image textures and the dynamic changes of multi-view matching errors is proposed. Through joint modeling of image structure perception and reconstruction consistency measurement, a covariance degradation risk probability map is generated, providing accurate target region annotations for subsequent Gaussian parameter tuning and optimization strategies.
[0081] Specifically, step 101 may include the following sub-steps S11-S19:
[0082] Step S11: Preprocess the multi-view image, output the target multi-view image, and convert the target image into a grayscale image.
[0083] It should be noted that multi-view images of the 3D scene to be modeled were acquired. Image acquisition utilized a hexacopter UAV platform equipped with a high-resolution visible light camera, which autonomously flew within a set trajectory range, uniformly capturing images from five fixed perspectives: front view, left front oblique view, right front oblique view, side view, and top view, ensuring that each 3D spatial point was covered by at least three different perspectives. The flight trajectory was planned using a serpentine scanning path, with the shooting height controlled between 3 and 10 meters above the target object, and the flight speed controlled within 1.5 meters per second to avoid image blurring. The image sensor employed a global shutter structure with a resolution of 4000×3000 pixels and a capture frequency of 3 frames per second. After acquisition, the multi-view images underwent preprocessing, including edge denoising, histogram equalization, and illumination normalization, removing ghost frames and low-contrast frames, and retaining high-quality image sequences for subsequent analysis.
[0084] Step S12: Use the Sobel operator to calculate the gradient of the grayscale image, output the first-order reciprocal image in the horizontal and vertical directions, and calculate the sum of squared gradient magnitudes corresponding to the multi-view image of the target based on the first-order reciprocal image in the horizontal and vertical directions.
[0085] Step S13: Average the sum of squared gradient magnitudes within a local area of the grayscale image to generate a texture gradient intensity map.
[0086] Step S14: Divide the texture gradient intensity map to generate multiple equal-sized grids corresponding to the target multi-view image, and calculate the texture response rate of each equal-sized grid to determine the texture response rate of each equal-sized grid.
[0087] It should be noted that a texture structure perception operation based on local gradients is performed on the obtained multi-view images. Specifically, after converting the color image (multi-view image) to grayscale, the Sobel operator is used to calculate the gradient of each pixel in both the horizontal and vertical directions, resulting in first-order derivative maps (i.e., first-order reciprocal maps in the horizontal and vertical directions). By calculating the sum of squared gradient magnitudes of each pixel, i.e., calculating the sum of squared gradient magnitudes of the multi-view image using the first-order reciprocal maps in the horizontal and vertical directions, and averaging the sum of squared gradient magnitudes within a local region of 11×11 (i.e., a local region of the grayscale image), a complete texture gradient intensity map is obtained. The texture gradient intensity map is divided into 64×48 equally sized grid regions (equal-size grids). The mean and standard deviation of the texture intensity map are calculated within each grid to quantify the texture response rate of that region, thus obtaining the texture response rate corresponding to each equal-size grid.
[0088] Furthermore, a lower threshold for the response rate is set to 0.05 (normalized scale). Mesh areas below this threshold are marked as sparse texture regions; areas above 0.2 are marked as rich texture regions; and the remaining areas are defined as medium texture regions. The texture level label of each mesh is preserved and projected into the point cloud space as a texture confidence reference for the 3D points.
[0089] It's worth noting that the process of using the Sobel operator to calculate the gradient for each pixel in both the horizontal and vertical directions essentially involves performing local convolution operations on the image's grayscale values to extract the edge change rates along these two directions, forming first-order derivative maps in both directions. First, the color image is converted to grayscale to uniformly process light intensity information. Then, two fixed 3×3 convolution kernels are defined: one for detecting brightness changes in the horizontal direction (often called the Gx kernel), and the other for detecting brightness changes in the vertical direction (called the Gy kernel). For example, the Gx kernel is... Gy core is These two kernels are convolved with the grayscale image respectively, yielding gradient images for each pixel in the horizontal (x-axis) and vertical (y-axis) directions. The values of each pixel in these derivative maps represent the rate of brightness change in that direction; larger values indicate more pronounced edge or structural changes in that direction. These first-order derivative maps in these two directions can be further used to calculate gradient magnitude maps and gradient direction maps, forming a fundamental step in texture extraction, edge detection, and feature analysis.
[0090] Step S15: Select the target region in the reference view image, and use the mutual information weighted block matching algorithm to search for the maximum luminous consistency pixel block based on the target region and each equal-sized grid, generating multiple matching points corresponding to each equal-sized grid.
[0091] Step S16: Perform multiple linear triangulations on multiple matching points, output multiple three-dimensional points corresponding to each matching point, and construct a triangulation consistency vector corresponding to each matching point based on the multiple three-dimensional points corresponding to each matching point.
[0092] Step S17: Standardize the multiple triangulated consistency vectors, output the stability evaluation score of each 3D point corresponding to each matching point, and perform linear interpolation combination of the stability evaluation score of each 3D point and the texture response rate associated with the equal-sized mesh corresponding to each 3D point to calculate the degradation probability score corresponding to each 3D point.
[0093] Step S18: Construct a three-dimensional space based on each three-dimensional point, and divide the three-dimensional space to generate multiple cubic voxel units.
[0094] Step S19: Calculate the mean of the degradation probability scores of multiple 3D points in each cubic voxel unit, output the covariance degradation risk level of each cubic voxel unit, and construct a covariance degradation risk probability map based on the covariance degradation risk level of each cubic voxel unit.
[0095] It should be noted that a multi-view matching error consistency evaluation operation is performed to measure the reliability of triangulation. First, a region (target region) with a feature density greater than a threshold is selected from the reference view image. Using the PatchMatch algorithm (mutual information weighted block matching algorithm), the corresponding pixel block with the maximum luminance consistency is searched in all other images (i.e., the same-sized grid regions corresponding to the multi-view images), and the matching corresponding points are extracted. Linear triangulation is performed on the matching corresponding points using camera intrinsic and extrinsic parameters to obtain the three-dimensional position (3D point) of each matching point. This process is repeated for each reference point to perform N triangulation experiments (N is 5), and the variance of the triangulation depth value and its disparity variation range are calculated for each triangulation. At the same time, multiple sets of view angles and the projection consistency residuals of the reconstructed points between each view are recorded. The depth fluctuation, angle range, and residual standard deviation are jointly constructed into a triangulation consistency vector, that is, the triangulation consistency vector corresponding to each matching corresponding point is constructed based on the multiple 3D points corresponding to each matching corresponding point. This vector, after being standardized, is used to represent the stability evaluation score of each 3D point in the geometric reconstruction process. That is, the stability evaluation score of each 3D point corresponding to each matching point is output after the standardization of multiple triangulation consistency vectors.
[0096] The "Multi-view Matching Error Consistency Evaluation Operation" refers to performing feature matching and triangulation reconstruction on the same spatial point in images taken from multiple different angles, and evaluating the point's stable and reliable geometric positioning capability in 3D space by comparing the error variation trends between the reconstruction results from different viewpoints. The core purpose of this operation is to detect whether the matching of a spatial point in multi-view images is consistent and whether the reconstruction error is controllable. Specifically, for each candidate pixel in a reference image, corresponding points with similar luminance and texture features are found in other viewpoint images. Triangulation is performed through pairwise or multi-view matching pairs to calculate its 3D position. Then, the geometric stability of the point is quantified by comparing its depth value, position change, disparity residual, and other indicators through multiple sets of triangulation results. If a point exhibits large positional shifts and drastic error fluctuations in different viewpoint combinations, it indicates inconsistent multi-view matching errors and is a geometrically unstable point; conversely, if the error is stable and the projection residual is low, it indicates high reliability in 3D reconstruction. This evaluation operation is crucial in the entire point cloud generation process. It can effectively identify points in the image that are prone to covariance degradation due to sparse texture, abnormal reflection, or occlusion, providing a reliable screening basis for subsequent robust modeling, covariance adjustment, and point cloud optimization.
[0097] The "PatchMatch algorithm based on mutual information weighted block matching" refers to replacing the pixel similarity measurement method in the traditional PatchMatch algorithm with a weighted block matching strategy driven by mutual information (Mutual Information), thereby improving the robustness of matching in complex lighting variations and sparse texture regions. The PatchMatch algorithm is an efficient image patch matching method, initially proposed by Barnes et al. in image editing. Its core idea is to quickly estimate the optimal patch correspondence between images through random initialization and propagation mechanisms. Traditional PatchMatch relies on Euclidean distance or color difference as matching criteria, which is easily affected by lighting changes or surface reflections. Introducing mutual information as a matching metric measures the degree of information sharing between two image patches, is insensitive to brightness changes, and is suitable for robust matching of images across different viewpoints. In this step, the method is used to find the corresponding positions of the same spatial point in different images, including a reference image and multiple other viewpoint images. Specifically, a fixed-size image patch is extracted centered on a candidate point in the reference image, and the image patch with the maximum mutual information value is searched in other viewpoint images as the matching target, thus obtaining a more reliable matching pair. This method can effectively improve the point cloud initialization accuracy in areas with weak texture, heavy occlusion, or drastic brightness changes. It is an extended version of common PatchMatch variants for multi-view stereo reconstruction, such as PatchMatchStereo, COLMAP (Camera Localization and Mapping), and MVSNet (Depth Inference for Unstructured Multi-view Stereo). Its advantage lies in achieving high-quality dense matching, providing high-confidence matching input for subsequent triangulation reconstruction.
[0098] Furthermore, the image texture responsivity and triangulation consistency vector are aligned and fused in 3D space to generate a covariance degradation risk probability map. Specifically, for each 3D point, based on the texture responsivity of its corresponding image region and its own triangulation consistency score, a linear interpolation function is used to combine the two to calculate a degradation probability score, ranging from 0 to 1. The 3D space is further divided into cubic voxel units with a side length of 0.2 meters. The degradation probability values of all points within each voxel are averaged and assigned as the covariance degradation risk level of that voxel. The risk levels of all voxels are combined into a 3D covariance degradation risk probability map (i.e., the covariance degradation risk probability map), which will be used as input for screening high-risk point cloud regions in subsequent steps.
[0099] Linear interpolation is a method used to estimate the value at a certain intermediate position between two known values, based on the relative proportion of the intermediate position. It assumes that the change between the two known values is uniform and linear, thus allowing the value at any intermediate point to be calculated through a simple proportional relationship. In this invention, the linear interpolation function integrates two different sources and scales—image texture response rate and multi-view triangulation consistency score—into a unified covariance degradation probability score. Specifically, each 3D point or image mesh region corresponds to both a texture score and a geometric stability score. The linear interpolation function generates a continuous risk value based on the relative weights of these two indicators, reflecting whether the point is likely to degrade in subsequent covariance modeling. This method is simple, computationally efficient, and suitable for rapid scoring processing in large-scale point cloud scenarios. Common implementations can refer to the image interpolation functions in Python's NumPy library (Numerical Python, a Python library for scientific computing) or OpenCV (Open Source Computer Vision Library), which have wide applications in image processing, rendering fusion, and multi-source data mapping.
[0100] In this embodiment, by fusing multi-source information, high-risk areas where covariance values may degrade during 3D point cloud modeling are accurately identified. This provides a preliminary risk perception basis and input screening mechanism for subsequent Gaussian parameter initialization, morphological adjustment, and adaptive covariance optimization. In point cloud modeling based on 3D Gaussian sputtering, the covariance matrix is used to describe the diffusion characteristics and directionality of each point in 3D space. Once numerical degradation occurs (such as eigenvalues approaching zero or abnormal amplification), it will lead to a series of problems such as point cloud morphological distortion, rendering occlusion errors, and non-convergence of photometric optimization. Therefore, this step collects multi-view images and constructs a covariance robustness identification model (i.e., the functional general term for the processing flow of extracting texture responsivity, triangulation consistency parameters, and generating a covariance degradation risk probability map). Starting from the two perspectives of texture structure changes in the image itself and error distribution of geometric reconstruction, it extracts spatially meaningful texture responsivity and triangulation consistency parameters, and then fuses them through a linear interpolation function to generate a covariance degradation risk probability map, forming a mechanism that can warn of potential problem points in the early stages of modeling. This image can not only be used as a point cloud filter to select usable point sets, but also for targeted modeling and dynamic adjustment of high-risk areas, greatly improving the numerical stability and modeling accuracy of the entire scene point cloud, and has important control significance and engineering application value.
[0101] Step 102: Generate multiple Gaussian 3D points to be controlled based on the covariance degradation risk probability map.
[0102] It should be noted that, based on the covariance degradation risk probability map, a Gaussian ellipsoid initialization stability discrimination mechanism is established. This mechanism jointly analyzes the point cloud aggregation density and luminosity consistency deviation in corresponding regions of the image to determine the diffusion principal axis direction and structural deviation degree of high-risk points, thus constructing a set of Gaussian parameters to be adjusted. Specifically, to further improve the stability of the Gaussian ellipsoid initialization process in 3D point clouds and avoid subsequent rendering distortion and optimization failure due to covariance abnormalities, a high-precision and highly targeted Gaussian ellipsoid initialization stability discrimination process is constructed based on the covariance degradation risk probability map. This process identifies 3D points prone to covariance degradation through joint analysis of image spatial structure distribution and point cloud morphology information, and selects the parameter set that requires subsequent covariance adjustment.
[0103] Specifically, step 102 may include the following sub-steps S21-S211:
[0104] Step S21: In the covariance degradation risk probability map, any cube voxel unit corresponding to a covariance degradation risk level greater than the preset level threshold is taken as a high-risk voxel unit, and each high-risk voxel unit is divided to output multiple three-dimensional cube meshes.
[0105] Step S22: Count the number of three-dimensional points in each three-dimensional cube grid, and compare the number of three-dimensional points in each three-dimensional cube grid with the preset point number threshold.
[0106] Step S23: Take any three-dimensional cube mesh corresponding to the number of three-dimensional points that is less than the preset first point number threshold as a potential degenerate mesh.
[0107] Step S24: Calculate the cluster density coefficient of the three-dimensional cube mesh corresponding to the number of any three-dimensional points that is greater than or equal to the preset second point quantity threshold, and compare it with the preset density coefficient threshold.
[0108] Step S25: Take any three-dimensional cubic mesh corresponding to an aggregation density coefficient greater than the preset density coefficient threshold as a structurally unstable mesh.
[0109] The risk level of covariance degradation is the risk probability value.
[0110] It should be noted that all voxel units at the high-risk level (covariance degradation risk level greater than a preset threshold) are extracted from the covariance degradation risk probability map and treated as high-risk voxel units. The 3D points within each high-risk voxel unit are then traversed. Each point contains its 3D coordinates (X, Y, Z), the previously extracted texture responsivity score (between 0 and 1), the triangulation consistency score (normalized value), the initial covariance matrix (symmetric positive definite 3×3 matrix), and the corresponding image number and pixel position. To further analyze the spatial distribution density, the entire high-risk region is divided into a 3D cubic mesh with a side length of 0.1 meters. For each 3D cube mesh, the number of points (i.e., the number of 3D points) is counted. If the number of points is less than 10 (the first threshold for the number of points is preset to 10), the region is considered sparse, with insufficient covariance data samples, constituting a potential degenerate point. That is, any 3D cube mesh with a number of 3D points less than the preset threshold is considered a potentially degenerate mesh. Conversely, if the number of points exceeds 60 (the second threshold for the number of points is preset to 60), its clustering density coefficient is calculated based on the variance of the distance between points. If the density is too high, it may indicate abnormal duplicate points due to occlusion or viewpoint overlap, and it should also be marked as a structurally unstable region. That is, any 3D cube mesh with a clustering density coefficient greater than the preset threshold is considered a structurally unstable mesh. For 3D cube meshes with a number of 3D points greater than or equal to the preset first threshold, or with a number of 3D points less than the preset second threshold, no further processing steps are performed.
[0111] Step S26: Perform back projection operations on multiple 3D points in the structurally unstable mesh and multiple 3D points in the potentially degenerate mesh, and output the grayscale values of each 3D point in the structurally unstable mesh and the grayscale values of each 3D point in the potentially degenerate mesh from multiple perspectives.
[0112] Step S27: Based on the gray values of each three-dimensional point in the structurally unstable grid and the gray values of each three-dimensional point in the potentially degenerate grid under multiple views, calculate the sum of squares of the photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squares of the photometric residuals of each three-dimensional point in the potentially degenerate grid.
[0113] Step S28: Normalize the sum of squared photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squared photometric residuals of each three-dimensional point in the potentially degenerate grid to determine the photometric consistency deviation values of each three-dimensional point in the structurally unstable grid and the photometric consistency deviation values of each three-dimensional point in the potentially degenerate grid, and compare them with preset deviation thresholds respectively.
[0114] Step S29: Take any three-dimensional point corresponding to a photometric consistency deviation value greater than the preset deviation threshold as a three-dimensional reconstruction imbalance point.
[0115] It should be noted that for the points marked as sparse or abnormally dense (i.e., multiple 3D points in structurally unstable grids and multiple 3D points in potentially degenerate grids), a robustness analysis based on image photometric consistency is performed. For each 3D point, a backprojection operation is performed on all observation view images according to its corresponding image pixel coordinates to obtain its pixel grayscale value at each view. Using the grayscale value in the reference view image as a benchmark, the sum of squared photometric residuals in all auxiliary view images is calculated and further normalized to obtain the photometric consistency deviation value of that point. When this deviation value exceeds 40 gray levels (taking an 8-bit grayscale image as an example, the value range is 0 to 255), it is considered that the point is affected by local reflection, shadows, or occlusion interference, resulting in poor image consistency and unstable initial 3D reconstruction results. That is, any 3D point corresponding to a photometric consistency deviation value greater than a preset deviation threshold is regarded as a 3D reconstruction imbalance point; for 3D points corresponding to photometric consistency deviation values less than or equal to the preset deviation threshold, no subsequent processing steps are performed.
[0116] "Robustness analysis based on image photometric consistency" refers to assessing the reliability of photometric response stability of a point in multi-view image reconstruction by detecting whether the brightness of the same 3D point is consistent across different viewpoints. This helps determine the reconstruction quality and the reliability of covariance estimation. Its purpose is to identify anomalous points with significant brightness differences due to factors such as illumination variations, reflection interference, occlusion shadows, or texture loss, thus avoiding including these inconsistent points in the direct modeling process of Gaussian parameters and preventing distortion in covariance matrix estimation. In the specific analysis, a 3D point and its corresponding pixel positions in all observed images are first selected, and the grayscale value or brightness channel value of that pixel in each image is extracted. Then, using a primary viewpoint image (such as a reference image) as a benchmark, the photometric residual of that point in other images is calculated, i.e., the difference between the brightness value of each viewpoint and the benchmark brightness. Next, all residuals are averaged or varianced to form a photometric consistency deviation index. If this index exceeds a set threshold (e.g., more than 20 gray levels), the point is considered to have a photometric inconsistency problem. This analysis process is robust, meaning it can tolerate a certain degree of brightness fluctuation, avoiding misjudgment of anomalies due to slight changes in illumination. At the same time, it has a strong ability to identify anomalies in areas with high brightness reflection or extremely weak texture, making it a key auxiliary indicator in point cloud stability assessment.
[0117] Step S210: Perform eigenvalue decomposition on the covariance matrix corresponding to each 3D reconstruction imbalance point to determine the principal axis eigenvectors and eigenvalues corresponding to each 3D reconstruction imbalance point.
[0118] Step S211: Based on the principal axis direction feature vector and feature value corresponding to each 3D reconstruction imbalance point, filter each 3D reconstruction imbalance point and output multiple Gaussian 3D points to be adjusted.
[0119] It should be noted that structural stability analysis is performed on the covariance matrix of points exhibiting joint anomalies in spatial density and photometric consistency (i.e., points of 3D reconstruction imbalance). Specifically, eigenvalue decomposition is performed on the covariance matrix of each point to obtain the eigenvectors along the three principal axes and the corresponding eigenvalues. In other words, eigenvalue decomposition is performed on the covariance matrix of each 3D reconstruction imbalance point to determine the eigenvectors and eigenvalues along the principal axes. The ratio between the largest and smallest eigenvalues is calculated. If this ratio is greater than 40, it indicates that the ellipsoid has been overstretched or compressed in a certain direction, losing normal spatial diffusion equilibrium. Meanwhile, by comparing the angle between the principal axis direction of the point and the principal axis direction of other points in its grid, if the average angle exceeds 30 degrees, it indicates that its diffusion direction deviates structurally from the neighboring points, and is suspected to be an anomaly point or a boundary distortion region. Such points will be recorded as structural direction deviation points. That is, based on the principal axis direction feature vector and feature value corresponding to each 3D reconstruction imbalance point, each 3D reconstruction imbalance point is screened, and multiple Gaussian 3D points to be adjusted are output.
[0120] Among them, the eigenvectors of the three principal axes are obtained by performing eigenvalue decomposition on the covariance matrix of the three-dimensional point. They essentially reflect the dominant direction of the distribution of data around the point in space, and can be understood as three mutually perpendicular directions of the "natural extension" of the data.
[0121] It is worth mentioning that this invention selects points that simultaneously meet the following three conditions as targets for adjustment (Gaussian 3D points to be adjusted): their location belongs to a high-risk area for covariance degradation, their spatial density or photometric consistency is abnormal, and the ratio of eigenvalues of the covariance matrix is greater than a threshold and the principal axis direction deviates from the neighborhood direction. For these target points, all their parameter information is extracted, including 3D coordinates, original values of the covariance matrix, principal axis direction vector, image number, pixel position index, texture score, and photometric deviation, etc., and uniformly saved as a Gaussian parameter adjustment set. This set will serve as the input point set for the subsequent covariance adaptive adjustment and reconstruction mechanism.
[0122] In this embodiment, in the early stages of 3D point cloud modeling, a detailed analysis of the structural state and observational stability of high-risk areas is conducted to identify Gaussian points that may cause numerical degradation of the covariance matrix. This forms a set of Gaussian parameters to be adjusted (i.e., a set composed of all the Gaussian 3D points to be adjusted), which is then used for subsequent adaptive covariance adjustment and Gaussian ellipsoid reconstruction. Since the previous step generated a covariance degradation risk probability map, identifying potential problem areas under conditions such as sparse image texture and unstable triangulation errors, this risk map itself does not directly reveal the actual stability and local structural deviations of each point. Therefore, this step further analyzes the point cloud cluster density and photometric consistency deviation of each 3D point in these areas based on the risk map, judging its reliability as the basis of the Gaussian distribution from two key dimensions. Point cloud cluster density reflects whether there are abnormal distributions in the space surrounding the point that are too sparse or too dense. The former may lead to insufficient geometric constraints, while the latter may cause duplicate data due to occlusion or redundant perspectives; both can interfere with covariance estimation. Simultaneously, the photometric consistency deviation assessment evaluates the brightness retention of this point across multiple viewpoints. Excessive brightness variation indicates potentially unstable matching relationships across different images, further increasing the uncertainty of the covariance matrix. After identifying these unstable points, this step also assesses the structural deviation of the diffusion principal axis of their covariance matrix, detecting significant inconsistencies with neighboring directions and identifying morphological anomalies such as principal axis eigenvalue imbalance or extreme stretching. This ensures that the final constructed Gaussian parameter set accurately locates points with genuine numerical instability risks. This process not only effectively implements the transition from probabilistic maps to target point selection but also provides clear input for subsequent morphological adjustments, serving as a crucial intermediate step in the robust modeling process of Gaussian point clouds.
[0123] Step 103: Based on each Gaussian 3D point to be adjusted, perform 3D point screening and calculate fitting confidence score, output multiple secondary-adjusted 3D points and the fitting confidence score corresponding to each secondary-adjusted 3D point, and construct an initial 3D point cloud model based on the multiple secondary-adjusted 3D points.
[0124] It should be noted that, for the set of Gaussian parameters to be adjusted, a covariance adaptive adjustment mechanism based on normal entropy constraints is introduced. Based on the principal axis eigenvalue ratio compression rule and the spatial consistency entropy value evaluation criterion, the scale distribution of the Gaussian ellipsoid in each direction is dynamically reconstructed, generating a Gaussian distribution template with enhanced structural stability, and outputting the corresponding fitting confidence mapping tensor. Specifically, to address the problems of principal axis imbalance, directional distortion, or scale distortion that easily occur at high-risk points of covariance degradation in the Gaussian ellipsoid, a point-by-point analysis is performed on the selected set of Gaussian parameters to be adjusted, and a covariance adaptive adjustment mechanism based on normal entropy constraints is introduced. This covariance adaptive adjustment mechanism based on normal entropy constraints refers to the joint modeling of eigenvalue ratio control and spatial direction entropy value judgment to complete the dynamic reconstruction of the covariance matrix, generate a Gaussian distribution template with enhanced structural stability, and output a fitting confidence mapping tensor describing its geometric confidence level.
[0125] Specifically, step 103 may include the following sub-steps S31-S39:
[0126] Step S31: Based on the eigenvalues corresponding to each Gaussian 3D point to be adjusted, calculate the eigenvalue ratio of each Gaussian 3D point to be adjusted, and compare the eigenvalue ratio of each Gaussian 3D point to be adjusted with the preset ratio threshold.
[0127] Step S32: Take any Gaussian 3D point to be adjusted corresponding to a feature value ratio greater than a preset ratio threshold as a morphological anomaly point, adjust the feature values of each morphological anomaly point, and output the adjusted feature values corresponding to each morphological anomaly point.
[0128] It should be noted that for each 3D point in the set of Gaussian parameters to be adjusted (i.e., the 3D Gaussian point to be adjusted), its original covariance matrix is extracted, and a 3D eigenvalue decomposition operation is performed to obtain three eigenvalues (denoted as λ1, λ2, and λ3, sorted from largest to smallest) and the corresponding three orthogonal principal axis direction vectors. This set of eigenvalues represents the diffusion scale of the point in the X, Y, and Z directions. Then, the ratio between the largest and smallest eigenvalues (eigenvalue ratio), i.e., λ1 / λ3, is calculated. If the eigenvalue ratio exceeds 50 (the preset ratio threshold is 50), it indicates that the scale of the 3D Gaussian point to be adjusted is excessively stretched or compressed in a certain direction, which is a morphological anomaly, and it is regarded as a morphological anomaly point. At this time, the eigenvalues of the morphological anomaly point are adjusted according to the principal axis compression strategy: the maximum allowable ratio threshold is set to 10, λ1 is compressed to within 10 times λ3, while λ2 is kept to meet the condition of constant eigenvalue sum after adjustment. After the eigenvalues are adjusted, the three eigenvalues are recombined, and the corrected covariance matrix is restored using a matrix reconstruction algorithm to ensure that it is in a symmetric positive definite form, which facilitates stable rendering in the future.
[0129] Step S33: Construct a spherical neighborhood for each morphological anomaly point and output the spherical neighborhood corresponding to each morphological anomaly point.
[0130] Step S34: Take the three-dimensional Gaussian points to be controlled in each spherical neighborhood as the nearest neighbors of each morphological anomaly point, and divide the spherical neighborhood corresponding to each morphological anomaly point based on the polar coordinate angle associated with the principal axis direction feature vector of the nearest neighbors of each morphological anomaly point, and determine multiple equally spaced grids corresponding to each morphological anomaly point.
[0131] Step S35: Count the number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each morphological anomaly point, and calculate the compressed neighborhood space entropy value corresponding to each morphological anomaly point based on the number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each morphological anomaly point.
[0132] Step S36: Based on the compressed neighborhood space entropy value corresponding to each morphological anomaly point, calculate the entropy change rate corresponding to each morphological anomaly point, and compare the entropy change rate corresponding to each morphological anomaly point with the preset change rate threshold.
[0133] Step S37: Take any morphological anomaly point corresponding to an entropy change rate greater than a preset change rate threshold as a three-dimensional point to be fine-tuned, and fine-tune the adjustment feature value, principal axis direction feature vector, and covariance matrix corresponding to each three-dimensional point to be fine-tuned to determine multiple secondary adjustment three-dimensional points.
[0134] It should be noted that after eigenvalue adjustment, a spatial consistency entropy evaluation mechanism is further introduced to assess whether the morphological adjustment of morphological anomalies damages their spatial neighborhood. The specific method is as follows: A spherical neighborhood with a radius of 0.3 meters is constructed centered on the current 3D point (morphological anomaly). At least 20 nearest Gaussian points are extracted from this neighborhood (i.e., the 3D Gaussian points to be adjusted within each spherical neighborhood are considered the nearest neighbors of each morphological anomaly). The covariance matrix of these points is subjected to the same eigenvalue decomposition, and the distribution density of all principal axis directions on the sphere is calculated. The polar coordinate angle of each principal axis direction is discretized into equally spaced grids (e.g., every 10 degrees is a partition), the number of principal axes in each partition is counted, and the Shannon entropy value of this angular distribution is calculated as the spatial entropy of the neighborhood structure direction distribution (compressed neighborhood spatial entropy value). By comparing the changes in the entropy value of the neighborhood caused by the principal axis direction of the morphological anomaly point before and after compression (i.e., comparing the entropy value of the neighborhood space before and after compression of the morphological anomaly point, the entropy value of the neighborhood space before compression is determined based on the morphological anomaly point without eigenvalue adjustment), if the entropy value (entropy change rate) increases by more than 30% (the preset change rate threshold is 30%), it is considered that the compression has caused a structural orientation shift, and the eigenvalue of that point needs to be fine-tuned (i.e., any morphological anomaly point with an entropy change rate greater than the preset change rate threshold is taken as the three-dimensional point to be fine-tuned, and the adjustment eigenvalue, principal axis direction eigenvector, and covariance matrix corresponding to each three-dimensional point to be fine-tuned are fine-tuned to determine multiple secondary adjustment three-dimensional points). The weighted mean is used to approximate the average tilt angle of the neighborhood principal axis direction, and a soft constraint term is introduced to reduce its abruptness in the local structure.
[0135] The "spatial consistency entropy evaluation mechanism" refers to an analytical method that quantifies the consistency of structural orientation in a region by calculating the dispersion of the principal axis directions of points within a spatial neighborhood of a 3D point cloud. This allows for the assessment of whether adjusting the shape of a Gaussian point would disrupt the continuity of the local structure. Its core function is to determine whether the compressed principal axis direction of a point deviates significantly from the principal axis directions of its neighboring points when adaptively adjusting the covariance matrix of the Gaussian ellipsoid, thus avoiding spatial structural distortion or abrupt changes in orientation caused by single-point adjustments. In practice, covariance principal axis vectors of all adjacent Gaussian points are collected within a set radius, centered on the target point. The distribution of these directions in spherical coordinates is statistically analyzed, and the orientation angles are discretized into fixed partitions. The distribution entropy of the number of principal axes in each partition is calculated. A smaller entropy value indicates a more concentrated orientation and a more consistent structure; a larger entropy value indicates a more divergent orientation and a chaotic structure. If adjusting the covariance of the target point significantly increases the entropy value of the neighborhood, it is considered to have disrupted local consistency, and a reversal or smoothing adjustment should be implemented. Therefore, this mechanism plays a crucial role in maintaining the continuity of the local morphology and the integrity of the overall structure of the point cloud, and is the core judgment basis for achieving constrained optimization in the covariance adjustment stage.
[0136] Step S38: Reproject each of the two-dimensional adjustment points, output the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points, and calculate the area ratio corresponding to each of the two-dimensional adjustment points based on the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points.
[0137] Step S39: Based on the compressed neighborhood space entropy value corresponding to each of the two-dimensional adjustment points, calculate the neighborhood entropy value change ratio corresponding to each of the two-dimensional adjustment points, and use a weighted average strategy to output the fitting confidence score corresponding to each of the two-dimensional adjustment points based on the area ratio and neighborhood entropy value change ratio corresponding to each of the two-dimensional adjustment points.
[0138] It should be noted that after the aforementioned eigenvalue adjustment and structural orientation smoothing are completed, a new Gaussian distribution template is generated for each point (a 3D point undergoing secondary adjustment). This template includes: the reconstructed covariance matrix, the corresponding three principal axis direction vectors and their normalized eigenvalues, the 3D center coordinates of the point, and its pixel projection position in the original image. Subsequently, a fitting reliability score is calculated for each template. The score consists of two specific quantitative indicators: one is the ratio of the area of the reprojected region of the point in all viewpoint images to the area before adjustment (area ratio), used to measure the stability of its representation at the image level after adjustment; the other is the ratio of the change in neighborhood entropy, reflecting the degree of influence of the adjustment on the continuity of the surrounding structure. After normalization of both indicators, a weighted average strategy is used to fuse them into a fitting reliability score, with the range limited to between 0 and 1. The reliability scores of all points are output as tensors according to the 3D mesh mapping, called the fitting reliability mapping tensor, for use in subsequent multi-view consistency verification and re-optimization stages.
[0139] To ensure repeatable tracking and periodic checks of the adjustment process, key parameters of all adjusted points are packaged and recorded, including the original covariance matrix, the adjusted covariance matrix, eigenvalue ratio changes, principal axis direction changes, neighborhood entropy rate of change, fit confidence score, 3D coordinates, original image number, and pixel position index. This data is then uniformly written into the Gaussian distribution template update log for backtracking control and version comparison in subsequent covariance optimization processes.
[0140] In this embodiment, by precisely adjusting the shape and reconstructing the covariance of the selected set of Gaussian parameters to be adjusted, the ill-conditioned phenomenon of the covariance matrix caused by problems such as sparse texture and unstable parallax during the initial modeling process is solved, thereby enhancing the structural stability and spatial consistency of the Gaussian ellipsoid in the 3D point cloud. In actual point cloud reconstruction, some points have extremely unbalanced eigenvalues in their covariance matrix due to the presence of highlights, shadows, or missing textures in the input image, resulting in abnormal expansion along the principal axis direction. This leads to error propagation and shape distortion during rendering or photometric optimization. To avoid such problems, this step introduces an adaptive adjustment mechanism based on normal entropy constraints. The eigenvalue ratio of the covariance matrix of each Gaussian point is compressed, and the scale ratio between the maximum and minimum eigenvalues is forcibly controlled to prevent the ellipsoid from being infinitely stretched or collapsed in a certain direction. Simultaneously, to ensure that the adjustment does not disrupt the local point cloud structure, a spatial consistency entropy evaluation mechanism is introduced. By analyzing the distribution concentration of points along the principal axis in their neighborhood, it is determined whether the adjustment causes abrupt changes in direction or structural fragmentation. If a significant increase in entropy is detected, the adjustment intensity is rolled back or smoothed. After completing the above optimization, a new covariance matrix and principal axis combination are generated based on the adjustment results to construct a Gaussian distribution template with enhanced structural stability. A fitting reliability mapping tensor reflecting the morphological stability and reprojection adaptability of each point is output, providing strong numerical evidence for subsequent consistency verification, local resampling, and global optimization. This step not only improves the geometric continuity and robustness of the point cloud model but also lays the foundation for accurate and high-fidelity spatial measurement, representing a crucial link in achieving the core technology closed loop of this invention.
[0141] Step 104: Optimize each of the two-dimensional adjusted 3D points based on the multi-view images and the fitting confidence scores corresponding to each of the two-dimensional adjusted 3D points, and output multiple target Gaussian 3D points.
[0142] It should be noted that, using the fitted confidence mapping tensor, multi-view image reprojection consistency verification is performed on the Gaussian distribution template. Each Gaussian ellipsoid is mapped to all acquired image planes, and the error distribution between the photometric rendered image and the actual acquired image is compared pixel-by-pixel. Regions of abrupt error residual changes are located, and local resampling of the point cloud is triggered in these regions. Specifically, to improve the reconstruction accuracy and structural consistency of the 3D Gaussian point cloud model in the actual image space, after adjusting the structural stability of the Gaussian distribution template, multi-view image reprojection consistency verification is performed on the point cloud based on the fitted confidence mapping tensor. The core purpose of this step is to detect the photometric matching performance of the adjusted Gaussian ellipsoid under all observation views, locate regions of abrupt photometric residual changes, and reconstruct and enhance these error abrupt change regions through local resampling, thereby constructing a more continuous and accurate point cloud representation structure.
[0143] Specifically, step 104 may include the following sub-steps S41-S4:
[0144] Step S41: Based on the fitting confidence score corresponding to each secondary adjusted 3D point, screen each secondary adjusted 3D point to determine multiple medium confidence 3D points.
[0145] Step S42: Construct a three-dimensional Gaussian ellipsoid for each moderately reliable three-dimensional point, output the Gaussian ellipsoid corresponding to each moderately reliable three-dimensional point, and map each Gaussian ellipsoid based on multiple perspectives corresponding to the multi-view image, output the two-dimensional elliptical region of each moderately reliable three-dimensional point in each perspective.
[0146] Step S43: Generate a photometric rendering map of each medium-confidence 3D point in each viewpoint based on the contribution weight and opacity value of multiple pixels in the two-dimensional elliptical region of each medium-confidence 3D point.
[0147] It should be noted that the confidence score (fit confidence score) of each Gaussian point (second-order adjusted 3D point) in the fitting confidence mapping tensor is read. Based on the fitting confidence score corresponding to each second-order adjusted 3D point, each second-order adjusted 3D point is screened. Specifically, the process is as follows: moderately confident points (moderately confident 3D points) with fitting confidence scores between 0.4 and 0.8 are selected as the key verification objects in this stage. These points are neither completely confident nor excluded, and have a certain adjustment space and observation value. For each Gaussian point (moderately confident 3D point), it is represented as an ellipsoid (Gaussian ellipsoid) according to its center position, covariance matrix, and principal axis direction, and is mapped onto the image plane corresponding to all acquisition viewpoints using perspective projection. During projection, the intrinsic parameters (focal length, principal point position) and extrinsic parameters (camera pose) of the image are considered, and the 3D Gaussian ellipsoid is mapped into a 2D Gaussian distributed elliptical region (2D elliptical region). At the same time, the contribution weight and opacity value of each pixel are recorded to generate a differentiable photometric rendering map.
[0148] Step S44: Based on the photometric rendering maps of each moderately reliable 3D point at each viewpoint, generate residual maps of each moderately reliable 3D point at each viewpoint, and perform standard deviation analysis on the residual maps of each moderately reliable 3D point at each viewpoint to generate error anomaly target regions of each moderately reliable 3D point at each viewpoint.
[0149] It should be noted that the photometric rendering of each moderately reliable 3D point in each viewpoint of the multi-view image is compared pixel-by-pixel with the brightness value of the same pixel location in the real image to calculate the residual map. To eliminate the interference of camera noise and exposure differences, the real image is first linearly normalized to map its pixel brightness range to the [0,1] interval. The brightness value of each pixel in the photometric rendering is calculated by weighting the opacity of a Gaussian ellipsoid, and the residual value is defined as the absolute difference between the rendered image and the actual image for that pixel. Subsequently, standard deviation analysis is performed on each residual map to extract the average residual region and the high residual abrupt change region. Regions with residuals exceeding twice the average value and showing a continuous distribution are specifically marked as error anomaly target regions.
[0150] Step S45: Based on the image pixel coordinates associated with the error anomaly target region of each moderately confident 3D point in each viewpoint, generate multiple candidate 3D point locations for each moderately confident 3D point in each viewpoint.
[0151] Step S46: Perform multi-view consistency verification on multiple candidate 3D points of each moderately confident 3D point in each view, and output the estimated position of multiple real 3D mutation points of each moderately confident 3D point in each view.
[0152] Step S47: Perform local resampling on the estimated positions of multiple real 3D mutation points of each moderately confident 3D point from each viewpoint to generate multiple initial Gaussian 3D points.
[0153] Step S48: Filter multiple initial Gaussian 3D points and output multiple target Gaussian 3D points.
[0154] It should be noted that a backprojection operation is performed on the image pixel coordinates corresponding to all high residual regions to restore their approximate positions in 3D space. During backprojection, image intrinsic and extrinsic parameter information are combined with the depth estimation results of the current viewpoint to generate multiple candidate 3D points. Points with relatively consistent spatial positions are selected as the estimates of true 3D mutation points through multi-view consistency verification. These points are marked as locally inconsistent regions, and a spatial search grid is constructed within a radius of 0.3 meters in their vicinity to identify the original Gaussian point set within this region. A local resampling operation is performed on these point sets, that is, the estimated positions of multiple true 3D mutation points of each moderately confident 3D point in each viewpoint are locally resampled. This process is as follows: more unused pixel blocks from the image of more viewpoints are introduced into the current region, and new 3D points (initial Gaussian 3D points) are extracted by mutual information weighted block matching. Their spatial positions, covariance matrices, and color features are re-estimated. Specifically, after completing the local resampling, the fitting confidence score of each Gaussian point (initial Gaussian 3D point) in the current point cloud is updated. For points in the local resampling region, the photometric map is regenerated and the residual distribution is recalculated. If the average residual decreases by more than 30%, its confidence level is increased by 0.2 and it is marked as a high-confidence point (target Gaussian 3D point). If the decrease is not significant, the original confidence level is retained and it is recorded as an optimization failure sample. All adjustment results and resampling information are recorded in the update log file, including the source viewpoint, matching confidence level, photometric consistency score, spatial location, and corresponding image index for each newly added point.
[0155] The multi-view consistency verification screening refers to screening multiple candidate 3D points obtained by back-projecting the pixel coordinates of high residual area images from different acquisition perspectives by checking the consistency of these candidate points in spatial location. Specifically, the spatial positional deviation between the candidate 3D points obtained by back-projection from different perspectives is calculated, and points with deviations less than a preset threshold are retained as estimates of true 3D mutation points.
[0156] In this embodiment, the performance of the Gaussian distribution template in multi-view images is refined by fitting the point cloud morphological stability and photometric consistency indices provided by the confidence mapping tensor. This identifies and corrects potential error concentration areas in the model, enhancing the reprojection consistency and structural integrity of the point cloud in image space. Even after covariance adjustment and morphological reconstruction during the construction of the Gaussian point cloud model, geometric errors or observational biases may still exist in certain areas. These problems often manifest as residual abrupt changes in specific pixel regions between the rendered photometric map and the actual image. Through the multi-view image reprojection consistency verification mechanism introduced in this step, each Gaussian ellipsoid is projected from multiple camera perspectives, and pixel-by-pixel brightness differences are calculated with the real image to generate a residual map, accurately capturing subtle inconsistencies between the model and observations. Based on this, by analyzing the location and distribution of high-brightness abrupt change regions in the residual map, the corresponding error abrupt change regions in three-dimensional space are further located. These regions are often located under conditions of extremely weak texture, occluded boundaries, or complex lighting, and are considered high-error, high-uncertainty areas. Therefore, this step triggers local resampling of the point cloud at these locations, that is, re-extracting local image patches from the original multi-view images for 3D point reconstruction, supplementing or replacing the original point cloud structure, thereby improving the model's expressiveness and reprojection accuracy in key areas. This mechanism not only achieves closed-loop control of residuals during point cloud modeling but also effectively prevents errors from accumulating and amplifying in subsequent optimization processes. It is crucial for ensuring the spatial continuity, rendering quality, and measurement accuracy of the overall 3D model and is an important supporting element for achieving high-fidelity point cloud reconstruction capabilities in this invention.
[0157] Step 105: Update the initial 3D point cloud model using multiple target Gaussian 3D points to determine the intermediate 3D point cloud model.
[0158] It should be noted that the obtained multiple target Gaussian 3D points are incorporated into the current point cloud model (initial 3D point cloud model) for completion and replacement. This ultimately forms a new Gaussian point cloud structure (intermediate 3D point cloud model), which possesses higher image consistency and spatial structure accuracy, providing a more stable geometric foundation for subsequent covariance convergence control and overall modeling optimization.
[0159] Step 106: The intermediate 3D point cloud model is updated using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and the target 3D point cloud model is output.
[0160] It should be noted that, based on the consistency verification of multi-view image reprojection, a covariance update gating mechanism based on gradient convergence dynamic monitoring is established. This mechanism analyzes in real time the gradient magnitude changes and local error curvature changes of Gaussian parameters during backpropagation, dynamically adjusts the parameter update magnitude and optimization step rhythm, and records the historical evolution trajectory of the covariance parameters of each Gaussian point. Specifically, to ensure the numerical stability and morphological convergence consistency of the covariance matrix update during the backpropagation optimization of the Gaussian point cloud, and to avoid covariance divergence or getting trapped in local minima due to gradient oscillations or misleading updates, a covariance update gating mechanism based on gradient convergence dynamic monitoring is proposed. This mechanism refers to the real-time tracking and adjustment of the gradient magnitude changes, local photometric error curvature trends, and their evolution laws of each Gaussian point during the optimization iteration process. This allows for adaptive adjustment of the update magnitude and step rhythm, and records the evolution trajectory of the covariance parameters of each Gaussian point, providing a reference for subsequent overall modeling quality assessment and control.
[0161] Specifically, step 106 may include the following sub-steps S61-S62:
[0162] Step S61: Using a gradient convergence dynamic monitoring-based covariance update gating mechanism, determine the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model.
[0163] It should be noted that the process of determining the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model is as follows: Before each backpropagation optimization iteration, the covariance matrix of all Gaussian points to be optimized (Gaussian 3D points in the intermediate 3D point cloud model) is extracted, and the gradient value of the covariance parameters in the current gradient iteration is calculated. The gradient values are obtained based on the partial derivative of the loss function with respect to the covariance parameters, and the gradient components and their magnitudes on the three principal axes are recorded to form a three-dimensional gradient vector set. Subsequently, the maximum, minimum, and average rate of change of the gradient history window (e.g., the last 5 iterations) of each point are statistically analyzed to construct a "gradient stability criterion." If the gradient magnitude changes by more than 50% within the window, the point is considered to be in a non-convergent state. In this state, the update gating control flag is activated to prepare for the next update rhythm adjustment.
[0164] Furthermore, by combining the image reprojection residual map of the current round, local error curvature analysis is performed on each point. By performing a secondary fitting on the pixel error of the projection region of that point in the residual map, its curvature feature value is extracted to measure the error change gradient of that point in the image domain, i.e., to determine whether it is in a high error fluctuation region. If the residual curvature value exceeds a preset threshold (e.g., the second derivative of the average error fitting curve is greater than 0.02), the point is determined to be an "error unstable point." This curvature feature is combined with the aforementioned gradient change trend, and a "covariance update suppression factor" is constructed through logical judgment to guide whether the next round of update strategy should be weighted lower.
[0165] For Gaussian points marked as "non-convergent" or "error unstable," a dynamic step rate adjustment mechanism is initiated. This mechanism suppresses update divergence by slowing down the learning rate and introduces a historical direction momentum factor to enhance the smoothness of the optimization path. Specifically, the standard learning rate is multiplied by a dynamic scaling factor, which is weighted by both the current gradient change magnitude and the error curvature, ranging from 0.1 to 0.9. If a point steadily decreases and the error curvature flattens within three consecutive iterations, the original learning rate is gradually restored to maintain global optimization consistency. The updated covariance matrix undergoes positive definiteness verification. If negative eigenvalues appear or the matrix deviates from symmetry, matrix regularization is performed, replacing the smallest eigenvalue with a set threshold and reconstructing the principal axis direction to ensure that subsequent covariance expressions still possess physical rationality.
[0166] To ensure the traceability and transparency of the entire optimization process, the covariance parameter changes of each Gaussian point throughout the optimization cycle are recorded, and a "covariance evolution trajectory table" is constructed. This table includes: the 3D feature values of each optimization round, the corresponding principal axis angle changes, gradient magnitude, residual curvature value, learning rate adjustment coefficient, and gating state label. All information is stored in a structured log file in chronological order for reference in the final modeling quality diagnosis and possible re-optimization path restart. Simultaneously, all updated Gaussian points will synchronously output their new covariance matrix and fitting confidence score, feeding back into the next step of the point cloud overall morphology convergence control process, forming a closed-loop coordination mechanism between numerical optimization and structural modeling.
[0167] In this embodiment, during the optimization process based on multi-view image error backpropagation, the covariance parameter update of each point in the Gaussian point cloud (i.e., the Gaussian 3D point in the intermediate 3D point cloud model) is both stable and effective, thus avoiding overall model distortion, abnormal ellipsoidal shape, or structural continuity disruption caused by numerical non-convergence, abrupt updates, or incorrect optimization directions. Since the covariance matrix of the 3D Gaussian points directly determines the spatial diffusion pattern and projection performance of each point, its update process must be precisely adjusted while maintaining numerical stability. However, in actual optimization, the gradient may fluctuate drastically due to factors such as illumination interference, image noise, sparse texture, or local occlusion, leading to unexpected oscillations or getting trapped in local minima in the optimization path. To address this issue, this step introduces a gradient convergence dynamic monitoring mechanism based on reprojection consistency verification. By continuously tracking the gradient magnitude change of the Gaussian point covariance parameter in each round of backpropagation and combining it with the error curvature change trend in the pixel residual map, it determines whether the point is in a "stable convergence" or "abnormal fluctuation" state. Once an unstable gradient or abrupt error region is detected, the learning rate and update step rhythm at that point will be automatically adjusted, such as by reducing the step size, introducing historical momentum, or delaying updates, thereby mitigating divergence and enhancing numerical convergence. Simultaneously, to ensure model controllability and backtracking of the optimization process, this step also records the covariance evolution trajectory of each Gaussian point throughout the entire training process, including eigenvalue changes, principal axis rotation, gradient magnitude, and adjustment coefficients, constructing a structured log. This provides crucial reference for subsequent model repair, error backtracking, and global structure evaluation. In summary, this step not only provides a dynamic steady-state control mechanism for the efficient optimization of Gaussian point cloud covariance parameters but also establishes an important guarantee for the continuous improvement and consistency maintenance of the model's global quality, making it an indispensable supporting component for achieving a high-precision 3D reconstruction system.
[0168] Step S62: Based on the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model, optimize the intermediate 3D point cloud model and output the target 3D point cloud model.
[0169] It should be noted that, based on the evolution trajectory of the covariance parameters of each Gaussian point, the fitted confidence mapping tensor, the error distribution between the rendered photometric map and the actual sampled map, and the local resampling operation records, a global Gaussian ellipsoid dynamic adjustment model is constructed. This model performs iterative covariance calibration and ellipsoid morphology convergence control on the point cloud of the entire scene, thereby enhancing the continuity of the point cloud structure, improving image rendering accuracy, and ensuring the reliability of spatial measurements. Specifically, to further improve the overall performance of the 3D point cloud model in terms of structural coherence, image rendering accuracy, and spatial measurement stability, a global Gaussian ellipsoid dynamic adjustment model is proposed for the point cloud data that has already undergone local optimization, Gaussian covariance adjustment, multi-view error verification, and resampling. This model uses the historical evolution trajectory of the covariance parameters of each Gaussian point, the fitted confidence mapping tensor, the error distribution information between the rendered photometric map and the actual sampled map, and the local resampling records as inputs to construct an iterative covariance calibration mechanism across the entire scene, completing the gradual convergence control of the Gaussian ellipsoid morphology globally.
[0170] Specifically, when initializing the global dynamic adjustment model, i.e., when optimizing the intermediate 3D point cloud model, the evolution trajectory of the covariance parameters of all Gaussian points recorded during continuous backpropagation iterations is uniformly loaded (the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model). The trajectory information of each Gaussian point includes the principal axis eigenvalue sequence (λ1, λ2, λ3) in each iteration, the principal axis direction unit vector, the gradient magnitude change rate, the convergence judgment label (whether it has entered the stable descent interval), the learning rate adjustment factor, and whether the gating constraint flag has been triggered due to instability. At the same time, the confidence values of all Gaussian points in the current frame are extracted from the fitting confidence mapping tensor, ranging from 0 to 1. Points that meet the following two conditions: ① the covariance gradient oscillation cycle is greater than 5 times; ② the current confidence value is less than 0.4, are marked as "high adjustment priority points". This marking will directly affect the adjustment amplitude and iteration frequency in subsequent global morphological adjustments to avoid potential numerical instability points interfering with the overall structure.
[0171] For each Gaussian point, the residual distribution between the actual image pixel values and the rendered values is extracted from all the rendered photometric maps it participated in generating. The point is projected onto the corresponding image plane through viewpoint number matching, and a residual map sub-window (e.g., an 11×11 pixel range) of its affected area is collected. The average photometric residual value and standard deviation of the point under each viewpoint are calculated. The average is taken across all viewpoints to construct the "multi-view reprojection error vector" for that point. Points with error values greater than 0.15 (unit: normalized grayscale) are recorded individually, and combined with a label indicating whether reconstruction failed during local resampling operations, a "structural instability indicator matrix" is constructed. This matrix is subsequently used to guide whether a penalty term should be added or intensity control should be implemented in regional covariance adjustment.
[0172] Initiate an iterative covariance calibration process for the entire Gaussian point cloud. Divide the entire point cloud space into a cubic grid with a side length of 0.5 meters. For the Gaussian point set within each grid cell, calculate the following three metrics: ① polar coordinate distribution entropy along the principal axis, assessing the consistency of the ellipsoidal orientation in that region; ② eigenvalue standard deviation, representing the degree of scale difference; ③ mean of fit confidence. If the entropy value is greater than 0.85, the scale standard deviation exceeds 0.1, and the mean confidence is less than 0.6, then the voxel cell is determined to be a "locally non-convergent region." Within these regions, a joint covariance adjustment strategy is adopted: the mean vector of the principal axis of the local point set is used as the adjustment reference for the principal axis orientation of each point, a micro-rotation operation is performed, and the maximum eigenvalue is limited to the 80th percentile of the eigenvalues of all points within the grid to prevent occlusion errors caused by local extremum expansion. After this adjustment, recalculate the covariance matrix and perform positive definiteness verification. If a non-positive definite case occurs, a minimum eigenvalue truncation strategy is introduced to restore it to an invertible matrix.
[0173] Global consistency verification is performed on the adjusted point cloud set, and the decision to proceed to the next round of global fine-tuning is based on the verification results. The verification process involves regenerating the Gaussian point cloud rendering photometric map for each viewpoint and comparing it pixel-by-pixel with the original captured image. The average residual, standard deviation, number of high residual pixels, and their proportion of the total pixels are calculated. If the residual decreases by more than 15% and the area of high residual regions decreases by more than 20%, the adjustment is considered effective, and the process proceeds to the next round of fine-tuning, updating the fitting confidence score of all Gaussian points. Conversely, if residuals rebound or structural distortion increases in some regions after adjustment, the process rolls back to the previous state based on the covariance evolution trajectory, reallocating mesh adjustment weights or delaying convergence speed. Finally, if the global average residual is below 0.08, the residual standard deviation is below 0.03, and no new unstable regions are added in two consecutive rounds of adjustment, the iteration terminates, the covariance matrix and principal axis direction of all Gaussian points are frozen, and the finally converged point cloud model (i.e., the target 3D point cloud model) is output.
[0174] In this embodiment, the present invention integrates local adjustment with global consistency to construct a Gaussian ellipsoid dynamic adjustment model covering the entire 3D scene point cloud range. This addresses issues such as morphological divergence, principal axis imbalance, and scale inconsistency that arise during the iterative optimization of the Gaussian point covariance matrix, thereby achieving an overall improvement in the point cloud model's structural continuity, image rendering quality, and measurement reliability. In complex scenes, due to factors such as texture loss, occlusion, and viewpoint changes, local point clouds are prone to forming ellipsoidal structures with inconsistent orientations and abrupt scale changes during covariance updates. These anomalies not only affect the visual naturalness and fidelity of the 3D model but also cause error concentration during image reprojection, thus affecting the measurement accuracy and stability of the point cloud. This step integrates dynamic information from multiple sources, including the evolution trajectory of covariance parameters of each Gaussian point during continuous iteration (such as principal axis eigenvalues, orientation rotation angles, and gradient change trends), the fitting reliability mapping tensor reflecting morphological reliability, the pixel-by-pixel error distribution between the rendered photometric map and the actual image, and records of previously triggered local resampling behaviors, to construct a comprehensive state judgment basis across scales and time dimensions. Building upon this foundation, this step guides a unified morphological calibration operation on the global covariance matrix. Through spatial region segmentation and index screening mechanisms, it identifies structurally inconsistent regions and performs operations such as local principal axis alignment, scale difference smoothing, and matrix positive definiteness repair. This promotes the co-evolution of all Gaussian point cloud units, ultimately achieving morphological convergence of the model at a global scale. Simultaneously, through multi-round rendering residual analysis and credibility feedback, this step also implements a dynamic self-feedback mechanism, ensuring that each morphological adjustment is based on real image observations and possesses optimization loop capabilities. This process effectively avoids structural breakage issues caused by single-point adjustments, improves the overall geometric stability of the point cloud, image reprojection accuracy, and spatial measurement consistency, and is a key link in achieving engineering reliability and accuracy control within the entire 3D modeling scheme.
[0175] For comparison of technical effects, existing technologies can be used as a reference. In the reconstruction of 3D Gaussian sputtered point clouds, a covariance matrix needs to be defined for the Gaussian points to characterize their spatial distribution. However, in areas with sparse textures such as metallic highlights and solid-color backgrounds, traditional strategies are prone to causing ill-conditioned covariance matrices (abnormal eigenvalues): this leads to distorted rendering ellipsoids, causing occlusion and boundary distortion; and gradient anomalies interfere with optimization convergence, compromising model stability and measurement accuracy. For more information on these issues, please refer to [link to relevant documentation]. Figure 2This invention proposes a point cloud reconstruction method based on 3D Gaussian sputtering. Guided by a covariance degradation risk probability map, it introduces a covariance health assessment mechanism to identify unstable regions. It achieves Gaussian ellipsoidal directional reconstruction through normal entropy constraints and eigenvalue compression, avoiding ellipsoidal distortion common in traditional methods. A dynamic resampling and optimization gating mechanism, combined with fitting reliability, is constructed to ensure stable parameter updates. Global Gaussian ellipsoid dynamic adjustment achieves coordinated convergence of the entire point cloud morphology, forming an error closed-loop adjustment. This method improves the 3D reconstruction accuracy of complex structural regions and enhances the engineering usability of non-contact measurement. Especially in scenes with scarce textures or limited viewpoints, it significantly enhances the structural stability and measurement reliability of point cloud models, making it suitable for non-contact measurement in complex scenes. It possesses practical value and technological advancement.
[0176] In this embodiment of the invention, a point cloud reconstruction method based on three-dimensional Gaussian sputtering is provided. The method acquires multi-view images and a reference view image, and constructs a covariance degradation risk probability map based on these images. Multiple Gaussian three-dimensional points to be adjusted are generated based on the covariance degradation risk probability map. Three-dimensional point selection and fitting confidence score calculation are performed on each of the Gaussian three-dimensional points to be adjusted, outputting multiple second-order adjusted three-dimensional points and their corresponding fitting confidence scores. An initial three-dimensional point cloud model is constructed based on these multiple second-order adjusted three-dimensional points. The second-order adjusted three-dimensional points are optimized based on the multi-view images and their corresponding fitting confidence scores, outputting multiple target Gaussian three-dimensional points. The invention employs multiple target Gaussian 3D points to update the initial 3D point cloud model, determining an intermediate 3D point cloud model. A covariance update gating mechanism based on gradient convergence dynamic monitoring is then used to update the intermediate 3D point cloud model, outputting the target 3D point cloud model. Based on this scheme, the invention pre-judges covariance anomaly risks through a covariance degradation risk probability map, reducing the generation of ill-conditioned covariance matrices from the source. By screening 3D points based on each Gaussian 3D point to be controlled, it can specifically screen and correct 3D points prone to gradient anomalies, avoiding the optimization process from getting trapped in local minima. This effectively avoids gradient propagation anomalies, optimization non-convergence, or local minima, significantly improving the structural stability of the 3D point cloud model.
[0177] Please see Figure 3 , Figure 3 This is a structural block diagram of a point cloud reconstruction system based on three-dimensional Gaussian sputtering, provided in Embodiment 2 of the present invention.
[0178] This invention provides a point cloud reconstruction system based on three-dimensional Gaussian sputtering, comprising:
[0179] The acquisition module 301 is used to acquire multi-view images and reference view images, and to construct a covariance degradation risk probability map based on the multi-view images and reference view images.
[0180] The generation module 302 is used to generate multiple Gaussian three-dimensional points to be controlled based on the covariance degradation risk probability map.
[0181] The filtering module 303 is used to filter three-dimensional points and calculate the fitting confidence score based on each Gaussian three-dimensional point to be adjusted, output multiple secondary-adjusted three-dimensional points and the fitting confidence score corresponding to each secondary-adjusted three-dimensional point, and construct an initial three-dimensional point cloud model based on the multiple secondary-adjusted three-dimensional points.
[0182] The optimization module 304 is used to optimize each of the two-dimensional adjustment points based on the multi-view images and the fitting confidence scores corresponding to each two-dimensional adjustment point, and output multiple target Gaussian three-dimensional points.
[0183] The first update module 305 is used to update the initial three-dimensional point cloud model using multiple target Gaussian three-dimensional points to determine the intermediate three-dimensional point cloud model.
[0184] The second update module 306 is used to update the intermediate 3D point cloud model using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and output the target 3D point cloud model.
[0185] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and modules described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0186] This invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the above embodiments.
[0187] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the above embodiments.
[0188] This invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any of the above embodiments.
[0189] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A point cloud reconstruction method based on three-dimensional Gaussian sputtering, characterized in that, include: Acquire multi-view images and reference view images, and construct a covariance degradation risk probability map based on the multi-view images and the reference view images; Based on the covariance degradation risk probability map, multiple Gaussian three-dimensional points to be controlled are generated. Based on the three-dimensional Gaussian points to be adjusted, three-dimensional points are screened and fitting confidence scores are calculated. Multiple secondary-adjusted three-dimensional points and the fitting confidence scores corresponding to each secondary-adjusted three-dimensional point are output. An initial three-dimensional point cloud model is constructed based on the multiple secondary-adjusted three-dimensional points. Based on the multi-view images and the fitting confidence scores corresponding to each of the secondary-adjusted 3D points, the secondary-adjusted 3D points are optimized to output multiple target Gaussian 3D points. The initial three-dimensional point cloud model is updated using multiple target Gaussian three-dimensional points to determine the intermediate three-dimensional point cloud model; The intermediate 3D point cloud model is updated using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and the target 3D point cloud model is output.
2. The point cloud reconstruction method based on three-dimensional Gaussian sputtering according to claim 1, characterized in that, The step of constructing a covariance degradation risk probability map based on the multi-view image and the reference view image includes: The multi-view image is preprocessed to output a target multi-view image, and the target image is converted into a grayscale image. The Sobel operator is used to calculate the gradient of the grayscale image, and the first-order reciprocal image in the horizontal and vertical directions is output. Based on the first-order reciprocal image in the horizontal and vertical directions, the sum of squared gradient magnitudes corresponding to the target multi-view image is calculated. The average of the squared gradient magnitudes is calculated within a local area of the grayscale image to generate a texture gradient intensity map. The texture gradient intensity map is divided to generate multiple equal-sized grids corresponding to the target multi-view image, and the texture response rate of each equal-sized grid is calculated to determine the texture response rate of each equal-sized grid. In the reference view image, a target region is selected, and a mutual information weighted block matching algorithm is used to search for the maximum luminous consistency pixel block based on the target region and each of the equal-sized grids, generating multiple matching points corresponding to each of the equal-sized grids; Multiple linear triangulations are performed on the multiple matching corresponding points to output multiple three-dimensional points corresponding to each matching corresponding point. Based on the multiple three-dimensional points corresponding to each matching corresponding point, a triangulation consistency vector corresponding to each matching corresponding point is constructed. The multiple triangulation consistency vectors are standardized to output the stability evaluation scores of each three-dimensional point corresponding to each matching point. The stability evaluation scores of each three-dimensional point and the texture response rate associated with the equal-sized mesh corresponding to each three-dimensional point are linearly interpolated and combined to calculate the degradation probability score of each three-dimensional point. Based on each of the three-dimensional points, a three-dimensional space is constructed, and the three-dimensional space is divided to generate multiple cubic voxel units; The degradation probability scores corresponding to multiple three-dimensional points in each cubic voxel are averaged to output the covariance degradation risk level corresponding to each cubic voxel. Based on the covariance degradation risk level corresponding to each cubic voxel, a covariance degradation risk probability map is constructed.
3. The point cloud reconstruction method based on three-dimensional Gaussian sputtering according to claim 2, characterized in that, The step of generating multiple Gaussian three-dimensional points to be adjusted based on the covariance degradation risk probability map includes: In the covariance degradation risk probability map, any cube voxel unit corresponding to a covariance degradation risk level greater than a preset level threshold is taken as a high-risk voxel unit, and each of the high-risk voxel units is divided to output multiple three-dimensional cube meshes. Count the number of 3D points in each of the 3D cube meshes; The three-dimensional cube mesh corresponding to any number of three-dimensional points less than the preset first point number threshold is taken as a potential degenerate mesh; Calculate the cluster density coefficient of the 3D cube mesh corresponding to the number of any 3D points that is greater than or equal to the preset second point number threshold, and compare it with the preset density coefficient threshold. Any three-dimensional cubic mesh corresponding to an aggregation density coefficient greater than the preset density coefficient threshold is considered a structurally unstable mesh. Perform back projection operations on multiple 3D points in the structurally unstable mesh and multiple 3D points in the potentially degenerate mesh to output the grayscale values of each 3D point in the structurally unstable mesh and the grayscale values of each 3D point in the potentially degenerate mesh from multiple perspectives. Based on the gray values of each three-dimensional point in the structurally unstable grid and the gray values of each three-dimensional point in the potentially degenerate grid under multiple views, calculate the sum of squares of the photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squares of the photometric residuals of each three-dimensional point in the potentially degenerate grid. The sum of squared photometric residuals of each three-dimensional point in the structurally unstable grid and the sum of squared photometric residuals of each three-dimensional point in the potentially degenerate grid are normalized to determine the photometric consistency deviation values of each three-dimensional point in the structurally unstable grid and the photometric consistency deviation values of each three-dimensional point in the potentially degenerate grid, and then compared with preset deviation thresholds respectively. Any three-dimensional point corresponding to a photometric consistency deviation value greater than the preset deviation threshold is taken as a three-dimensional reconstruction imbalance point. Perform eigenvalue decomposition on the covariance matrix corresponding to each of the three-dimensional reconstruction imbalance points to determine the principal axis eigenvectors and eigenvalues corresponding to each of the three-dimensional reconstruction imbalance points. Based on the principal axis direction feature vector and feature value corresponding to each of the three-dimensional reconstruction imbalance points, the three-dimensional reconstruction imbalance points are filtered to output multiple Gaussian three-dimensional points to be adjusted.
4. The point cloud reconstruction method based on three-dimensional Gaussian sputtering according to claim 3, characterized in that, The process of selecting 3D points and calculating fitting confidence scores based on the Gaussian 3D points to be adjusted, and outputting multiple secondary-adjusted 3D points and their corresponding fitting confidence scores, includes: Based on the feature values corresponding to each of the three-dimensional Gaussian points to be adjusted, the feature value ratio of each of the three-dimensional Gaussian points to be adjusted is calculated, and the feature value ratio of each of the three-dimensional Gaussian points to be adjusted is compared with a preset ratio threshold. Take any Gaussian 3D point to be adjusted that corresponds to a feature value ratio greater than the preset ratio threshold as a morphological anomaly point, adjust the feature value of each morphological anomaly point, and output the adjusted feature value corresponding to each morphological anomaly point. A spherical neighborhood is constructed for each of the aforementioned morphological anomalies, and the spherical neighborhood corresponding to each of the aforementioned morphological anomalies is output. The three-dimensional Gaussian points to be controlled in each of the spherical neighborhoods are taken as the nearest neighbors of each of the morphological anomalies. Based on the polar coordinate angle associated with the principal axis direction feature vectors of the nearest neighbors of each of the morphological anomalies, the spherical neighborhoods corresponding to each of the morphological anomalies are divided to determine multiple equally spaced grids corresponding to each of the morphological anomalies. The number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each of the morphological anomalies is counted, and the compressed neighborhood space entropy value corresponding to each of the morphological anomalies is calculated based on the number of principal axis direction feature vectors in multiple equally spaced grids corresponding to each of the morphological anomalies. Based on the compressed neighborhood space entropy value corresponding to each of the morphological anomalies, the entropy change rate corresponding to each of the morphological anomalies is calculated, and the entropy change rate corresponding to each of the morphological anomalies is compared with a preset change rate threshold. Any morphological anomaly point corresponding to an entropy change rate greater than the preset change rate threshold is taken as a three-dimensional point that needs fine-tuning. The adjustment feature value, principal axis direction feature vector, and covariance matrix corresponding to each three-dimensional point that needs fine-tuning are then fine-tuned to determine multiple secondary adjustment three-dimensional points. Reproject each of the two-dimensional adjustment points, output the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points, and calculate the area ratio corresponding to each of the two-dimensional adjustment points based on the area of the image reprojection coverage region corresponding to each of the two-dimensional adjustment points. Based on the compressed neighborhood space entropy value corresponding to each of the two-dimensional adjustment points, the change ratio of the neighborhood entropy value corresponding to each of the two-dimensional adjustment points is calculated. Then, a weighted average strategy is used to output the fitting confidence score corresponding to each of the two-dimensional adjustment points based on the area ratio and the change ratio of the neighborhood entropy value.
5. The point cloud reconstruction method based on three-dimensional Gaussian sputtering according to claim 4, characterized in that, The optimization of each of the secondary-adjusted 3D points is based on the multi-view images and the fitting confidence scores corresponding to each of the secondary-adjusted 3D points, outputting multiple target Gaussian 3D points, including: Based on the fitting confidence scores corresponding to each of the aforementioned secondary-adjusted 3D points, multiple moderately confident 3D points are selected. A three-dimensional Gaussian ellipsoid is constructed for each of the moderately reliable three-dimensional points, and the Gaussian ellipsoid corresponding to each of the moderately reliable three-dimensional points is output. Based on the multiple views corresponding to the multi-view image, the Gaussian ellipsoid is mapped to output the two-dimensional elliptical region of each of the moderately reliable three-dimensional points in each of the views. Based on the contribution weights and opacity values of multiple pixels in the two-dimensional elliptical regions of each medium-confidence 3D point in each viewpoint, a photometric rendering map of each medium-confidence 3D point in each viewpoint is generated. Based on the photometric rendering maps of each moderately reliable 3D point at each viewpoint, a residual map of each moderately reliable 3D point at each viewpoint is generated, and standard deviation analysis is performed on the residual maps of each moderately reliable 3D point at each viewpoint to generate the error anomaly target region of each moderately reliable 3D point at each viewpoint. Based on the image pixel coordinates associated with the error anomaly target region of each of the medium-confidence 3D points in each of the viewpoints, multiple candidate 3D point positions of each of the medium-confidence 3D points in each of the viewpoints are generated. Perform multi-view consistency verification on multiple candidate 3D points of each moderately confident 3D point in each viewpoint, and output the estimated positions of multiple real 3D mutation points of each moderately confident 3D point in each viewpoint. Local resampling is performed on the estimated positions of multiple real three-dimensional mutation points of each of the medium-confidence three-dimensional points in each of the aforementioned viewpoints to generate multiple initial Gaussian three-dimensional points; Multiple initial Gaussian 3D points are filtered to output multiple target Gaussian 3D points.
6. The point cloud reconstruction method based on three-dimensional Gaussian sputtering according to claim 1, characterized in that, The method employs a covariance update gating mechanism based on gradient convergence dynamic monitoring to update the intermediate 3D point cloud model and output the target 3D point cloud model, including: The covariance update gating mechanism based on gradient convergence dynamic monitoring is used to determine the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model according to the intermediate 3D point cloud model. Based on the historical evolution trajectory of the covariance parameters of each Gaussian 3D point in the intermediate 3D point cloud model, the intermediate 3D point cloud model is optimized to output the target 3D point cloud model.
7. A point cloud reconstruction system based on three-dimensional Gaussian sputtering, characterized in that, include: The acquisition module is used to acquire multi-view images and reference view images, and construct a covariance degradation risk probability map based on the multi-view images and the reference view images; The generation module is used to generate multiple Gaussian three-dimensional points to be controlled based on the covariance degradation risk probability map. The filtering module is used to filter three-dimensional points and calculate fitting confidence scores based on each of the Gaussian three-dimensional points to be adjusted, output multiple secondary-adjusted three-dimensional points and the fitting confidence scores corresponding to each of the secondary-adjusted three-dimensional points, and construct an initial three-dimensional point cloud model based on the multiple secondary-adjusted three-dimensional points. The optimization module is used to optimize each of the secondary adjusted 3D points based on the multi-view image and the fitting confidence score corresponding to each of the secondary adjusted 3D points, and output multiple target Gaussian 3D points. The first update module is used to update the initial three-dimensional point cloud model using multiple target Gaussian three-dimensional points to determine the intermediate three-dimensional point cloud model. The second update module is used to update the intermediate 3D point cloud model using a covariance update gating mechanism based on gradient convergence dynamic monitoring, and output the target 3D point cloud model.
8. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer performs the point cloud reconstruction method based on three-dimensional Gaussian sputtering as described in any one of claims 1-6.
Citation Information
Cited By
Historic building point cloud model completion method based on multi-modal feedback and feature purification
CN121482299A
Jaw crusher eccentric shaft anomaly detection method and system
CN121837281A
A method and system for detecting an abnormality of an eccentric shaft of a jaw crusher
CN121837281B