An underwater three-dimensional Gaussian scene reconstruction method and system based on ray-Gaussian intersection depth constraint

CN122841671APending Publication Date: 2026-09-29UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610945827.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

NeRF 利用编码在神经网络中的隐式表示建模场景的几何和外观,但计算复杂度高,难以在资源受限的水下平台上进行实时推理

Benefits of technology

本发明通过引入射线-高斯交点深度约束,摒弃了传统方法中对聚合深度施加衰减的近似策略,直接在主高斯原语的一维响应峰值深度处解析施加物理衰减,显著提升了水下场景重建的几何精度与颜色解耦能力。结合多层次细节稀疏网格结构和前景-背景掩码引导的深度平滑正则化,有效抑制了远景区域深度不连续与伪影现象,避免了过度平滑对前景细节的破坏。联合优化三维高斯原语属性与介质参数预测网络,能够自适应地模拟水下波长依赖的吸收、散射与后向散射效应,在保持实时渲染效率的同时,在多个真实与合成水下数据集上获得了更高的峰值信噪比与结构相似性指标,且在不同浑浊程度环境中均展现出良好的泛化性与几何一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841671A_ABST
    Figure CN122841671A_ABST
Patent Text Reader

Abstract

The application discloses a ray-Gaussian intersection depth constraint-based underwater three-dimensional Gaussian scene reconstruction method and system. By introducing the ray-Gaussian intersection depth constraint, the application discloses a method for directly applying physical attenuation at the one-dimensional response peak depth of the main Gaussian primitive, thereby abandoning the approximate strategy of applying attenuation to the aggregated depth in the traditional method, and significantly improving the geometric accuracy and color decoupling capability of underwater scene reconstruction. In combination with a multi-level detail sparse grid structure and foreground-background mask guided depth smoothing regularization, the application effectively suppresses the depth discontinuity and artifact phenomenon in the long-range area, and avoids the damage of excessive smoothing to the foreground details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer vision, differentiable rendering, participating medium imaging and 3D reconstruction technology, and specifically relates to an underwater 3D Gaussian scene reconstruction method and system based on ray-Gaussian intersection depth constraints. Background Technology

[0002] 3D scene reconstruction is a key technology in underwater robot vision, marine scientific research, and seabed resource exploration. It primarily utilizes imaging systems mounted on underwater vehicles to acquire multi-view images, thereby reconstructing the scene's 3D geometry and appearance. However, the underwater environment presents unique visual perception challenges. Light is absorbed and scattered by water molecules and suspended particles during propagation, leading to color distortion, reduced contrast, and decreased visibility. Furthermore, wavelength-dependent attenuation gives underwater images a distinctive blue-green hue. In addition, backscattering produces occlusion effects, introducing ghosting artifacts in multi-view reconstruction, severely impacting reconstruction quality.

[0003] In recent years, two mainstream methods have emerged in the field of 3D scene reconstruction: Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). NeRF uses implicit representations encoded in neural networks to model the geometry and appearance of a scene, but it has high computational complexity and is difficult to perform real-time inference on resource-constrained underwater platforms. 3DGS provides faster rendering speeds through explicit Gaussian primitive representations, but the standard implementation assumes uniform atmospheric conditions and does not consider the depth-dependent attenuation and wavelength-specific absorption characteristics of the underwater environment.

[0004] For underwater scenes, existing methods attempt to combine physical underwater imaging models with the aforementioned reconstruction frameworks. The SeaThru-NeRF method integrates an underwater image formation model into a neural radiation field, achieving direction-dependent sampling of medium parameters. However, based on implicit NeRF representations, it requires extensive sampling queries during training and rendering, resulting in low inference efficiency. The Water-Splatting method fuses 3DGS with volume rendering, using 3DGS to represent explicit geometry and a separate volume field to capture the scattering medium. However, the effective depth of Water-Splatting is only the projected depth of the Gaussian mean, resulting in a loose coupling between geometry and physics. The optimizer can trade off pixel colors by shifting the Gaussian position and varying the medium parameters for each ray, often leading to a large solution space with incorrect geometry. Furthermore, due to the nonlinear characteristics of exponential decay, applying attenuation or integral attenuation to the aggregate depth introduces bias. The SeaSpla method combines a physical underwater image formation model with 3DGS, but its generalization ability in complex turbid environments is limited, and it lacks depth constraints for infinite background regions. Summary of the Invention

[0005] To address the technical problems mentioned above, this invention provides a method for reconstructing an underwater 3D Gaussian scene based on ray-Gaussian intersection depth constraints, comprising the following steps: S1. Obtain a multi-view image sequence of the target underwater scene, and obtain a preprocessed multi-view image sequence based on the multi-view image sequence; S2. Based on the preprocessed multi-view image sequence, obtain the initial sparse point cloud, camera parameters, and initial set of 3D Gaussian primitives. S3. Based on the direction of each pixel ray, obtain the participating medium parameters corresponding to that ray; S4. Based on the set of three-dimensional Gaussian primitives ordered along the ray and the participating medium parameters, obtain the depth of the ray-Gaussian intersection point, and obtain the final rendering color of the pixel based on the depth of the ray-Gaussian intersection point. S5. Based on the multi-level sparse mesh structure, obtain the foreground-background binary mask, and based on the foreground-background binary mask, obtain the depth smoothing regularization loss of the background region; S6. Based on the reconstruction loss between the final rendered color and the real color, as well as the depth smoothing regularization loss, obtain the optimized 3D Gaussian primitive properties and medium parameters to predict the network parameters. S7. Based on the target viewpoint parameters and the optimized underwater scene reconstruction model, obtain the rendering results from the target viewpoint.

[0006] Preferably, S3 includes: obtaining a spherical harmonic coding feature vector through a spherical harmonic coding layer based on the direction of each pixel ray; obtaining the participating medium parameters through a medium parameter prediction MLP based on the spherical harmonic coding feature vector; wherein the participating medium parameters include medium color, backscattering coefficient, and attenuation coefficient.

[0007] Preferably, in step S4, the step of obtaining the ray-Gaussian intersection depth based on the set of three-dimensional Gaussian primitives ordered along the ray and the participating medium parameters includes: Based on the set of three-dimensional Gaussian primitives ordered along the ray, each three-dimensional Gaussian primitive is converted into a one-dimensional probability density function along the ray; the cumulative transmittance is calculated based on the contribution weight of the one-dimensional response peak of each Gaussian primitive; the principal Gaussian primitive is determined based on the condition that the cumulative transmittance first drops below a preset threshold; and the depth of the ray-Gaussian intersection is obtained by analytical formula based on the principal Gaussian primitive.

[0008] Preferably, step S4, obtaining the final rendered color of the pixel based on the ray-Gaussian intersection depth, includes: The object color component is obtained based on the attenuation coefficient and the ray-Gaussian intersection depth in the participating medium parameters; the medium color component is obtained based on the backscattering coefficient, the medium color, and the ray-Gaussian intersection depth in the participating medium parameters; the object color component and the medium color component are added together to obtain the final rendered color of the pixel.

[0009] Preferably, S5 includes: based on the multi-level sparse mesh structure, counting the number of intersections between each ray and the finest LOD level Gaussian primitive; generating a foreground-background binary mask based on the comparison result of the number of intersections and the density threshold; and calculating the depth smoothing regularization loss for the background region identified by the foreground-background binary mask.

[0010] Preferably, S6 includes: Based on the rendered color and the real color, obtain the HDR reconstruction loss; based on the distribution of Gaussian primitives along the ray, obtain the depth distortion loss; based on the weighted sum of the HDR reconstruction loss, depth distortion loss, and depth smoothing regularization loss, obtain the total loss; based on the total loss, use the Adam optimizer to obtain the optimized 3D Gaussian primitive properties and medium parameters to predict the network parameters.

[0011] Preferably, S6 further includes: During the optimization process, cloning or splitting operations are performed on Gaussian primitives whose position gradient exceeds a preset threshold; pruning operations are performed on Gaussian primitives whose opacity is below a preset threshold, whose scale exceeds a preset range, or whose depth exceeds the effective range.

[0012] Preferably, S7 includes: Receive user-specified target viewpoint parameters, including camera pose and intrinsic parameters; input the target viewpoint parameters into the trained underwater scene reconstruction model, execute the rendering pipeline, and obtain underwater rendered images, depth maps, object color maps, descattered images, 3D Gaussian scene models, and / or video sequences from the target viewpoint.

[0013] This invention also provides an underwater 3D Gaussian scene reconstruction system based on ray-Gaussian intersection depth constraints. The system is used to implement the above method and includes: a data preprocessing and initialization module, a medium parameter prediction module, a ray-Gaussian surface depth calculation module, a physically guided rendering module, a LOD regularization module, a joint optimization module, and a new perspective output module.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention introduces ray-Gaussian intersection depth constraints, abandoning the approximate strategy of applying attenuation to the aggregation depth in traditional methods. Instead, it directly applies physical attenuation analytically at the one-dimensional response peak depth of the main Gaussian primitive, significantly improving the geometric accuracy and color decoupling capability of underwater scene reconstruction. Combined with a multi-level detailed sparse mesh structure and foreground-background mask-guided depth smoothing regularization, it effectively suppresses depth discontinuities and artifacts in distant regions, avoiding the destruction of foreground details by excessive smoothing. The jointly optimized 3D Gaussian primitive attribute and medium parameter prediction network can adaptively simulate underwater wavelength-dependent absorption, scattering, and backscattering effects. While maintaining real-time rendering efficiency, it achieves higher peak signal-to-noise ratio and structural similarity indices on multiple real and synthetic underwater datasets, and exhibits good generalization and geometric consistency in environments with varying turbidity levels. Attached Figure Description

[0015] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating the overall framework of the underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraints provided in this embodiment of the invention. Figure 2 This is a schematic diagram of ray-Gaussian intersection geometry calculation according to an embodiment of the present invention, where (a) shows the relationship between Gaussian primitives and rays in the traditional method, and (b) shows the transformation of Gaussian primitives along rays into one-dimensional forms and constraint to surface depth t in the method of the present invention. The illustration; Figure 3 This is a schematic diagram of the LOD-aware regularization mechanism in an embodiment of the present invention, illustrating the process of distinguishing the foreground surface and sparse background medium using hierarchical scaling of Gaussian primitives. Figure 4 This is a visual comparison of the embodiments of the present invention and existing methods on the SeaThru-NeRF dataset; Figure 5 This is a comparison chart of the ablation experiment results in the embodiments of the present invention; Figure 6 A schematic diagram of the system module structure provided in this embodiment of the invention; Figure 7 A schematic diagram of the new perspective rendering and result output process in an embodiment of the present invention; Figure 8 A schematic diagram of the electronic device structure in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Example 1 The overall framework and flow of the reconstruction method in this embodiment are as follows: Figure 1 As shown, its specific implementation method is as follows: S1. Obtain a multi-view image sequence of the target underwater scene, and obtain a preprocessed multi-view image sequence based on the multi-view image sequence.

[0020] S101. Acquire a multi-view image sequence of the target underwater scene. The images can be acquired by a camera mounted on an underwater vehicle, and the number of images is usually tens to hundreds.

[0021] S102. Perform data preprocessing on the acquired multi-view image sequence. Specifically, perform quality screening, distortion correction, noise filtering, resolution unification, and color normalization on the multi-view images, and convert the images from the sRGB color space to the linear RGB color space to eliminate the influence of gamma correction.

[0022] S103. Input the preprocessed image sequence into SfM tools such as COLMAP for processing. Obtain the correspondence between images through feature extraction and matching. Optimize with Bundle Adjustment to obtain the initial sparse point cloud and the camera intrinsic and extrinsic parameters (pose) corresponding to each image.

[0023] S104. Organize the initial sparse point cloud into a multi-level sparse grid structure according to spatial location. Each grid cell stores a list of Gaussian primitive indices corresponding to that level, supporting fast lookup. Calculate the LOD level of each Gaussian primitive using the following formula: in, H Total hierarchical depth (in this embodiment) H =4), tmax Represented as the preset percentile (preferably the 95th percentile) of the scene's visible depth. tIt represents the statistical value of the distance from the Gaussian primitive to the current camera center or to the centers of multiple training cameras. A higher LOD level indicates finer foreground details, while a lower LOD level indicates sparser background areas.

[0024] S2. Based on the preprocessed multi-view image sequence, obtain the initial sparse point cloud, camera parameters, and initial set of 3D Gaussian primitives.

[0025] S201, Using each 3D point in the initial sparse point cloud as the position p of the Gaussian primitive. k .

[0026] S202. For each point, estimate the local point cloud density from its k nearest neighbors, and initialize the Gaussian scaling parameter S accordingly. k The scale parameter is a three-dimensional diagonal matrix, with each dimension reflecting the degree of spatial expansion in the corresponding direction.

[0027] S203, Rotation parameter R k Initialize to the rotation matrix corresponding to the unit quaternion.

[0028] S204, Opacity αk Initialize to a preset initial value (e.g., 0.1).

[0029] S205, the spherical harmonic color coefficients are obtained by sampling the pixel colors of the corresponding images at each viewpoint.

[0030] Specifically, each Gaussian primitive is projected onto each visible viewpoint, and the linear RGB values ​​of the corresponding pixels are sampled to initialize the zeroth-order coefficients of the spherical harmonic function.

[0031] S3. Based on the direction of each pixel ray, obtain the participating medium parameters corresponding to that ray.

[0032] S301. For each pixel under a given camera viewpoint, calculate its corresponding ray direction v.

[0033] S302. Encode the normalized direction of the ray using a 4th-order spherical harmonic function to obtain the spherical harmonic encoded feature vector.

[0034] S303. Input the spherical harmonic encoded feature vector into the medium parameter prediction MLP (containing 32 hidden units), and output the three participating medium parameters corresponding to the ray: medium color. cmed The backscattering coefficient is activated by the Sigmoid function and has a range of [0,1]. σbs The non-negativity is ensured by using the Softplus activation function; the decay coefficient... σattn The Softplus activation function ensures non-negativity. σattn and σbsIt can be a scalar, an RGB three-channel vector, or a parameter vector corresponding to a color channel; when it is a color channel vector, the exponentiation operation is performed on a channel-by-channel basis.

[0035] The design treats the medium properties as constant along a single camera ray, and the depth-dependent effect is introduced analytically.

[0036] S4. Based on the set of three-dimensional Gaussian primitives ordered along the ray and the participating medium parameters, obtain the ray-Gaussian intersection depth, and based on the ray-Gaussian intersection depth, obtain the final rendering color of the pixel.

[0037] S401, reference Figure 2 For each pixel ray Obtain all three-dimensional Gaussian primitives sorted along this ray. .

[0038] S402. Transform the spatial probability density function of each three-dimensional Gaussian primitive into a one-dimensional probability density function along the ray. The quadratic coefficients in the world coordinate system are: Through rotational decomposition, it is equivalently represented in the local Gaussian coordinate system as: in, ; Bk Gaussian primitives k The coefficient of the first-order term after being converted to one-dimensional quadratic exponential form along the ray; Ck represents the constant term in the form of a one-dimensional quadratic exponent; o represents the camera center or the origin of the ray; Sk The three-dimensional scaling matrix of the Gaussian primitive k is represented by T; T represents the matrix or vector transpose; v represents the direction vector of the pixel ray in the world coordinate system. k d represents the local ray direction of v after Gaussian rotation matrix transformation. k p represents the camera center relative to the Gaussian center. k The local displacement vector after rotation transformation.

[0039] because It is a positive definite diagonal matrix. Heng was established.

[0040] S403. Calculate the one-dimensional peak contribution weight of each Gaussian primitive along the ray. ωq : Right now ωqThe occlusion contribution is obtained by normalizing or clipping the one-dimensional Gaussian response after opacity normalization, satisfying... .

[0041] S404. Traverse the Gaussian primitives along the ray in depth order and calculate the cumulative transmittance. Find the first time the cumulative transmittance drops to the threshold. The following are Gaussian primitives: in, j This indicates the candidate Gaussian primitive number after being sorted by depth along the ray; q This indicates the Gaussian primitive number in the cumulative transmittance multiplication term.

[0042] S405, Calculate surface depth using analytical formulas .because ,in, t This represents the depth of the ray-Gaussian intersection point obtained analytically. k Indicates the index of the main Gaussian primitive. Ak , Bk Let represent the coefficients of the quadratic and linear terms of the one-dimensional quadratic function corresponding to the principal Gaussian primitive, respectively. This formula gives the minimum point of the quadratic function, i.e., the peak depth of the one-dimensional response of the Gaussian primitive along the ray. In this embodiment, this peak depth serves as an effective representation of the depth of the ray-Gaussian intersection point. Further... Cut off to the effective depth range The depth is differentiable with respect to the position, rotation, and scale parameters of the Gaussian primitive, allowing gradient backpropagation.

[0043] Step 406: Based on the participating medium parameters obtained in S3 and the surface depth calculated in step 405. The underwater rendering equation, guided by physics, decomposes the rendered color into object color components: Where Cobj(r) represents the object color component corresponding to ray r, and e represents the base of the natural exponential function. Indicates the first i The cumulative transmittance before object sampling or Gaussian primitives. Indicates the first i The object color is sampled or derived from a Gaussian primitive. i This indicates the object sample or Gaussian primitive number ordered along the ray. σobj This represents the density or opacity attenuation coefficient corresponding to the color component of an object.

[0044] and medium color components: Here, cmed represents the medium color or backscattered color vector output by the network that predicts the medium parameters.

[0045] The final underwater rendering colors are: Physical attenuation at analytically calculated surface depth Apply at *, rather than at the aggregation depth or the desired depth.

[0046] S407. The above rendering process is executed through a CUDA-accelerated differential rasterizer to generate the final underwater rendered image.

[0047] S5. Based on the multi-level sparse mesh structure, obtain the foreground-background binary mask, and based on the foreground-background binary mask, obtain the depth smoothing regularization loss of the background region.

[0048] like Figure 3 The diagram shown illustrates the LOD-aware depth regularization process in this step.

[0049] Step 501: Using the multi-level sparse mesh structure constructed in S1, count the number of intersections between each ray and the Gaussian primitive at the finest LOD level (highest LOD). A foreground-background binary mask is generated based on a density threshold τ (preferably 5, but can also be an adjustable parameter): in, Nh (r) represents the number of intersections between ray r and the finest LOD level Gaussian primitive.

[0050] S502. Calculate the depth smoothing regularization loss for the background region of the mask identifier to constrain the depth continuity of the distant view: in, Represents pixels ( i , j The foreground-background binary mask at position () is used, where ∇x represents the differential gradient operator along the horizontal direction of the image. This represents the depth value at pixel (i,j).

[0051] The background smoothing loss described above is applied only to the background region to avoid over-smoothing of foreground geometric details.

[0052] S6. Based on the reconstruction loss between the final rendered color and the real color, as well as the depth smoothing regularization loss, obtain the optimized 3D Gaussian primitive properties and medium parameters to predict the network parameters.

[0053] S601, Calculate the HDR reconstruction loss between rendered colors and real colors. The loss is calculated in the tone mapping space.

[0054] S602, Calculate Depth Distortion Loss This loss penalizes the dispersion of Gaussian primitives along the ray, causing Gaussian primitives to concentrate near the real geometric surface.

[0055] S603, Calculate the total loss In one exemplary embodiment, it is preferable to However, the above values ​​do not constitute a limitation on the scope of protection.

[0056] S604 uses the Adam optimizer for backpropagation to jointly optimize the 3D Gaussian primitive properties (position, rotation, scale, opacity, spherical harmonic color coefficient) and medium parameters to predict network parameters.

[0057] Step 605: Periodically perform adaptive densification operations during the optimization process: for location gradients exceeding a preset threshold τ The `gradient` Gaussian primitive is cloned or split. Cloning replicates a new Gaussian primitive near its original location, while splitting replaces a single Gaussian primitive with two smaller-scale Gaussian primitives. In one exemplary embodiment, the gradient threshold... The intensive operation is performed once every 500 iterations, within the first 15,000 iterations.

[0058] S606. Periodically perform pruning operations during the optimization process: remove Gaussian primitives with opacity below a preset threshold, and remove Gaussian primitives with excessively large scales or low long-term contributions. In an exemplary embodiment, the opacity pruning threshold is... Remove the corresponding Gaussian primitive; Gaussian primitives with a scale exceeding 1 / 10 of the scene diagonal are considered too large and removed; for depth values... Exceeding the effective depth range of the current scene Gaussian primitives are identified as depth anomalies and removed. This depth anomaly pruning operation prevents individual Gaussian primitives from drifting to physically unreasonable positions due to gradient instability during optimization, forming a complementary constraint with the depth truncation operation in S4. The pruning operation is executed synchronously with the compaction operation, also once every 500 iterations.

[0059] S607. After each intensive and pruning operation, reset the moving average of the relevant parameters in the Adam optimizer to avoid historical gradient information interfering with newly introduced or reset Gaussian primitives. The above thresholds and period values ​​do not constitute a limitation on the protection range.

[0060] S7. Based on the target viewpoint parameters and the optimized underwater scene reconstruction model, obtain the rendering results from the target viewpoint.

[0061] S701: Receives user-specified target view parameters, including camera pose and intrinsic parameters.

[0062] S702. Input the target viewpoint parameters into the trained underwater scene reconstruction model, execute the rendering pipeline of stage three to stage four, and generate the rendering result under the target viewpoint.

[0063] S703, output one or more of the following results: underwater rendered image from the target's perspective, depth map, object color map, descattered image, 3D Gaussian scene model, or video sequence generated from multiple consecutive frames.

[0064] Example 2 Reference Figure 6 The underwater 3D Gaussian scene reconstruction system based on ray-Gaussian intersection depth constraint provided in this embodiment of the invention includes the following modules: The data preprocessing and initialization module receives multi-view image sequences of underwater scenes as input. It performs preprocessing operations such as quality screening, distortion correction, noise filtering, resolution unification, color normalization, and color space conversion. It then uses the SfM tool to obtain initial sparse point clouds and camera parameters. Based on these initial sparse point clouds, it initializes a set of 3D Gaussian primitives and organizes the point clouds into a multi-level sparse mesh structure according to their spatial location. This module outputs the initial point cloud and mesh structure to the ray-Gaussian surface depth calculation module and the LOD regularization module.

[0065] The medium parameter prediction module comprises a spherical harmonic coding layer and an MLP inference unit. The spherical harmonic coding layer encodes the direction of each pixel ray using a 4-level spherical harmonic function. The MLP inference unit contains 32 hidden units, outputting the medium color, backscattering coefficient, and attenuation coefficient. σ attn and σ bs can be a scalar or a color channel vector. This module will participate in the output of media parameters to the physically guided rendering module.

[0066] The ray-Gaussian surface depth calculation module includes a Gaussian projection unit, a ray sorting unit, a one-dimensional probability density transformation unit, an opacity contribution calculation unit, a principal Gaussian determination unit, and a surface depth analysis unit. The Gaussian projection unit projects a 3D Gaussian primitive onto the current viewpoint; the ray sorting unit sorts the Gaussian primitives by depth; and the one-dimensional probability density transformation unit converts the 3D Gaussian primitive into a one-dimensional quadratic exponential form along the ray, explicitly calculating the quadratic coefficients. Ak , Bk , Ck The opacity contribution calculation unit calculates the contribution through cropping. ωqThe principal Gaussian determinant is determined based on the cumulative transmittance threshold to determine the principal Gaussian primitive. k *; Surface depth analytical unit through analytical formula Calculate the surface depth and truncate it to the effective range. This module will calculate the surface depth. t * Output to the Physically Guided Rendering module.

[0067] Physically Guided Rendering Module: This module includes an object component calculation unit, a medium component calculation unit, and a final color compositing unit. The object component calculation unit is based on surface depth. t * and attenuation coefficient are used to calculate the object's color components; the medium component calculation unit is based on surface depth. t * The backscattering coefficient is used to calculate the medium color component; the final color synthesis unit adds the object color component and the medium color component to obtain the final rendered color. .

[0068] The LOD regularization module includes an LOD level calculation unit, a ray cross-statistics unit, a mask generation unit, and a background depth smoothing loss calculation unit. The LOD level calculation unit calculates the LOD level based on the distance from the Gaussian primitive to the camera; the ray cross-statistics unit counts the number of intersections between each ray and the highest LOD level Gaussian primitive; the mask generation unit generates a foreground-background binary mask based on a density threshold; and the background depth smoothing loss calculation unit calculates the depth smoothing regularization loss for the background region. This module outputs the regularization loss to the joint optimization module.

[0069] Joint optimization module: Based on the total loss function (weighted sum of HDR reconstruction loss, depth distortion loss and LOD smoothing loss), the Adam optimizer performs end-to-end joint optimization of the network parameters for predicting 3D Gaussian primitive attributes and medium parameters, and periodically performs densification and pruning operations during the optimization process.

[0070] New Perspective Output Module: Receives target perspective parameters, calls the ray-Gaussian surface depth calculation module and the physical guided rendering module to generate rendering results under the target perspective, and outputs underwater rendered images, depth maps, object color maps, descattered images, 3D Gaussian scene models and / or video sequences.

[0071] Example 3 To verify the effectiveness of the present invention, experimental evaluations were conducted on multiple benchmark datasets.

[0072] The datasets include: (1) the SeaThru-NeRF dataset, containing real underwater scenes from locations such as the IUI3 Red Sea, Curaçao, Japanese Gardens Red Sea, and Panama; (2) the Mip-NeRF-360 dataset, which synthesizes fog effects using a standard atmospheric scattering model; and (3) the NeRF-Synthetic dataset, which synthesizes fog effects in synthetic environments using estimated depth maps. Comparison methods include 3DGS, SeaThru-NeRF, Water-Splatting, and SeaSplat. Evaluation metrics are PSNR, SSIM, and LPIPS.

[0073] Reference Figure 4 On the SeaThru-NeRF dataset, this invention achieved good performance in multiple scenarios. In the IUI3 scenario, it achieved PSNR=29.99, SSIM=0.895, and LPIPS=0.182; in the Curaçao scenario, it achieved PSNR=32.42, SSIM=0.933, and LPIPS=0.109; in the Japanese Gardens scenario, it achieved PSNR=24.57, SSIM=0.901, and LPIPS=0.114; and in the Panama scenario, it achieved PSNR=32.03, SSIM=0.933, and LPIPS=0.071.

[0074] On the Mip-NeRF-360 fog dataset, this invention achieves PSNR=16.23, SSIM=0.390, and LPIPS=0.550 in simple fog scenarios, and PSNR=13.85, SSIM=0.480, and LPIPS=0.510 in difficult fog scenarios.

[0075] On the NeRF-Synthetic dataset, this invention achieves PSNR=15.66, SSIM=0.511, and LPIPS=0.581 in simple fog scenes, and PSNR=14.35, SSIM=0.475, and LPIPS=0.598 in difficult fog scenes.

[0076] Reference Figure 5The ablation experiments verified the effectiveness of each component. (1) Effectiveness of LOD regularization: In the IUI3 scenario, the PSNR increased from 29.83 to 29.99 after regularization; in the Curaçao scenario, the PSNR increased from 31.53 to 32.42; in the Panama scenario, the PSNR increased from 31.16 to 32.03. (2) Comparison of integration strategies: The complete pipeline proposed in this invention showed good performance in multiple scenarios. (3) Parameter sensitivity: In the Curaçao scenario, the total hierarchical depth H=4 and the density threshold τ=5 are the optimal configurations.

[0077] Example 4 Reference Figure 7 This embodiment describes the process of rendering and outputting a new perspective based on a trained underwater scene reconstruction model.

[0078] S701: Receive the target view parameters specified by the user, including camera pose (rotation matrix and translation vector) and camera intrinsic parameters (focal length, principal point, distortion coefficient).

[0079] S702. Input the target viewpoint parameters into the trained underwater scene reconstruction model and execute the following rendering pipeline: For each pixel in the target viewpoint, construct the corresponding camera ray; use the medium parameter prediction network to obtain the participating medium parameters of the ray; convert the three-dimensional Gaussian primitive along the ray into a one-dimensional quadratic exponential form, determine the principal Gaussian primitive, and calculate the surface depth analytically. t *; Physics-guided underwater rendering equations in t * Decouples the object color and medium color at the * location to generate the final rendered color for that pixel.

[0080] S703, output one or more of the following results: underwater rendered image from the target's perspective, depth map, object color map (object surface color after removing medium attenuation), descattered image, 3D Gaussian scene model, or video sequence generated from multiple consecutive frames.

[0081] Example 5 Reference Figure 8 This invention also provides an electronic device, including at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed, enable the at least one processor to perform the method described in Embodiment 1. The processor may be one or more of a central processing unit (CPU), a graphics processing unit (GPU), or a tensor processing unit (TPU). The memory may include high-speed random access memory and / or non-volatile memory. The electronic device may also include an input interface and an output interface for receiving external data and outputting processing results.

[0082] Example 6 This invention also provides a non-transient computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1. The computer-readable storage medium includes, but is not limited to, read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, USB flash drive, or solid-state drive.

[0083] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for reconstructing an underwater 3D Gaussian scene based on ray-Gaussian intersection depth constraints, characterized in that, Includes the following steps: S1. Obtain a multi-view image sequence of the target underwater scene, and obtain a preprocessed multi-view image sequence based on the multi-view image sequence; S2. Based on the preprocessed multi-view image sequence, obtain the initial sparse point cloud, camera parameters, and initial set of 3D Gaussian primitives. S3. Based on the direction of each pixel ray, obtain the participating medium parameters corresponding to that ray; S4. Based on the set of three-dimensional Gaussian primitives ordered along the ray and the participating medium parameters, obtain the depth of the ray-Gaussian intersection point, and obtain the final rendering color of the pixel based on the depth of the ray-Gaussian intersection point. S5. Based on the multi-level sparse mesh structure, obtain the foreground-background binary mask, and based on the foreground-background binary mask, obtain the depth smoothing regularization loss of the background region; S6. Based on the reconstruction loss between the final rendered color and the real color, as well as the depth smoothing regularization loss, obtain the optimized 3D Gaussian primitive properties and medium parameters to predict the network parameters. S7. Based on the target viewpoint parameters and the optimized underwater scene reconstruction model, obtain the rendering results from the target viewpoint.

2. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, S3 includes: obtaining a spherical harmonic coding feature vector through a spherical harmonic coding layer based on the direction of each pixel ray; obtaining the participating medium parameters through a medium parameter prediction MLP based on the spherical harmonic coding feature vector; wherein the participating medium parameters include medium color, backscattering coefficient, and attenuation coefficient.

3. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, In step S4, the step of obtaining the ray-Gaussian intersection depth based on the set of three-dimensional Gaussian primitives ordered along the ray and the participating medium parameters includes: Based on the set of three-dimensional Gaussian primitives ordered along the ray, each three-dimensional Gaussian primitive is converted into a one-dimensional probability density function along the ray; the cumulative transmittance is calculated based on the contribution weight of the one-dimensional response peak of each Gaussian primitive; the principal Gaussian primitive is determined based on the condition that the cumulative transmittance first drops below a preset threshold; and the ray-Gaussian intersection depth is obtained by analytical formula based on the principal Gaussian primitive.

4. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 3, characterized in that, Step S4, which involves obtaining the final rendered color of the pixel based on the ray-Gaussian intersection depth, includes: The object color component is obtained based on the attenuation coefficient and the ray-Gaussian intersection depth in the participating medium parameters; the medium color component is obtained based on the backscattering coefficient, the medium color, and the ray-Gaussian intersection depth in the participating medium parameters; the object color component and the medium color component are added together to obtain the final rendered color of the pixel.

5. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, S5 includes: based on the multi-level sparse mesh structure, counting the number of intersections between each ray and the finest LOD level Gaussian primitive; generating a foreground-background binary mask based on the comparison result of the number of intersections and the density threshold; and calculating the depth smoothing regularization loss for the background region identified by the foreground-background binary mask.

6. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, S6 includes: Based on the rendered color and the real color, obtain the HDR reconstruction loss; based on the distribution of Gaussian primitives along the ray, obtain the depth distortion loss; based on the weighted sum of the HDR reconstruction loss, depth distortion loss, and depth smoothing regularization loss, obtain the total loss; based on the total loss, use the Adam optimizer to obtain the optimized 3D Gaussian primitive properties and medium parameters to predict the network parameters.

7. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, S6 further includes: During the optimization process, cloning or splitting operations are performed on Gaussian primitives whose position gradient exceeds a preset threshold; pruning operations are performed on Gaussian primitives whose opacity is below a preset threshold, whose scale exceeds a preset range, or whose depth exceeds the effective range.

8. The underwater 3D Gaussian scene reconstruction method based on ray-Gaussian intersection depth constraint according to claim 1, characterized in that, S7 includes: Receive user-specified target viewpoint parameters, including camera pose and intrinsic parameters; input the target viewpoint parameters into the trained underwater scene reconstruction model, execute the rendering pipeline, and obtain underwater rendered images, depth maps, object color maps, descattered images, 3D Gaussian scene models, and / or video sequences from the target viewpoint.

9. An underwater 3D Gaussian scene reconstruction system based on ray-Gaussian intersection depth constraints, the system being used to implement the method described in any one of claims 1-8, characterized in that, include: The module includes a data preprocessing and initialization module, a medium parameter prediction module, a ray-Gaussian surface depth calculation module, a physically guided rendering module, a LOD regularization module, a joint optimization module, and a new perspective output module.