Underwater scene 3D characterization method based on three-dimensional Gaussian splashing

Through the 3D characterization method of underwater scenes based on three-dimensional Gaussian splashing, the water body effect and the real appearance color and geometric information of the scene in the underwater image are decoupled, and the problem of blurred details and color artifacts in the reconstruction scene in the prior art is solved, achieving high-quality underwater scene representation and support for a variety of downstream tasks.

CN120219664AActive Publication Date: 2025-06-27BEIHANG UNIV

Patent Information

Application Number
CN202510697121.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively decouple the water body effect and the real appearance color and geometric information of the scene in the 3D representation of underwater scenes, resulting in the reconstructed scenes having problems such as blurred details and color artifacts.

Method used

Using a 3D characterization method for underwater scenes based on three-dimensional Gaussian splashing, the underwater image set is obtained and the camera parameters are calibrated using the COLMAP algorithm, the pseudo-depth map is derived, the properties of the water body Gaussian and scene Gaussian are initialized, and iterative optimization is performed to decouple the water body properties, scene appearance color and geometric properties.

Benefits of technology

It realizes the effective decoupling of complex water effects and the appearance color and geometric information of the scene itself from underwater images, and supports tasks such as underwater scene restoration, new perspective synthesis and image restoration dataset construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219664A_ABST
    Figure CN120219664A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater scene 3D characterization method based on three-dimensional Gaussian splashing, and the method comprises the steps: obtaining an original underwater image set shot by a single underwater scene, calibrating an internal parameter matrix and an external parameter matrix of a camera corresponding to each underwater image in the original underwater image set through employing a COLMAP algorithm, and outputting a sparse point cloud at the same time; utilizing a pre-trained monocular depth estimator to export a pseudo depth map corresponding to each underwater image in the original underwater image set; based on the sparse point cloud, initializing water body Gaussian attributes to obtain a corresponding water body Gaussian set; based on the sparse point cloud and the pseudo-depth map, initializing attributes of scene gauss to obtain a corresponding scene gauss set; iterative optimization is carried out on attribute parameters in the water body Gaussian set and the scene Gaussian set, and corresponding water body attributes, scene appearance colors and geometric attributes are decoupled and captured. According to the method, the complex water body effect and the appearance color and geometric information of the scene can be decoupled from the underwater image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional scene representation, and more specifically, to a method for 3D representation of underwater scenes based on three-dimensional Gaussian splatting. Background Art

[0002] The 3D representation of underwater scenes plays a crucial role in many application fields, including autonomous underwater vehicles (AUVs), marine ecological research, and underwater virtual reality systems. An effective underwater 3D representation should effectively model both the water body and the scene itself (i.e., the appearance color and geometric structure of the underwater scene itself when there is no water body). However, due to the significant differences between underwater imaging and imaging in the atmospheric environment, existing vision-based 3D representation methods still face severe challenges in effectively representing underwater 3D scenes.

[0003] Next, the background and technologies related to the present invention will be introduced separately from four aspects: the characteristics of underwater imaging, traditional 3D representation technologies for underwater scenes, 3D representation technologies for underwater scenes based on Neural Radiance Field (NeRF), and three-dimensional Gaussian splatting (3DGS) technology.

[0004] Underwater imaging has three main characteristics compared to imaging in the atmospheric environment: The first characteristic is that light is affected by the attenuation effect during its propagation in water. This is because the light reflected by an object is partially absorbed by water molecules during its journey to the camera, and the attenuation of red-band light is more significant than that of blue-green band light, which makes the appearance color of the scene in the underwater image show a color deviation towards blue-green. The second characteristic is that light is affected by the scattering effect during its propagation in water, especially the backscattering effect. This is because the underwater ambient light is scattered by suspended particles in the water into the camera, resulting in degradation phenomena such as a haze appearance and low contrast in the appearance color of the scene in the underwater image. The third characteristic is that both the aforementioned attenuation effect and the backscattering effect are highly correlated with the line of sight distance (LOS Distance) and increase significantly with the increase of the line of sight distance. This means that the same object photographed from different angles and distances will present different degrees of degraded appearance colors, no longer faithfully reflecting the true appearance color of the scene itself, nor having the cross-view content consistency possessed by imaging in the atmospheric environment. Generally speaking, the captured underwater image couples the true appearance color of the scene and complex water body effects, bringing great challenges to the effective 3D representation of underwater scenes.

[0005] Traditional 3D representation methods for underwater scenes are mainly implemented based on stereo vision technology. Such methods usually adopt a two-stage implementation approach. According to different technical routes, these methods can be further divided into 2 subclasses. The methods in the first subclass adopt the technical route of "removing water first - then reconstructing". Such methods first use image restoration algorithms to remove the water body effect in each underwater image one by one and restore the "true" appearance of the scene, and then use technologies such as stereo matching and Structure from Motion (SfM) to reconstruct the underwater 3D scene and its "true" appearance. The methods in the second subclass adopt the technical route of "reconstructing first - then removing water". Such methods first reconstruct the geometric structure of the scene and the degraded appearance with the water body effect through technologies such as stereo matching, and then introduce an underwater imaging model to superimpose the water body effect, so as to decouple the "true" appearance of the scene from the underwater image through iterative optimization.

[0006] The 3D representation technology for underwater scenes based on neural radiance fields is an implicit 3D representation technology for underwater scenes. Neural radiance fields (NeRF) are the core technology of such methods. Its principle is to use a deep neural network to implicitly model the scene. The network takes the three-dimensional coordinates and viewing directions in a scene as inputs and outputs the color and transparency of that point, and then realizes high-quality three-dimensional scene rendering through the ray marching strategy and volumetric rendering mechanism. Correspondingly, the 3D representation technology for underwater scenes based on NeRF models the coupling mechanism between the water body effect and the true appearance color of the scene in the underwater imaging process by expanding the volumetric rendering framework of NeRF, and then decouples the true appearance color, geometric information, and water body attributes of the scene. Taking SeaThru-NeRF as an example, this method assigns the transparency and color attributes belonging to the object and the transparency and color attributes belonging to the water body to each position in space at the same time, and modifies the volumetric rendering mechanism of NeRF that originally only models opaque objects to be compatible with the semi-transparent water body between the camera and the object, so as to model the attenuation and backscattering effects in the underwater imaging process, and thus can directly learn the true appearance color, geometry, and water body parameters of the scene from underwater images and realize the 3D representation of the underwater scene.

[0007] The three-dimensional Gaussian splashing technology is an explicit 3D representation technology for scenes. This technology uses a set of 3D Gaussian basis elements with learnable parameters to explicitly represent the scene, and at the same time introduces a differentiable rasterization pipeline, which is superior to the neural radiance field technology in multiple dimensions such as training time consumption, rendering quality, and rendering speed. Next, the specific implementation of the 3DGS technology is briefly introduced. In the 3DGS technology, each 3D Gaussian basis element has a mean , covariance , opacity o, color attributes. In fact, the covariance attribute of each 3D Gaussian basis element is calculated through an element-wise learnable scale matrix and a rotation matrix according to the formula to ensure the positive semi-definite property of the covariance matrix . The 3D Gaussian distribution corresponding to each 3D Gaussian basis element is expressed as , where refers to the coordinates of any point in the world coordinate system. In the setting of the 3DGS technology, to model non-Lambertian surfaces, a set of parameter-learnable spherical harmonic functions is used to calculate the final color attributes involved in rendering. During the rendering process, the 3DGS technology introduces a tile-based differentiable rasterization pipeline. First, given the camera view to be rendered, each 3D Gaussian basis element is projected onto the corresponding image coordinate system to become a 2D Gaussian basis element ; then, these 2D Gaussian basis elements are sorted according to the depth values of the means of the original 3D Gaussian basis elements in the corresponding camera coordinate system; then, the sorted 2D Gaussian basis elements are used to render the color of each pixel through the algorithm, and the specific rendering formula is referred to as , where represents the coordinates of the given pixel in the image coordinate system, S represents the set composed of all Gaussian basis elements participating in the rendering of the given pixel, represents the opacity contribution of the i-th Gaussian basis element participating in the rendering at the current given pixel position. By traversing all the pixels corresponding to the camera view to be rendered and completing the rendering, the rendered image can be obtained. The 3DGS technology usually uses and such as color reconstruction losses to calculate the difference between the rendered image and the original image , and optimizes the learnable parameters of all Gaussian basis elements through the differentiable rasterization pipeline to achieve an accurate 3D representation of the scene.

[0008] However, traditional 3D representation techniques for underwater scenes, 3D representation techniques for underwater scenes based on neural radiance fields, and 3D Gaussian splash techniques still have some defects.

[0009] This section introduces the problems existing in traditional 3D representation methods for underwater scenes. First, the common problem of these two-stage methods is that in the front and back stages, the decoupling of scene appearance and the reconstruction of scene geometry are carried out separately, and the reconstruction effect of the latter stage depends heavily on the reconstruction quality of the former stage. Specifically, for the "dewatering first - reconstruction later" type of methods, the geometric reconstruction quality depends on the quality of image restoration in the previous stage. However, since image restoration processes each image separately and the restoration results of different images are independent of each other, the restored images cannot guarantee the consistency between contents, which will affect the accuracy of geometric reconstruction. For the "reconstruction first - dewatering later" type of methods, the appearance restoration quality depends on the quality of the scene geometry output by stereo matching in the previous stage. However, due to the geometric structure reconstructed directly from underwater images, especially in the distant areas, there are problems such as inaccurate reconstruction and large depth map errors, which will affect the accuracy of water effect superposition and the effectiveness of decoupling the true appearance of the scene. Second, these methods adopt data-driven image restoration algorithms or predefined water parameter settings, which means that these methods lack detailed modeling of water parameters, making the learned 3D representation of the scene unable to well explain the water property information in the scene, thus making it difficult to carry out downstream tasks such as underwater novel view synthesis.

[0010] This section introduces the problems existing in the 3D representation technology for underwater scenes based on neural radiance fields. First, the performance of these methods is limited by the inherent defects of their underlying technology, neural radiance fields - the volume rendering framework and ray marching strategy they adopt result in huge computational overhead, long training and inference speeds, restricting the application of these methods in more efficiently trained and real-time rendering scenes. Second, the implicit modeling ideas for both water bodies and the scene itself in these methods often lead to incomplete decoupling between water effects and scene appearance geometry, and the reconstructed scene itself often has problems such as blurred details, color artifacts, and floaters.

[0011] This section describes the problems faced by the 3D Gaussian Splatter (3DGS) technique in performing 3D characterization of underwater scenes. First, the original 3DGS technique assumes that images are captured in an atmospheric environment. In other words, the image set used to train 3DGS needs to have cross-view content consistency. However, the distance-related water body effects during underwater imaging cause the collected image set to no longer meet the cross-view content consistency. Therefore, directly using 3DGS for 3D characterization of underwater scenes will learn incorrect geometric structures and severe color artifacts, and it is impossible to achieve 3D characterization of underwater scenes. Second, recently, there have been some works that extend the 3DGS technique to underwater scenes. For example, WaterSpaltting introduces a volume rendering framework to model water body effects while explicitly modeling the scene itself. Seasplat introduces a physics-based underwater imaging model to superimpose water body effects while using 3DGS to model the scene appearance and geometry. UW-GS introduces a physics-based color model for each Gaussian to model the influence of water body effects. However, these methods have not been able to directly and explicitly model the massive water body medium. The modeling quality of water body effects highly depends on the accuracy of scene modeling itself, and it is difficult to accurately and effectively decouple the true appearance, geometry, and water body properties of underwater scenes.

[0012] Therefore, how to decouple the complex water body effects, as well as the appearance color and geometric information of the scene itself, from underwater images, and achieve the rendering of the appearance color of the scene itself and the rendering of underwater images with superimposed water body effects, so as to support the development of various downstream tasks such as underwater scene restoration, underwater novel view synthesis, and construction of underwater image restoration datasets, is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0013] In view of the above problems, the present invention provides a 3D characterization method for underwater scenes based on 3D Gaussian Splatter to at least solve some of the technical problems mentioned in the above background technology.

[0014] To achieve the above object, the present invention adopts the following technical solutions:

[0015] The present invention provides a 3D characterization method for underwater scenes based on 3D Gaussian Splatter, including:

[0016] Obtain the original underwater image set captured from a single underwater scene, and use the COLMAP algorithm to calibrate the internal parameter matrix and external parameter matrix of the camera corresponding to each underwater image in the original underwater image set, and at the same time output a sparse point cloud;

[0017] Use a pre-trained monocular depth estimator to derive the pseudo-depth map corresponding to each underwater image in the original underwater image set;

[0018] Based on the sparse point cloud, initialize the attributes of the water body Gaussians to obtain the corresponding water body Gaussian set;

[0019] Initialize the properties of the scene Gaussian based on the sparse point cloud and the pseudo-depth map to obtain the corresponding set of scene Gaussians;

[0020] Iteratively optimize the property parameters in the water Gaussian set and the scene Gaussian set to decouple and capture the corresponding water properties, the appearance color and geometric properties of the scene itself.

[0021] Furthermore, the properties of the water Gaussian specifically include:

[0022] (1) Mean property of the water Gaussian:

[0023] For each water Gaussian, its mean property is denoted as:

[0024]

[0025] where is the mean property of the p-th water Gaussian, representing the point coordinates in the world coordinate system; represents the x-axis coordinate of the p-th water Gaussian in the world coordinate system; represents the y-axis coordinate of the p-th water Gaussian in the world coordinate system; represents the z-axis coordinate of the p-th water Gaussian in the world coordinate system; The parameters , and are non-learnable parameters and remain fixed after initialization;

[0026] (2) Covariance property of the water Gaussian:

[0027] Model the covariance property of the water Gaussian as isotropic corresponding to a spherical distribution; For each water Gaussian, its covariance property is denoted as:

[0028]

[0029] where represents the covariance property of the p-th water Gaussian; represents the shared parameter of the covariance property of all water Gaussians, which is a non-learnable parameter and remains fixed after initialization; represents the identity matrix of shape 3×3;

[0030] (3) Opacity property of the water Gaussian:

[0031] For each water Gaussian, its opacity property is denoted as:

[0032]

[0033] where Denotes the opacity attribute of the p-th water Gaussian, which is a vector of length 3 corresponding to the three RGB channels to model the wavelength selectivity of the water effect; Denotes the shared parameter of the opacity attributes of all water Gaussians; Denotes the opacity offset corresponding to the current p-th water Gaussian; and Are both vectors of length 3; The function is used to ensure that The value range of each channel is (0, 1); and Are both learnable parameters;

[0034] (4)Color attributes of water Gaussians:

[0035] For each water Gaussian, its color attribute is denoted as:

[0036]

[0037] Among them, Denotes the color attribute of the p-th water Gaussian; Denotes the shared parameter of the color attributes of water Gaussians in the entire space, which is a vector of length 3 corresponding to the three RGB channels; Denotes the color offset corresponding to the p-th water Gaussian, which is a vector of length 3; The function is used to ensure that The value range of each channel is (0, 1); and Are learnable parameters.

[0038] Furthermore, the attributes of the scene Gaussian specifically include:

[0039] (1)Mean attribute of the scene Gaussian:

[0040] For each scene Gaussian, its mean attribute is denoted as:

[0041]

[0042] Among them, Is the mean attribute of the r-th scene Gaussian, representing the point coordinates in the world coordinate system; Represents the x-axis coordinate of the r-th scene Gaussian in the world coordinate system; Represents the y-axis coordinate of the r-th scene Gaussian in the world coordinate system; Represents the z-axis coordinate of the r-th scene Gaussian in the world coordinate system; 、 and Are learnable parameters;

[0043] (2) Covariance property of scene Gaussian:

[0044] Model the covariance property of scene Gaussian as the anisotropy of the corresponding ellipsoidal distribution; for each scene Gaussian, its covariance property is denoted as:

[0045]

[0046] where represents the covariance property of the r-th scene Gaussian; represents the scale matrix of the ellipsoidal distribution; represents the rotation matrix of the ellipsoidal distribution; and are both learnable parameters;

[0047] (3) Opacity property of scene Gaussian:

[0048] For each scene Gaussian, its opacity property is denoted as ; represents the opacity property of the r-th scene Gaussian, which is a scalar value and a learnable parameter with a value range of (0, 1);

[0049] (4) Color property of scene Gaussian:

[0050] For each scene Gaussian, its color property is denoted as:

[0051]

[0052] where represents the color property of the r-th scene Gaussian; , and correspond to the three RGB channels respectively, and are learnable parameters with a value range of (0, 1).

[0053] Furthermore, based on the sparse point cloud, initialize the properties of the water body Gaussian to obtain the corresponding water body Gaussian set; specifically including:

[0054] (1) Initialize the mean property of the water body Gaussian:

[0055] According to the distribution range of the sparse point cloud, calculate and set a cuboid bounding box as the scene boundary to limit the scene range;

[0056] Inside the scene boundary, construct a regular grid in the XYZ three-axis directions with a fixed step size, and place a water body Gaussian at each grid node, taking the position of the grid node in the world coordinate system as the mean property of the corresponding water body Gaussian;

[0057] (2) Initialize the covariance attributes of the water body Gaussians:

[0058] Share the parameters of the covariance attributes of all water body Gaussians , and initialize them according to the formula ; where represents the scale factor; represents the fixed step size;

[0059] (3) Initialize the opacity attributes of the water body Gaussians:

[0060] For each water body Gaussian, initialize its opacity offset to [0, 0, 0];

[0061] Share the parameters of the opacity attributes of all water body Gaussians and initialize them to ;

[0062] (4) Initialize the color attributes of the water body Gaussians:

[0063] For each water body Gaussian, initialize its color offset to [0, 0, 0];

[0064] Share the parameters of the color attributes of all water body Gaussians and initialize them to ;

[0065] (5) Denote all the water body Gaussians with the mean attributes, covariance attributes, opacity attributes, and color attributes all initialized as the water body Gaussian set.

[0066] Further, initialize the attributes of the scene Gaussians based on the sparse point cloud and the pseudo-depth map to obtain the corresponding scene Gaussian set; including:

[0067] Augment the sparse point cloud according to the pseudo-depth map to obtain the augmented point cloud;

[0068] Initialize the attributes of the scene Gaussians based on the augmented point cloud to obtain the corresponding scene Gaussian set.

[0069] Further, the augmenting the sparse point cloud according to the pseudo-depth map to obtain the augmented point cloud specifically includes:

[0070] (1) According to the pseudo-depth map corresponding to each underwater image in the original underwater image set, divide each underwater image into multiple regions at a preset depth value interval, and each region corresponds to a different depth interval on the pseudo-depth map;

[0071] (2)Extract the sparse point cloud subset observable from the camera view corresponding to each underwater image in the original underwater image set, calculate the depth values of all points in the sparse point cloud subset in the corresponding camera coordinate system, and obtain the projection points of all points in the sparse point cloud subset on the corresponding image coordinate system, denoted as the original projection points;

[0072] (3)For each region of each underwater image in the original underwater image set, calculate the number of pixels covered by the current region, determine the subset of original projection points of the sparse point cloud subset corresponding to the current underwater image in the current region, and determine the number of original projection points included in the current subset of original projection points. Calculate the number of projection points to be newly added in the current region, and randomly add the corresponding number of projection points in the current region, denoted as the newly added projection points;

[0073] The number of projection points to be newly added in the current region is expressed as:

[0074]

[0075] Among them, represents the original underwater image set the set composed of the projection points to be newly added in the t-th region of the h-th underwater image; represents the set the number of projection points to be newly added in; represents the minimum projection point density of each region; represents the t-th region of the h-th underwater image in the original underwater image set; represents the number of pixels included in; represents the subset of original projection points of the h-th underwater image in the original underwater image set in the t-th region; represents the subset of original projection points the number of original projection points included in;

[0076] (4)For each newly added projection point in each region of each underwater image in the original underwater image set, the depth value and color attribute of the corresponding newly added spatial point in the camera coordinate system are expressed as:

[0077]

[0078] Among them, m represents the m-th newly added projection point; represents the set composed of the k nearest original projection points of the m-th newly added projection point found from the set of original projection points in the current region using the K-nearest neighbor algorithm; represents The depth value of the original spatial point corresponding to the nth original projection point in the current camera coordinate system; denote The color value of the original spatial point corresponding to the nth original projection point in the current camera coordinate system; denote the weighting weight;

[0079] where, denote the coordinates of the mth newly added projection point on the image coordinate system as , and denote the coordinates of the nth original projection point on the image coordinate system as , then the weighting weight is expressed as:

[0080]

[0081] (5) According to the camera projection transformation and the external parameter matrix of the camera, obtain the position of the newly added spatial point corresponding to each newly added projection point in the world coordinate system;

[0082] (6) Calculate the positions and color attributes of all newly added spatial points in all regions of all underwater images in the original underwater image set. The union of the set composed of these newly added spatial points and the set of original spatial points corresponding to the sparse point cloud is the augmented point cloud.

[0083] Furthermore, in each iteration optimization process of the attribute parameters in the water body Gaussian set and the scene Gaussian set, select an underwater image and its corresponding pseudo-depth map from the original underwater image set. Use the water body Gaussian basis element and the scene Gaussian basis element of the current iteration through the per-channel multi-output rasterization rendering module to obtain the underwater image , the image of the scene itself , the pure water background image and the depth map under the same camera view; realize the optimization of the water body Gaussian attribute and the scene Gaussian attribute by calculating the loss function and backpropagating the gradient;

[0084] The per-channel multi-output rasterization rendering module includes:

[0085] (1) Underwater image rendering branch:

[0086] ① Denote the set composed of all water body Gaussians as , and denote the set composed of all scene Gaussians as ;

[0087] ② For the pth water body Gaussian, re-express its opacity attribute as , where , and are The three opacity components, which are independent of each other and respectively correspond to the three RGB channels;

[0088] ③ For the r-th scene Gaussian, re-express its opacity property as , where , and are the three opacity components of , corresponding to the three RGB channels respectively, and

[0089] ④ Given the camera view corresponding to the underwater image to be rendered, that is, the internal and external parameter matrices of the camera, obtain the union of the water body Gaussian set and the scene Gaussian set , and then sort all the Gaussian basis elements in the set according to their depth values in the given camera coordinate system for tile-based -blending rendering;

[0090] ⑤ The color of each channel of any pixel of the underwater image to be rendered is expressed as:

[0091]

[0092] where represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element during the rendering process. Here, the color channel of the i-th Gaussian basis element corresponds to the color channel of represents the opacity contribution of the i-th Gaussian basis element to the corresponding color channel during the rendering of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel;

[0093] ⑥ Render each color channel of any pixel of the underwater image to be rendered independently;

[0094] ⑦ Traverse all pixels of the underwater image to be rendered and complete the rendering, and output the final underwater image ;

[0095] (2) Scene itself image rendering branch:

[0096] ①Extract all Gaussian basis elements in the scene Gaussian set ;

[0097] ②For the r-th scene Gaussian, re-denote its opacity attribute as a scalar ;

[0098] ③Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, use the rasterization rendering pipeline in the original 3DGS technology to process all Gaussian basis elements in the scene Gaussian set and output the final image of the scene itself ;

[0099] (3)Pure water background image rendering branch:

[0100] ①Extract all Gaussian basis elements in the water body Gaussian set ;

[0101] ②For the p-th water body Gaussian, re-express its opacity attribute as , where the 3 opacity components are independent of each other;

[0102] ③Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, sort all Gaussian basis elements in the water body Gaussian set according to their depth values in the given camera coordinate system for tile-based rendering;

[0103] ④The color of each channel of any pixel of the pure water background image to be rendered is expressed as:

[0104]

[0105] where, represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element during the rendering process, where the color channel of the i-th Gaussian basis element corresponds to the color channel of; represents the opacity contribution of the i-th Gaussian basis element in the corresponding color channel during the rendering process of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel during the rendering process; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel;

[0106] ⑤ Independently render the colors of each channel of any pixel for rendering the pure water background image;

[0107] ⑥ Traverse all pixels of the pure water background image to be rendered and complete the rendering, and output the final pure water background image ;

[0108] (4) Depth map rendering branch:

[0109] ① Extract all Gaussian primitives in the scene Gaussian set and participate in rasterization and rendering;

[0110] ② For the r-th scene Gaussian, re-record its opacity attribute as a scalar ;

[0111] ③ Given the camera view corresponding to the image to be rendered, that is, the intrinsic matrix and extrinsic matrix of the camera, sort all Gaussian primitives in the scene Gaussian set according to their depth values in the given camera coordinate system for tile-based rendering;

[0112] ④ The depth value of any pixel of the depth map to be rendered is expressed as:

[0113]

[0114] where represents the position coordinates of the pixel in the image coordinate system, represents the opacity contribution of the i-th Gaussian primitive in the current pixel rendering process during the rendering process, and ; represents the value of the 2D Gaussian distribution of the i-th Gaussian primitive in the current image coordinate system at the current pixel; represents the depth value of the i-th Gaussian primitive in the current camera coordinate system during the rendering process;

[0115] ⑤ When all pixels of the image to be rendered are completed, the desired depth map is obtained .

[0116] Further, in the process of iteratively optimizing the attribute parameters in the water body Gaussian set, the attribute optimization strategy of the water body Gaussian includes:

[0117] (1) The mean attribute and covariance attribute of each water body Gaussian are frozen after initialization and no longer participate in the optimization process;

[0118] (2) According to the number of parameter optimization iteration rounds, the optimization process of the opacity attribute and color attribute of each water body Gaussian is divided into two stages:

[0119] During the optimization process of the first stage, the opacity offset of each water body Gaussian and the color offset are kept frozen and not involved in the optimization. Only the opacity attribute sharing parameters of all water body Gaussians and the color attribute sharing parameters are optimized;

[0120] During the optimization process of the second stage, the opacity offset of each water body Gaussian and the color offset are thawed and optimized together with the opacity attribute sharing parameters shared by all water body Gaussians and the color attribute sharing parameters

[0121] Furthermore, during the process of iteratively optimizing the attribute parameters in the water body Gaussian set and the scene Gaussian set, the total loss function is expressed as:

[0122] In the first stage of optimization, ;

[0123] In the second stage of optimization, ;

[0124] Among them, represents the loss function based on image content; represents the coarse-grained depth loss function; represents the local smoothness regularization term of the water body Gaussian opacity attribute; represents the local smoothness regularization term of the water body Gaussian color attribute; represents the coarse-grained depth loss function weight; represents the local smoothness regularization term of the water body Gaussian opacity attribute weight; represents the local smoothness regularization term of the water body Gaussian color attribute weight.

[0125] Furthermore:

[0126] (1) The construction method of the loss function based on image content specifically includes:

[0127] ① Using the pseudo-depth map of the current iteration round to create a background mask ; Given a predefined depth threshold ; The area composed of pixels with pseudo-depth values greater than is regarded as the foreground area, and the corresponding background mask The pixel value at the corresponding position is 0; the pseudo-depth value is less than or equal to The area composed of pixels is regarded as the background area, and the corresponding background mask The pixel value at the corresponding position is set to 1;

[0128] ② Denote the underwater image in the original underwater image set as the true underwater image, and the symbol is Construct a weight map with the same shape as the pseudo-depth map Specifically, the value of the l-th pixel of the weight map is referred to the following requirements:

[0129] When the value of the background mask corresponding to the l-th pixel is 0, the l-th pixel of the weight map corresponds to the foreground area. At this time, the weight map corresponding to the l-th pixel is expressed as:

[0130]

[0131] Among them, represents the gradient cut-off operator; represents the color value corresponding to the l-th pixel of the rendered underwater image;

[0132] When the value of the background mask corresponding to the l-th pixel is 1, the l-th pixel of

[0133]

[0134] corresponds to the background area. At this time, the weight map corresponding to the l-th pixel is expressed as: Then, the foreground-enhanced underwater image and the true foreground-enhanced underwater image

[0135]

[0136] Among them, represents the element-wise multiplication operator;

[0137] ③ Extract all the pixels with a value of 1 on the background mask to obtain the set of background pixels, denoted as , and all the pixels belonging to this set correspond to the area in the picture where there is no underwater scene and only pure water exists;

[0138] ④ The loss function based on the image content is expressed as:

[0139]

[0140] Among them, represents the weight of the loss function term ; represents the weight of the loss function term ; represents the weight of the loss function term ; represents the loss function in the original 3DGS technology; represents the pure water background image at the th pixel of the region; represents the color value of the underwater image ground truth at the th pixel of the region;

[0141] (2) Coarse-grained depth loss function is constructed as follows:

[0142] Convert the rendered depth map into an approximate disparity map , where represents the scale parameter;

[0143] Use the Pearson correlation coefficient with scale invariance characteristics as a similarity metric;

[0144] The coarse-grained depth loss function is expressed as:

[0145]

[0146] Among them, represents the Pearson correlation coefficient calculation function; represents the pseudo-depth map;

[0147] (3) Local smoothing regularization term for water body Gaussian opacity attribute , is expressed as:

[0148]

[0149] (4) Local smoothing regularization term for water body Gaussian color attribute , is expressed as:

[0150]

[0151] Among them, represents the set of water body Gaussians participating in rendering; Denote the set composed of the k nearest neighbor water body Gaussians of the p-th water body Gaussian in the world coordinate system obtained by the KNN algorithm; Denote the weight factor, and Denote as ; Denote the scale factor; Denote the mean attribute of the p-th water body Gaussian in the world coordinate system; Denote the mean attribute of the q-th water body Gaussian in the world coordinate system.

[0152] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a 3D representation method for underwater scenes based on three-dimensional Gaussian splashing, which has the following beneficial effects:

[0153] The present invention can decouple the complex water body effects, as well as the appearance color and geometric information of the scene itself, from underwater images. Furthermore, it can realize the rendering of the appearance color of the scene itself and the rendering of underwater images with superimposed water body effects, thereby supporting the development of various downstream tasks such as underwater scene restoration, underwater new view synthesis, and underwater image restoration dataset construction.

[0154] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0155] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained according to the provided accompanying drawings without creative efforts.

[0156] Figure 1 It is a schematic diagram of the input and output of the per-channel multi-output rasterization rendering module provided by the embodiment of the present invention.

[0157] Figure 2 It is a schematic diagram of the framework of the 3D representation method for underwater scenes based on three-dimensional Gaussian splashing provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0158] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0159] An embodiment of the present invention discloses an underwater scene 3D characterization method based on three-dimensional Gaussian splashing, including:

[0160] Obtain the original underwater image set captured from a single underwater scene, use the COLMAP algorithm to calibrate the internal parameter matrix and external parameter matrix of the camera corresponding to each underwater image in the original underwater image set, and simultaneously output a sparse point cloud;

[0161] Use a pre-trained monocular depth estimator to derive the pseudo-depth map corresponding to each underwater image in the original underwater image set;

[0162] Based on the sparse point cloud, initialize the attributes of the water Gaussian to obtain the corresponding water Gaussian set;

[0163] Based on the sparse point cloud and the pseudo-depth map, initialize the attributes of the scene Gaussian to obtain the corresponding scene Gaussian set;

[0164] Iteratively optimize the attribute parameters in the water Gaussian set and the scene Gaussian set to decouple and capture the corresponding water attributes, the appearance color and geometric attributes of the scene itself.

[0165] An embodiment of the present invention provides an underwater scene 3D characterization method based on three-dimensional Gaussian splashing (3DGS), which adopts different 3D Gaussian basis element setting strategies to explicitly represent the water body and the scene itself in a complex underwater scene respectively, and introduces a per-channel multi-output rasterization rendering module to implement the rendering of different types of outputs to meet the requirements of parameter optimization and various downstream tasks. Next, first, the 3D Gaussian basis element setting strategy proposed by the embodiment of the present invention is described, then the proposed per-channel multi-output rasterization rendering module is introduced in detail, and then the implementation process of the proposed underwater scene 3D characterization method is introduced in detail.

[0166] 1. 3D Gaussian basis element setting strategy for water body and scene itself

[0167] In view of the characteristics of the semi-transparent water medium and the opaque scene itself in the underwater scene, an embodiment of the present invention designs two different types of 3D Gaussian basis elements to represent the water body and the scene itself respectively. Specifically, the 3D Gaussian basis element representing the water medium is simply referred to as the water Gaussian, and the 3D Gaussian basis element representing the scene itself is simply referred to as the scene Gaussian. The attribute characteristics of these two different types of Gaussian basis elements are introduced below.

[0168] 1.1. Attribute characteristics of water Gaussian basis element

[0169] 1.1.1. Mean attribute of water Gaussian

[0170] The mean attribute setting of the 3D Gaussian basis element in the original 3DGS technology is the same. For each water Gaussian, taking the p-th water Gaussian as an example, its mean attribute is denoted as:

[0171]

[0172] Among them, is the mean attribute of the p-th Gaussian of the water body, representing the point coordinates in the world coordinate system; represents the x-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; represents the y-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; represents the z-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; The parameters , and are non-learnable parameters and remain fixed after initialization.

[0173] 1.1.2. Covariance Attribute of Gaussian of Water Body

[0174] Considering that the overall water body medium in the underwater environment is relatively uniform, and the main differences lie in the spatial non-uniformity caused by local components, light direction, etc., such as light transmittance and color attributes, therefore, different from the anisotropic Gaussian distribution adopted by the original 3DGS technology, the distribution of the Gaussian of the water body in the embodiments of the present invention is isotropic, corresponding to a spherical distribution. For each Gaussian of the water body, taking the p-th Gaussian of the water body as an example, its covariance attribute is denoted as:

[0175]

[0176] Among them, represents the covariance attribute of the p-th Gaussian of the water body; represents the shared parameter of the covariance attributes of all Gaussians of the water body, which is a non-learnable parameter and remains fixed after initialization; represents the identity matrix with a shape of 3×3;

[0177] 1.1.3. Opacity Attribute of Gaussian of Water Body

[0178] Considering the spatial non-uniform characteristics of the water body medium attributes and the wavelength selectivity of the water body effect, for each Gaussian of the water body, taking the p-th Gaussian of the water body as an example, its opacity attribute is denoted as:

[0179]

[0180] Different from the opacity attribute of each Gaussian basis element in the original 3DGS technology being a scalar, the opacity attribute of the Gaussian of the water body proposed in the embodiments of the present invention is a vector with a length of 3, corresponding to the three RGB channels, to model the wavelength selectivity of the water body effect. represents the shared parameter of the opacity attributes of all Gaussians of the water body in the entire space; Represents the opacity offset corresponding to the p-th water body Gaussian; and are both vectors of length 3; The function can ensure that the value range of each channel of is (0, 1). and are both learnable parameters.

[0181] 1.1.4. Color Attributes of Water Body Gaussians

[0182] Considering the spatially non-uniform characteristics of the water body medium properties, different from the color attribute settings of each Gaussian basis element in the original 3DGS technology, for each water body Gaussian, taking the p-th water body Gaussian as an example, its color attributes are denoted as:

[0183]

[0184] where represents the color attributes of the p-th water body Gaussian; represents the color attribute sharing parameter of the water body Gaussian in the entire space, which is a vector of length 3, corresponding to the three RGB channels; represents the color offset corresponding to the p-th water body Gaussian, which is a vector of length 3; The function is used to ensure that the value range of each channel is (0, 1); and are learnable parameters.

[0185] 1.2. Attribute Characteristics of Scene Gaussian Basis Elements

[0186] 1.2.1. Mean Attribute of Scene Gaussians

[0187] The same as the mean attribute setting of the 3D Gaussian basis element in the original 3DGS technology, for each scene Gaussian, taking the r-th scene Gaussian as an example, its mean attribute is denoted as:

[0188]

[0189] where is the mean attribute of the r-th scene Gaussian, representing the point coordinates in the world coordinate system; represents the x-axis coordinate of the r-th scene Gaussian in the world coordinate system; represents the y-axis coordinate of the r-th scene Gaussian in the world coordinate system; represents the z-axis coordinate of the r-th scene Gaussian in the world coordinate system; , and are learnable parameters.

[0190] 1.2.2. Covariance Attributes of Scene Gaussians

[0191] Similar to the covariance attribute settings of 3D Gaussian basis elements in the original 3DGS technology, the covariance attributes of scene Gaussians in the embodiments of the present invention are anisotropic and correspond to an ellipsoidal distribution. For each scene Gaussian, taking the r-th scene Gaussian as an example, its covariance attribute is denoted as:

[0192]

[0193] where, represents the covariance attribute of the r-th scene Gaussian; represents the scale matrix of the ellipsoidal distribution; represents the rotation matrix of the ellipsoidal distribution; and are both learnable parameters.

[0194] 1.2.3. Opacity Attributes of Scene Gaussians

[0195] Similar to the opacity attribute settings of 3D Gaussian basis elements in the original 3DGS technology, for each scene Gaussian, taking the r-th scene Gaussian as an example, its opacity attribute is denoted as , which is a scalar value and a learnable parameter with a value range of (0, 1).

[0196] 1.2.4. Color Attributes of Scene Gaussians

[0197] Considering that underwater scenes usually exhibit Lambertian Surface characteristics, different from the color attribute settings of each Gaussian basis element in the original 3DGS technology that model color attributes through spherical harmonics, the color attributes of the scene Gaussian basis elements proposed in the embodiments of the present invention are directly modeled by a vector of length 3. For each scene Gaussian, taking the r-th scene Gaussian as an example, its color attribute is denoted as:

[0198]

[0199] where, represents the color attribute of the r-th scene Gaussian; , and correspond to the three RGB channels respectively and are learnable parameters with a value range of (0, 1).

[0200] 2. Per-Channel Multi-Output Rasterization Rendering Module

[0201] To be compatible with the wavelength selectivity of the water body effect in the underwater imaging process, solve the problem that the opacity components of different channels of the water body Gaussian are independent of each other and are not compatible with the differentiable rasterization pipeline in the original 3DGS technology, and at the same time output the results to meet the needs of optimization and application, the embodiment of the present invention designs a per-channel multi-output rasterization rendering module. As Figure 1 shown, this module includes 4 rendering branches: the underwater image rendering branch, the scene itself image rendering branch, the pure water background image rendering branch, and the depth map rendering branch, which can realize the rendering of underwater images , the scene itself image , the pure water background image , the depth map and other four kinds of results. The rendering process of each rendering branch is introduced below.

[0202] 2.1. Underwater Image Rendering Branch

[0203] ① Denote the set composed of all water body Gaussians as , and the set composed of all scene Gaussians as ;

[0204] ② For each water body Gaussian, take the p-th water body Gaussian as an example, and re-express its opacity attribute as , where , and are the 3 opacity components of , these 3 opacity components are independent of each other and do not affect each other, and respectively correspond to the 3 RGB channels;

[0205] ③ For the r-th scene Gaussian, re-express its opacity attribute as , where , and are the 3 opacity components of , corresponding to the 3 RGB channels respectively, and ;

[0206] ④ Given the camera view corresponding to the underwater image to be rendered, that is, the internal parameter matrix and external parameter matrix of the camera, obtain the union of the water body Gaussian set and the scene Gaussian set , and then sort all the Gaussian primitives in the set according to their depth values in the given camera coordinate system, so as to carry out -blending rendering based on tiles;

[0207] ⑤For each channel color of any pixel in the underwater image to be rendered (the position coordinates of this pixel in the image coordinate system are denoted as x), it can be obtained according to the following formula:

[0208]

[0209] where, represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element in the rendering process. Here, the color channel of the i-th Gaussian basis element corresponds to the color channel of; represents the opacity contribution of the i-th Gaussian basis element in the corresponding color channel during the rendering of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel. It should be noted that the i-th Gaussian basis element here refers to the i-th Gaussian basis element in the set of Gaussian basis elements participating in the rendering during the entire rendering process, which is different from the p-th water body Gaussian basis element or the r-th scene Gaussian basis element mentioned before;

[0210] ⑥It should be noted that different from the original 3DGS technology where the same set of opacity contributions is used for the rendering of all color channels of any pixel. In order to model the wavelength selectivity of the water body effect and solve the problem that the opacity components corresponding to different color channels of the water body Gaussian are independent and different from each other, in the "per-channel multi-output rasterization rendering" module proposed in the embodiments of the present invention, the opacity contributions used for the rendering of different color channels of the same pixel are different. For each color channel of any pixel in the underwater image to be rendered, independent rendering is performed. Only when all color channels are completed with rendering, the color rendering process of the given pixel is considered to end.

[0211] ⑦By traversing all pixels in the underwater image to be rendered and completing the rendering, the final underwater image can be output .

[0212] 2.2. Scene itself image rendering branch

[0213] ①Extract all Gaussian basis elements in the scene Gaussian set ;

[0214] ②For each scene Gaussian, taking the r-th scene Gaussian as an example, rewrite its opacity attribute as a scalar ;

[0215] ③ Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, all the Gaussian basis elements in the set are processed using the rasterization rendering pipeline in the original 3DGS technology, and the final image of the scene itself can be output .

[0216] 2.3. Pure water background image rendering branch

[0217] ① Extract all the Gaussian basis elements in the water Gaussian set ;

[0218] ② Similar to the "underwater image rendering branch", for each water Gaussian, taking the p-th water Gaussian as an example, its opacity attribute is re-expressed as , where the three opacity components of

[0219] ③ Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, all the Gaussian basis elements in the water Gaussian set are sorted according to their depth values in the given camera coordinate system for Tile-Based -blending rendering.

[0220] ④ The color of each channel of any pixel of the pure water background image to be rendered (the position coordinates of this pixel in the image coordinate system are denoted as ) can be obtained according to the following formula:

[0221]

[0222] where represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element during the rendering process, where the color channel of the i-th Gaussian basis element here corresponds to the color channel of represents the opacity contribution of the i-th Gaussian basis element in the corresponding color channel during the rendering of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel during the rendering process; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel during the rendering process.

[0223] ​​⑤ Similar to the "underwater image rendering branch", for any pixel of the pure water background image to be rendered (the position coordinates of this pixel in the image coordinate system are denoted as ), the colors of each channel are rendered independently. Only when all the color channels of this pixel are completed, the color rendering process of this pixel is considered truly ended.

[0224] ⑥ Traverse all the pixels of the pure water background image to be rendered and complete the rendering, then the final pure water background image can be output .

[0225] 2.4. Depth map rendering branch

[0226] ① Extract all the Gaussian basis elements in the scene Gaussian set and participate in rasterization and rendering.

[0227] ② Similar to the "scene itself image rendering branch", for each scene Gaussian, taking the r-th scene Gaussian as an example, its opacity attribute is re-denoted as a scalar .

[0228] ③ Given the camera view corresponding to the image to be rendered, that is, the internal parameter matrix and external parameter matrix of the camera, sort all the Gaussian basis elements in the scene Gaussian set according to their depth values in the given camera coordinate system, so as to carry out tile-based -blending rendering.

[0229] ④ The depth value of any pixel of the depth map to be rendered (the position coordinates of this pixel in the image coordinate system are denoted as ) can be obtained according to the following formula:

[0230]

[0231] where, represents the position coordinates of the pixel in the image coordinate system, represents the opacity contribution of the i-th Gaussian basis element in the current pixel rendering process during the rendering process, and ; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel during the rendering process; represents the depth value of the i-th Gaussian basis element in the current camera coordinate system during the rendering process;

[0232] ⑤ When all the pixels of the image to be rendered are completed, the expected depth map can be obtained.

[0233] 3. Implementation process of the underwater scene 3D characterization method based on three-dimensional Gaussian splash provided by the embodiments of the present invention

[0234] Figure 2 The flowchart of the underwater scene 3D characterization method based on three-dimensional Gaussian splash provided by the embodiments of the present invention is shown. Given a set of original underwater images taken underwater, first, carry out the "image preprocessing stage based on COLMAP" and the "pseudo-depth map derivation stage" based on this set of original underwater images. Through the "image preprocessing stage based on COLMAP", the internal parameter matrix, external parameter matrix and sparse point cloud of the camera required in the subsequent stages can be output; the "pseudo-depth map derivation stage" outputs the pseudo-depth map corresponding to each underwater image in the original underwater image set required in the subsequent stages. Then, through the "water body Gaussian initialization stage" and the "scene Gaussian initialization stage", the corresponding water body Gaussian set and scene Gaussian set can be obtained. Then, through the parameter iteration optimization of the "Gaussian basis element learnable parameter optimization stage", the corresponding water body attributes, the appearance color and geometric attributes of the scene itself are decoupled and captured to meet the needs of the "downstream task application stage". The specific content of each stage is introduced in detail below.

[0235] 3.1. Image preprocessing stage based on COLMAP

[0236] Given a set of original underwater images taken of a single underwater scene , use the COLMAP algorithm to calibrate the internal parameter matrix and external parameter matrix of the camera corresponding to each underwater image in this image set, and at the same time output a sparse point cloud (Sparse Point Cloud), denoted as , for the needs of subsequent Gaussian basis element initialization and view rendering.

[0237] 3.2. Pseudo-depth map derivation stage

[0238] The coupled water body effect in underwater images poses a huge challenge to the development of 3D characterization. Therefore, in the embodiments of the present invention, an attempt is made to use the powerful geometric prior ability of a pre-trained monocular depth estimator to guide the initialization and optimization process of the underwater scene 3D characterization method.

[0239] Specifically, use a pre-trained monocular depth estimator, such as the DepthAnything algorithm, to derive the pseudo-depth map corresponding to each underwater image in the original underwater image set . Usually, the pseudo-depth map derived by the monocular depth estimator reflects relative parallax information, and the value range of the pseudo-depth value is [0,1]. The pseudo-depth value corresponding to the distant area is smaller, and the pseudo-depth value corresponding to the near area is larger.

[0240] 3.3. Water body Gaussian initialization stage

[0241] Different from the attribute initialization strategy of the 3D Gaussian basis elements in the original 3DGS technology, the initialization methods of various attributes of the water body Gaussian proposed in the embodiments of the present invention are introduced below respectively.

[0242] (1) Initialize the mean attribute of the water body Gaussian:

[0243] According to the distribution range of the sparse point cloud exported by COLMAP, calculate and set a cuboid bounding box (AABB, axis-aligned bounding box) as the scene boundary to limit the scene range;

[0244] Inside the scene boundary, construct a regular grid in the XYZ three coordinate axis directions with a fixed step size, and place a water body Gaussian at each grid node, and use the position of the grid node in the world coordinate system as the mean attribute of the corresponding water body Gaussian;

[0245] (2) Initialize the covariance attribute of the water body Gaussian:

[0246] The covariance attributes of all water body Gaussians share the parameter , and are initialized according to the formula ; where, represents the scale factor; represents the fixed step size;

[0247] (3) Initialize the opacity attribute of the water body Gaussian:

[0248] For each water body Gaussian, taking the p-th water body Gaussian as an example, initialize its opacity offset to [0, 0, 0];

[0249] Initialize the opacity attribute sharing parameter of all water body Gaussians to to ensure that the opacity attribute of the water body Gaussian is initialized to [0.1, 0.1, 0.1];

[0250] (4) Initialize the color attribute of the water body Gaussian:

[0251] For each water body Gaussian, taking the p-th water body Gaussian as an example, initialize its color offset to [0, 0, 0];

[0252] Initialize the color attribute sharing parameter of all water body Gaussians to to ensure that the color attribute of the water body Gaussian is initialized to [0.1, 0.1, 0.1].

[0253] (5) Denote all the water body Gaussians with the mean attribute, covariance attribute, opacity attribute, and color attribute initialized as the water body Gaussian set.

[0254] 3.4. Scene Gaussian Initialization Phase

[0255] The water body effect existing in the underwater image causes a serious sparsity problem in the sparse point cloud output by COLMAP. Especially, the imaging results in the area far from the camera are severely affected by the water body effect, and the number of spatial points corresponding to the distant area is extremely small, which is not conducive to the initialization and optimization of the 3D Gaussian basis elements. To solve this problem, in the embodiments of the present invention, first, a "pseudo-depth map-guided sparse point cloud amplification phase" is used to perform amplification and densification processing on the overly sparse point cloud, and then a "scene Gaussian initialization phase based on the amplified point cloud" is used to initialize the attributes of the scene Gaussian to obtain the scene Gaussian set. The specific implementations of these two phases are introduced below.

[0256] 3.4.1. Pseudo-depth map-guided sparse point cloud amplification phase

[0257] In the embodiments of the present invention, in the pseudo-depth map-guided sparse point cloud amplification phase, it aims to make up for the problem that the COLMAP point cloud is too sparse caused by the serious degradation of the underwater image and the great difficulty in feature matching. Specifically:

[0258] (1) According to the pseudo-depth map corresponding to each underwater image in the original underwater image set, each underwater image is divided into multiple regions at a preset depth value interval, and each region corresponds to a different depth interval on the pseudo-depth map;

[0259] For example, assume the original underwater image set has images. Taking the h-th underwater image as an example, according to its corresponding pseudo-depth map, the image is divided into 10 regions at a depth value interval of 0.1 , and each region corresponds to a different depth interval on the pseudo-depth map. For example, the pseudo-depth value of the pixels corresponding to the t-th region of the h-th underwater image is in the range of [0.1*(t - 1), 0.1*t).

[0260] (2) Extract the subset of the sparse point cloud that can be observed from the camera view corresponding to each underwater image in the original underwater image set, calculate the depth values of all points in the sparse point cloud subset in the corresponding camera coordinate system, and obtain the projection points of all points in the sparse point cloud subset on the corresponding image coordinate system, denoted as the original projection points;

[0261] Repeat the same operation for each original underwater image: Taking the h-th underwater image as an example, extract the subset of the COLMAP sparse point cloud observable from the corresponding camera view , and calculate the depth values of all points in the subset of the sparse point cloud in the corresponding h-th camera coordinate system, as well as the projection points of these points on the corresponding h-th image coordinate system (denoted as the original projection points). Denote the set composed of all original projection points as the original projection point set ;

[0262] (3) For each region of each underwater image in the original underwater image set, calculate the number of pixels covered by the current region, determine the subset of the original projection points of the subset of the sparse point cloud corresponding to the current underwater image in the current region, and determine the number of original projection points included in the current subset of the original projection points. Calculate the number of projection points to be newly added in the current region, and randomly add the corresponding number of projection points in the current region, denoted as the newly added projection points;

[0263] Repeat the same operation for each region of each original underwater image: Taking the t-th region of the h-th underwater image as an example, calculate the number of pixels covered by the current region , determine the original projection point set of the current underwater image in the t-th region, , and determine the number of original projection points included in the subset of the original projection points . Specify that the minimum projection point density included in each region is . Denote the set composed of the projection points to be newly added in the t-th region as . Then the number of projection points to be newly added in the t-th region is calculated by the following formula:

[0264]

[0265] where represents the set composed of the projection points to be newly added in the t-th region of the h-th underwater image in the original underwater image set ; represents the number of projection points to be newly added in the set ; If takes the value of 0, it means that no new projection points need to be added; If , then the number of projection points to be randomly added in the t-th region is .

[0266] (4) Repeat the same operation for each newly added projection point in each region of each underwater image in the original underwater image set: Taking the t-th region of the h-th underwater image The mth newly added projection point in For example, the corresponding new spatial point Depth value in camera coordinate system and color attributes It is expressed as:

[0267]

[0268] Wherein, m in the formula represents the mth newly added projection point; It means from the current area The original projection point set The set of k original projection points closest to the mth newly added projection point found by using the K-Nearest Neighbor (KNN) algorithm; express The depth value of the original space point corresponding to the nth original projection point in the current camera coordinate system; express The color value of the original space point corresponding to the nth original projection point in the current camera coordinate system; represents weighted weight;

[0269] Among them, the coordinates of the mth newly added projection point in the image coordinate system are marked as , the coordinate of the nth original projection point in the image coordinate system is denoted as , then the weighted weight It is expressed as:

[0270]

[0271] (5) According to the camera projection transformation and the camera's external parameter matrix, the position of each newly added spatial point corresponding to the newly added projection point in the world coordinate system is obtained; specifically:

[0272] Let the width of the image taken by the current camera be w, the height be h, and the focal length be f. The corresponding transformation matrix of the camera coordinate system is (This matrix is ​​the inverse matrix of the camera's external parameter matrix). According to the basic principle of projection transformation, the mth newly added projection point The corresponding new space point The coordinates in the current camera coordinate system are , add space points The homogeneous coordinates in the world coordinate system can be obtained by the formula: get;

[0273] (6) Calculate the position and color attributes of all newly added spatial points in all regions of all underwater images in the original underwater image set. The union of the set composed of these newly added spatial points and the original spatial point set corresponding to the sparse point cloud is the desired amplified point cloud.

[0274] 3.4.2. Scene Gaussian Initialization Stage Based on Amplified Point Cloud

[0275] In the embodiment of the present invention, in the scene Gaussian initialization stage based on the amplified point cloud, the same Gaussian basis element initialization strategy as the original 3DGS is adopted, that is, the initialization methods of the mean, covariance, color, and opacity attributes of each scene Gaussian basis element adopt the default initialization methods of 3DGS. The only difference is that the point cloud used to initialize the scene Gaussian is no longer the sparse point cloud output by COLMAP, but the "pseudo-depth map-guided amplified point cloud" provided by the embodiment of the present invention.

[0276] 3.5. Optimization Stage of Learnable Parameters of Gaussian Basis Elements

[0277] In the embodiment of the present invention, the learnable parameters of the Gaussian basis elements are gradually optimized through multiple iterations to achieve accurate 3D representation of the underwater scene. Specifically, in each iteration process, a single underwater image is randomly selected from the original underwater image set and its corresponding pseudo-depth map , and the underwater image, scene image , pure water background image , and depth map in the same camera view are obtained by using the water body Gaussian basis elements and scene Gaussian basis elements of the current iteration through the proposed "per-channel multi-output rasterization rendering module". By calculating the loss function and backpropagating the gradient, the optimization of the water body Gaussian attributes and scene Gaussian attributes is achieved. First, the attribute optimization strategies of the water body Gaussian and the scene Gaussian are introduced below, and then the loss function used in the embodiment of the present invention is introduced.

[0278] 3.5.1. Attribute Optimization Strategy of Water Body Gaussian

[0279] During the optimization process, the mean and covariance attributes of each water body Gaussian are frozen after initialization and no longer participate in the optimization process. The optimization processes of the opacity and color attributes of each water body Gaussian are divided into two stages according to different iteration rounds. The optimization strategies of the opacity and color attributes of each water body Gaussian are introduced in detail below.

[0280] Divide the total number of iterations into two stages. For example, if the total number of iterations is recorded as 30,000 times, the first 15,000 iterations are the first stage, and the last 15,000 iterations are the second stage.

[0281] During the optimization process of the first stage, the opacity offset of each water Gaussian and the color offset are kept frozen and do not participate in the optimization. Only the opacity attribute sharing parameters of all water Gaussians and the color attribute sharing parameters are optimized.

[0282] During the optimization process of the second stage, the opacity offset of each water Gaussian and the color offset are thawed and optimized together with the opacity attribute sharing parameters shared by all water Gaussians and the color attribute sharing parameters.

[0283] During the optimization process of all water Gaussians, the adaptive density control strategy in the original 3DGS technology is no longer performed on the water Gaussians.

[0284] 3.5.2. Attribute Optimization Strategy of Scene Gaussians

[0285] During the optimization process, the attribute optimization strategy of all scene Gaussians adopts the conventional Gaussian basis element attribute optimization strategy used in the original 3DGS technology, and still adopts the adaptive density control strategy used in the original 3DGS technology.

[0286] 3.5.3. Loss Function

[0287] In the embodiments of the present invention, the loss function adopted in the optimization process includes 4 types, namely, the loss function based on image content, the coarse-grained depth loss function, the local smooth regularization term of the water Gaussian opacity attribute, and the local smooth regularization term of the water Gaussian color attribute. The specific principles and specific implementations of each loss function are introduced in detail below.

[0288] 3.5.3.1. Loss Function Based on Image Content

[0289] The loss and loss and other color reconstruction loss functions used in the original 3DGS technology face the following problems in the 3D representation of underwater scenes: The first problem is that the attenuation effect of water makes the brightness of the seabed area in the underwater image decrease, which will generate insufficient gradients for the 3D optimization of the corresponding scene area; the second problem is that these color reconstruction loss functions are only calculated between the original underwater image and the underwater image to be rendered , lacking direct supervision for the optimization of water Gaussians, which will lead to suboptimal representation results. Therefore, in the embodiments of the present invention, a loss function based on image content is designed , on the one hand, strengthen the representation of the seabed area (foreground) where the attenuation effect is severely affected, and on the other hand, use the color information of the pure water background area (background) to directly supervise the properties of the water body Gaussian. Next, the loss function based on image content will be introduced in detail. Calculation method.

[0290] ① Use the pseudo-depth map of the current iteration to create a background mask . Given a predefined depth threshold , the area composed of pixels with pseudo-depth values greater than is regarded as the foreground region, and the pixels at the corresponding positions of the corresponding background mask take the value of 0; the area composed of pixels with pseudo-depth values less than or equal to is regarded as the background region, and the pixels at the corresponding positions of the corresponding background mask are set to 1.

[0291] ② Denote the underwater images in the original underwater image set as the true underwater images, with the symbol , and construct a weight map with the same shape as the pseudo-depth map; specifically, the value of the l-th pixel of the weight map is referred to the following requirements:

[0292] When the value of the background mask corresponding to the l-th pixel is 0, the l-th pixel of the weight map corresponds to the foreground region. At this time, the l-th pixel of the corresponding weight map is expressed as:

[0293]

[0294] where represents the gradient cut-off operator; represents the color value corresponding to the l-th pixel of the rendered underwater image ;

[0295] When the value of the background mask corresponding to the l-th pixel is 1, the l-th pixel of the corresponding weight map corresponds to the background region. At this time, the l-th pixel of the corresponding weight map is expressed as:

[0296]

[0297] Then, obtain the relationship between the foreground-enhanced underwater image and the foreground-enhanced true underwater image :

[0298]

[0299] Among them, represents an element-wise multiplication operator.

[0300] ③ Extract the background mask All pixels with a value of 1 in the above are extracted to obtain a set of background pixels, denoted as , and all pixels belonging to this set correspond to the area in the picture where there is no underwater scene and only pure water exists.

[0301] ④ Loss function based on image content It is expressed as:

[0302]

[0303] Among them, represents the weight of the loss function term ; represents the weight of the loss function term ; represents the weight of the loss function term ; These three weights are used to balance the influence of different loss function terms on the final optimization result, and the values are ; represents the loss function in the original 3DGS technology; represents the pure water background image at the th pixel in the region ; represents the color value of the underwater image ground truth at the th pixel in the region ;

[0304] 3.5.3.2. Coarse-grained depth loss function

[0305] The inconsistent viewpoints among the images in the original underwater image set will cause the geometric constraint ability in the optimization process to decline, making it difficult for the optimized 3D representation to accurately reflect the geometric information of the scene. To solve this problem, the embodiment of the present invention designs a coarse-grained depth loss function , and uses the geometric prior of the pseudo-depth map as additional coarse-grained geometric supervision.

[0306] First of all, since the pseudo-depth map reflects relative parallax information rather than absolute depth information, the embodiment of the present invention converts the rendered depth map into an approximate parallax map , where is a scale parameter; the approximate disparity map has a value characteristic of "objects closer appear larger and objects farther appear smaller" similar to the disparity space.

[0307] Secondly, to solve the scale ambiguity problem between the pseudo-depth map and the approximate disparity map , the embodiment of the present invention uses the Pearson correlation coefficient with scale invariance characteristics as a similarity metric.

[0308] Finally, the calculation method of the coarse-grained depth loss function is , and this formula is the intermediate variable after converting the rendered depth map into the disparity space; where represents the Pearson correlation coefficient calculation function.

[0309] 3.5.3.3. Local smoothing regularization term for the Gaussian opacity attribute of water bodies , and the local smoothing regularization term for the Gaussian color attribute of water bodies ;

[0310] In the second stage of optimizing the Gaussian attributes of water bodies, to ensure the optimization smoothness of the opacity offset attribute and the color offset attribute of each water body Gaussian, the embodiment of the present invention designs the local smoothing regularization term for the Gaussian opacity attribute of water bodies and the local smoothing regularization term for the Gaussian color attribute of water bodies respectively. The specific calculation methods of these two regularization terms are as follows:

[0311]

[0312]

[0313] Among them, represents the set of water body Gaussians participating in rendering; represents the set of water body Gaussians the p-th water body Gaussian's opacity offset in represents the set of water body Gaussians the p-th water body Gaussian's color offset in represents the -th water body Gaussian's set of its nearest neighbor water body Gaussians in the world coordinate system obtained by the KNN algorithm; mentioned here is a weight factor. The calculation formula of is is a scaling factor, and They are respectively the th Gaussian of water body and the mean attribute of its th neighboring Gaussian of water body in the world coordinate system.

[0314] 3.5.3.4. Final total loss function

[0315] The total loss function used in the final training is calculated as follows:

[0316] In the first stage of optimization, , .

[0317] In the second stage of optimization, , , , ; represents the weight of the coarse-grained depth loss function ; represents the weight of the local smooth regularization term of the water body Gaussian opacity attribute ; represents the weight of the local smooth regularization term of the water body Gaussian color attribute ; , and are used to balance the influence of different loss function terms on the final optimization result.

[0318] The Adam optimizer is used to optimize the learnable parameters of all Gaussian basis elements throughout the optimization process.

[0319] 3.6. Downstream task application stage

[0320] The 3D representation method of the underwater scene proposed in the embodiment of the present invention can meet the requirements of various downstream tasks. The following are examples of its applications in three downstream tasks: underwater scene restoration, underwater novel view synthesis, and underwater image restoration dataset construction.

[0321] 3.6.1. Example of underwater scene restoration task

[0322] First step, collect a set of corresponding original underwater image sets for an underwater scene, and use the method proposed in the embodiment of the present invention to learn the 3D representation of the corresponding underwater scene.

[0323] Second step, given the camera view corresponding to each underwater image in the original underwater image set , enable the image rendering branch of the scene itself, and render the scene itself image in the scene Gaussian set . Since the scene itself image During the rendering process, the Gaussian set of water bodies is not used, so it can reflect the true appearance color of the scene when there is no water body effect. Therefore, the rendered can be regarded as an underwater image image restoration result.

[0324] 3.6.2. Underwater New Viewpoint Synthesis Task Example

[0325] First step, collect a corresponding set of original underwater images for an underwater scene, and use the method proposed in the embodiments of the present invention to learn the 3D representation of the corresponding underwater scene.

[0326] Second step, given any camera viewpoint, enable the underwater image rendering branch, and render all the Gaussian basis elements in the union of the Gaussian set of water bodies and the Gaussian set of the scene, then the underwater image rendering branch can be used to render the underwater image from the corresponding easy viewpoint . Since the underwater image rendering branch uses all the Gaussian basis elements in the union of the Gaussian set of water bodies and the Gaussian set of the scene for rendering, the rendered underwater image is superimposed with the influence of the water body effect, reflecting the degraded appearance of the scene under the influence of the water body effect. Therefore, given any camera viewpoint, the 3D representation learned in the embodiments of the present invention can be used to render the underwater image from the corresponding viewpoint, that is, the synthesis of the underwater new viewpoint is realized.

[0327] 3.6.3. Underwater Image Restoration Dataset Construction Task Example

[0328] First step, collect a corresponding set of original underwater images for an underwater scene, and use the method proposed in the embodiments of the present invention to learn the 3D representation of the corresponding underwater scene.

[0329] Second step, given a set of obtained camera viewpoint sets, for each camera viewpoint in this set, respectively enable the scene itself image rendering branch and the underwater image rendering branch to render the scene itself image without the water body effect and the underwater image affected by the water body effect . For the and rendered for each camera viewpoint, the image content of the two reflects the content of the same underwater scene photographed from the same camera viewpoint. The difference is that reflects the true appearance color of the scene itself, while reflects the scene appearance with color degradation under the influence of the water body effect. Therefore, for the and rendered for each camera viewpoint, they can be regarded as the water-free ground truth image and the water-containing degraded image in the underwater image restoration dataset respectively, and an image pair for supervised learning is constructed.

[0330] Compared with the existing method, the embodiment of the present invention has the following advantages:

[0331] (1) Joint optimization of water parameters, scene geometry, and appearance color: In order to address the problem faced by traditional underwater scene 3D representation methods that the staged optimization is difficult to achieve optimal representation of both scene geometry and appearance color, the embodiments of the present invention simultaneously set Gaussian primitives representing the water medium and Gaussian primitives representing the scene itself in space, and render underwater images through a customized differentiable rasterization pipeline. In this way, the parameters of the water medium and the color geometry parameters of the scene itself can be learned simultaneously in the end-to-end optimization process, thereby achieving better scene representation quality.

[0332] (2) Explicitly characterizing the water body and the scene itself to achieve high-speed, high-quality characterization: In order to address the problem that the NeRF-based underwater scene 3D characterization method implicitly represents the water body and the scene itself, which results in huge training and rendering time and computing power overhead, and the quality of rendering results is difficult to satisfy, the embodiment of the present invention designs two types of attribute-customized Gaussian primitives, namely water body Gaussian and scene Gaussian, based on the characteristics of the water environment and the scene itself during underwater imaging, to explicitly represent the water body and the scene itself in the scene, respectively. This not only achieves faster training and rendering speeds, but also achieves fine-grained scene 3D characterization and achieves richer rendering quality.

[0333] (3) Directly and explicitly modeling water bodies to avoid dependence on the modeling quality of the scene itself: In response to the existing method of using 3DGS technology to perform 3D representation of underwater scenes, the embodiments of the present invention directly model the influence of water effects in the underwater imaging process by designing water body Gaussian basis elements and distributing them in large numbers in space. The optimization process of water body parameters is changed from relying on the reconstruction quality of the Gaussian of the scene itself as adopted by the existing method to being parallel to the optimization of the Gaussian of the scene itself, thereby being able to more accurately and effectively decouple the real appearance, geometry and water properties of the underwater scene in the end-to-end joint optimization process.

[0334] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0335] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An underwater scene 3D representation method based on three-dimensional Gaussian splashing, characterized in that Including: Obtain the original underwater image set captured in a single underwater scene, and use the COLMAP algorithm to calibrate the intrinsic matrix and extrinsic matrix of the camera corresponding to each underwater image in the original underwater image set, and at the same time output a sparse point cloud; Use a pre-trained monocular depth estimator to derive the pseudo-depth map corresponding to each underwater image in the original underwater image set; Based on the sparse point cloud, initialize the attributes of the water body Gaussian to obtain the corresponding water body Gaussian set; Based on the sparse point cloud and the pseudo-depth map, initialize the attributes of the scene Gaussian to obtain the corresponding scene Gaussian set; Iteratively optimize the attribute parameters in the water body Gaussian set and the scene Gaussian set to decouple and capture the corresponding water body attributes, the appearance color and geometric attributes of the scene itself.

2. The 3D characterization method of underwater scenes based on three-dimensional Gaussian splashing according to claim 1, characterized in that, The attributes of the water body Gaussian specifically include: (1) The mean attribute of the water body Gaussian: For each water body Gaussian, its mean attribute is denoted as: ; Among them, is the mean attribute of the p-th Gaussian of the water body, representing the point coordinates in the world coordinate system; represents the x-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; represents the y-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; represents the z-axis coordinate of the p-th Gaussian of the water body in the world coordinate system; The parameters , and are non-learnable parameters and remain fixed after initialization; (2) The covariance attribute of the water body Gaussian: Model the covariance attribute of the water body Gaussian as isotropic corresponding to a spherical distribution; for each water body Gaussian, its covariance attribute is denoted as: ; Among them, represents the covariance attribute of the p-th water body Gaussian; represents the covariance attribute shared parameter of all water body Gaussians, which is a non-learnable parameter and remains fixed after initialization; represents an identity matrix with a shape of 3×3; (3) The opacity attribute of the water body Gaussian: For each water body Gaussian, its opacity attribute is denoted as: ; Among them, represents the opacity attribute of the p-th water body Gaussian, which is a vector of length 3 corresponding to the three RGB channels to model the wavelength selectivity of the water body effect; represents the shared parameter of the opacity attributes of all water body Gaussians; represents the opacity offset corresponding to the current p-th water body Gaussian; and are both vectors of length 3; The function is used to ensure that the value range of each channel is (0, 1); and are both learnable parameters; (4) The color attribute of the water body Gaussian: For each water body Gaussian, its color attribute is denoted as: ; Among them, represents the color attribute of the p-th water body Gaussian; represents the color attribute sharing parameter of the full-space water body Gaussian, which is a vector of length 3 corresponding to the three channels of RGB; represents the color offset corresponding to the p-th water body Gaussian, which is a vector of length 3; The function is used to ensure that the value range of each channel is (0, 1); and are learnable parameters.

3. A 3D characterization method for underwater scenes based on three-dimensional Gaussian splashing according to claim 1, characterized in that The attributes of the scene Gaussian specifically include: (1) The mean attribute of the scene Gaussian: For each scene Gaussian, its mean attribute is denoted as: ; Among them, is the mean property of the r-th scene Gaussian, representing the point coordinates in the world coordinate system; represents the x-axis coordinate of the r-th scene Gaussian in the world coordinate system; represents the y-axis coordinate of the r-th scene Gaussian in the world coordinate system; represents the z-axis coordinate of the r-th scene Gaussian in the world coordinate system; , and are learnable parameters; (2) The covariance attribute of the scene Gaussian: Model the covariance attribute of the scene Gaussian as anisotropic corresponding to an ellipsoidal distribution; for each scene Gaussian, its covariance attribute is denoted as: ; Among them, represents the covariance attribute of the r-th scene Gaussian; represents the scale matrix of the ellipsoidal distribution; represents the rotation matrix of the ellipsoidal distribution; and are both learnable parameters; (3) The opacity attribute of the scene Gaussian: For each scene Gaussian, its opacity attribute is denoted as ; denotes the opacity attribute of the r-th scene Gaussian, which is a scalar value and a learnable parameter with a value range of (0, 1); (4) The color attribute of the scene Gaussian: For each scene Gaussian, its color attribute is denoted as: ; Among them, represents the color attribute of the r-th scene Gaussian; , and respectively correspond to the three RGB channels and are learnable parameters with a value range of (0, 1).

4. A 3D characterization method for underwater scenes based on three-dimensional Gaussian splashing according to claim 1, characterized in that, The initialization of the attributes of the water body Gaussian based on the sparse point cloud to obtain the corresponding water body Gaussian set specifically includes: (1) Initialize the mean attribute of the water body Gaussian: According to the distribution range of the sparse point cloud, calculate and set a cuboid bounding box as the scene boundary to limit the scene range; Inside the scene boundary, construct a regular grid in the XYZ three coordinate axes directions with a fixed step size, and place a water body Gaussian at each grid node, and use the position of the grid node in the world coordinate system as the mean attribute of the corresponding water body Gaussian; (2) Initialize the covariance attribute of the water body Gaussian: All Gaussian covariance property sharing parameters of water bodies , are initialized according to the formula ; among them, represents the scale factor; represents the fixed step size; (3) Initialize the opacity attribute of the water body Gaussian: For each water Gaussian, its opacity offset is initialized to [0, 0, 0]; Share the opacity attribute shared parameter of all water bodies Gaussian Initialize to ; (4) Initialize the color attribute of the water body Gaussian: For each water body Gaussian, initialize its color offset to [0, 0, 0]; Share the color attribute sharing parameters of all Gaussian water bodies is initialized to ; (5) Denote all the water body Gaussians whose mean attribute, covariance attribute, opacity attribute, and color attribute are all initialized as the water body Gaussian set.

5. A 3D characterization method for underwater scenes based on three-dimensional Gaussian splash according to claim 1, characterized in that The initialization of the attributes of the scene Gaussian based on the sparse point cloud and the pseudo-depth map to obtain the corresponding scene Gaussian set includes: Augment the sparse point cloud according to the pseudo-depth map to obtain an augmented point cloud; Based on the augmented point cloud, initialize the attributes of the scene Gaussian to obtain the corresponding scene Gaussian set.

6. The 3D characterization method of underwater scenes based on three-dimensional Gaussian splashing according to claim 5, wherein, The augmentation processing of the sparse point cloud according to the pseudo-depth map to obtain the augmented point cloud specifically includes: (1)According to the pseudo-depth map corresponding to each underwater image in the original underwater image set, each underwater image is divided into multiple regions at preset depth value intervals, and each region corresponds to a different depth interval on the pseudo-depth map; (2)Extract the sparse point cloud subset observable from the camera view corresponding to each underwater image in the original underwater image set, calculate the depth values of all points in the sparse point cloud subset in the corresponding camera coordinate system, and obtain the projection points of all points in the sparse point cloud subset on the corresponding image coordinate system, denoted as the original projection points; (3)For each region of each underwater image in the original underwater image set, calculate the number of pixels covered by the current region, determine the subset of original projection points of the sparse point cloud subset corresponding to the current underwater image in the current region, and determine the number of original projection points included in the current subset of original projection points, calculate the number of projection points to be newly added in the current region, and randomly add the corresponding number of projection points in the current region, denoted as the newly added projection points; The number of projection points to be newly added in the current region is expressed as: ; Among them, represents the set of projection points to be newly added in the t-th region of the h-th underwater image in the original underwater image set ; represents the number of projection points to be newly added in the set ; represents the minimum projection point density of each region; represents the t-th region of the h-th underwater image in the original underwater image set; represents the number of pixels contained in ; represents the original projection point subset in the t-th region of the h-th underwater image in the original underwater image set; represents the number of original projection points contained in the original projection point subset (4)For each newly added projection point in each region of each underwater image in the original underwater image set, the depth value and color attribute of the corresponding newly added spatial point in the camera coordinate system are expressed as: For each newly added projection point in each region of each underwater image in the original underwater image set, the depth value and color attribute of the corresponding newly added spatial point in the camera coordinate system ; Among them, m represents the m-th newly added projection point; represents the set composed of the k nearest original projection points of the m-th newly added projection point found by using the K-nearest neighbor algorithm from the set of original projection points in the current area; represents the depth value of the original space point corresponding to the n-th original projection point in represents the color value of the original space point corresponding to the n-th original projection point in represents the weighted weight; Among them, denote the coordinates of the m-th newly added projection point in the image coordinate system as , and denote the coordinates of the n-th original projection point in the image coordinate system as , then the weighted weight is expressed as: ; (5)According to the camera projection transformation and the external parameter matrix of the camera, obtain the position of the newly added spatial point corresponding to each newly added projection point in the world coordinate system; (6)Calculate the positions and color attributes of all newly added spatial points in all regions of all underwater images in the original underwater image set. The union of the set composed of these newly added spatial points and the set of original spatial points corresponding to the sparse point cloud is the augmented point cloud.

7. A method for 3D characterization of underwater scenes based on three-dimensional Gaussian splashing according to claim 5, characterized in that During each iterative optimization of the attribute parameters in the water body Gaussian set and the scene Gaussian set, an underwater image is selected from the original underwater image set. and its corresponding pseudo depth map , using the current iteration of the water body Gaussian primitives and the scene Gaussian primitives through the channel-by-channel multi-output rasterization rendering module, to obtain the underwater image under the same camera perspective , the scene image itself , pure water background image and depth map ; By calculating the loss function and returning the gradient, the Gaussian properties of the water body and the scene are optimized; The per-channel multi-output rasterization rendering module includes: (1)Underwater image rendering branch: ① Denote the set composed of all water body Gaussians as , and denote the set composed of all scenario Gaussians as ; ②For the p-th water Gaussian, its opacity attribute is reformulated as , where , and are 's three opacity components. These three opacity components are independent of each other and respectively correspond to the three RGB channels; ③ For the r-th scene Gaussian, re-express its opacity property as , where , and are 's three opacity components, corresponding to the three RGB channels respectively, and ; ④ Given the camera perspective corresponding to the underwater image to be rendered, that is, the camera's intrinsic parameter matrix and extrinsic parameter matrix, find the water body Gaussian set and scene Gaussian set The union of , then for the set All Gaussian primitives in are sorted by their depth values ​​in a given camera coordinate system for tile-based - Blending rendering; ⑤ The color of each channel of any pixel of the underwater image to be rendered , is expressed as: ; Among them, represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element in the rendering process. Here, the color channel of the i-th Gaussian basis element corresponds to the color channel of ; represents the opacity contribution of the i-th Gaussian basis element in the corresponding color channel during the rendering of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel during the rendering process; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel during the rendering process; ⑥ Independently render each color channel of any pixel of the underwater image to be rendered; ⑦Traverse all pixels of the underwater image to be rendered and complete the rendering, and output the final underwater image ; (2)Scene itself image rendering branch: ①Extract the Gaussian set of the scene All Gaussian basis elements in it; ②For the r-th scene Gaussian, its opacity attribute is re-denoted as a scalar ; ③ Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, the rasterization rendering pipeline in the original 3DGS technology is used to process all the Gaussian basis elements in the scene Gaussian set and output the final image of the scene itself ; (3)Pure water background image rendering branch: ①Extract all Gaussian basis elements in the Gaussian set of the water body ; ②For the p-th water body Gaussian, its opacity attribute is reformulated as , where The three opacity components of are independent of each other; ③ Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, for the Gaussian set of water bodies sort all the Gaussian basis elements in it according to their depth values in the given camera coordinate system, so as to carry out tile-based rendering; ④ The color of each channel of any pixel of the pure water background image to be rendered is expressed as: ; Among them, represents the position coordinates of the pixel in the image coordinate system; corresponds to different color channels; represents the color value of the color channel of the i-th Gaussian basis element in the rendering process. Here, the color channel of the i-th Gaussian basis element corresponds to the color channel; represents the opacity contribution of the i-th Gaussian basis element in the corresponding color channel during the rendering of the current pixel, and ; represents the opacity component of the i-th Gaussian basis element in the corresponding color channel during the rendering process; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel. ⑤ Independently render the color of each channel of any pixel of the pure water background image to be rendered; ⑥Traverse all pixels of the pure water background image to be rendered and complete the rendering, and output the final pure water background image ; (4)Depth map rendering branch: ①Extract the Gaussian set of the scene All Gaussian basis elements in it are involved in rasterization and rendering; ②For the r-th scene Gaussian, its opacity attribute is re-denoted as a scalar ; ③ Given the camera view corresponding to the image to be rendered, i.e., the intrinsic matrix and extrinsic matrix of the camera, for the scene Gaussian set sort all the Gaussian basis elements in it according to their depth values in the given camera coordinate system, so as to carry out tile-based rendering; ④ The depth value of any pixel of the depth map to be rendered is expressed as: ; Among them, represents the position coordinates of the pixel in the image coordinate system, represents the opacity contribution of the i-th Gaussian basis element in the current pixel rendering process during the rendering process, and ; represents the value of the 2D Gaussian distribution of the i-th Gaussian basis element in the current image coordinate system at the current pixel during the rendering process; represents the depth value of the i-th Gaussian basis element in the current camera coordinate system during the rendering process; ⑤ After all the pixels of the image to be rendered are completed, the desired depth map is obtained. 。 8. A method for 3D characterization of underwater scenes based on three-dimensional Gaussian splashing according to claim 1, characterized in that In the process of iteratively optimizing the attribute parameters of the water body Gaussian, the attribute optimization strategy of the water body Gaussian includes: (1)The mean attribute and covariance attribute of each water body Gaussian remain frozen after initialization and no longer participate in the optimization process; (2)According to the number of parameter optimization iteration rounds, the optimization process of the opacity attribute and color attribute of each water body Gaussian is divided into two stages: During the optimization process of the first stage, the opacity offset of each water body Gaussian and the color offset are kept frozen and do not participate in the optimization. Only the opacity attribute sharing parameters of all water body Gaussians and the color attribute sharing parameters are optimized; During the optimization process of the second stage, thaw the opacity offset of each water body Gaussian and the color offset , and optimize them together with the opacity attribute sharing parameters and the color attribute sharing parameters shared by all water body Gaussians.

9. A method for 3D characterization of underwater scenes based on three-dimensional Gaussian splashing according to claim 7, characterized in that During the process of iteratively optimizing the attribute parameters in the water body Gaussian set and the scene Gaussian set, the total loss function used is expressed as: In the first stage of optimization, ; In the second stage of optimization, ; Among them, represents the loss function based on image content; represents the coarse-grained depth loss function; represents the local smoothing regularization term of the Gaussian opacity attribute of water bodies; represents the local smoothing regularization term of the Gaussian color attribute of water bodies; represents the coarse-grained depth loss function weight; represents the local smoothing regularization term of the Gaussian opacity attribute of water bodies weight; represents the local smoothing regularization term of the Gaussian color attribute of water bodies weight.

10. The 3D representation method of an underwater scene based on three-dimensional Gaussian splash according to claim 9, characterized in that: (1) Construction method of loss function based on image content Specifically includes: ①Create a background mask using the pseudo-depth map of the current iteration round ; Given a predefined depth threshold ; The region composed of pixels with pseudo-depth values greater than is regarded as the foreground region, and the pixels at the corresponding positions of the corresponding background mask take the value 0; The region composed of pixels with pseudo-depth values less than or equal to is regarded as the background region, and the pixels at the corresponding positions of the corresponding background mask are set to 1; ​ ② Denote the underwater images in the original underwater image set as the underwater image ground truth, with the symbol , and construct a weight map with the same shape as the pseudo-depth map ; specifically, the value of the l-th pixel of the weight map is determined according to the following requirements: When the background mask corresponding to the l-th pixel has a value of 0, the l-th pixel of the weight map corresponds to the foreground region. At this time, the weight map corresponding to the l-th pixel is expressed as: ; Among them, represents a gradient cut-off operator; represents the color value corresponding to the l-th pixel of the rendered underwater image; When the background mask corresponding to the l-th pixel has a value of 1, the l-th pixel corresponds to the background area. At this time, the weight map corresponding to the l-th pixel is expressed as: ; Then, an underwater image with enhanced foreground is obtained and the ground truth of the underwater image with enhanced foreground The relationship is as follows: ; Among them, represents an element-wise multiplication operator; ③Extract the background mask All pixels with a value of 1 are taken to obtain the set of background pixels, denoted as , and all pixels belonging to this set correspond to areas in the picture where there is no underwater scene and only pure water exists; ④ Loss function based on image content It is expressed as: ; Among them, represents the weight of the loss function term ; represents the weight of the loss function term ; represents the weight of the loss function term ; represents the loss function in the original 3DGS technology; represents the color value of the pure water background image at the th pixel in the region; represents the color value of the underwater image ground truth at the th pixel in the region; (2)Coarse-grained depth loss function is constructed in the following specific ways: Convert the rendered depth map into an approximate disparity map , where represents the scale parameter; Adopt the Pearson correlation coefficient with scale invariance characteristics as a similarity metric; Coarse-grained depth loss function Expressed as: ; Among them, represents the Pearson correlation coefficient calculation function; represents the pseudo-depth map; (3)Local smoothing regularization term for the Gaussian opacity property of water bodies , expressed as: ; (4)Local smoothing regularization term for Gaussian color attributes of water bodies , expressed as: ; Among them, represents the Gaussian set of water bodies participating in rendering; represents the set composed of the k nearest neighbor water body Gaussians of the p-th water body Gaussian in the world coordinate system obtained by the KNN algorithm; represents a weight factor, and is expressed as ; represents a scale factor; represents the mean attribute of the p-th water body Gaussian in the world coordinate system; represents the mean attribute of the q-th water body Gaussian in the world coordinate system.

Citation Information

Patent Citations

  • Sparse visual angle three-dimensional reconstruction method based on depth prior information

    CN118657888A

  • Efficient 3D Gaussian scene reconstruction method based on multi-modal depth distribution supervision

    CN119648925A

  • Structural perception three-dimensional scene reconstruction method and device

    CN119888133A

  • Three-dimensional scene-oriented real-time rendering method based on spatial perception

    CN119941950A

  • Underwater three-dimensional scene reconstruction method and device based on three-dimensional Gaussian splashing and underwater imaging model

    CN120014164A

Cited By

  • Multi-model collaborative two-dimensional Gaussian splash three-dimensional reconstruction method

    CN121392157A

  • Three-dimensional Gaussian splash reconstruction method for underwater scene

    CN121639946A