End-to-end underwater three-dimensional reconstruction method and system based on underwater imaging model
Through an end-to-end underwater 3D reconstruction method, combined with a pose estimation network and an underwater imaging model, the accuracy and robustness issues of underwater 3D reconstruction methods in dynamic and complex environments are solved, and high-precision 3D reconstruction results are achieved, which are suitable for ocean exploration and underwater robot navigation.
Patent Information
- Application Number
- CN202511120632.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing underwater 3D reconstruction methods are highly dependent on underwater image quality and camera pose estimation accuracy, have poor adaptability, and lack robustness, especially in dynamic and complex underwater environments.
An end-to-end underwater 3D reconstruction method is adopted, combining the pose estimation network and underwater imaging model. Through deep learning, underwater light propagation and 3D volume structure are simulated, the pose estimation and 3D reconstruction processes are integrated, and the light attenuation and scattering blur in the underwater environment are adaptively simulated, reducing dependence on traditional methods.
It improves the accuracy and robustness of underwater 3D reconstruction, enhances the system's adaptability to different water qualities and environmental conditions, simplifies the operation process, and generates high-precision, detail-rich 3D reconstruction results, which are suitable for fields such as ocean exploration and underwater robot navigation.
Smart Images

Figure CN120635333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, image processing and three-dimensional reconstruction, and in particular to an end-to-end underwater three-dimensional reconstruction method and system based on an underwater imaging model. Background Art
[0002] The rapid development of fields such as marine resource development, seabed exploration, marine engineering, and underwater robotics has placed higher demands on underwater environment perception and 3D modeling. As a key component of underwater perception, underwater 3D reconstruction technology, which restores the geometric structure of target objects from image information, is crucial for ensuring underwater operational safety and improving automated operations.
[0003] However, compared to terrestrial environments, underwater scenes are characterized by complex lighting conditions, variable medium properties, and significant imaging degradation, including rapid illumination decay, color casts, scattering blur, and overall contrast loss. These degradation factors severely degrade underwater image quality, further impacting key steps such as image feature extraction, cross-view matching, and camera pose estimation, significantly limiting the accuracy and robustness of 3D reconstruction.
[0004] Traditional underwater 3D reconstruction methods are typically based on technologies such as structured light, laser scanning, stereo vision, or SfM (Structure-from-Motion). However, these methods have the following major shortcomings: 1. Lack of modeling of the physical mechanisms of underwater imaging: This approach cannot adapt to the imaging degradation characteristics in different water quality environments, affecting image quality restoration and subsequent modeling accuracy. 2. Highly dependent on camera pose estimation accuracy: Traditional methods rely on high-quality image features and precise camera calibration, which can lead to positioning deviations or reconstruction failures in dynamic scenes or areas with weak textures. 3. Complex processing flow and poor robustness: Usually requires tedious image preprocessing steps (such as color correction, white balance, scatter removal, etc.), resulting in insufficient system stability and adaptability; 4. Insufficient end-to-end capabilities: Most methods are modular in design, and the various stages cannot be jointly optimized, which easily leads to error transmission and information loss.
[0005] In recent years, the rapid development of deep learning technology, particularly methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting, has provided new insights into underwater 3D reconstruction. These methods enable researchers to more precisely simulate underwater light propagation and 3D volumetric structure, resulting in high-quality 3D reconstructions. However, these methods typically rely on traditional camera pose estimation techniques (such as COLMAP) to obtain accurate camera pose information. This poses significant challenges in dynamic underwater environments, especially for motion devices and complex scenes requiring real-time performance. Summary of the Invention
[0006] The purpose of the present invention is to propose an end-to-end underwater 3D reconstruction method and system based on an underwater imaging model to solve the problems of existing underwater 3D reconstruction methods, such as strong dependence on underwater image quality and camera pose estimation accuracy, poor adaptability, and insufficient robustness.
[0007] To achieve the above object, the present invention is achieved through the following technical solutions: An end-to-end underwater 3D reconstruction method based on an underwater imaging model comprises the following steps: S1: Collect multiple underwater image sequences to obtain raw image data containing complex underwater environment features such as light attenuation, scattering blur, and color shift; S2: Construct an end-to-end learnable pose estimation network to perform spatiotemporal modeling and feature extraction on the input image sequence and predict the relative six-degree-of-freedom pose between image frames. The pose estimation network structure includes: an input module, an image feature extraction module, a fusion encoding module, a Transformer encoding module, and a pose regression module. S3: Build an underwater imaging model to simulate the underwater image degradation process caused by the medium and embed it into the 3D rendering process to enhance color reproduction and physical consistency of imaging; S4: Based on the pose estimation network and underwater imaging model, a dense reconstruction module based on 3D Gaussian distribution is constructed to perform 3D reconstruction and output a 3D reconstructed image.
[0008] Furthermore, the pose estimation network in S2 includes the following modules: (1) Input module: Receives multiple multi-view images from underwater scenes as observation input for pose estimation, providing basic data for subsequent feature extraction and inference; (2) Image feature extraction module: The deep visual encoding network is used to extract high-dimensional spatial features of the input image, obtaining a feature map containing rich spatial structural information. The feature map is then flattened and positionally encoded to form an image token with spatial position information, which facilitates subsequent Transformer processing and global context modeling. (3) Fusion encoding module: combines the image token with a set of learnable camera tokens, which serve as implicit camera pose representations. This module integrates visual information with camera priors, constructs a joint context representation, and provides rich input semantics and geometric information to the Transformer, assisting pose inference. (4) Transformer Encoding Module: This module consists of a multi-layer Transformer architecture, including standard components such as Multi-Head Self Attention, Multi-Head Cross Attention, a FeedForward network, residual connections, and layer normalization. This module deeply models the joint sequence of image tokens and camera tokens, mining the spatiotemporal dependencies and structural matching cues between multiple viewpoints to achieve context-aware pose estimation.
[0009] (5) Pose regression module: Based on the joint features encoded by Transformer, the relative rotation (usually in the form of quaternions) and translation parameters corresponding to the image are predicted through a fully connected regression network to achieve high-precision camera pose estimation.
[0010] Furthermore, in S2, in order to achieve data compatibility and seamless connection between the pose estimation network and the 3D reconstruction module, a standardized data conversion process is further designed to uniformly convert the camera pose, intrinsic parameters and sparse point cloud obtained by the pose estimation network into structured input that can be used by the subsequent reconstruction module; specifically, the following steps are included: (1) Pose conversion and format normalization: The relative pose sequence predicted by the network is converted into the global camera pose of each frame (external parameters in the world coordinate system) through matrix multiplication, and then the rotation matrix is transformed into a quaternion to adapt to the image-pose expression format required by standard 3D reconstruction systems such as COLMAP.
[0011] (2) Camera intrinsic parameter initialization and normalization: If there is no explicit calibration data, the default parameters are used (such as focal length fx=fy=500, the principal point is located at the center of the image), and the size and intrinsic parameter format of all images are unified to generate the PINHOLE model parameters corresponding to each image.
[0012] (3) Synchronous packaging of images and poses: Traverse the image sequence, correspond the image file names to the predicted poses one by one, and write them into cameras.txt and images.txt to describe the camera model and external parameter file name information of each frame image, ensuring that the image frame timing is consistent with the pose transformation.
[0013] (4) Sparse point cloud extraction and color addition: Read the spatial point position information from the generated initial point cloud (such as .ply format), extract the color of each point (if available), and generate a sparse point set with color, and write it into the standard points3D.txt format as the initial scene structure representation.
[0014] (5) Unified output standard ternary data format: The final output files are cameras.bin, images.txt and points3D.txt, which constitute the sparse reconstruction input format that meets the COLMAP / 3DGS input requirements and are loaded and used by the subsequent differentiable rendering module based on Gaussian representation.
[0015] Through structured conversion and format standardization, the above processing flow enables the feedforward network to directly generate sparse model input data for 3D reconstruction without the need for manual calibration or the involvement of third-party SfM tools, supporting the complete closed-loop operation of the SfM-Free 3D reconstruction system.
[0016] Furthermore, the underwater imaging model includes a medium modeling module (MLP) that models the underwater light propagation and imaging process in an end-to-end learnable manner: ; : is the pixel color finally observed on channel c∈{R,G,B} in the camera image; : is the true reflection color of the 3D Gaussian point corresponding to the pixel, generated by the 3D reconstruction module; d(x): represents the depth value corresponding to the three-dimensional point, that is, the distance from the point to the center of the camera; : Absorption coefficient, used to simulate the exponential decay of light intensity during propagation; : Scattering coefficient, used to simulate the superposition contribution of water background light to color; : Medium background light color (or inherent water color), reflecting the basic offset of image color caused by water quality; The underwater imaging model takes three-dimensional spatial coordinates (x, y, z) as input and outputs the above parameters through a multi-layer perceptron structure to achieve modeling of light propagation characteristics in different water environments.
[0017] Furthermore, the medium modeling module is implemented in the form of a multi-layer perceptron (MLP), which receives a three-dimensional spatial position (x, y, z)(x, y, z)(x, y, z) as input and outputs underwater imaging parameters of the position in each color channel, including: a. Absorption coefficient : Used to describe the energy attenuation of incident light during its propagation in water; b. Scattering coefficient : Used to model the color superposition and blurring caused by backscattering; c. Background light color : Indicates the color of the water itself or the background light intensity.
[0018] Furthermore, the specific steps of S4 are as follows: S4-1: Initialize Gaussian Point Set: Initialize a sparse 3D point set based on the number and pose information of the input image. Each point is represented by a learnable 3D Gaussian distribution, which includes parameters such as spatial position, covariance, color attributes, and opacity.
[0019] S4-2: Construct an underwater projection mechanism: Use the camera pose and the underwater imaging model parameters to project the three-dimensional Gaussian points onto the image plane, and construct a differentiable projection rendering model that conforms to the laws of physical propagation.
[0020] S4-3: Gaussian Differentiable Rendering: Based on the 3D Gaussian Differentiable Rendering framework, color, transparency, and coverage are fused in the image space to generate a composite image and the following joint loss function is calculated.
[0021] S4-4: Color consistency loss: Minimize the pixel error between the rendered image and the ground truth image; S4-5: Transparency regularization: constrains point redundancy and improves efficiency;
[0022] S4-6: Depth consistency loss (optional): constrains the point cloud geometry and improves 3D restoration accuracy.
[0023] S4-7: Hierarchical optimization strategy: A coarse-to-fine optimization process is used to increase the number of Gaussian points and image resolution in successive rounds, supplemented by a redundant elimination mechanism to retain only points that contribute significantly to the reconstruction quality.
[0024] S4-8: The final output is a high-fidelity dense point cloud with complete geometric structure, texture, and boundary details. This module can be trained collaboratively with the pose estimation network and imaging modeling module to build an end-to-end 3D reconstruction system that does not require manual calibration or third-party SfM tools.
[0025] An end-to-end underwater 3D reconstruction system based on an underwater imaging model, including the following modules: Image data acquisition module: This module is used to collect multiple frames of underwater image sequences and obtain original image data; Pose estimation network module: This module performs spatiotemporal modeling and feature extraction on the input image sequence to predict the pose between image frames; Underwater Imaging Model Module: This module simulates the underwater image degradation process caused by the medium and is embedded into the 3D rendering process to enhance color reproduction and physical consistency of imaging. 3D reconstruction module: performs 3D reconstruction based on dense reconstruction of 3D Gaussian distribution and outputs 3D reconstructed images.
[0026] Compared with the prior art, the advantages and technical effects of the present invention are:
[0027] The present invention has several significant advantages. First, through an end-to-end optimization framework, the pose estimation and three-dimensional reconstruction processes are integrated into the same deep neural network architecture, achieving joint optimization and reducing the dependence of traditional methods on pose estimation and image preprocessing, thereby improving the reconstruction accuracy and robustness. Secondly, the introduction of the adaptive underwater imaging model enables the present invention to dynamically simulate imaging degradation factors such as light attenuation, scattering blur and color shift in the underwater environment, enhancing the system's adaptability to different water qualities and environmental conditions, thereby improving the realism and accuracy of the reconstruction results. In addition, the pose estimation module without external calibration requirements can directly derive the camera's pose from the underwater image sequence, simplifying the system's setup and operation and improving ease of use. In terms of robustness, the system optimizes pose estimation and underwater imaging modeling, successfully solving problems such as low image quality and scattering blur in complex underwater environments, and enhancing the robustness and stability of the system in dynamic and low-contrast environments. Finally, by combining 3D Gaussian Splatting with volume rendering technology, the generated 3D reconstruction results are highly accurate and rich in details, and can accurately restore the geometric structure and texture of underwater scenes. It provides high-precision 3D data support for fields such as ocean exploration, underwater robot navigation, and structural health inspection, significantly improving the feasibility and reliability of practical applications.
[0028] This invention will help improve the accuracy and robustness of underwater 3D reconstruction, especially in complex dynamic environments, and promote its implementation and development in practical applications such as offshore wind power, marine aquaculture, and dock construction. Therefore, this invention not only has important theoretical significance but also has a positive impact on practical applications such as underwater data acquisition, environmental monitoring, and resource management. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is the overall flow chart of the present invention.
[0030] Figure 2 Underwater multi-viewing Figure 3 Schematic diagram of the framework of the dimensional reconstruction method.
[0031] Figure 3 This is the pose estimation network architecture diagram.
[0032] Figure 4 This is the flow chart of the 3D reconstruction system.
[0033] Figure 5 This is the result of the three-dimensional reconstruction of the present invention. DETAILED DESCRIPTION
[0034] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described below in detail with reference to specific embodiments and the accompanying drawings. It is apparent that the embodiments described are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments disclosed herein without inventive effort are intended to fall within the scope of protection of the present invention.
[0035] Example 1: This embodiment proposes an end-to-end underwater 3D reconstruction method based on an underwater imaging model. By building a unified deep learning framework, it integrates camera pose estimation, underwater image modeling, and 3D reconstruction processes. This method can effectively adapt to different underwater environmental changes and improve reconstruction accuracy and system robustness. This method does not require traditional external calibration or multi-stage processing, and can automatically recover camera motion trajectories and dense 3D structures directly from multi-frame underwater image sequences. It is suitable for high-quality 3D modeling tasks in complex underwater environments. Figure 1 、 Figure 2 As shown, the three-dimensional reconstruction method includes the following steps: Step 1: Multi-frame image input and overall network construction: First, a multi-frame underwater image sequence is continuously collected through a multi-view underwater camera device to ensure the temporal continuity and spatial coverage of the image, and to obtain the original image data containing complex underwater environment features such as light attenuation, scattering blur and color shift, providing rich perceptual information for subsequent pose estimation and 3D reconstruction. The continuous multi-frame underwater image sequence is collected as input data, denoted as , where t represents the time frame number.
[0036] The input image is processed by a unified end-to-end deep neural network. This network architecture comprises two core functional submodules: a pose estimation module, which estimates the six-degree-of-freedom camera transformation parameters between adjacent image frames; and a 3D reconstruction module, which combines the estimated pose with the image content to generate a dense point cloud or voxel representation. The network employs a joint optimization mechanism to ensure that the reconstruction task's high dependency on the pose estimation results is co-optimized during end-to-end training, significantly improving overall accuracy and stability.
[0037] Step 2: If Figure 3 As shown, the implementation of the camera pose estimation module: To achieve a 3D reconstruction process without the need for external SfM tools, we designed a Transformer-based end-to-end camera pose estimation network to directly regress the camera pose from multiple frames of underwater imagery. Its core architecture, input and output definitions, and training mechanism are as follows: (1) Input and output definition: Input image sequence, continuous T frame images, size is (T, 3, H, W), resolution is H = 384, W = 512; finally output the global pose corresponding to each frame image, a total of T poses.
[0038] (2) Network structure: a. Image feature extraction module: This module uses a ResNet-18 encoder with the last fully connected layer removed to extract multi-scale deep visual features from the input image. The output feature map has 128 channels, and the feature map size corresponds to the downsampled spatial size (e.g., 48×64).
[0039] b. Image Patch Encoding + Transformer Encoder: Each image is divided into 16×16 image blocks. Each block undergoes a linear transformation and is mapped into a fixed-dimensional vector (e.g., 128). Positional encoding is then applied to form a sequence of tokens. Assuming each image is H×W in size, each frame is encoded into H / 16×H / 16 patch tokens. The token sequences of multiple frames are concatenated in chronological order to form the input sequence for the Transformer Encoder.
[0040] c. Cross-frame feature fusion and spatiotemporal modeling: The input sequence passes through a multi-layer (four layers in this example) standard Transformer encoder. Each layer incorporates multi-head self-attention, a multi-layer feedforward network, residual connections, and layer normalization. The Transformer module uses self-attention to implicitly model the spatiotemporal dependencies and structural correspondences between multiple frames, enabling cross-frame feature fusion. This mechanism effectively captures geometric and content matching under varying viewpoints without requiring explicit computation of matching maps or optical flow.
[0041] d. Pose Regression Module: Finally, the token sequence output by the Transformer is reshaped into a frame-wise descriptor for each frame. This descriptor represents a global semantic vector or token for each frame. This representation is input into the pose regression module (composed of two MLPs), which predicts the relative rotation parameters (quaternions) and translation vectors between each pair of image frames, thereby deriving a global pose sequence relative to the first frame.
[0042] e. Data structure conversion: Export to COLMAP-style format (cameras.txt / images.txt) for 3D reconstruction input. If no internal parameters are available, use the PINHOLE model parameters: fx=fy=500, cx=W / 2, cy=H / 2. Finally, package the pose and image sequence in order and output a scene representation suitable for 3D Gaussian Splatting.
[0043] Step S3: Figure 4 As shown, underwater imaging model modeling and optimization: To enhance the realism and adaptability of 3D reconstructions in diverse underwater environments, a learnable underwater imaging model was introduced to simulate the propagation and attenuation of light in water. This model, deeply integrated with the 3D differentiable rendering module, provides a physically accurate imaging foundation for image synthesis, effectively improving the system's robustness to factors such as water quality variations, depth attenuation, and color shift.
[0044] The underwater imaging model includes a medium modeling module (Medium MLP): This module is implemented in the form of a multi-layer perceptron (MLP), receives a three-dimensional spatial position (x, y, z)(x, y, z)(x, y, z) as input, and outputs underwater imaging parameters of the position in each color channel, including: d. Absorption coefficient : Used to describe the energy attenuation of incident light during its propagation in water; e. Scattering coefficient : Used to model the color superposition and blurring caused by backscattering; f. Background light color : Indicates the color of the water itself or the background light intensity.
[0045] These parameters are generated by Medium MLP in the form of continuous spatial functions and embedded into the rendering process as physical fields, which can be optimized during training to adapt to the real image characteristics under different water conditions.
[0046] To realistically simulate the image degradation process in underwater environments and improve the adaptability and robustness of the reconstruction system under different water quality conditions, an adaptive underwater imaging model is introduced. This model models the underwater light propagation and imaging process in an end-to-end learnable manner: ; : is the pixel color finally observed on channel c∈{R,G,B} in the camera image; : is the true reflection color of the 3D Gaussian point corresponding to the pixel, generated by the 3D reconstruction module; d(x): represents the depth value corresponding to the three-dimensional point, that is, the distance from the point to the center of the camera; : Absorption coefficient, used to simulate the exponential decay of light intensity during propagation; : Scattering coefficient, used to simulate the superposition contribution of water background light to color; : Medium background light color (or inherent water color), reflecting the basic offset of image color caused by water quality.
[0047] The above medium parameters 、 and These are generated by a learnable medium modeling module (Medium MLP). This module takes three-dimensional spatial coordinates (x, y, z) as input and outputs the above parameters through a multi-layer perceptron structure, enabling modeling of light propagation characteristics in different water environments.
[0048] The underwater imaging model adopts a physical parameterized expression method, mainly including imaging parameters such as light attenuation coefficient, medium background brightness, and optional scattering factor. The imaging parameters can be estimated through image content statistics, or obtained as optimizable variables through end-to-end learning during training, and are encapsulated and called in a one-to-one correspondence with the image frame index. The imaging parameters are integrated with the projection mechanism of three-dimensional Gaussian points during the three-dimensional rendering process, and differentiable physical rendering of colors is achieved through a color modulation function based on the medium transmittance, thereby improving the consistency and authenticity between the rendered image and the real underwater image. This module significantly improves the robustness and accuracy of the system's three-dimensional reconstruction under different water quality, depth and lighting conditions, providing strong support for subsequent end-to-end optimization.
[0049] Step S4: 3D reconstruction based on Gaussian rendering: After completing image input, pose estimation, and adaptive underwater imaging modeling, the system enters the dense reconstruction phase based on a 3D Gaussian distribution. The core of this phase is to build a physically consistent and differentiably optimized image synthesis mechanism, enabling the entire process from sparse point initialization to high-quality point cloud recovery.
[0050] First, a set of Gaussian points is initialized in 3D space. Each Gaussian point is represented by learnable parameters, including: 3D position coordinates μ = (x, y, z), covariance matrix Σ, color vector c = (r, g, b), opacity α, and then a set of point clouds is initialized. Next, the rendering process begins: (1) Differentiable rendering process based on physical modeling: In the rendering stage, the system projects the Gaussian points onto the image plane and integrates the medium absorption and scattering parameters obtained in the previous stage (S3) to complete the physically consistent rendering based on the light propagation model.
[0051] (2) Joint loss function design: To achieve end-to-end optimization and image reconstruction consistency, this system uses a joint loss function for optimization during training, including reconstruction error (L1) and structural similarity (SSIM) loss to measure the color and structural consistency between the synthetic image and the real image. The L1 loss is used to ensure the accuracy of color restoration, while the SSIM loss helps maintain the stability of the image edge and texture structure. The weighted combination of the two constitutes the final loss function, guiding the system to achieve the optimal balance between three-dimensional structure, image appearance and physical consistency. The following loss function composition: a. Reconstruction loss (L1 Loss): ; in is the system rendered image, I(x) is the real image, and N is the total number of pixels.
[0052] b. Structural Similarity Loss (SSIM Loss): To enhance structural fidelity, the SSIM loss term is introduced to measure the similarity of image structures in local areas: ,SSIM comprehensively evaluates image consistency through brightness, contrast and structural ,similarity.
[0053] c. Total loss function (weighted combination): Combining the above two indicators, the final optimization goal is: ,in , is the weight coefficient, usually set to =1.0, = 0.1∼0.5 to balance the reconstruction accuracy and structural fidelity.
[0054] This embodiment uses 3D Gaussian Splatting or volume rendering-based methods to perform high-quality 3D reconstruction. By combining image content with estimated pose information, a dense point cloud, mesh, or voxelized 3D representation is generated, effectively restoring the geometry and texture information of the underwater scene. The system supports real-time visualization and various post-processing operations, such as filtering and noise reduction, texture mapping, and model optimization, improving the model's usability and expressiveness in complex engineering environments. The resulting output is a complete and detailed 3D model, widely applicable to various underwater mission scenarios, including marine resource exploration, underwater robot navigation, structural health inspection, and emergency operation analysis.
[0055] Example 2: In order to verify the effectiveness of the end-to-end underwater 3D reconstruction method based on the underwater imaging model proposed in this paper, a set of typical experiments were designed, covering 3D modeling tasks under different water conditions.
[0056] This embodiment uses some public data sets and some self-taken data sets. This example shows a public underwater data set, which is processed into an image with an image resolution of 512×384 as input. In the experiment, the system adopts the complete end-to-end network architecture described in Example 1, without the need for external camera calibration, depth map or structured light assistance, and directly inputs the image sequence, and performs the following processes in sequence: using the Transformer pose estimation network to regress the relative camera pose between consecutive image frames and convert it into a global pose sequence; constructing an adaptive underwater imaging module to estimate light attenuation, scattering and background light parameters based on the three-dimensional point coordinates; performing differentiable three-dimensional image synthesis based on the physical model in the Gaussian rendering module; and jointly optimizing the pose and Gaussian point attributes in the network through pixel-level L1 loss and SSIM loss, and finally outputting a high-fidelity dense three-dimensional point cloud.
[0057] To further verify the effectiveness of the proposed physical modeling module, Figure 5 A comparison of image rendering results in a typical underwater scene is shown: the left image shows the rendering result without underwater imaging modeling (Render), and the right image shows the rendering result after incorporating water optical modeling (With Water). It can be observed that the introduction of physical modeling allows the system to more accurately reproduce the light attenuation and scattering phenomena in underwater images, enhancing the realism and visual consistency of the synthesized images, fully demonstrating the effectiveness and advantages of the proposed method in complex water environments.
[0058] Figure 5This figure shows a comparison of image rendering results for a typical underwater scene: the left image shows the result (render) without considering the underwater imaging model, and the right image shows the rendering result (with_water) after introducing physical modeling. It can be seen that the inclusion of the model more accurately reproduces underwater light attenuation and scattering, enhancing the realism and consistency of the image.
[0059] In order to verify the effectiveness of the present invention, a comparison was made with some existing methods, and the results are shown in the following table: ; The table above clearly shows that the proposed method has significant advantages over existing technologies in multiple key metrics. In terms of PSNR, the proposed method achieved 26.1, 23.7, and 30.8 dB in the Coral, JAN, and Curasao scenes, respectively, outperforming both Seasplat and CF-GS. The improvement was particularly significant in the complex underwater environment of Curasao, demonstrating higher image fidelity. In terms of SSIM, the proposed method was slightly lower than Seasplat in the Coral and JAN scenes but far exceeded CF-GS, while significantly outperforming both in Curasao, demonstrating more stable structure preservation and detail restoration capabilities. Furthermore, the LPIPS metric shows that the proposed method significantly outperformed the comparison methods in all scenes, with significant improvements in perceptual quality and realism.
[0060] In summary, the method of the present invention achieves comprehensive optimization in the three dimensions of image fidelity, structural restoration, and visual realism, especially showing stronger robustness and adaptability in complex underwater environments, fully verifying its advanced nature and practical value in the field of underwater 3D reconstruction.
[0061] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects disclosed in the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An end-to-end underwater 3D reconstruction method based on an underwater imaging model, characterized in that: The following steps are involved: S1: Collect multiple underwater image sequences to obtain original image data containing complex underwater environment features such as illumination attenuation, scattering blur, and color shift; S2: Build an end-to-end learnable pose estimation network to perform spatiotemporal modeling and feature extraction on the input image sequence, and predict the relative 6DOF pose between image frames; The pose estimation network structure includes: an input module, an image feature extraction module, a fusion encoding module, a Transformer encoding module and a pose regression module; S3: Build an underwater imaging model to simulate the underwater image degradation process caused by the medium and embed it into the 3D rendering process to enhance color reproduction and physical consistency of imaging; S4: Based on the pose estimation network and underwater imaging model, a dense reconstruction module based on 3D Gaussian distribution is constructed to perform 3D reconstruction and output a 3D reconstructed image.
2. The end-to-end underwater 3D reconstruction method according to claim 1, characterized in that: The pose estimation network in S2 includes the following modules: (1) Input module: receives underwater multi-view images as observation input for pose estimation; (2) Image feature extraction module: The high-dimensional spatial features of the input image are extracted through a deep visual encoding network to obtain a feature map containing spatial structure information; the feature map is then flattened and position encoded to form an image token with spatial position information, which facilitates subsequent Transformer processing and global context modeling; (3) Fusion encoding module: combines image tokens with learnable camera tokens, which serve as implicit camera pose representations. This module integrates visual information with camera priors, constructs a joint context representation, and provides rich input semantics and geometric information to the Transformer, assisting pose inference. (4) Transformer encoding module: It consists of a multi-layer Transformer, including multi-head self-attention, multi-head cross-attention, feedforward network, residual connection and layer normalization; this module deeply models the joint sequence of image tokens and camera tokens to achieve context-aware estimation of pose; (5) Pose regression module: Based on the joint features encoded by Transformer, the relative rotation and translation parameters corresponding to the image are predicted through a fully connected regression network to achieve high-precision camera pose estimation.
3. The end-to-end underwater 3D reconstruction method according to claim 1, wherein: The S2 also includes standardized data conversion, which specifically includes the following steps: (1) Pose conversion and format normalization: The relative pose sequence predicted by the pose estimation network is converted into the global camera pose of each frame through matrix multiplication, and then the rotation matrix is transformed into a quaternion; (2) Camera intrinsic parameter initialization and normalization: If there is no explicit calibration data, the default parameters are used, and the size and intrinsic parameter format of all images are unified to generate the PINHOLE model parameters corresponding to each image; (3) Synchronous packaging of images and poses: traverse the image sequence and correspond the image file names to the predicted poses one by one; (4) Sparse point cloud extraction and color addition: Read the spatial point position information from the generated initial point cloud, extract the color of each point, and generate a sparse point set with color, write it into a standard format, and use it as the initial scene structure representation; (5) Unified output standard ternary data format: The final output files are cameras.bin, images.txt and points3D.txt, which constitute the sparse reconstruction input format that meets the COLMAP / 3DGS input requirements and are loaded and used by the subsequent differentiable rendering module based on Gaussian representation.
4. The end-to-end underwater 3D reconstruction method according to claim 1, wherein: The underwater imaging model includes a medium modeling module (MLP) that models the underwater light propagation and imaging process in an end-to-end learnable manner: ; : is the pixel color finally observed on channel c∈{R,G,B} in the camera image; : is the true reflection color of the 3D Gaussian point corresponding to the pixel, generated by the 3D reconstruction module; d(x): represents the depth value corresponding to the three-dimensional point, that is, the distance from the point to the center of the camera; : Absorption coefficient, used to simulate the exponential decay of light intensity during propagation; : Scattering coefficient, used to simulate the superposition contribution of water background light to color; : Medium background light color, reflecting the basic offset of water quality on image color; The underwater imaging model takes three-dimensional spatial coordinates (x, y, z) as input and outputs the above parameters through a multi-layer perceptron (MLP) to achieve modeling of light propagation characteristics in different water environments.
5. The end-to-end underwater 3D reconstruction method according to claim 4, characterized in that: The medium modeling module is implemented in the form of a multi-layer perceptron (MLP), which receives a three-dimensional spatial position (x, y, z)(x, y, z)(x, y, z) as input and outputs underwater imaging parameters of the position in each color channel, including: a. Absorption coefficient : Used to describe the energy attenuation of incident light during its propagation in water; b. Scattering coefficient : Used to model the color superposition and blurring caused by backscattering; c. Background light color : Indicates the color of the water itself or the background light intensity.
6. The end-to-end underwater 3D reconstruction method according to claim 4, characterized in that: The specific steps of S4 are as follows: S4-1: Initialize Gaussian point set: Initialize a sparse 3D point set based on the number and pose information of the input images; each point is represented by a learnable 3D Gaussian distribution, including spatial position, covariance, color attributes, and opacity parameters; S4-2: Construct underwater projection mechanism: Use the camera pose and the underwater imaging model parameters to project the 3D Gaussian points onto the image plane and construct a differentiable projection rendering model; S4-3: Gaussian Differentiable Rendering: Based on the 3D Gaussian Differentiable Rendering framework, color, transparency, and coverage are fused in image space to generate a composite image and the following joint loss function is calculated. S4-4: Color consistency loss: Minimize the pixel error between the rendered image and the ground truth image; S4-5: Transparency regularization: constrains point redundancy and improves efficiency; S4-6: Depth consistency loss: constrains the point cloud geometry and improves 3D restoration accuracy; S4-7: Hierarchical optimization strategy: A coarse-to-fine optimization process is used to increase the number of Gaussian points and image resolution in successive rounds, supplemented by a redundant elimination mechanism; S4-8: The final output is a high-fidelity dense point cloud with complete geometric structure, texture and boundary details.
7. An end-to-end underwater 3D reconstruction system based on an underwater imaging model, characterized in that: Includes the following modules: Image data acquisition module: This module is used to collect multiple frames of underwater image sequences and obtain original image data; Pose Estimation Network Module: This module performs spatiotemporal modeling and feature extraction on the input image sequence to predict the relative six-degree-of-freedom pose between image frames; Underwater Imaging Model Module: This module simulates the underwater image degradation process caused by the medium and is embedded into the 3D rendering process to enhance color reproduction and physical consistency of imaging. 3D reconstruction module: performs 3D reconstruction based on dense reconstruction of 3D Gaussian distribution and outputs 3D reconstructed images.
Citation Information
Patent Citations
Camera image quality improvement method based on neural radiation field
CN116957931A
Three-dimensional entity reconstruction method and system based on end-to-end three-dimensional entity reconstruction network
CN119206047A
Underwater scene nerve implicit three-dimensional reconstruction method
CN119863569A
End-to-end monocular visual odometer method fusing space-time semantic information
CN120088332A
Forest and fruit pose estimation method and system suitable for depth information missing scene
CN120451265A
Cited By
Three-dimensional reconstruction method based on pulse camera and electronic equipment
CN121053335A
Underwater target volume measurement method, device and equipment based on three-dimensional reconstruction and medium
CN121353377A
Tower three-dimensional reconstruction method, storage medium and electronic equipment
CN121414990A
Underwater range gating dynamic imaging method based on meta-heuristic attitude estimation
CN121522662A
Forecasting method and system for observing dense space-time physical field based on atmosphere and ocean points
CN121542655A