An underwater multi-view three-dimensional reconstruction method
By introducing physical perception-based underwater image enhancement and a neural symbolic distance function (SDF) reconstruction method with hybrid geometric priors, the problems of high equipment cost, system complexity, and insufficient reconstruction accuracy in underwater 3D reconstruction are solved, achieving efficient and accurate underwater 3D reconstruction.
Patent Information
- Application Number
- CN202411024889.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-07-29
AI Technical Summary
Existing underwater 3D reconstruction technologies have many shortcomings, including high equipment costs, system complexity, difficulty in data fusion, strong environmental dependence, limited image enhancement effects, and insufficient reconstruction accuracy. In particular, the application of multi-sensor fusion and deep learning methods in underwater environments has not yielded ideal results.
The Physically Aware Underwater Image Enhancement (PUIE) method is used for image enhancement. It combines a few-shot, multi-view target segmentation strategy to generate a foreground mask and a 3D geometric prior. 3D reconstruction is performed using the Neural Symbolic Distance Function (SDF). The training of the neural network is optimized using a hybrid geometric prior and a joint loss function.
It significantly improves the image quality and accuracy of underwater 3D reconstruction, reduces equipment costs and system complexity, reduces reliance on large amounts of labeled data, and provides an efficient and accurate 3D reconstruction solution.
Smart Images

Figure CN119295645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to three-dimensional reconstruction technology, and in particular to an underwater multi-view three-dimensional reconstruction method. Background Technology
[0002] In underwater environments, acquiring high-quality images and accurate 3D reconstructions is an extremely challenging task due to the unique characteristics of light absorption and scattering. Existing 3D reconstruction technologies mainly rely on multi-sensor fusion and deep learning methods, each with its own advantages and disadvantages, but still have many shortcomings.
[0003] Multi-sensor fusion techniques typically combine sonar and optical sensors to acquire geometric information about underwater objects. For example, sonar sensors can penetrate turbid water, but their low resolution makes it difficult to provide fine geometric details. Optical sensors (such as underwater cameras) can capture high-resolution images, but their image quality is easily affected by light absorption and scattering due to the optical characteristics of the underwater environment. To overcome their respective limitations, researchers have attempted to fuse data from sonar and optical sensors to improve overall reconstruction accuracy.
[0004] However, multi-sensor fusion methods have the following shortcomings:
[0005] 1. High cost and complexity: Multi-sensor systems have high equipment costs and complex system integration and calibration, which increases the difficulty of deployment and maintenance.
[0006] 2. Difficulty in data synchronization and fusion: The different working principles of sonar and optical sensors make it difficult to synchronize data in time and space, which affects the accuracy of fusion.
[0007] 3. High dependence on the environment: Multi-sensor systems are highly dependent on the underwater environment, especially in turbid water and low light conditions, where data quality will significantly decrease.
[0008] In recent years, deep learning technology has made significant progress in computer vision and image processing. In underwater image processing, deep learning models such as convolutional neural networks (CNNs) have been widely applied to tasks such as image enhancement, dehazing, and object detection. For example, some studies have utilized generative adversarial networks (GANs) to enhance underwater images and improve image quality. However, in the field of 3D reconstruction, especially in underwater environments, deep learning technology still faces challenges.
[0009] Specifically, existing deep learning methods have the following shortcomings:
[0010] 1. Limited image enhancement effect: Although deep learning models can improve the quality of underwater images to some extent, the enhancement effect is limited in cases of severe light absorption and scattering.
[0011] 2. Insufficient reconstruction accuracy: When deep learning-based 3D reconstruction methods, such as Neural Radiation Field (NeRF), are applied in underwater environments, the reconstruction results are often not accurate enough and the details are not fully restored due to the lack of effective geometric priors and lighting models.
[0012] 3. High demand for training data: Deep learning models typically require a large amount of labeled data for training, but in an underwater environment, obtaining and labeling high-quality training data is costly and difficult.
[0013] A typical underwater 3D reconstruction technique is based on the implicit representation method of Neural Radiation Field (NeRF). NeRF uses multi-view images to train a neural network to implicitly represent the voxel color and density of the scene. However, directly applying NeRF to underwater scenes yields unsatisfactory results. The NeRF method proposed by Mildenhall et al. performs excellently in aerial environments, but in underwater environments, due to optical characteristics and a lack of geometric priors, the reconstruction results are blurry and inaccurate. The UWSLAM method proposed by Bi et al. combines underwater image enhancement and 3D reconstruction techniques, but still relies on multi-sensor data, resulting in a complex system and limited reconstruction accuracy.
[0014] In summary, existing multi-sensor fusion technologies and deep learning methods still have many shortcomings in underwater 3D reconstruction, especially in terms of equipment cost, system complexity, data fusion, environmental dependence, image enhancement effect, and reconstruction accuracy.
[0015] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0016] The main objective of this invention is to solve the problems existing in the above-mentioned background technology and provide an underwater multi-view three-dimensional reconstruction method.
[0017] To achieve the above objectives, the present invention adopts the following technical solution:
[0018] In a first aspect of the present invention, an underwater multi-view three-dimensional reconstruction method includes the following steps:
[0019] S1. Acquire multi-view underwater images and perform image enhancement using the Physically Aware Underwater Image Enhancement (PUIE) method.
[0020] S2. Initialize the neural symbolic distance function SDF. The initialized SDF is represented by a neural network that parameterizes an implicit three-dimensional field to describe the geometry of the object.
[0021] S3. Based on image enhancement, a foreground mask is generated through a few-sample, multi-view target segmentation strategy as a two-dimensional geometric prior, and three-dimensional geometric information is obtained through monocular depth estimation and normal vector prediction as a three-dimensional geometric prior.
[0022] S4. Combining image enhancement results and geometric prior information, perform neural symbolic distance function (SDF) training. Specifically, a joint loss function is constructed using a hybrid geometric prior of two-dimensional foreground mask and three-dimensional depth and normal vector information. The enhanced multi-view images and their corresponding geometric prior information are used to train the neural network to minimize the loss function between reconstruction error and geometric prior constraints. The trained neural network is then used for three-dimensional reconstruction.
[0023] In a second aspect of the invention, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the underwater multi-view three-dimensional reconstruction method.
[0024] In a third aspect of the invention, a computer program product includes a computer program that, when executed by a processor, implements the underwater multi-view three-dimensional reconstruction method.
[0025] The present invention has the following beneficial effects:
[0026] This invention provides a more efficient and accurate underwater 3D reconstruction solution by introducing hybrid geometric priors and optimizing the reconstruction process using the symbolic distance function (SDF).
[0027] This invention's underwater multi-view 3D reconstruction method significantly improves the quality and efficiency of underwater environment 3D reconstruction through a series of innovative technical steps. First, it employs the Physically Based Image Enhancement (PUIE) method to enhance underwater images, effectively improving image clarity and contrast and resolving color distortion and blurring issues caused by complex underwater lighting conditions. Second, it generates a high-quality foreground mask as a 2D geometric prior through a few-sample, multi-view target segmentation strategy. This mask, combined with 3D geometric information obtained from monocular depth estimation and normal vector prediction, provides rich geometric constraints for neural network training. Furthermore, this invention introduces a hybrid geometric prior and optimizes the reconstruction process of the symbolic distance function (SDF) through a joint loss function, significantly improving reconstruction accuracy and detail recovery. Compared to traditional multi-sensor fusion technologies and deep learning methods, this invention reduces equipment costs and system complexity, decreases reliance on large amounts of labeled data, and increases automation, providing an efficient and accurate 3D reconstruction solution for fields such as marine scientific research, underwater engineering maintenance, underwater archaeology, and autonomous underwater vehicles.
[0028] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the framework of the underwater multi-view three-dimensional reconstruction method according to an embodiment of the present invention. Detailed Implementation
[0030] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.
[0031] This invention optimizes the reconstruction process of the Neural Symbolic Distance Function (SDF) by combining hybrid geometric prior information. The Physically Aware Underwater Image Enhancement (PUIE) method estimates illumination attenuation and scattering parameters based on a physical model for image enhancement. The SDF reconstruction technique with joint geometric priors constructs a joint loss function by combining 2D foreground masks and 3D geometric information to optimize neural network training. A few-shot, multi-view target segmentation strategy achieves automatic multi-view target segmentation based on a limited number of annotations through data augmentation and transfer learning.
[0032] The method of this invention mainly includes: Image acquisition: acquiring underwater environment images using a multi-view optical camera. Image enhancement: enhancing the images using the PUIE method to improve clarity and contrast. Specifically, this includes: Illumination distribution estimation: estimating illumination attenuation and scattering parameters in the image using a physical model. Color distortion correction: correcting color distortion in the image based on the estimated parameters. Geometric prior acquisition: Two-dimensional foreground mask: initially labeling a small number of images, training a segmentation model, and automatically segmenting the remaining images. Three-dimensional geometric information: obtaining depth and normal vector information through monocular depth estimation and normal vector prediction. SDF initialization and training: initializing the SDF model and parameterizing it into a neural network. Combining geometric prior information, constructing a joint loss function, optimizing neural network training, and generating high-precision three-dimensional reconstruction results.
[0033] See Figure 1 This invention provides an underwater multi-view 3D reconstruction method, comprising the following steps:
[0034] S1. Image Acquisition and Enhancement: Acquire underwater images from multiple perspectives and use the Physically Aware Underwater Image Enhancement (PUIE) method to enhance the images, providing high-quality image data for subsequent geometric prior information acquisition and 3D reconstruction.
[0035] S2. Initialize Neural SDF: The neural symbolic distance function SDF is initialized through pre-training or randomization. The initialized SDF is represented by a neural network that parameterizes an implicit three-dimensional field to describe the geometry of the object.
[0036] S3. Geometric Prior Acquisition: Based on image enhancement, a foreground mask is generated through a few-sample, multi-view target segmentation strategy as a two-dimensional geometric prior, and three-dimensional geometric information is obtained through monocular depth estimation and normal vector prediction as a three-dimensional geometric prior.
[0037] S4. Neural SDF Reconstruction: The neural symbolic distance function (SDF) is trained by combining the image enhancement results and geometric prior information with a neural network. The trained neural network model is then used for 3D reconstruction. A joint loss function is constructed by using a hybrid geometric prior of 2D foreground mask and 3D depth and normal vector information. The enhanced multi-view images and their corresponding geometric prior information are used to train the neural network to minimize the loss function between reconstruction error and geometric prior constraints.
[0038] In a preferred embodiment, step S1, image enhancement using the PUIE method, specifically includes: Distribution estimation: Modeling the underwater illumination and imaging process using a physical model to estimate the illumination distribution and color distortion in the image; calculating illumination attenuation and scattering parameters by analyzing the image's color histogram and illumination distribution. Consensus process: Correcting and enhancing the image based on the estimated illumination and color distortion parameters; optimizing the image enhancement effect through information consistency among multi-view images to ensure the enhanced image has higher clarity and contrast.
[0039] In a preferred embodiment, step S3 specifically includes the following steps: Initial annotation: Manually annotating images from a small number of selected viewpoints to generate an initial target mask, which serves as few-sample data for training the segmentation model; Segmentation model training: Training the segmentation model on the few-sample data, using data augmentation and transfer learning techniques to improve the model's segmentation capability; Automatic segmentation: Automatically segmenting the remaining viewpoint images to generate a foreground mask, which serves as a two-dimensional geometric prior to constrain the training process of the neural SDF.
[0040] In a preferred embodiment, in step S4, the multi-view images used for SDF modeling are first preprocessed, including image alignment, denoising, and normalization, to ensure the consistency and quality of the input data. Further, in step S4, the foreground mask generated using a few-sample multi-view target segmentation strategy is used as a two-dimensional geometric prior to constrain the training process of the neural network; additional three-dimensional geometric information obtained through monocular depth estimation and normal vector prediction is used to further optimize the SDF training.
[0041] See Figure 1In a preferred embodiment, step S4 specifically includes: extracting depth maps and normal maps from the enhanced multi-view RGB image, these feature maps being used to capture the geometric properties of the object surface; generating a foreground mask using a few-sample multi-view object segmentation strategy to distinguish the object surface from other background elements; combining the two-dimensional foreground mask with the three-dimensional geometric information obtained through monocular depth estimation and normal vector prediction to form a complete geometric prior; learning the SDF representation of the object from the extracted features and geometric prior through a multilayer perceptron (MLP) network; wherein, the definition of the loss function combines the constraints of reconstruction error and geometric prior and is used as the optimization objective during the training process; during training, the MLP network is trained using ray sampling and volumetric rendering techniques to minimize the loss function, thereby learning an accurate SDF representation; through iterative optimization, the network parameters are adjusted until a high-precision three-dimensional reconstruction effect is achieved.
[0042] In some embodiments, data is acquired using a multispectral camera, and image enhancement is performed by combining multispectral information to improve image quality under different spectra.
[0043] In some embodiments, multi-sensor fusion can be performed: combining sonar or lidar data with optical images for 3D reconstruction to further improve reconstruction accuracy.
[0044] In some embodiments, the real-time image enhancement and 3D reconstruction system obtained using the present invention is applied to autonomous underwater vehicles (AUVs) or remotely operated vehicles (ROVs) to achieve real-time environmental perception and navigation.
[0045] In some embodiments, the present invention can be used in augmented reality (AR) applications: applying the reconstructed 3D model to an AR system to assist underwater engineering or archaeological research, improving work efficiency and accuracy.
[0046] The significant advantages of this invention are:
[0047] Significantly Improved Image Quality: The PUIE method, based on a physical model, can more accurately correct color distortion and improve image clarity. High Reconstruction Accuracy: The SDF reconstruction technique combined with geometric priors significantly improves the geometric details and overall accuracy of 3D reconstruction. High Efficiency and Automation: The few-sample, multi-view target segmentation strategy reduces the need for manual annotation and improves the efficiency of foreground mask acquisition. Low System Cost: Relying on a single optical sensor reduces equipment cost and system complexity. Advantages Compared to Existing Solutions: Multi-sensor fusion technologies have high equipment costs and system complexity; this invention reduces cost and system complexity. Deep learning methods require a large amount of labeled data; this invention reduces data requirements through a few-sample strategy.
[0048] The following describes specific embodiments of the present invention.
[0049] This invention proposes a neural symbolic distance function (SDF) reconstruction method based on hybrid geometric priors, which is particularly suitable for 3D reconstruction in underwater environments. The invention proposes a framework called UW-SDF, which optimizes the reconstruction process of the neural symbolic distance function (SDF) by introducing hybrid geometric priors, thereby significantly improving the quality and efficiency of underwater 3D reconstruction.
[0050] The method specifically includes:
[0051] 1. Physically-Aware Underwater Image Enhancement (PUIE)
[0052] Illumination distribution estimation: The underwater illumination and imaging process is modeled using a physical model to estimate the illumination attenuation and scattering parameters in the image.
[0053] Color distortion correction: Based on estimated illumination and color distortion parameters, correct color distortion in an image to improve color fidelity.
[0054] Multi-view information consistency optimization: By analyzing the information consistency between multi-view images, the image enhancement effect is optimized to ensure that the enhanced image has higher clarity and contrast.
[0055] 2. Reconstruction of Neural Symbolic Distance Function (SDF)
[0056] SDF initialization method: The initial SDF is represented by a neural network that parameterizes an implicit three-dimensional field to describe the geometry of the object.
[0057] Design of loss function based on joint geometric prior: By combining two-dimensional foreground mask and three-dimensional geometric information (depth and normal vector), a joint loss function is constructed to optimize the training process of neural network.
[0058] As a concrete example, the joint loss function includes the following RGB reconstruction loss, Eikonal loss, 2D masking loss, and 3D geometric information loss:
[0059]
[0060] RGB reconstruction loss:
[0061]
[0062] The SDF value can be normalized using the Eikonal term. The Eikonal loss function is as follows:
[0063]
[0064] The binary cross-entropyloss (BCELoss) is used to supervise the prediction of the two-dimensional foreground mask. The two-dimensional mask loss function is as follows:
[0065]
[0066] When handling shape radiative blur, depth and normal priors are used to supervise the network. The 3D geometric information loss function is as follows: Depth and surface normals are calculated and supervised using L1 loss and angleL1 loss respectively.
[0067]
[0068] Where x is a batch of points uniformly sampled in the three-dimensional space and near the surface, C represents the RGB image, M represents the mask, D represents the depth image, N represents the normal image, and the symbols above the letters ^ represent the predicted value and - represent the prior.
[0069] Neural network training strategies: Utilize incremental learning and transfer learning techniques to improve the training efficiency and reconstruction accuracy of neural networks.
[0070] 3. Few-sample, multi-view target segmentation strategy
[0071] Initial annotation method: Manually annotate a small number of viewpoint images to generate an initial target mask.
[0072] Data augmentation and transfer learning: Expand the training dataset using data augmentation techniques and improve the generalization ability of the segmentation model through transfer learning.
[0073] Automatic segmentation technology: Using a trained segmentation model, multi-view images are automatically segmented to generate high-quality foreground masks.
[0074] 4. Utilization of mixed geometric priors
[0075] Acquisition and application of two-dimensional geometric priors: The foreground mask generated by the few-sample, multi-view target segmentation strategy is used as a two-dimensional geometric prior to constrain the training process of the neural network.
[0076] Acquisition and application of 3D geometric priors: Obtaining... (The sentence is incomplete and requires more context to translate accurately.) Additional 3D geometric information further optimizes SDF training.
[0077] The geometric prior joint constraint method combines two-dimensional foreground mask and three-dimensional geometric information to jointly constrain the training process of neural networks, thereby improving reconstruction accuracy and detail recovery.
[0078] The underwater 3D reconstruction system architecture of this invention includes: an image acquisition and enhancement module, comprising a multi-view camera system and image enhancement technology based on the PUIE method; a geometric prior acquisition module, comprising a few-sample multi-view target segmentation strategy and 3D geometric information acquisition technology; a neural SDF reconstruction module, comprising SDF initialization, loss function design based on joint geometric priors, and neural network training strategy; and a reconstruction result display module, used to display and analyze the final high-precision 3D reconstruction results.
[0079] This invention enables high-quality 3D reconstruction in underwater environments, providing an efficient and accurate solution and offering technical support for multiple fields such as marine scientific research, underwater engineering maintenance, underwater archaeology, and autonomous underwater vehicles.
[0080] In some specific embodiments, the following main steps are included:
[0081] First, the captured multi-view underwater images are enhanced to improve image quality. We employ a Physically Aware Underwater Image Enhancement (PUIE) method, which consists of two parts: distribution estimation and consensus process. Distribution estimation: A physical model is used to model the underwater illumination and imaging process, estimating the illumination distribution and color distortion in the image. Illumination attenuation and scattering parameters are calculated by analyzing the image's color histogram and illumination distribution. Consensus process: Based on the estimated illumination and color distortion parameters, the image is corrected and enhanced. By ensuring information consistency among multi-view images, the image enhancement effect is optimized, ensuring the enhanced image has higher sharpness and contrast.
[0082] Next, based on the enhanced image, the Neural Symbolic Distance Function (SDF) is trained to represent the object surface using the concept of Neural Radiation Field (NeRF). The specific steps are as follows: SDF Initialization: The initial SDF is represented by a neural network that parameterizes an implicit 3D field to describe the object's geometry. Initial network parameters can be pre-trained or randomly initialized. Image Preprocessing: The multi-view images used for SDF neural network modeling are pre-processed, including image alignment, denoising, and normalization, to ensure the consistency and quality of the input data. Geometric Prior Introduction: During SDF training, 2D and 3D geometric prior information is incorporated for optimization; Obtaining a 2D Foreground Mask: A foreground mask generated using a few-sample multi-view object segmentation strategy is used as a 2D geometric prior to constrain the neural network training process; Obtaining 3D Depth and Normal Vectors: Additional 3D geometric information is obtained through monocular depth estimation and normal vector prediction to further optimize the SDF. of Training. Neural network training: using enhanced multi-view images and their corresponding geometric priors. The goal is to train a neural network by minimizing the loss function between the reconstruction error and the geometric prior constraints, thereby obtaining high-precision 3D reconstruction results.
[0083] To efficiently obtain foreground masks, this invention proposes a few-shot, multi-view target segmentation strategy, with the following steps: Initial annotation: Manual annotation is performed on a small number of viewpoint images to generate initial target masks. These annotations serve as few-shot data for training the segmentation model. Segmentation model training: A general segmentation model (such as SAM, SegmentAnythingModel) is trained on the few-shot data. Data augmentation and transfer learning techniques are used to improve the model's segmentation ability. Automatic segmentation: The remaining multi-view images are automatically segmented to generate foreground masks. These masks are used as two-dimensional geometric priors to constrain the training process of the neural SDF.
[0084] During the training of the neural SDF, a hybrid geometric prior is utilized to fully leverage 2D and 3D geometric prior information to improve reconstruction accuracy and detail recovery: 2D geometric prior: Using foreground masks as constraints, this ensures that the network can correctly identify and segment the foreground and background regions of the object during reconstruction, avoiding background noise interference. 3D geometric prior: Combining monocular depth estimation and normal vector prediction, this provides additional 3D geometric information, helping the network to more accurately recover the geometric details and surface texture of the object.
[0085] Figure 1 This is a technical flowchart of the UW-SDF framework according to an embodiment of the present invention, illustrating the workflow and main technical aspects of the entire system as follows:
[0086] Image acquisition and enhancement: Underwater images were acquired using a multi-view optical camera, and the PUIE method was used for image enhancement.
[0087] Geometric prior acquisition: A foreground mask is generated using a few-sample, multi-view target segmentation strategy, and 3D geometric prior is acquired through monocular depth estimation and normal vector prediction.
[0088] SDF Initialization and Training: Initialize the symbolic distance function (SDF), combine it with geometric prior information to train the neural network, and finally generate high-precision 3D reconstruction results.
[0089] The implementation of the PUIE method includes: Distribution estimation: estimating illumination attenuation and scattering parameters using optical models and image statistical properties. Consensus process: optimizing image enhancement effects based on information consistency among multi-view images.
[0090] The training strategy for neural SDF includes: constructing a joint loss function by combining foreground masks and 3D geometric priors to constrain the training process of the neural network; using incremental learning and transfer learning techniques to improve the training efficiency and reconstruction accuracy of the neural network; and employing a few-shot, multi-view target segmentation strategy: initial labeling and data augmentation: training on a small amount of labeled data and expanding the training dataset using data augmentation techniques; and automatic segmentation: using the trained segmentation model to automatically segment multi-view images and generate high-quality foreground masks.
[0091] Through the above technical solution, the present invention has achieved remarkable results in underwater three-dimensional reconstruction, especially in terms of image quality improvement, reconstruction accuracy and efficiency, it has obvious advantages.
[0092] Compared with existing technologies, this invention significantly improves the quality and efficiency of underwater 3D reconstruction by introducing hybrid geometric priors and optimizing the neural symbolic distance function (SDF) reconstruction process. Specific advantages are as follows:
[0093] 1. Improve image quality
[0094] Physically-aware underwater image enhancement (PUIE) method:
[0095] Distribution estimation: By estimating light attenuation and scattering parameters through a physical model, color distortion in images can be accurately corrected.
[0096] Consensus process: Optimize image enhancement effects by leveraging the consistency of information among multi-view images.
[0097] This method significantly improves the clarity and contrast of underwater images, enabling subsequent 3D reconstruction processes to be based on higher-quality image data.
[0098] 2. High-precision 3D reconstruction
[0099] Neural Symbolic Distance Function (SDF) Reconstruction Technique:
[0100] Combined with geometric priors: The training process of SDF is optimized by introducing two-dimensional foreground masks and three-dimensional geometric prior information (such as depth and normal vectors).
[0101] Neural network training: By utilizing a joint loss function and combining incremental learning and transfer learning techniques, the efficiency of network training and reconstruction accuracy can be improved.
[0102] Compared to traditional multi-sensor fusion technology and existing deep learning methods, this invention achieves higher accuracy in geometric detail recovery and overall reconstruction in underwater environments.
[0103] 3. Efficient target segmentation
[0104] Few-sample, multi-view target segmentation strategy:
[0105] Initial annotation and data augmentation: Manual annotation is performed on a small number of viewpoint images, and data augmentation techniques are combined to expand the training dataset.
[0106] Automatic segmentation: The system is trained using a general segmentation model (SAM) to automatically segment multi-view images and generate high-quality foreground masks.
[0107] This strategy significantly reduces the need for manual annotation, improves the efficiency of foreground mask acquisition, and provides a reliable geometric prior for training neural SDF.
[0108] Key technological advantages:
[0109] 1. Significantly improved image quality: The PUIE method significantly improves the clarity and contrast of underwater images, providing high-quality data input for 3D reconstruction.
[0110] 2. Excellent reconstruction accuracy and detail recovery: Combining SDF reconstruction technology with hybrid geometric priors, the actual... It achieves higher accuracy in geometric detail restoration and overall reconstruction.
[0111] 3. Improved efficiency: The few-sample, multi-view target segmentation strategy reduces the need for manual annotation and improves the efficiency of foreground mask acquisition, thereby accelerating the overall 3D reconstruction process.
[0112] Compared to existing technologies: Existing multi-sensor fusion technologies suffer from high equipment costs, system complexity, and difficulties in data fusion. In contrast, this invention relies on a single optical sensor, reducing costs and system complexity. Existing deep learning methods have limited image enhancement effects, insufficient reconstruction accuracy, and high training data requirements. This invention, by incorporating geometric priors and optimizing neural network training, significantly improves reconstruction accuracy and reduces training data requirements.
[0113] The application areas of this invention include: Underwater exploration: suitable for 3D reconstruction of underwater objects and environments in fields such as marine scientific research and ecological environment monitoring. Underwater engineering: applicable to 3D modeling and analysis of engineering projects such as underwater infrastructure inspection, maintenance, and repair. Underwater archaeology: providing high-precision technical support for the 3D reconstruction of underwater remains and artifacts, assisting in archaeological research and conservation work. Autonomous underwater vehicles (AUVs) and remotely operated vehicles (ROVs): achieving high-precision environmental perception and target reconstruction in autonomous navigation, target recognition, and operation tasks.
[0114] In summary, this invention achieves significant technical results in underwater 3D reconstruction, particularly in terms of image quality improvement, reconstruction accuracy, and efficiency. By introducing hybrid geometric priors and optimizing the SDF reconstruction process, this invention provides an efficient and accurate solution for 3D reconstruction of underwater environments.
[0115] To demonstrate the specific application effects of the present invention, several examples are provided below to illustrate the application and technical advantages of the present invention in different scenarios.
[0116] Example 1: 3D Reconstruction of Underwater Coral Reefs
[0117] Background: In marine ecological research, accurately reconstructing the three-dimensional structure of underwater coral reefs is of great significance for assessing their health status and ecological changes.
[0118] step:
[0119] 1. Image Acquisition: Use a multi-view optical camera to photograph the coral reef from different angles and distances to obtain multiple sets of underwater images.
[0120] 2. Image Enhancement: The acquired images are enhanced using the Physically Based Underwater Image Enhancement (PUIE) method to improve image clarity and contrast.
[0121] 3. Geometric prior acquisition:
[0122] Two-dimensional foreground masking: The foreground region of the coral reef is manually labeled on a small number of viewpoint images to train the segmentation model, and then the remaining images are automatically segmented.
[0123] 3D geometric information: Obtain the 3D geometry of the coral reef through monocular depth estimation and normal vector prediction. Geometric information.
[0124] 4. SDF Initialization and Training: Initialize the Neural Symbolic Distance Function (SDF), combine it with geometric prior information, train the neural network, optimize the loss function, and generate high-precision 3D reconstruction results of the coral reef.
[0125] Results: This invention achieves fine structural restoration and high-precision geometric detail reconstruction in the three-dimensional reconstruction of coral reefs, providing an accurate three-dimensional model that helps ecologists conduct detailed research and evaluation.
[0126] Example 2: Underwater Pipeline Inspection and Maintenance
[0127] Background: The inspection and maintenance of underwater pipelines require high-precision 3D models in order to identify potential damage and carry out accurate repair work.
[0128] step:
[0129] 1. Image acquisition: Using a multi-view camera system mounted on an underwater robot, multiple sets of images of the underwater pipe are captured.
[0130] 2. Image Enhancement: The PUIE method is used to enhance the acquired images, eliminating the effects of light attenuation and scattering, and improving image quality.
[0131] 3. Geometric prior acquisition:
[0132] Two-dimensional foreground mask: Initially label the pipe regions in a small number of images, train the segmentation model, and automatically segment the remaining images.
[0133] 3D geometric information: Using depth estimation techniques, the 3D geometric information of the pipeline is obtained, including depth and normal vector.
[0134] 4. SDF Initialization and Training: Initialize and train the SDF model, optimize it by combining geometric priors, and generate a high-precision 3D model of the pipeline.
[0135] Results: This invention provides a high-precision 3D model for underwater pipeline inspection, which can clearly show the structure and details of the pipeline, helping engineers to accurately locate and repair potential damage, and improving the efficiency and accuracy of maintenance work.
[0136] Example 3: 3D Reconstruction of Underwater Archaeological Sites
[0137] Background: In underwater archaeological research, accurately reconstructing the three-dimensional structure of a site is of great significance for its protection and study.
[0138] step:
[0139] 1. Image Acquisition: Using a multi-view camera system, images of the underwater site were captured from different angles.
[0140] 2. Image Enhancement: The PUIE method is applied to enhance image quality, eliminate the influence of the underwater environment on the image, and improve clarity.
[0141] 3. Geometric prior acquisition:
[0142] Two-dimensional foreground mask: Manually label the archaeological sites in a small number of images, train the segmentation model, and automatically generate the foreground mask for the remaining images.
[0143] Three-dimensional geometric information: The three-dimensional geometric information of the site is obtained through monocular depth estimation technology.
[0144] 4. SDF Initialization and Training: Initialize and train the SDF model, combine geometric priors to optimize the reconstruction process, and generate a high-precision 3D model of the site.
[0145] Results: This invention performs exceptionally well in the 3D reconstruction of underwater archaeological sites, providing detailed 3D models that showcase the complete structure and details of the sites, thus assisting archaeologists in conducting in-depth research and conservation work.
[0146] Example 4: Environmental Perception of Autonomous Underwater Vehicles (AUVs)
[0147] Background: Autonomous underwater vehicles (AUVs) require high-precision environmental perception capabilities to achieve autonomous navigation and target recognition.
[0148] step:
[0149] 1. Image Acquisition: The AUV is equipped with a multi-view camera system to acquire multiple images of the underwater environment during autonomous cruising.
[0150] 2. Image Enhancement: The PUIE method is used to enhance image quality and improve the clarity and contrast of environmental images.
[0151] 3. Geometric prior acquisition:
[0152] Two-dimensional foreground mask: A small number of key target regions are annotated from the AUV viewpoint to train the segmentation model and automatically segment the rest of the image.
[0153] 3D geometric information: Using depth estimation techniques, obtain the 3D geometric information of the environment.
[0154] 4. SDF Initialization and Training: Initialize and train the SDF model, optimize it by combining geometric priors, and generate a high-precision 3D environmental model.
[0155] Results: This invention provides a high-precision three-dimensional environment model for AUV environmental perception, improving the autonomous navigation and target recognition capabilities of AUVs and ensuring efficient operation in complex underwater environments.
[0156] Through the above examples, the present invention has demonstrated its superior technical effects in different scenarios, specifically including:
[0157] 1. High-quality image enhancement: The PUIE method significantly improves the clarity and contrast of underwater images.
[0158] 2. High-precision 3D reconstruction: By combining hybrid geometric priors and optimizing SDF model training, high-precision 3D reconstruction results are generated.
[0159] 3. Efficient target segmentation: The few-sample, multi-view target segmentation strategy reduces the need for manual annotation and improves the efficiency of foreground mask acquisition.
[0160] 4. Wide range of applications: It is applicable to a wide range of fields, including marine ecological research (e.g., 3D reconstruction and monitoring of marine organisms and ecological environment), underwater engineering maintenance (e.g., 3D modeling and inspection of underwater structures such as pipelines and platforms), underwater archaeology (e.g., 3D reconstruction and protection of underwater sites and cultural relics), and autonomous underwater vehicles (e.g., environmental perception, navigation and target recognition), demonstrating its broad application potential and technological advantages.
[0161] These embodiments fully demonstrate the significant technical effects of the present invention in underwater three-dimensional reconstruction, providing effective technical support for the exploration and application of underwater environments.
[0162] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.
[0163] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.
[0164] This invention also provides a processor that executes a computer program, at least performing the methods described above.
[0165] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk storage device or a magnetic tape storage device. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0166] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the various components shown or discussed are interconnected. The coupling, or direct coupling, or communication connection between them can be through some interfaces, devices, or units. Indirect coupling or communication connection can be electrical, mechanical, or other forms.
[0167] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0168] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0169] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0170] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
[0171] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0172] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0173] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0174] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. Regarding the technology to which this invention pertains... For technical personnel in the field, without getting rid of Based on the concept of this invention, several other methods can also be made. Any product that is a substitute or obvious modification, and has the same performance or use, should be considered to fall within the scope of protection of this invention.
Claims
1. An underwater multi-view three-dimensional reconstruction method, characterized in that, Includes the following steps: S1. Acquire multi-view underwater images and perform image enhancement using the Physically Aware Underwater Image Enhancement (PUIE) method. S2. Initialize the neural symbolic distance function SDF. The initialized SDF is represented by a neural network that parameterizes an implicit three-dimensional field to describe the geometry of the object. S3. Based on image enhancement, a foreground mask is generated through a few-sample, multi-view target segmentation strategy as a two-dimensional geometric prior, and three-dimensional geometric information is obtained through monocular depth estimation and normal vector prediction as a three-dimensional geometric prior. S4. Combine the image enhancement results and geometric prior information to train the neural symbolic distance function SDF. Specifically, a joint loss function is constructed using a hybrid geometric prior of two-dimensional foreground mask and three-dimensional depth and normal vector information. The enhanced multi-view images and their corresponding geometric prior information are used to train the neural network to minimize the loss function between reconstruction error and geometric prior constraints. The trained neural network is then used for three-dimensional reconstruction.
2. The underwater multi-view three-dimensional reconstruction method as described in claim 1, characterized in that, In step S1, image enhancement using the PUIE method specifically includes: Distribution estimation: The underwater illumination and imaging process is modeled using a physical model to estimate the illumination distribution and color distortion in the image; by analyzing the color histogram and illumination distribution of the image, the illumination attenuation and scattering parameters are calculated. Consensus process: Based on the estimated illumination and color distortion parameters, the image is corrected and enhanced; the image enhancement effect is optimized by ensuring information consistency among multi-view images.
3. The underwater multi-view three-dimensional reconstruction method as described in any one of claims 1 to 2, characterized in that, In step S2, the initial network parameters of the neural network are initialized either through pre-training or randomization.
4. The underwater multi-view three-dimensional reconstruction method as described in any one of claims 1 to 2, characterized in that, In step S3, the few-sample, multi-view target segmentation strategy specifically includes: Manual annotations are performed on a small number of selected viewpoints to generate an initial target mask. These annotations serve as few-sample data and are used to train the segmentation model. The segmentation model is trained on a small number of samples, and its segmentation ability is improved through data augmentation and transfer learning techniques. The remaining viewpoint images are automatically segmented to generate a foreground mask, which is used as a two-dimensional geometric prior to constrain the training process of the neural symbolic distance function SDF.
5. The underwater multi-view three-dimensional reconstruction method as described in any one of claims 1 to 2, characterized in that, In step S4, the multi-view images used for SDF modeling are first preprocessed, including image alignment, denoising, and normalization.
6. The underwater multi-view three-dimensional reconstruction method as described in any one of claims 1 to 2, characterized in that, In step S4, the foreground mask generated by the few-sample multi-view target segmentation strategy is used as a two-dimensional geometric prior to constrain the training process of the neural network; the training of SDF is further optimized by using the additional three-dimensional geometric information obtained through monocular depth estimation and normal vector prediction.
7. The underwater multi-view three-dimensional reconstruction method as described in any one of claims 1 to 2, characterized in that, Step S4 specifically includes: Depth maps and normal maps are extracted from the enhanced multi-view RGB images to capture the geometric properties of the object's surface; A foreground mask is generated using a few-sample, multi-view target segmentation strategy to distinguish between object surfaces and background elements; By combining the two-dimensional foreground mask with the three-dimensional geometric information obtained through monocular depth estimation and normal vector prediction, a complete geometric prior is formed. The SDF representation of the object is learned from the extracted features and geometric priors through a multilayer perceptron (MLP) network. The loss function is defined by combining the constraints of reconstruction error and geometric prior, and is used as the optimization objective during the training process. During training, the MLP network is trained using ray sampling and volumetric rendering techniques to minimize the loss function, thereby learning an accurate SDF representation. Through iterative optimization, network parameters are adjusted until the desired 3D reconstruction accuracy is achieved.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the underwater multi-view three-dimensional reconstruction method as described in any one of claims 1-7.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the underwater multi-view three-dimensional reconstruction method as described in any one of claims 1-7.
Citation Information
Patent Citations
Systems and methods for end to end scene reconstruction from multiview images
US20210279943A1
Techniques for training a machine learning model to reconstruct different three-dimensional scenes
US20240161404A1