A method and system for single-view three-dimensional reconstruction of deep-sea polymetallic nodules

Through the combination of multi-view generation model and DreamBooth tool, the difficulty in data acquisition in the three-dimensional reconstruction of deep-sea multi-metal nodules is solved, and efficient and low-cost three-dimensional reconstruction and evaluation are achieved. The generated multi-metal nodules image is in line with the actual characteristics.

CN120088412BActive Publication Date: 2025-07-22SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510569940.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-22
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and economically realize the three-dimensional reconstruction of deep-sea polymetallic nodules through single-view images, which is limited by the difficulty in obtaining underwater image quality and viewing angle parameters, resulting in difficulty in obtaining three-dimensional data and large evaluation errors.

Method used

The multi-view generation model is used as a 2D prior, and 3D Gaussian is optimized through fractional distillation sampling loss, combined with DreamBooth tool for model fine-tuning, and the Unet noise prediction module and generative Gaussian splatter are used to perform three-dimensional reconstruction of multi-metal nodules to avoid dependence on real three-dimensional data.

Benefits of technology

It realizes efficient and low-cost three-dimensional reconstruction of polymetallic nodules under single-view conditions, improves the accuracy and completeness of reconstruction, reduces the cost of data acquisition, and the generated polymetallic nodules images conform to actual characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088412B_ABST
    Figure CN120088412B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical fields of image processing and deep-sea mining, and discloses a method and system for single-view three-dimensional reconstruction of deep-sea polymetallic nodules. Aiming at the problem of difficult acquisition of real three-dimensional data of underwater polymetallic nodules, the present invention uses a multi-view generation model as a 2D prior, optimizes 3D Gaussian through a fractional distillation sampling loss, and does not use any three-dimensional data for training, thereby breaking through the problem of three-dimensional data acquisition and realizing the three-dimensional reconstruction of polymetallic nodules. In addition, the present invention proposes a method of fine-tuning only through images, realizing the functional decoupling of the noise prediction module (Unet) and the conditional control network (multi-view model) based on DreamBooth, and still maintaining the stability of multi-view generation through pre-trained view parameters under the input of an empty condition; avoiding the problem of underwater camera parameter acquisition through a conditional blanking mechanism, enabling the model to complete fine-tuning only relying on single-view images, and reducing the data acquisition cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of image processing and deep-sea mining, and particularly relates to a method and system for single-view three-dimensional reconstruction of deep-sea polymetallic nodules. Figure 3 Background Art

[0002] Deep-sea deposits have rich reserves of polymetallic nodules. Polymetallic nodules are one of the most widely studied and geographically distributed deep-sea mineral resources. The metals highly enriched in these polymetallic nodules play a crucial role in many high-tech, green-tech, emerging-tech, and energy applications. The volume estimation of traditional deep-sea polymetallic nodules mainly relies on equipment such as underwater mining vehicles and underwater buckets to carry out sampling work. However, deep-sea polymetallic nodules are generally located on the seabed thousands of meters deep. Such a volume estimation method that simply relies on sampling has the following defects. The deep-sea sampling operation is not only extremely difficult but also extremely costly. Each sampling requires a large amount of human, material, and time costs. In addition, due to the limited number of sampling points, it is difficult to comprehensively and accurately reflect the distribution state and actual volume of polymetallic nodules in the entire deep-sea area, ultimately resulting in a large error in the evaluation results.

[0003] In view of this, three-dimensional reconstruction of polymetallic nodules based on underwater images, especially single-view images, to evaluate the volume of polymetallic nodules and then achieve abundance estimation has become the most economical and efficient volume evaluation scheme at present. However, there are many problems with underwater images. For example, there are problems such as color distortion and uneven illumination. Coupled with the low shooting resolution of AUVs, it is difficult to extract effective features from the images. At the same time, due to the limitations of the underwater AUV shooting conditions, it is difficult to obtain or evaluate key information such as perspective parameters of the images. In addition, given the high difficulty of underwater mining, it is impossible to obtain the true three-dimensional data of polymetallic nodules corresponding to the underwater images, making many methods based on neural network learning of three-dimensional features ineffective, making the three-dimensional reconstruction task of underwater polymetallic nodules extremely difficult, and making it difficult for previous technologies to achieve volume evaluation of underwater polymetallic nodules based on single-view images.

[0004] With the proposal of 3D Gaussian Splatting (3DGS), 3D reconstruction based on underwater images has become possible. 3DGS mainly relies on data such as multi-view images and poses, accurately stitches and fuses multi-view images in 3D space, thereby constructing an accurate 3D model, which undoubtedly breaks the dependence on real 3D data. At the same time, relying on the powerful image generation ability of the diffusion model, it makes it possible to infer images from other perspectives. The diffusion model generates images that conform to a specific distribution by gradually adding noise to the data distribution and then learning the process of removing noise in reverse. In the reconstruction of underwater polymetallic nodules, this characteristic makes it possible to infer images from other perspectives starting from existing multi-view underwater images. This image perspective inference based on the diffusion model not only helps to make up for the data missing due to perspective limitations during the actual shooting process, but also provides more comprehensive image information for 3DGS, thereby improving the accuracy and integrity of the 3D reconstruction of polymetallic nodules and providing a more solid foundation for subsequent volume estimation and abundance assessment.

[0005] However, there are problems such as color distortion and uneven illumination in deep-sea polymetallic nodule images. Coupled with the low shooting resolution of AUVs, it is difficult to extract effective features from the images. Limited by the underwater AUV shooting conditions, it is difficult to obtain or evaluate key information such as perspective parameters, which are the necessary data information for fine-tuning multi-view generation models such as Zero-1-to-3 and One-2-3-45. Therefore, how to innovate the fine-tuning means of multi-view generation models so that the multi-view generation models can be fine-tuned only through the existing image data and then adapt to the special characteristics of polymetallic nodule images has become an urgent problem to be solved. Summary of the Invention

[0006] Aiming at the problem of difficult acquisition of real 3D data of underwater polymetallic nodules, the present invention proposes a single-view Figure 3 3D reconstruction method for deep-sea polymetallic nodules, which uses a multi-view generation model as a 2D prior, optimizes 3D Gaussian through the fractional diffusion sampling loss, does not use any 3D data for training, breaks through the problem of 3D data acquisition, and realizes the 3D reconstruction of polymetallic nodules.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A single-view Figure 3 3D reconstruction method for deep-sea polymetallic nodules, comprising the following steps:

[0009] Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to segment individual nodules in the polymetallic nodule image, save the segmented nodule images and use them for subsequent model training;

[0010] Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model only with the nodule images segmented in Step 1, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules;

[0011] Among them, the process of fine-tuning the multi-view generation model is as follows:

[0012] Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constrained fine-tuning environment;

[0013] Use the nodule images segmented in Step 1 as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model;

[0014] Step 3. Embed the multi-view generation model trained in Step 2 into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for nodule reconstruction, realizing the three-dimensional reconstruction of polymetallic nodules.

[0015] In addition, based on the above single-view Figure 3 dimensional reconstruction method of deep-sea polymetallic nodules, the present invention also proposes a corresponding single-view Figure 3 dimensional reconstruction system of deep-sea polymetallic nodules, which adopts the following technical solutions:

[0016] A single-view Figure 3 dimensional reconstruction system of deep-sea polymetallic nodules, including the following modules:

[0017] Nodule segmentation module, used to input the polymetallic nodule images, use the pre-trained segmentation large model SAM to segment individual nodules in the polymetallic nodule images, save the segmented nodule images and use them for subsequent model training;

[0018] Model fine-tuning module, used to use the DreamBooth tool to fine-tune the multi-view generation model only with the segmented nodule images, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules;

[0019] Among them, the process of fine-tuning the multi-view generation model is as follows:

[0020] Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constrained fine-tuning environment;

[0021] Use the segmented nodule images as the only fine-tuning data source to train the Unet noise prediction module, and then transfer the trained Unet noise prediction module back to the multi-view generation model;

[0022] And the multi-metal nodule single-view Figure 3 dimensional reconstruction module, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for the reconstruction of nodules, so as to realize the three-dimensional reconstruction of multi-metal nodules.

[0023] In addition, based on the above-mentioned deep-sea multi-metal nodule single-view Figure 3 dimensional reconstruction method, the present invention also proposes a computer device, which includes a memory and one or more processors.

[0024] The memory stores executable code, and when the processor executes the executable code, it is used to implement the steps of the above-mentioned deep-sea multi-metal nodule single-view Figure 3 dimensional reconstruction method.

[0025] In addition, based on the above-mentioned deep-sea multi-metal nodule single-view Figure 3 dimensional reconstruction method, the present invention also proposes a computer-readable storage medium, on which a program is stored. When the program is executed by the processor, it is used to implement the steps of the above-mentioned deep-sea multi-metal nodule single-view Figure 3 dimensional reconstruction method.

[0026] The present invention has the following advantages:

[0027] As described above, the present invention describes a deep-sea multi-metal nodule single-view Figure 3 dimensional reconstruction method and system. Aiming at the problem of difficulty in obtaining real three-dimensional data of underwater multi-metal nodules, the present invention uses a multi-view generation model as a 2D prior, optimizes the 3D Gaussian through the fractional distillation sampling loss, and does not use any three-dimensional data for training, breaking through the problem of three-dimensional data acquisition and realizing the three-dimensional reconstruction of multi-metal nodules. Aiming at the problem of difficulty in fine-tuning the multi-view generation model, the present invention proposes a method of only fine-tuning through images to solve the problem that the multi-view generation model cannot adapt to underwater multi-metal nodule images. Specifically, the present invention proposes a method for fine-tuning a multi-view generation model based on DreamBooth, which subversively transforms the original design goal of DreamBooth, sacrifices the model generalization ability actively in exchange for the improvement of the generation quality of the target object (multi-metal nodules); realizes the functional decoupling of the generation network (Unet noise prediction module) and the conditional control network (view control module), and can still maintain the multi-view generation stability through the pre-trained view parameters under the input of an empty condition; avoids the problem of underwater camera parameter acquisition through the conditional blanking mechanism, enabling the model to complete fine-tuning only relying on single-view images, and significantly reducing the data acquisition cost. Brief Description of the Drawings

[0028] Figure 1 The flowchart of the single-view Figure 3 3D reconstruction method of deep-sea polymetallic nodules in the embodiment of the present invention;

[0029] Figure 2 The flowchart of the fine-tuning scheme of the multi-view generation model based on DreamBooth in the embodiment of the present invention;

[0030] Figure 3 The flowchart of the nodule reconstruction method based on the 3D Gaussian model in the embodiment of the present invention;

[0031] Figure 4 The schematic diagram of the image of the reconstructed nodule in the specific example of the present invention;

[0032] Figure 5 The experimental result comparison chart of the present invention's specific example of the One-2-3-45 and SF3D and other current advanced single-image three-dimensional reconstruction large models, the basic model before fine-tuning by the method of the present invention, and the model after fine-tuning by the method of the present invention. Detailed Embodiments

[0033] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0034] Embodiment 1

[0035] Aiming at the problem of difficult acquisition of real 3D data of underwater polymetallic nodules, the present invention proposes a single-view Figure 3 3D reconstruction method of deep-sea polymetallic nodules, which uses a multi-view generation model as a 2D prior, optimizes the 3D Gaussian through the fractional distillation sampling loss, does not use any 3D data for training, breaks through the problem of 3D data acquisition, and realizes the 3D reconstruction of polymetallic nodules.

[0036] As Figure 1 shown, the single-view Figure 3 3D reconstruction method of deep-sea polymetallic nodules in this embodiment includes the following steps:

[0037] Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to segment the polymetallic nodule image, save the segmented nodule image, and use it for subsequent model training.

[0038] Step 1.1. Use the ViT model pre-trained by MAE to process the input polymetallic nodule image, and further process it through 1×1 convolution and 3×3 convolution to output an image embedding feature map with a spatial resolution of 64×64 :

[0039] 。

[0040] Among them is the output feature of ViT, and two convolutions are used for dimensionality reduction and feature fusion, is layer normalization.

[0041] Step 1.2. Given the center point of a single tubercle in the image as a prompt, the coordinates of the center point of a single tubercle in the polymetallic nodule image are converted into a point embedding vector through a pre-trained position encoder 。

[0042] Step 1.3. Mask decoding is implemented using a dual Transformer decoder layer, and the image features in Step 1.1 and the point embedding vector in Step 1.2 are deeply interacted, and the formula is expressed as follows:

[0043] 。

[0044] Finally, the output feature is multiplied element-wise by the weight W, and then a mask is generated through sigmoid activation, and the derived segmentation mask is visualized to achieve the segmentation of a single tubercle.

[0045] Among them, represents self-attention operation, represents cross-attention operation, represents multi-layer perceptron operation, represents the output feature after operation; where W is pre-trained through the segmentation large model SAM.

[0046] Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model only through the tubercle images segmented in Step 1, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules.

[0047] The process of the multi-view generation model generating multi-view images is as follows:

[0048] For a given single RGB image , set the image of the required generated view and the given view image The view difference is , and an image is generated under the camera transformation through a diffusion model

[0049] 。

[0050] Select a pre-trained diffusion model with an image codec, a Unet noise prediction module, and a conditional encoder, and further train it on this basis.

[0051] Set the diffusion time step , and train the model by minimizing the following objective function:

[0052] ; where represents the loss function of minimizing the neural network, is random noise, is the predicted noise of the Unet noise prediction module, is the condition of the multi-view generation model, is the noise map after adding t-step noise to the latent feature map of the image at the p-th perspective, represents the square of the L2 norm, represents the expected value of the random variable .

[0053] After the multi-view generation model is trained, the inference model can generate multi-view images by iteratively denoising a Gaussian noise image conditioned on . Among them, the fine-tuning process of the multi-view generation model is as follows:

[0054] Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture. In the hybrid model architecture, the text condition input is set to an empty string to construct an unconditional constraint fine-tuning environment;

[0055] Use the single nodule image segmented in step 1 as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model. After training the multi-view generation model, the multi-view generation model already has excellent view control ability. However, limited by the natural distribution of images fitted by the multi-view generation model, there are no multi-metal nodule-related images, resulting in its lack of ability to generate multi-metal nodule images. Therefore, the present invention innovatively proposes a fine-tuning method, freezes the parameters of the conditional encoder of the multi-view generation model, migrates the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model, uses DreamBooth for fine-tuning, and sets the condition to an empty character to enhance the ability of the diffusion model to generate multi-metal nodule images.

[0056] The DreamBooth method can achieve optimized generation of specific objects while maintaining the generalization ability of the text-to-image model by establishing a strong correlation mapping between images and text conditions, effectively suppressing the text drift phenomenon (i.e., the problem of generating similar images for different text inputs). However, the conditional binding mechanism relies on the text-image pairing mode and cannot adapt to the "image + camera parameter" composite conditional input paradigm required by the multi-view generation model. At the same time, there are technical obstacles in the precise acquisition of camera parameters in the underwater operation scenario, resulting in a serious lack of conditional data required for the traditional fine-tuning method of the multi-view generation model, that is, fine-tuning based on the "image + camera parameter" composite conditional input paradigm. To address the above technical bottlenecks, the present invention achieves a breakthrough in the multi-view generation ability of polymetallic nodules through the following techniques.

[0057] DreamBooth is a method for fine-tuning a text-to-image model. This method is used to fine-tune the text-to-image model, taking the category text of the object in the image as a condition and strongly correlating it with the generated image to generate the desired image. Without affecting the general graphics generation ability of the large model, this method changes its generation of specific objects to achieve customized text-to-image generation. Its biggest feature is subject-driven. That is, for a specific subject, some images of the subject (and the corresponding class names) are used as inputs, and a fine-tuned text-to-image model is returned. This model encodes a unique identifier pointing to the subject. The identifier and the class to which the subject belongs are implanted into the existing "dictionary" of the diffusion model, which can save the overhead of writing detailed image descriptions. The original method uses a very simple prompt design: "a [identifier] [classnoun]". [identifier] represents the unique identifier, and [class noun] represents the class to which the subject belongs. During the training process, the part of the loss that supervises the model is called the prior preservation loss PPL (prior preservation loss), which preserves the category knowledge. Combining the original training loss with the prior preservation loss gives the new loss function, that is:

[0058] 。

[0059] Among them, represents the input image, represents the random noise, t is the number of random noise steps, represents the noise map after adding t steps of noise to the encoded input image, c is the text condition with a unique identifier, is the text condition without a unique identifier, is the picture generated by the text-to-image model, is the noise map after adding t steps of noise to the encoded picture generated by the text-to-image model, , respectively, the noise predicted by the Unet noise prediction module in the text-to-image model , and the noise in . Through the calculation of the loss function, the parameters of the Unet noise prediction module in the text-to-image model are optimized.

[0060] The present invention uses DreamBooth to fine-tune the multi-view generation model, which is actually the reverse use of the core purpose of DreamBooth. The original purpose of DreamBooth is to avoid text deviation and not affect the generalization ability of the text-to-image model. However, the purpose of the present invention is to enable the multi-view generation model to generate as many multi-views of polymetallic nodules as possible, and the lack of generalization ability has no impact on the present invention. By setting the conditions to be empty, that is, under any conditions, the Unet will be guided to generate polymetallic nodule images, supplemented by the powerful perspective control ability of the multi-view generation model, to improve the model's ability to generate multi-view images of polymetallic nodules.

[0061] As Figure 2 shown, the process of using DreamBooth to fine-tune the multi-view generation model is as follows:

[0062] Step 2.1. Freeze the conditional processing module of the multi-view generation model ( Figure 2 the right dashed box CLIP embedding module and the fully connected layer), the image encoder and the decoder, and implement the parameter freezing strategy to block the interference path of external conditional input.

[0063] Step 2.2. Migrate the core network component of the multi-view generation model, the Unet noise prediction module, from the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid generation architecture.

[0064] Among them, the image encoder, the decoder, and the text conditional encoder of the hybrid generation architecture come from the text-to-image diffusion model, and the Unet noise prediction module of the hybrid generation architecture comes from the multi-view generation model.

[0065] After the hybrid generation architecture is constructed, network training of the hybrid generation architecture is carried out. As Figure 2 shown in the conditional box of the hybrid architecture in , set the text conditional input in DreamBooth to an empty string, and the image input is the nodule image segmented in step 1; this operation constructs a fine-tuning environment without conditional constraints, and at the same time, the input of the polymetallic nodule image forces the Unet network to autonomously learn the inherent feature distribution of the polymetallic nodule image, breaking through the strong dependence of the traditional method on conditional information.

[0066] Step 2.4. After the segmented nodule images in Step 1 are input, the VAE encoder of the hybrid generation architecture is used to achieve the conversion between the pixel space and the latent space, and the polymetallic nodule images are encoded into the latent space;

[0067] After the empty string is input, it is encoded by the text conditional encoder in the hybrid generation architecture.

[0068] Step 2.5. Add random noise to the encoded images , and under the guidance of the empty character conditional encoding, the Unet noise prediction module predicts the added noise, and calculates the loss between the predicted noise and the actually added noise. The calculation formula of the loss is as follows:

[0069] .

[0070] Among them, represents the segmented nodule image, represents the random noise, t is the random noise step, represents the noise map after adding t steps of noise to the encoded segmented nodule image, s is the conditional guidance, that is, the empty string, is the picture generated by the hybrid generation architecture, is the noise map after adding t steps of noise to the encoded picture generated by the hybrid generation architecture, , are respectively the , noises predicted by the Unet noise prediction module in the hybrid generation architecture; represents the expectation calculation of the joint distribution of , represents the expectation calculation of the joint distribution of .

[0071] Through the calculation of the loss function, the parameters of the Unet noise prediction module in the hybrid generation architecture are optimized. Through the calculation of the loss function, the parameters of the Unet noise prediction module in the hybrid generation architecture are optimized.

[0072] Step 2.6. After training is iterated M times, the training stops and the fine-tuning is completed. In this embodiment, M is set to 800 for example.

[0073] Step 2.7. Migrate the fine-tuned Unet noise prediction module back to the original multi-view generation model architecture. At this time, the multi-view generation model already has the ability to generate polymetallic nodule images. At the same time, since the perspective conditional processing module, image encoder, and decoder parameters of the original multi-view generation model are frozen except for the Unet network, it does not affect its perspective control ability.

[0074] Step 3. Embed the multi-view generation model trained in Step 2 into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for the reconstruction of nodules, thereby achieving the three-dimensional reconstruction of polymetallic nodules.

[0075] In the present invention, by jointly optimizing the latent space distribution and view control ability of the multi-view generation model, at the cost of sacrificing the generalization ability of the multi-view generation model (since this model is not used to generate images of other objects, this cost is acceptable), the problem of the lack of polymetallic nodule representation in the natural image distribution of the multi-view generation model is specifically corrected.

[0076] The present invention utilizes the view control parameters pre-trained by the multi-view generation model to guide the generation of geometric consistency of new views; the present invention strengthens the modeling ability of the Unet network for the texture features of polymetallic nodules through empty-condition fine-tuning.

[0077] In the nodule reconstruction method based on the 3D Gaussian model, each Gaussian voxel is fully characterized by a parameter group where the elements are respectively the position center coordinates, scaling factor, rotation quaternion, opacity, and color features of each Gaussian sphere.

[0078] As Figure 3 shown, the nodule reconstruction method based on the 3D Gaussian model is specifically as follows:

[0079] Step 3.1. Initialize the 3D Gaussian in the preset spatial domain. Each Gaussian voxel follows a uniform random distribution in the preset spatial domain, and its initial scaling parameter is set to the unit value and in the zero rotation state.

[0080] Through subsequent iterative optimization processes, dynamically adjust the spatial distribution density of the voxels to achieve a progressive improvement in the reconstruction accuracy.

[0081] Step 3.2. Render it into a two-dimensional image through the volume rendering of the Gaussian splash.

[0082] Take a random view P. At the P view, adopt the segmented accumulation method to sample the ray N times , where represents the position parameters of N sampling points along the ray path, is the starting point of the ray, is the direction vector of the ray, represents the extension distance of the ray in the direction . The pixels of the two-dimensional image are calculated as: where represents the transmittance from the ray starting point to the sampling point represents the sampling point Opacity at, indicating Sampling point Color at.

[0083] Step 3.3. Render the two-dimensional image Encode it through an image encoder to obtain a feature map in the latent space to extract key features and compress the image size, .

[0084] Step 3.4. Add t steps of random noise to the feature map in the latent space , with the formula as follows: .

[0085] .

[0086] Among them, represents the noise map after adding t steps of noise. , where , where Value extracted from the linear variance table, that is: , where T is the maximum number of steps 1000.

[0087] Step 3.5. Predict the added t steps of random noise through the Unet noise prediction module fine-tuned in Step 2 , where , are the default view single view and view difference.

[0088] Step 3.6. Optimize the 3D Gaussian through the score distillation sampling loss , with the formula expressed as follows:

[0089] .

[0090] Among them, represents the score distillation sampling loss, is the weighted function generated by the multi-view generation model, represents taking the mathematical expectation of the joint distribution of the random variable , represents the parameter group of all Gaussian nuclides.

[0091] Step 3.7. Repeat the above Steps 3.2 - 3.6, randomly sample the view P and time step t, and perform iterative optimization.

[0092] Step 3.8. Scale the reconstructed nodules according to the pixel width of the two-dimensional rendering image of the reconstructed nodules and the actual width between the red dots in the polymetallic nodule image (for example, the distance between two red dots in the image is 300 pixels, which corresponds to 120 cm in real space), and then obtain the estimated volume of the reconstructed nodules.

[0093] In addition, extensive experiments are performed to conduct quantitative and qualitative evaluations to verify the effectiveness of the proposed method.

[0094] The single-image 3D reconstruction model uses the nodule images segmented by the nodule classification and segmentation model, including sampled nodule images and underwater nodule images, to fine-tune the large model. Images that do not overlap with the training dataset are selected for testing.

[0095] The experiments of this paper were implemented using Pytorch 1.12.1 and NVIDIA RTX 3090 card on Ubuntu 18.04 server. Figure 3 During fine-tuning of the dimensional reconstruction model, the batch size was set to 4, the initial learning rate was set to 5*10-6, the learning rate remained unchanged during training, and the training stopped when the number of training steps reached 800.

[0096] The present invention compares One-2-3-45 with the current advanced single-image 3D reconstruction large model such as SF3D, the basic model before the fine-tuning method in step 2, and the model after fine-tuning. The experimental results are as follows Figure 5 As shown in the figure, three common shapes of polymetallic nodules (disc-shaped: experimental Figure 1 , 2 , ellipsoidal: experimental Figure 3 , conjoined body: Experimental Figure 4 ) to compare the generation effects. Through the experimental results, it can be intuitively observed that the current advanced single-image 3D reconstruction large models such as One-2-3-45 and SF3D have serious shape distortion in the 3D reconstruction of deep-sea polymetallic nodules, and even generate a two-dimensional flat plate shape, and cannot generate a usable 3D shape. Although the model before the fine-tuning method in step 2 can generate the approximate shape of the nodules, there are a lot of gullies and bumps on its surface, which does not conform to the physical characteristics of polymetallic nodules. At the same time, this defect will cause serious errors in volume estimation; the model after the fine-tuning method in step 2 generates nodules without distortion, and its shape is round and smooth, which conforms to the characteristics of nodules.

[0097] In addition, the volume evaluation of polymetallic nodules was also performed based on the method of the present invention, as shown in the following Table 1

[0098] Table 1

[0099]

[0100] It can be obtained from Table 1 above that according to the pixel width of the reconstructed two-dimensional image of the nodule and the true distance between the red dots in the image (as shown below), the reconstructed nodule is scaled, and then its estimated volume is accurately obtained. Figure 4 as shown

[0101] Example 2

[0102] This Example 2 describes a single-view Figure 3 dimensional reconstruction system for deep-sea polymetallic nodules. This single-view Figure 3 dimensional reconstruction system for deep-sea polymetallic nodules is based on the same inventive concept as the single-view Figure 3 dimensional reconstruction method in Example 1 above. Specifically, the single-view Figure 3 dimensional reconstruction system for deep-sea polymetallic nodules in this example includes the following modules:

[0103] A single-view Figure 3 dimensional reconstruction system for deep-sea polymetallic nodules includes the following modules:

[0104] A nodule segmentation module, which is used to segment individual nodules from the input polymetallic nodule image using the pre-trained segmentation large model SAM, and save the segmented nodule images for subsequent model training;

[0105] A model fine-tuning module, which is used to fine-tune the multi-view generation model only through the segmented nodule images using the DreamBooth tool, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules;

[0106] Among them, the fine-tuning process of the multi-view generation model is as follows:

[0107] Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constraint fine-tuning environment;

[0108] Use the segmented nodule images as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model;

[0109] And a single-view Figure 3 dimensional reconstruction module, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash to reconstruct the nodules, and realize the three-dimensional reconstruction of the deep-sea polymetallic nodules.

[0110] It should be noted that the single-view Figure 3In the three-dimensional reconstruction system, the implementation processes of the functions and roles of each functional module are specifically described in the corresponding steps of the method in the above-mentioned Embodiment 1, and will not be elaborated here.

[0111] Embodiment 3

[0112] This Embodiment 3 describes a computer device, which includes a memory and one or more processors. An executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of the three-dimensional reconstruction method for deep-sea polymetallic nodules in the above-mentioned Embodiment 1. Figure 3 steps.

[0113] In this embodiment, the computer device is any device or apparatus with data processing capabilities, which will not be elaborated here.

[0114] Embodiment 4

[0115] This Embodiment 4 describes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the three-dimensional reconstruction method for deep-sea polymetallic nodules in the above-mentioned Embodiment 1. Figure 3 steps.

[0116] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device.

[0117] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A three-dimensional reconstruction method for a single view of deep-sea polymetallic nodules, characterized in that, It includes the following steps: Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to segment individual nodules in the polymetallic nodule image, save the segmented nodule images and use them for subsequent model training; Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model only with the nodule images segmented in Step 1, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture. Set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constraint fine-tuning environment; Use the nodule images segmented in Step 1 as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model; Step 3. Embed the multi-view generation model trained in Step 2 into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for nodule reconstruction, realizing the three-dimensional reconstruction of polymetallic nodules; In the said Step 3, the nodule reconstruction method based on the 3D Gaussian model is specifically: Step 3.

1. Initialize the 3D Gaussian in the preset spatial domain. Each Gaussian voxel obeys a uniform random distribution in the preset spatial domain, and its initial scaling parameter is set to the unit value and in a zero rotation state; Step 3.

2. Render a two-dimensional image at an arbitrary view angle p through volume rendering with Gaussian splashing Step 3.

3. Render the two-dimensional image Encode it through an image encoder to obtain the feature map z in the latent space P ; Step 3.

4. Add t-step random noise ε to the feature map z in the latent space, and the formula is as follows: P ​ Among them, represents the noise map after adding noise for t steps, where α s = 1 - β s , where β s is the value extracted from the linear variance table, that is: where T is the maximum number of steps; Step 3.

5. Predict the added t-step random noise using the Unet noise prediction module fine-tuned in Step 2 where Δp is the difference between the default perspective single view and the perspective difference; Step 3.

6. Optimize the 3D Gaussian Θ through the fractional diffusion sampling loss, and the formula is expressed as follows: Among them, represents the fractional distillation sampling loss, is a weighted function generated by the multi-view generation model, represents taking the mathematical expectation of the joint distribution of the random variables t, p, ε, and Θ represents the parameter group of all Gaussian nuclides; Step 3.

7. Repeat the above Steps 3.2 - 3.6, randomly sample the view angle P and the time step t, and perform iterative optimization; Step 3.

8. Scale the reconstructed nodule according to the pixel width of the two-dimensional rendered image of the reconstructed nodule and the actual width between the red dots in the polymetallic nodule image, so as to obtain the estimated volume of the reconstructed nodule.

2. The single-view three-dimensional reconstruction method of deep-sea polymetallic nodules according to claim 1, wherein The said Step 1 is specifically: Step 1.

1. Process the input polymetallic nodule image using the ViT model pre-trained with MAE, and further process it through 1×1 convolution and 3×3 convolution to output an image embedding feature map with a spatial resolution of 64×64 F' = Conv 3×3 (LayerNorm(Conv 1×1 (F vit ))); Among which F vit is the output feature of ViT. Two convolutions are used for dimensionality reduction and feature fusion, and LayerNorm is layer normalization; Step 1.

2. Given the center point of a single nodule in the image as a prompt, convert the coordinates of the center point of a single nodule in the polymetallic nodule image into a point embedding vector V through a pre-trained position encoder p ; Step 1.

3. Mask decoding is implemented using a dual Transformer decoder layer, which deeply interacts the image features F' in Step 1.1 with the point embedding vector V in Step 1.2, and is expressed by the following formula: p The depth interaction is as follows: F″ = CrossAttn(F′, MLP(CrossAttn(SelfAttn(V P ), F′))); Finally, multiply the output feature F″ and the weight W element by element through the dynamic linear classifier, then generate a mask through sigmoid activation, and visualize the derived segmentation mask to realize the segmentation of individual nodules; Among them, SelfAttn represents the self-attention operation, CrossAttn represents the cross-attention operation, MLP represents the multi-layer perceptron operation, and F″ represents the output feature after the operation; W is obtained through model pre-training.

3. The single-view three-dimensional reconstruction method of deep-sea polymetallic nodules according to claim 1, wherein The said Step 2 is specifically: Step 2.

1. Freeze the visual condition processing module, image encoder and decoder of the multi-view generation model, implement the parameter freezing strategy, and block the interference path of external condition input; Step 2.

2. Migrate the core network component Unet noise prediction module of the multi-view generation model from the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid generation architecture; Step 2.

3. After the hybrid generation framework is constructed, perform network training on the hybrid generation architecture; set the text condition input in DreamBooth to an empty string, and the image input is the tuberculosis image segmented in Step 1. Step 2.

4. After the segmented tuberculosis image is input, the VAE encoder of the hybrid generation architecture is used to realize the conversion between the pixel space and the latent space, and encode a single tuberculosis image into the latent space. After the empty string is input, the text condition encoder in the hybrid generation architecture encodes it. Step 2.

5. Add random noise ε to the encoded image. The Unet noise prediction module of the hybrid generation architecture predicts the added noise under the guidance of the empty character condition encoding, and calculates the loss between the predicted noise and the actual added noise. Step 2.

6. After M training iterations, the training stops and the fine-tuning is completed. Step 2.

7. Migrate the fine-tuned Unet noise prediction module back to the original multi-view generation model architecture. At this time, the multi-view generation model already has the ability to generate the segmented tuberculosis image.

4. The method for single-view three-dimensional reconstruction of deep-sea polymetallic nodules according to claim 3, characterized in that in Step 2, the image encoder, decoder and text condition encoder of the hybrid generation architecture are from the text-to-image diffusion model, and the Unet noise prediction module of the hybrid generation architecture is from the multi-view generation model. In Step 2.5, the calculation formula of the loss is as follows: Among them, represents the segmented tuberculosis image, ε represents random noise, t is the number of steps of random noise, and z t represents the noise map after adding t steps of noise to the encoded segmented tuberculosis image, s is the conditional guidance, that is, an empty string, and x prior is the image generated by the hybrid generation architecture, is the noise map after adding t steps of noise to the encoded image generated by the hybrid generation architecture, ε φ (z t , t, s), are the noises predicted by the Unet noise prediction module in the hybrid generation architecture for z t and respectively; represents the expectation calculation of the joint distribution of ε and t, represents the expectation calculation of the joint distribution of x prior , ε, and t; Through the calculation of the loss function, optimize the parameters of the Unet noise prediction module in the hybrid generation architecture.

5. The method for single-view three-dimensional reconstruction of deep-sea polymetallic nodules according to claim 1, characterized in that in Step 3, using the fine-tuned multi-view generation model, after inputting the polymetallic nodule image and the camera parameter difference between the desired view and the input image, the desired view image of the polymetallic nodule can be generated.

6. The method for single-view three-dimensional reconstruction of deep-sea polymetallic nodules according to claim 1, characterized in that In step 3, in the tuberculosis reconstruction method based on the 3D Gaussian model, each Gaussian voxel is completely characterized by the parameter set Θ i ={x i , s i , q i , α i , c i}, where the elements x i , s i , q i , α i , and c are the position center coordinates, scaling factor, rotation quaternion, opacity, and color characteristics of each Gaussian sphere, respectively.

7. A deep-sea polymetallic nodule single-view three-dimensional reconstruction system for implementing the deep-sea polymetallic nodule single-view three-dimensional reconstruction method according to claim 1, characterized in that the deep-sea polymetallic nodule single-view three-dimensional reconstruction system includes the following modules: A tuberculosis segmentation module, which is used to perform single tuberculosis segmentation on the input polymetallic nodule image using the pre-trained segmentation large model SAM, save the segmented tuberculosis image and use it for subsequent model training. A model fine-tuning module, which is used to fine-tune the multi-view generation model only through the segmented tuberculosis image using the DreamBooth tool, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules. Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constraint fine-tuning environment. Use the segmented tubercle images as the only fine-tuning data source to train the Unet noise prediction module, and then transfer the trained Unet noise prediction module back to the multi-view generation model; And a multi-metal tubercle single-view three-dimensional reconstruction module, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash to reconstruct the tubercles, realizing the three-dimensional reconstruction of multi-metal tubercles.

8. A computer device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements the steps of the deep-sea multi-metal tubercle single-view three-dimensional reconstruction method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it is used to implement the steps of the deep-sea multi-metal tubercle single-view three-dimensional reconstruction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Ship single-view three-dimensional reconstruction method

    CN118570382A

  • Underwater multi-view three-dimensional reconstruction method

    CN119295645A