Deep sea polymetallic nodule single view three-dimensional reconstruction method and system
By adopting a combination method of multi-view generation model and DreamBooth tool in the field of deep-sea multi-metal nodules, the problem of three-dimensional reconstruction of multi-metal nodules under single-view images is solved, and three-dimensional reconstruction and efficient volume evaluation are achieved without real three-dimensional data.
Patent Information
- Application Number
- CN202510569940.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The three-dimensional reconstruction task of deep-sea polymetallic nodules is difficult, and the prior art is difficult to achieve the volume evaluation of underwater polymetallic nodules based on single-view images, and the multi-view generation model cannot adapt to the special characteristics of underwater polymetallic nodules images.
The multi-view generation model is used as a 2D prior, and 3D Gaussian is optimized through fractional distillation sampling loss. Combined with the DreamBooth tool, the multi-view generation model is fine-tuned only through a single-view image, so that the model can generate multi-metallic nodule multi-view image, thereby realizing three-dimensional reconstruction.
The three-dimensional reconstruction of polymetallic nodules without real three-dimensional data training is achieved, which reduces the cost of data acquisition, improves the accuracy and completeness of reconstruction, and provides a more solid foundation for volume estimation and abundance evaluation.
Smart Images

Figure CN120088412A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and deep-sea mining, and particularly relates to a method and system for three-dimensional reconstruction of single-view deep-sea polymetallic nodules. Figure 3 Background Art
[0002] Deep-sea deposits have rich reserves of polymetallic nodules, which are one of the most widely studied and geographically distributed deep-sea mineral resources. The metals highly enriched in these polymetallic nodules play a crucial role in many high-tech, green-tech, emerging-tech, and energy applications. The volume estimation of traditional deep-sea polymetallic nodules mainly relies on sampling work carried out by equipment such as underwater mining vehicles and underwater buckets. However, deep-sea polymetallic nodules are generally located at the bottom of the sea thousands of meters deep. This simple sampling-dependent volume estimation method has the following defects: the deep-sea sampling operation is not only extremely difficult but also extremely costly, and each sampling requires a large amount of human, material, and time costs; in addition, due to the limited number of sampling points, it is difficult to comprehensively and accurately reflect the distribution state and actual volume of polymetallic nodules in the entire deep-sea area, ultimately resulting in a large error in the evaluation results.
[0003] In view of this, three-dimensional reconstruction of polymetallic nodules based on underwater images, especially single-view images, to evaluate the volume of polymetallic nodules and then achieve abundance estimation has become the most economical and efficient volume evaluation scheme at present. However, there are many problems with underwater images. For example, there are problems such as color distortion and uneven illumination. Coupled with the low shooting resolution of AUVs, it is difficult to extract effective features from the images. At the same time, due to the limitations of the underwater AUV shooting conditions, it is difficult to obtain or evaluate key information such as perspective parameters of the images. In addition, given the high difficulty of underwater mining, it is impossible to obtain the true three-dimensional data of polymetallic nodules corresponding to the underwater images, making many methods based on neural network learning of three-dimensional features ineffective, making the three-dimensional reconstruction task of underwater polymetallic nodules extremely difficult, and making it difficult for previous technologies to achieve the volume evaluation of underwater polymetallic nodules based on single-view images.
[0004] With the proposal of 3D Gaussian Splatting (3DGS), three-dimensional reconstruction based on underwater images has become possible. 3DGS mainly relies on data such as multi-view images and poses, accurately stitches and fuses multi-view images in three-dimensional space to construct an accurate three-dimensional model, which undoubtedly breaks the dependence on real three-dimensional data. At the same time, relying on the powerful image generation ability of the diffusion model, it makes it possible to infer images from other perspectives. The diffusion model generates images that conform to a specific distribution by gradually adding noise to the data distribution and then reversely learning the process of removing noise. In the reconstruction of underwater polymetallic nodules, this characteristic makes it possible to infer images from other perspectives starting from the existing multi-view underwater images. This image perspective inference based on the diffusion model not only helps to make up for the data missing due to perspective limitations in the actual shooting process, but also provides more comprehensive image information for 3DGS, thereby improving the accuracy and integrity of the three-dimensional reconstruction of polymetallic nodules and providing a more solid foundation for subsequent volume estimation and abundance assessment.
[0005] However, there are problems such as color distortion and uneven illumination in deep-sea polymetallic nodule images. Coupled with the low shooting resolution of AUVs, it is difficult to extract effective features from the images. Limited by the underwater AUV shooting conditions, it is difficult to obtain or evaluate key information such as perspective parameters, which are the data information necessary for fine-tuning multi-view generation models such as Zero-1-to-3 and One-2-3-45. Therefore, how to innovate the fine-tuning means of multi-view generation models so that the multi-view generation models can be fine-tuned only through the existing image data and then adapt to the special characteristics of polymetallic nodule images has become an urgent problem to be solved. Summary of the Invention
[0006] Aiming at the problem of difficult acquisition of real three-dimensional data of underwater polymetallic nodules, the present invention proposes a single-view Figure 3 three-dimensional reconstruction method for deep-sea polymetallic nodules, which uses a multi-view generation model as a 2D prior, optimizes 3D Gaussian through fractional distillation sampling loss, does not use any three-dimensional data for training, breaks through the problem of three-dimensional data acquisition, and realizes the three-dimensional reconstruction of polymetallic nodules.
[0007] To achieve the above object, the present invention adopts the following technical solutions: A single-view Figure 3 three-dimensional reconstruction method for deep-sea polymetallic nodules, comprising the following steps: Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to segment individual nodules in the polymetallic nodule image, save the segmented nodule images and use them for subsequent model training; Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model only with the nodule images segmented in Step 1, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the process of fine-tuning the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constrained fine-tuning environment; Use the nodule images segmented in Step 1 as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model; Step 3. Embed the multi-view generation model trained in Step 2 into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for nodule reconstruction, so as to realize the three-dimensional reconstruction of polymetallic nodules.
[0008] In addition, based on the above single-view Figure 3 three-dimensional reconstruction method of deep-sea polymetallic nodules, the present invention also proposes a corresponding single-view Figure 3 three-dimensional reconstruction system of deep-sea polymetallic nodules, which adopts the following technical solutions: A single-view Figure 3 three-dimensional reconstruction system of deep-sea polymetallic nodules, including the following modules: A nodule segmentation module, which is used to segment individual nodules from the input polymetallic nodule images using the pre-trained segmentation large model SAM, save the segmented nodule images and use them for subsequent model training; A model fine-tuning module, which is used to fine-tune the multi-view generation model only with the segmented nodule images using the DreamBooth tool, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the process of fine-tuning the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constrained fine-tuning environment; Use the segmented nodule images as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model; And a single-view Figure 3 three-dimensional reconstruction module, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for nodule reconstruction, so as to realize the three-dimensional reconstruction of polymetallic nodules.
[0009] In addition, in the above-mentioned deep-sea polymetallic nodules Figure 3 Based on the dimensional reconstruction method, the present invention also proposes a computer device, which includes a memory and one or more processors.
[0010] The memory stores executable code, and when the processor executes the executable code, it is used to implement the above-mentioned deep-sea polymetallic nodule single-view Figure 3 Steps of the reconstruction method.
[0011] In addition, in the above-mentioned deep-sea polymetallic nodules Figure 3 Based on the 3D reconstruction method, the present invention also proposes a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned single-view deep-sea polymetallic nodule reconstruction method. Figure 3 Steps of the reconstruction method.
[0012] The present invention has the following advantages: As described above, the present invention relates to a deep-sea polymetallic nodule single-view Figure 3 3D reconstruction method and system. Aiming at the problem of difficulty in acquiring real 3D data of underwater polymetallic nodules, the present invention adopts a multi-view generation model as a 2D prior, optimizes 3D Gaussian by fractional distillation sampling loss, does not use any 3D data training, breaks through the difficulty of 3D data acquisition, and realizes 3D reconstruction of polymetallic nodules. Aiming at the problem of difficulty in fine-tuning the multi-view generation model, the present invention proposes a method of fine-tuning only by images, which is used to solve the problem that the multi-view generation model cannot adapt to underwater polymetallic nodule images. Specifically, the present invention proposes a multi-view generation model fine-tuning method based on DreamBooth, which subverts the original design goal of DreamBooth, and actively sacrifices the generalization ability of the model in exchange for the generation quality improvement of the target object (polymetallic nodule); realizes the functional decoupling of the generation network (Unet noise prediction module) and the conditional control network (view control module), and can still maintain the stability of multi-view generation through pre-trained view parameters under empty conditional input; circumvents the difficulty of underwater camera parameter acquisition through the conditional blanking mechanism, so that the model can complete fine-tuning only by relying on single-view images, significantly reducing the data acquisition cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 The deep-sea polymetallic nodule single-view Figure 3 Flowchart of the reconstruction method; Figure 2 It is a process framework diagram of a multi-view generation model fine-tuning solution based on DreamBooth in an embodiment of the present invention; Figure 3Flowchart of the tuberculosis reconstruction method based on the 3D Gaussian model in the embodiments of the present invention; Figure 4 Schematic diagram of the reconstructed tuberculosis image in the specific example of the present invention; Figure 5 Comparison chart of experimental results of One-2-3-45 and current advanced single-image 3D reconstruction large models such as SF3D, the basic model before fine-tuning by the method of the present invention, and the model after fine-tuning by the method of the present invention in the specific example of the present invention. Detailed implementation manners
[0014] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners: Embodiment 1 Aiming at the problem of difficult acquisition of real 3D data of underwater polymetallic nodules, the present invention proposes a single-view Figure 3 3D reconstruction method for deep-sea polymetallic nodules, which uses a multi-view generation model as a 2D prior, optimizes the 3D Gaussian through the fractional distillation sampling loss, does not use any 3D data for training, breaks through the problem of 3D data acquisition, and realizes the 3D reconstruction of polymetallic nodules.
[0015] As Figure 1 shown, the single-view Figure 3 3D reconstruction method for deep-sea polymetallic nodules in this embodiment includes the following steps: Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to segment the polymetallic nodule image for individual nodules, and save the segmented nodule images for subsequent model training.
[0016] Step 1.1. Use the ViT model pre-trained by MAE to process the input polymetallic nodule image, and further process it through 1×1 convolution and 3×3 convolution to output an image embedding feature map with a spatial resolution of 64×64 : .
[0017] Where is the ViT output feature, and the two convolutions are used for dimensionality reduction and feature fusion, is layer normalization.
[0018] Step 1.2. Given the center point of an individual nodule in the image as a prompt, convert the coordinates of the center point of an individual nodule in the polymetallic nodule image into a point embedding vector through a pre-trained position encoder .
[0019] Step 1.3. Mask decoding is implemented using a double Transformer decoder layer, and the image features in Step 1.1 The deep interaction with the point embedding vector in Step 1.2 is expressed by the following formula: .
[0020] Finally, the output features are multiplied element-wise with the weight W, and then a mask is generated through sigmoid activation, and the derived segmentation mask is visualized to achieve the segmentation of a single tuberculosis nodule.
[0021] Among them, represents the self-attention operation, represents the cross-attention operation, represents the multi-layer perceptron operation, represents the output features after the operation; where W is pre-trained through the segmentation large model SAM.
[0022] Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model only with the tuberculosis images segmented in Step 1, so that the fine-tuned model has the ability to generate multi-view images of multi-metal tuberculosis nodules.
[0023] The process of the multi-view generation model generating multi-view images is as follows: For a given single RGB image , set the image of the required generated view to have a view difference of from the given view image , and generate an image under the camera transformation through a diffusion model .
[0024] Select a diffusion model with an image codec, a Unet noise prediction module, and a conditional encoder, which has been pre-trained, and further train it on this basis.
[0025] Let the diffusion time step , and train the model by minimizing the following objective function: ; in the formula, represents minimizing the loss function of the neural network, is the random noise, is the predicted noise of the Unet noise prediction module, is the condition of the multi-view generation model, is the noise map after adding t-step noise to the latent feature map of the image at the p view, represents the square of the L2 norm, represents the expected value of the random variable .
[0026] After the multi-view generation model is trained, the inference model can generate multi-view images by iteratively denoising Gaussian noise images conditioned on . Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture. In the hybrid model architecture, the text condition input is set to an empty string to construct an unconditional constrained fine-tuning environment; Use the single nodule image segmented in step 1 as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model. After the training of the multi-view generation model, the multi-view generation model already has excellent view control ability. However, limited by the natural distribution of images fitted by the multi-view generation model, there are no polymetallic nodule-related images, resulting in a lack of the ability to generate polymetallic nodule images. Therefore, the present invention innovatively proposes a fine-tuning method, freezing the parameters of the conditional encoder of the multi-view generation model, transplanting the Unet noise prediction module of the multi-view generation model into the text-to-image diffusion model, using DreamBooth for fine-tuning, and setting the condition to an empty character to enhance the ability of the diffusion model to generate polymetallic nodule images.
[0027] The DreamBooth method can achieve the generation optimization of specific objects while maintaining the generalization ability of the text-to-image model by establishing a strong association mapping between images and text conditions, effectively suppressing the text drift phenomenon (i.e., the problem that similar images are generated from different text inputs). However, the conditional binding mechanism depends on the text-image pairing mode and cannot adapt to the "image + camera parameter" composite condition input paradigm required by the multi-view generation model. At the same time, the accurate acquisition of camera parameters in the underwater operation scenario faces technical obstacles, resulting in a serious lack of conditional data required for the traditional fine-tuning method of the multi-view generation model, that is, fine-tuning based on the "image + camera parameter" composite condition input paradigm. To address the above technical bottlenecks, the present invention achieves a breakthrough in the multi-view generation ability of polymetallic nodules through the following techniques.
[0028] DreamBooth is a method for fine-tuning text-to-image models. This method is used to fine-tune text-to-image models, taking the class text of objects in an image as a condition, strongly correlating it with the generated image to generate the desired image. Without affecting the general image generation ability of the large model, this method can change its generation of specific objects to achieve customized text-to-image generation. Its biggest feature is subject-driven. That is, for a specific subject, the subject (and some images with the corresponding class names) are used as inputs, and a fine-tuned text-to-image model is returned. This model encodes a unique identifier pointing to the subject. The identifier and the class to which the subject belongs are implanted into the existing "dictionary" of the diffusion model, which can save the overhead of writing detailed image descriptions. The original method uses a very simple prompt design: "a [identifier] [classnoun]". [identifier] represents the unique identifier, and [class noun] represents the class to which the subject belongs. During the training process, the loss part of the model supervision is called the prior preservation loss (PPL), which preserves the class knowledge. Combining the original training loss with the prior preservation loss gives the new loss function, that is: 。
[0029] Among them, represents the input image, represents the random noise, t is the number of random noise steps, represents the noise map after adding t steps of noise to the encoded input image, c is the text condition with a unique identifier, is the text condition without a unique identifier, is the image generated by the text-to-image model, is the noise map after adding t steps of noise to the encoded image generated by the text-to-image model, 、 respectively represent the 、 noise predicted by the Unet noise prediction module in the text-to-image model. By calculating the loss function, the parameters of the Unet noise prediction module in the text-to-image model are optimized.
[0030] The present invention uses DreamBooth to fine-tune the multi-view generation model, which is actually the reverse use of the core purpose of DreamBooth. The original purpose of DreamBooth is to avoid text drift and not affect the generalization ability of the text-to-image model. However, the purpose of the present invention is to enable the multi-view generation model to generate as many multi-views of polymetallic nodules as possible, and the lack of generalization ability has no impact on the present invention. By setting the conditions to be empty, that is, under any conditions, the Unet will be guided to generate polymetallic nodule images, supplemented by the powerful view control ability of the multi-view generation model, to improve the model's ability to generate multi-view images of polymetallic nodules.
[0031] As Figure 2 shown, the process of using DreamBooth to fine-tune the multi-view generation model is as follows: Step 2.1. Freeze the condition processing module of the multi-view generation model ( Figure 2 the right dotted box CLIP embedding module and the fully connected layer), the image encoder and the decoder, and implement the parameter freezing strategy to block the interference path of external condition input.
[0032] Step 2.2. Migrate the core network component of the multi-view generation model, the Unet noise prediction module, from the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid generation architecture.
[0033] Among them, the image encoder, decoder, and text condition encoder of the hybrid generation architecture come from the text-to-image diffusion model, and the Unet noise prediction module of the hybrid generation architecture comes from the multi-view generation model.
[0034] After the hybrid generation architecture is constructed, network training of the hybrid generation architecture is carried out. As Figure 2 shown in the condition box of the hybrid architecture, set the text condition input in DreamBooth to an empty string, and the image input is the nodule image segmented in step 1; this operation constructs a fine-tuning environment without conditional constraints, and at the same time, the input of the polymetallic nodule image forces the Unet network to autonomously learn the inherent feature distribution of the polymetallic nodule image, breaking through the strong dependence on conditional information in traditional methods.
[0035] After the nodule image segmented in step 1 is input, the pixel space and the latent space are converted through the VAE encoder of the hybrid generation architecture, and the polymetallic nodule image is encoded into the latent space; After the empty string is input, the text condition encoder in the hybrid generation architecture encodes it.
[0036] Step 2.5. Add random noise to the encoded image , and under the guidance of null character conditional encoding, the Unet noise prediction module predicts the added noise, and calculates the loss between the predicted noise and the actually added noise. The calculation formula of the loss is as follows: .
[0037] Among them, represents the segmented tuberculosis image, represents the random noise, t is the number of steps of random noise, represents the noise map after adding t steps of noise to the encoded segmented tuberculosis image, s is the conditional guidance, that is, an empty string, is the picture generated by the hybrid generation architecture, is the noise map after adding t steps of noise to the encoded picture generated by the hybrid generation architecture, , respectively are the , noises predicted by the Unet noise prediction module in the hybrid generation architecture; represents to calculate the expectation of the joint distribution of represents to calculate the expectation of the joint distribution of
[0038] Through the calculation of the loss function, the parameters of the Unet noise prediction module in the hybrid generation architecture are optimized. Through the calculation of the loss function, the parameters of the Unet noise prediction module in the hybrid generation architecture are optimized.
[0039] Step 2.6. After training and iterating M times, the training stops and the fine-tuning is completed. In this embodiment, M is set to 800 for example.
[0040] Step 2.7. Migrate the fine-tuned Unet noise prediction module back to the original multi-view generation model architecture. At this time, the multi-view generation model already has the ability to generate multi-metal tuberculosis images. At the same time, since except for the Unet network, the perspective conditional processing module, image encoder and decoder parameters of the original multi-view generation model are frozen, it does not affect its perspective control ability.
[0041] Step 3. Embed the multi-view generation model trained in Step 2 into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash for the reconstruction of tuberculosis, and realize the three-dimensional reconstruction of multi-metal tuberculosis.
[0042] The present invention jointly optimizes the latent space distribution and perspective control ability of the multi-view generation model, at the cost of sacrificing the generalization ability of the multi-view generation model (since this model is not used to generate images of other objects, this cost is acceptable), and specifically corrects the problem of the lack of multi-metal tuberculosis representation in the natural image distribution of the multi-view generation model.
[0043] The present invention utilizes the view control parameters pre-trained by the multi-view generation model to guide the geometrically consistent generation of new views; the present invention strengthens the modeling ability of the Unet network for the texture features of polymetallic nodules through zero-condition fine-tuning.
[0044] In the tuberculosis reconstruction method based on the 3D Gaussian model, each Gaussian voxel is characterized by a parameter group where the elements are the position center coordinates, scaling factor, rotation quaternion, opacity, and color features of each Gaussian sphere, respectively.
[0045] As Figure 3 shown, the tuberculosis reconstruction method based on the 3D Gaussian model is specifically as follows: Step 3.1. Initialize the 3D Gaussian in the preset spatial domain. Each Gaussian voxel follows a uniform random distribution in the preset spatial domain, and its initial scaling parameter is set to the unit value and in a zero-rotation state.
[0046] Through subsequent iterative optimization processes, dynamically adjust the spatial distribution density of the voxels to achieve a progressive improvement in the reconstruction accuracy.
[0047] Step 3.2. Render it into a two-dimensional image through volume rendering of Gaussian splashing.
[0048] Take a random view P. At the P view, adopt the segmented accumulation method to sample the light ray N times , representing the position parameters of N sampling points along the light ray path, being the starting point of the light ray, the direction vector of the light ray, representing the extension distance of the light ray in the direction . Each pixel of the two-dimensional image is calculated as: . represents the transmittance from the light ray starting point to the sampling point , represents the opacity at the sampling point , and represents the color at the sampling point .
[0049] Step 3.3. Encode the rendered two-dimensional image through an image encoder to obtain the feature map in the latent space to extract key features and compress the image size, .
[0050] Step 3.4. Feed the feature map in the latent space Add t-step random noise , and the formula is as follows: .
[0051] Among them, represents the noise map after adding t-step noise. , where , where The value extracted from the linear variance table, that is: , where T is the maximum number of steps, 1000.
[0052] Step 3.5. Predict the added t-step random noise through the Unet noise prediction module fine-tuned in Step 2 , where , are the default perspective single view and perspective difference.
[0053] Step 3.6. Optimize the 3D Gaussian through the score distillation sampling loss , and the formula is expressed as follows: .
[0054] Among them, represents the score distillation sampling loss, is the weighted function generated by the multi-view generation model, represents taking the mathematical expectation of the joint distribution of the random variable , represents the parameter group of all Gaussian nuclides.
[0055] Step 3.7. Repeat the above Steps 3.2 - 3.6, randomly sample the perspective P and time step t, and perform iterative optimization.
[0056] Step 3.8. Scale the reconstructed nodule according to the pixel width of the two-dimensional rendered image of the reconstructed nodule and the actual width between the red dots in the polymetallic nodule image (for example, a distance of 300 pixels between two red dots in the image corresponds to 120 cm in the real space), so as to obtain the estimated volume of the reconstructed nodule.
[0057] In addition, this method also conducts quantitative and qualitative evaluations through extensive experiments to verify the effectiveness of the proposed method.
[0058] For the single-image 3D reconstruction model, use the nodule images segmented by the nodule classification and segmentation model, including the sampled nodule images and underwater nodule images, to fine-tune the large model. And select images that have no intersection with the training dataset for testing.
[0059] Experiments of the present invention were implemented on an Ubuntu 18.04 server using Pytorch 1.12.1 and an NVIDIA RTX 3090 card. During the fine-tuning of the multi-metal nodule monocular Figure 3 dimensional reconstruction model, the batch size was set to 4, the initial learning rate was set to 5*10-6, and the learning rate remained unchanged during the training process. When the number of training steps reached 800 steps, the training stopped.
[0060] The present invention compared One-2-3-45 with current advanced single-image three-dimensional reconstruction large models such as SF3D, the basic model before the fine-tuning method in step 2, and the fine-tuned model. The experimental results are as Figure 5 shown. For three common-shaped multi-metal nodules (disc-shaped: experiment Figure 1 , 2 , ellipsoidal: experiment Figure 3 , conjoined: experiment Figure 4 ), the generation effects were compared. Through the experimental results, it can be intuitively observed that current advanced single-image three-dimensional reconstruction large models such as One-2-3-45 and SF3D have serious shape distortions in the three-dimensional reconstruction of deep-sea multi-metal nodules, and even generate two-dimensional-like flat shapes, and are unable to generate usable three-dimensional shapes. Although the model before the fine-tuning method in step 2 can generate the general shape of the nodule, there are a large number of grooves and unevenness on its surface, which does not conform to the physical characteristics of multi-metal nodules. At the same time, this defect will cause serious errors in volume estimation; the model after the fine-tuning method in step 2 generates nodules without distortion, and its shape is round and smooth, conforming to the characteristics of nodules.
[0061] In addition, the volume of multi-metal nodules was also evaluated based on the method of the present invention, as shown in Table 1 below Table 1
[0062] From Table 1 above, it can be concluded that the present invention scales the reconstructed nodule according to the pixel width of the two-dimensional image of the reconstructed nodule and the true distance between the red dots in the image (as shown Figure 4 below), and then accurately obtains its estimated volume.
[0063] Example 2 This Example 2 describes a deep-sea multi-metal nodule monocular Figure 3 dimensional reconstruction system, which is based on the same inventive concept as the deep-sea multi-metal nodule monocular Figure 3 dimensional reconstruction method in the above Example 1. Specifically, the deep-sea multi-metal nodule monocular Figure 3 dimensional reconstruction system in this example includes the following modules: Figure 3 A deep-sea multi-metal nodule monocular A deep-sea multi-metal nodule monocularFigure 3 3D reconstruction system, including the following modules: Tubercle segmentation module, which is used to segment individual tubercles from the input polymetallic nodule image using the pre-trained segmentation large model SAM, save the segmented tubercle images and use them for subsequent model training; Model fine-tuning module, which is used to fine-tune the multi-view generation model only through the segmented tubercle images using the DreamBooth tool, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-to-image diffusion model architecture to construct a hybrid model architecture, and set the text condition input of the hybrid model architecture to an empty string to construct an unconditional constraint fine-tuning environment; Use the segmented tubercle images as the only fine-tuning data source to train the Unet noise prediction module, and then migrate the trained Unet noise prediction module back to the multi-view generation model; And the single-view Figure 3 3D reconstruction module, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash to reconstruct the tubercles, so as to realize the three-dimensional reconstruction of polymetallic nodules.
[0064] It should be noted that in the single-view Figure 3 3D reconstruction system of deep-sea polymetallic nodules, the implementation processes of the functions and roles of each functional module are specifically described in the corresponding steps of the method in the above-mentioned Embodiment 1, and will not be elaborated here.
[0065] Embodiment 3 This Embodiment 3 describes a computer device, which includes a memory and one or more processors. An executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of the single-view Figure 3 3D reconstruction method of deep-sea polymetallic nodules in the above-mentioned Embodiment 1.
[0066] In this embodiment, the computer device is any device or apparatus with data processing capabilities, which will not be elaborated here.
[0067] Embodiment 4 This Embodiment 4 describes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the single-view Figure 3 3D reconstruction method of deep-sea polymetallic nodules in the above-mentioned Embodiment 1.
[0068] The computer-readable storage medium may be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device.
[0069] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. A single-view 3D reconstruction method for deep-sea polymetallic nodules, characterized in that: The steps include: Step 1. For the input polymetallic nodule image, use the pre-trained large segmentation model SAM to perform single nodule segmentation on the polymetallic nodule image, save the segmented nodule image and use it for subsequent model training; Step 2. Use the DreamBooth tool to fine-tune the multi-view generation model using only the nodule images segmented in step 1, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the text-based diffusion model architecture to build a hybrid model architecture. Set the text condition input of the hybrid model architecture to an empty string to build an unconditionally constrained fine-tuning environment. The nodule image segmented in step 1 is used as the only fine-tuning data source to train the Unet noise prediction module, and then the trained Unet noise prediction module is migrated back to the multi-view generation model; Step 3. The multi-view generative model trained in step 2 is embedded in the generative Gaussian splash as a 2D prior to guide the 3D Gaussian splash to reconstruct the nodules and achieve 3D reconstruction of polymetallic nodules.
2. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 1, characterized in that: The step 1 is specifically as follows: Step 1.
1. Use the ViT model pre-trained with MAE to process the input polymetallic nodule image and further process it through 1×1 convolution and 3×3 convolution to output an image embedding feature map with a spatial resolution of 64×64. : ; in Output features for ViT, two convolutions are used for dimensionality reduction and feature fusion. Normalize the layer; Step 1.
2. Given the center point of a single nodule in the image as a hint, the coordinates of the center point of a single nodule in the polymetallic nodule image are converted into a point embedding vector through a pre-trained position encoder ; Step 1.
3. Mask decoding is implemented using a double Transformer decoder layer, transforming the image features in step 1.1 The midpoint embedding vector in step 1.2 The deep interaction of , the formula is expressed as follows: ; Finally, the output features are transformed through the dynamic linear classifier Multiply the weight W element-wise, then generate a mask through sigmoid activation, and visualize the derived segmentation mask to achieve the segmentation of a single nodule; in, represents the self-attention operation, represents the cross attention operation, represents the multi-layer perceptron operation, Indicates the output features after operation; W is obtained through model pre-training.
3. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.
1. Freeze the visual condition processing module, image encoder and decoder of the multi-view generation model, implement parameter freezing strategy, and block the interference path of external condition input; Step 2.
2. Migrate the Unet noise prediction module, the core network component of the multi-view generation model, from the multi-view generation model to the Wensheng graph diffusion model architecture to build a hybrid generation architecture; Step 2.
3. After the hybrid generative architecture is constructed, the hybrid generative architecture network is trained; the text condition input in DreamBooth is set to an empty string, and the image input is the tuberculosis image segmented in step 1; Step 2.
4. After the segmented nodule image is input, the VAE encoder with a hybrid generative architecture is used to realize the conversion between the pixel space and the latent space, and the single nodule image is encoded into the latent space; After the empty string is input, the text conditional encoder in the hybrid generation architecture encodes it; Step 2.
5. Add random noise to the encoded image ,The Unet noise prediction module of the hybrid generative architecture predicts the added noise under the guidance of the empty character conditional encoding, and the loss is calculated by comparing the predicted noise with the actual added noise; Step 2.
6. After training iterations M times, the training stops and fine-tuning is completed; Step 2.
7. Migrate the fine-tuned Unet noise prediction module back to the original multi-view generation model architecture. At this time, the multi-view generation model has the ability to generate segmented nodule images.
4. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 3 is characterized in that: In step 2, the image encoder, decoder and text conditional encoder of the hybrid generation architecture are from the Wensheng graph diffusion model, and the Unet noise prediction module of the hybrid generation architecture is from the multi-view generation model; In step 2.5, the loss calculation formula is as follows: ; in, represents the segmented nodule image, represents random noise, t is the number of random noise steps, represents the noise image after the segmented tuberculosis image is encoded and t steps of noise are added. s is the conditional guide, that is, an empty string. is an image generated by a hybrid generation architecture. The noise map after adding t steps of noise to the image generated by the hybrid generation architecture after encoding, , They are predicted by the Unet noise prediction module in the hybrid generation architecture. , Noise in Express The joint distribution of is calculated as expected, Express Calculate the expectation of the joint distribution of By calculating the loss function, the parameters of the Unet noise prediction module in the hybrid generation architecture are optimized.
5. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 1, characterized in that: In step 3, by using the fine-tuned multi-view generation model, after inputting the polymetallic nodule image and the camera parameter difference between the desired viewing angle and the input image, the desired viewing angle image of the polymetallic nodule can be generated.
6. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 1, characterized in that: In the step 3, in the tuberculosis reconstruction method based on the 3D Gaussian model, each Gaussian nuclide body is composed of a parameter group Complete characterization, where the elements They are the position center coordinates, scaling factor, rotation quaternion, opacity and color characteristics of each Gaussian sphere.
7. The single-view 3D reconstruction method for deep-sea polymetallic nodules according to claim 6, characterized in that: In step 3, the tuberculosis reconstruction method based on the 3D Gaussian model is specifically as follows: Step 3.
1. Initialize the 3D Gaussian in the preset spatial domain. Each Gaussian nuclide body obeys a uniform random distribution in the preset spatial domain. Its initial scaling parameter is set to a unit value and is in a zero rotation state. Step 3.
2. Render a 2D image at any viewing angle p using Gaussian splatter volume rendering ; Step 3.
3. Render the 2D image Encode through the image encoder to obtain the feature map in the latent space ; Step 3.
4. Feature map to latent space Add t steps of random noise to , the formula is as follows: ; in, represents the noise map after adding t steps of noise, ,in ,in The values extracted from the linear variance table are: , where T is the maximum number of steps; Step 3.
5. Predict the added t-step random noise through the Unet noise prediction module fine-tuned in step 2 ,in , The default single view and the viewing angle difference; Step 3.
6. Optimize 3D Gaussian via fractional distillation sampling loss , the formula is as follows: ; in, represents the fractional distillation sampling loss, is the weighting function generated by the multi-view generation model, Represents a random variable Find the mathematical expectation of the joint distribution of represents the parameter set of all Gaussian nuclides; Step 3.
7. Repeat the above steps 3.2-3.6, randomly sample the view angle P and time step t, and iterate the optimization; Step 3.
8. Scale the reconstructed nodules according to the pixel width of the two-dimensional rendering image of the reconstructed nodules and the actual width between the red dots in the polymetallic nodule image, thereby obtaining the estimated volume of the reconstructed nodules.
8. A single-view 3D reconstruction system for deep-sea polymetallic nodules, characterized in that: Includes the following modules: The nodule segmentation module is used to perform single nodule segmentation on the input polymetallic nodule image using the pre-trained large segmentation model SAM, and save the segmented nodule image for subsequent model training; A model fine-tuning module is used to use the DreamBooth tool to fine-tune the multi-view generation model only through the segmented nodule images, so that the fine-tuned model has the ability to generate multi-view images of polymetallic nodules; Among them, the fine-tuning process of the multi-view generation model is as follows: Migrate the Unet noise prediction module of the multi-view generation model to the Wenshengtu diffusion model architecture to build a hybrid model architecture. The text condition input of the hybrid model architecture is set to an empty string to build an unconditionally constrained fine-tuning environment. The segmented nodule image is used as the only fine-tuning data source to train the Unet noise prediction module, and then the trained Unet noise prediction module is migrated back to the multi-view generation model; And a single-view 3D reconstruction module for polymetallic nodules, which is used to embed the trained multi-view generation model into the generative Gaussian splash as a 2D prior, guide the 3D Gaussian splash to reconstruct the nodules, and realize the 3D reconstruction of polymetallic nodules.
9. A computer device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, the steps of the single-view three-dimensional reconstruction method of deep-sea polymetallic nodules as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, it is used to implement the steps of the single-view three-dimensional reconstruction method of deep-sea polymetallic nodules as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Three-dimensional content generation method, system and equipment based on three-dimensional Gaussian, and medium
CN118411467A
Deep sea polymetallic nodule resource assessment method based on image processing
CN118505522A
Ship single-view three-dimensional reconstruction method
CN118570382A
Underwater multi-view three-dimensional reconstruction method
CN119295645A
Controllable generation method for text on image based on large diffusion model
CN119850792A
Cited By
Multi-view-angle-based three-dimensional model generation method and system
CN120599157A
Three-dimensional scene target blanking method and system based on 3DGS
CN120612260A
A 3D scene target culling method and system based on 3DGS
CN120612260B