A method for 3D reconstruction of a ship from a single view
Through the three-dimensional ship reconstruction method based on single-view pictures, the joint training of generator and discriminator is used to solve the problems of high data acquisition and processing, high computing resources and time costs, and low model accuracy in the prior art, and efficient and accurate ship reconstruction is achieved.
Patent Information
- Application Number
- CN202410699132.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-05-31
AI Technical Summary
The existing three-dimensional reconstruction technology of ships has problems such as difficult data acquisition and processing, high computing resources and time costs, and low model accuracy, especially in the case of unstable marine environment and variable lighting conditions.
A three-dimensional reconstruction method based on single-view pictures is adopted to realize the three-dimensional reconstruction of the ship through preprocessing of the image data set, simplified representation of the camera pose, and joint training of the generator and discriminator.
This method can realize three-dimensional reconstruction of the ship in the absence of perspective information, reducing time and economic costs, and obtaining more accurate reconstruction results.
Smart Images

Figure CN118570382B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of ship image processing, and particularly relates to a method for single-view Figure 3 3D reconstruction of ships. Background Art
[0002] Digital twin technology is a technology that realizes real-time monitoring, prediction, and optimization by creating an accurate virtual copy of a physical entity in the digital space. With the booming development of China's water transportation industry, ships, as important waterborne transportation tools, play a crucial role in multiple fields such as navigation safety and ship scheduling. The cost of three-dimensional reconstruction of each ship using traditional three-dimensional reconstruction methods is extremely high. Although professional devices such as lidar and RGB-D cameras can assist in obtaining accurate three-dimensional data, their high costs and application limitations in specific environments limit their widespread application in the water transportation field. Therefore, it has become crucial to develop a low-cost method for three-dimensional reconstruction of ships based on single-view images.
[0003] Existing three-dimensional reconstruction algorithms are mainly divided into traditional three-dimensional reconstruction algorithms and deep learning-based three-dimensional reconstruction algorithms. Traditional three-dimensional reconstruction algorithms mainly include the multi-view stereo (MVS) method and the structured light method. The MVS algorithm reconstructs the three-dimensional structure of a scene by analyzing multiple photos taken from different angles and using the parallax principle. This method relies on images from multiple perspectives and estimates depth information by matching corresponding points of the same scene in different images. The structured light method irradiates an object with a projected light pattern (such as stripes or dot matrices) and then reconstructs its three-dimensional shape by capturing the distorted light pattern on the object's surface. The structured light method is often more accurate than MVS, but it requires special lighting equipment and is limited in outdoor environments. In recent years, deep learning methods have been widely applied to the field of three-dimensional reconstruction. However, the existing methods for three-dimensional reconstruction of ships using deep learning methods still have the following defects:
[0004] (1) Difficulty in data acquisition and processing: Three-dimensional reconstruction of ships requires a large amount of high-quality data as a basis, including images taken from different angles and conditions. However, the instability of the actual marine environment (such as waves and climate change) and the variability of lighting conditions will seriously affect the image quality. In addition, for dynamic or distant ships, it is more difficult to obtain high-quality images, and the processing and cleaning of data are also a complex and time-consuming process.
[0005] (2) High computational resources and time costs: 3D reconstruction usually requires a large amount of computational resources and time. Especially when dealing with large ships or entire fleets, high-resolution images and complex algorithms require powerful computing capabilities and storage space. For application scenarios that require quick response, such as maritime emergency rescue, existing technologies are also difficult to meet the real-time requirements.
[0006] (3) Low model accuracy: For ships with complex structures and details, such as hull structures, texture materials, etc., it is very difficult to achieve highly accurate reconstruction through existing technologies. Summary of the Invention
[0007] The purpose of the present invention is to address the problems existing in the prior art and provide a single-view Figure 3 3D reconstruction method for ships, which can reduce the time and economic costs of ship 3D reconstruction.
[0008] The technical solution is as follows:
[0009] A single-view Figure 3 3D reconstruction method for ships, comprising the following steps:
[0010] Step 1: Preprocess the image dataset, obtain prior information with a specific probability distribution based on the image dataset, and then perform noise sampling from the specific probability distribution;
[0011] Step 2: Represent the orientation of the camera only using the horizontal angle and the vertical angle, define the camera on a sphere with a radius of 1 standard unit, connect the sampling point of the camera with the center of the sphere as the orientation d of the camera, and sample the coordinates p of several points on this direction vector to determine the camera pose;
[0012] Step 3: Input the noise sampling generated in Step 1 into the two-dimensional generative adversarial network generator, and reconstruct the hidden layer in the convolutional layer of the two-dimensional generative adversarial network generator into a three-dimensional tri-planar structure through the camera pose in Step 2 as the 3D generator-rendered picture;
[0013] Step 4: Obtain the corresponding masked binary map of the dataset through the semantic mask module of the rendered view, and use it to perform shape constraints on the picture rendered in Step 3;
[0014] Step 5: Use a convolutional neural network as a progressive discriminator to determine the authenticity of the picture after shape constraints in Step 4 and predict the camera pose in Step 2, prompting the generator to generate a more realistic 3D neural radiance field: Where θ is the pose sampled in the direction d during the generation process of the generator, I is the real picture, G represents the generator model, G(z,θ) is the picture generated by the generator, D represents the discriminator model, j gen 、j realis the discriminator's judgment on whether the picture is real or fake, is the predicted value of the pose by the discriminator;
[0015] Step 6: Use a convolutional neural network as an encoder to map the input pictures of the image dataset in a low-dimensional data space for inversion, so as to output a one-dimensional latent code with a length of 512 and a pose information: where E is the encoder model as the inversion module, is the latent code obtained by the encoder encoding, is the predicted value of the pose by the encoder;
[0016] Step 7: Jointly train the 3D generator in Step 3, the discriminator in Step 5, and the encoder in Step 6 to obtain the loss functions of the 3D generator, discriminator, and encoder respectively;
[0017] Step 8: Use the loss function to perform adversarial training on the 3D reconstruction model to obtain the weight file of the model. The trained weight file infers the corresponding 3D model from the single-view picture and performs inference verification on the validation set, and finally realizes the single-view Figure 3 3D reconstruction of the ship.
[0018] Furthermore, the preprocessing of the image dataset in Step 1 includes the following steps:
[0019] Step 11: Extract information from the red, green, and blue color channels of the image dataset, count the number of each channel for each pixel point in the image dataset, and finally obtain a probability distribution. For the image dataset, the probability of a certain color channel color value appearing is:
[0020] Step 12: Sample a color value from the probability distribution to construct a tensor with the same shape as the randomly generated noise, with a shape of (n, 3), where n is the length of the tensor, and each position is assigned a three-channel color combination. Define the SC function to sample n pixel values from the corresponding pixel value indices, denoted as z i = SC n=512 (k i , P i ), where k i respectively represent the pixel value indices of each color channel, and P i represents the probability of the color value of these pixel values appearing. i = red, green, blue, and n takes 512;
[0021] Step 13: Using a weight vector ([[0.299],[0.587],[0.114]]) that reflects the different importance of different colors in perceived brightness, convert the tensor (n, 3) to a one-dimensional representation (n, 1), and the sampled noise is: Normalize z 0 to obtain the noise sample z,
[0022] Furthermore, in Step 2, the camera pose consists of an extrinsic matrix that describes the position and orientation of the camera coordinate system relative to the world coordinate system and an intrinsic matrix that describes the properties of the camera itself, including focal length, principal point position, and pixel size. The extrinsic matrix is represented as a 4x4 matrix, and the intrinsic matrix is represented as a 3x3 matrix, totaling 25 parameters. The pose parameters are simplified to two parameters, the horizontal angle and the vertical angle, where is the vertical angle, θ is the horizontal angle, and x, y, z are the coordinates of the sampling point in the world coordinate system.
[0023] Furthermore, in Step 3, the 3D generator renders the picture through the following steps:
[0024] Step 31: Input the input noise or latent code into the StyleGAN 2D generator, and use the hidden layer in the generator's convolutional layer as the feature map;
[0025] Step 32: Reshape the feature map into three multi-channel planes to form a three-plane feature representation, and then project the camera coordinates p in Step 2 onto the three feature planes;
[0026] Step 33: Use trilinear interpolation to retrieve the corresponding feature vectors F xy (p), F xz (p), F yz (p), and use a lightweight neural decoder to convert them into the estimated density σ and color c at position p:
[0027] Step 34: Render through the volume rendering equation in the camera direction d in Step 2 to obtain the final color of the picture surface, where is the exponential of the negative volume density integral along the ray to distance t, indicating the attenuation experienced by the ray before reaching distance t. r represents the ray, t n and t f represent the near and far boundaries of the ray, respectively.
[0028] Furthermore, in Step 4, the processing steps of the semantic mask module include:
[0029] Step 41: Binarize the input picture;
[0030] Step 42: Set a target separation threshold to separate the foreground and background;
[0031] Step 43: Use contour detection technology to find the longest contour, ensuring that the mask of the main target area is extracted. Where is the image after shape constraint, I gen is the image generated by the generator, m I is the mask corresponding to the input image.
[0032] Furthermore, the loss of the generator in Step 7 is defined as:
[0033]
[0034] Where D j refers to the part of the discriminator's prediction of true or false, D p refers to the part of the discriminator's prediction of pose, f is the softplus function, and L recon is the reconstruction loss between the real image I and the reconstructed image G(E(I));
[0035] The loss of the discriminator is defined as:
[0036]
[0037] The loss of the encoder is defined as:
[0038]
[0039] Where E z and E p respectively refer to the latent code and the pose prediction part output by the encoder.
[0040] Furthermore, in Step 8, two-stage adversarial training is performed on the 3D reconstruction model. In the first stage, the discriminator and the encoder are trained synchronously, and the generator is set to the frozen state. That is, in this stage, the discriminator and the encoder adjust their parameters according to the gradients calculated by the loss function to reduce the error between the model output and the real label. The generator is used for inference without updating its network parameters. In the second stage, the generator enters the training state, and similarly, through backpropagation of errors, the parameters of the generator are adjusted to optimize the quality of the generated data. The discriminator and the encoder are set to the frozen state. In each training epoch, these two stages are executed once, ensuring that each part of the model can be optimized and updated at the appropriate time. This staged parameter update strategy helps to balance the efficiency of each part of the model and avoid interference and conflicts during the training process.
[0041] Beneficial effects:
[0042] 1) The present invention can achieve the 3D reconstruction of ships in the case of missing perspective information, greatly reducing the time and economic costs of ship 3D reconstruction.
[0043] 2) By using a generator, a discriminator, and an encoder to obtain the model weight file for 3D reconstruction, more accurate reconstruction results can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a schematic flowchart of the method of the present invention;
[0045] Figure 2 is a model architecture diagram of the present invention;
[0046] Figure 3 is a schematic diagram of the camera pose of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention:
[0048] As Figure 1 - Figure 2 shown, a single-view Figure 3 3D reconstruction method for ships includes the following steps:
[0049] Step 1: Preprocess the image dataset, obtain prior information with a specific probability distribution based on the image dataset, and then perform noise sampling from the specific probability distribution, including the following steps:
[0050] Step 11: Extract information from the red, green, and blue color channels of the image dataset, count the number of each channel for each pixel point in the image dataset, and finally obtain a probability distribution. For the image dataset, the probability of a certain color value in a color channel is:
[0051] Step 12: Sample a color value from the probability distribution, construct a tensor with a shape consistent with the randomly generated noise, with a shape of (n, 3), where n is the length of the tensor, and each position is assigned a three-channel color combination. Define the SC function to sample n pixel values from the corresponding pixel value indices, denoted as z i = SC n=512 (ki , P i ), where k i represents the pixel value index of each color channel respectively, and P i represents the probability of the color value of these pixel values occurring. i = red, green, blue, and n takes 512;
[0052] Step 13: Use the weight vector ([[0.299], [0.587], [0.114]]) that reflects the different importance of different colors in perceived brightness to convert the tensor (n, 3) into a one-dimensional representation (n, 1). The sampled noise is: For z 0 perform normalization to obtain the noise sample z,
[0053] Step 2: As Figure 3 shown, only use the horizontal angle and vertical angle to represent the orientation of the camera. Define the camera on a sphere with a radius of 1 standard unit. Connect the sampling point of the camera to the center of the sphere as the orientation d of the camera, and sample the coordinates p of several points on this direction vector to determine the camera pose. The camera pose consists of an external parameter matrix that describes the position and orientation of the camera coordinate system relative to the world coordinate system and an internal parameter matrix that describes the properties of the camera itself, including focal length, principal point position, and pixel size. The external parameter matrix is represented as a 4x4 matrix, and the internal parameter matrix is represented as a 3x3 matrix, totaling 25 parameters. Simplify the pose parameters to two parameters, the vertical angle θ and the horizontal angle. x, y, z are the coordinates of the sampling point in the world coordinate system; where is the vertical angle θ, is the horizontal angle, and x, y, z are the coordinates of the sampling point in the world coordinate system;
[0054] Step 3: Input the noise sample generated in Step 1 into the two-dimensional generative adversarial network generator, and reconstruct the hidden layer in the convolutional layer of the two-dimensional generative adversarial network generator into a three-dimensional three-plane structure through the camera pose in Step 2. Render the picture as a three-dimensional generator through the following steps:
[0055] Step 31: Input the input noise or latent code into the StyleGAN two-dimensional generator, and use the hidden layer in the convolutional layer of the generator as the feature map;
[0056] Step 32: Reshape the feature map into three multi-channel planes to form a three-plane feature representation, and then project the camera coordinates p in Step 2 onto the three feature planes;
[0057] Step 33: Use trilinear interpolation to retrieve the corresponding feature vectors F xy (p), F xz (p), F yz(p), which is converted into an estimated density σ and color c at position p using a lightweight neural decoder:
[0058] Step 34: Render along the camera direction d in Step 2 through the volume rendering equation to obtain the final color of the picture surface, where is the exponential of the negative volume density integral along the ray to distance t, representing the attenuation experienced by the ray before reaching distance t. r represents the ray, t n and t f represent the near and far boundaries of the ray respectively;
[0059] Step 4: Obtain the corresponding binary mask image of the dataset through the semantic mask module of the rendered view, and use it to perform shape constraints on the image rendered in Step 3, including the following steps:
[0060] Step 41: Binarize the input image;
[0061] Step 42: Set a target separation threshold to separate the foreground and background;
[0062] Step 43: Use contour detection technology to find the longest contour to ensure that the extracted mask is of the main target area, where is the image after shape constraint, I gen is the image generated by the generator, m I is the mask corresponding to the input image;
[0063] Step 5: Use a convolutional neural network as a progressive discriminator to determine the authenticity of the image after shape constraint in Step 4 and predict the camera pose in Step 2, prompting the generator to generate a more realistic three-dimensional neural radiance field: where θ is the pose sampled in the direction d during the generation process by the generator, I is the real image, G represents the generator model, G(z,θ) is the image generated by the generator, D represents the discriminator model, j gen 、j real is the discriminator's judgment on whether the image is true or false, is the predicted value of the discriminator for the pose;
[0064] Step 6: Use a convolutional neural network as an encoder to map the input image of the image dataset in a low-dimensional data space for inversion to output a one-dimensional latent code with a length of 512 and a pose information: where E is the encoder model as the inversion module, is the latent code obtained by the encoder encoding, is the predicted value of the encoder for the pose;
[0065] Step 7: Jointly train the 3D generator in Step 3, the discriminator in Step 5, and the encoder in Step 6 to obtain the loss functions of the 3D generator, discriminator, and encoder respectively. The loss of the generator is defined as:
[0066]
[0067] where D j refers to the prediction part of the discriminator for real or fake, D p refers to the prediction part of the discriminator for the pose, f is the softplus function, and L recon is the reconstruction loss between the real image I and the reconstructed image G(E(I));
[0068] The loss of the discriminator is defined as:
[0069]
[0070] The loss of the encoder is defined as:
[0071] where E z and E p respectively refer to the latent code and the pose prediction part output by the encoder;
[0072] Step 8: Use the loss function to perform adversarial training on the 3D reconstruction model to obtain the weight file of the model. The trained weight file infers the corresponding 3D model from a single-view image and performs inference verification on the validation set, finally realizing the single-view Figure 3 3D reconstruction of the ship. Perform two-stage adversarial training on the 3D reconstruction model. In the first stage, the discriminator and the encoder are trained synchronously, and the generator is set to the frozen state. That is, in this stage, the discriminator and the encoder adjust their parameters according to the gradients calculated by the loss function to reduce the error between the model output and the real label. The generator is used for inference without updating its network parameters; in the second stage, the generator enters the training state, and similarly, through error backpropagation, the parameters of the generator are adjusted to optimize the quality of the generated data, and the discriminator and the encoder are set to the frozen state; within each training epoch, these two stages are executed once, ensuring that each part of the model can be optimized and updated at the appropriate time. This staged parameter update strategy helps to balance the efficiency of each part of the model and avoid interference and conflicts during the training process.
[0073] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principle and spirit of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for 3D reconstruction of a ship from a single view, characterized by: The following steps are involved: Step 1: Preprocess the image data set, obtain prior information of a specific probability distribution based on the image data set, and then perform noise sampling from the specific probability distribution; Step 2: Only use horizontal angles and vertical angles to represent the direction of the camera. Define the camera on a sphere with a radius of 1 standard unit. Connect the camera's sampling point and the center of the sphere as the camera's direction d. Sample the coordinates p of several points on this direction vector to determine the camera's position and posture. Step 3: The noise samples generated in step 1 are input into the 2D GAN generator, and the hidden layer in the convolutional layer of the 2D GAN generator is reconstructed into a three-dimensional three-plane structure through the camera pose in step 2 as a 3D generator rendering picture; Step 4: Obtain the mask binary image corresponding to the dataset through the semantic mask module of the rendering view, and use it to constrain the shape of the image rendered in step 3; Step 5: Use a convolutional neural network as a progressive discriminator to determine the authenticity of the image after the shape constraint in step 4, and predict the camera pose in step 2, so that the generator can generate a more realistic three-dimensional neural radiation field: Where θ is the pose sampled in direction d during the generator generation process, I is the real image, G represents the generator model, G(z,θ) is the image generated by the generator, D represents the discriminator model, and j gen 、j real is the discriminator's judgment on whether the image is true or false, is the discriminator’s prediction of the pose; Step 6: Use a convolutional neural network as an encoder to map the input image of the image dataset into a low-dimensional data space for inversion to output a one-dimensional latent code of length 512 and a pose information: Where E is the encoder model as the inversion module, is the potential code obtained by the encoder, is the encoder's prediction of the pose; Step 7: jointly train the 3D generator in step 3, the discriminator in step 5, and the encoder in step 6 to obtain loss functions of the 3D generator, the discriminator, and the encoder, respectively; Step 8: Use the loss function to perform adversarial training on the 3D reconstruction model to obtain the model's weight file. The trained weight file is used to infer the corresponding 3D model from the single-view image, and the inference verification is performed on the validation set to finally achieve single-view 3D reconstruction of the ship.
2. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: The preprocessing of the image data set in step 1 includes the following steps: Step 11: Extract information from the red, green and blue color channels of the image data set, count the number of each channel for each pixel in the image data set, and finally get a probability distribution. For the image data set, the probability of a color channel color value appearing is: Step 12: Sample a color value from the probability distribution and construct a tensor with the same shape as the randomly generated noise, with a shape of (n, 3), where n is the length of the tensor. Each position is assigned a three-channel color combination. Define the n pixel values sampled by the SC function from the corresponding pixel value index, represented as z i =SC n=512 (k i ,P i ), where k i Represents the pixel value index of each color channel, P i The probability of occurrence of the color values representing these pixel values, i = red, green, blue, n is 512; Step 13: Using the weight vector ([[0.299], [0.587], [0.114]]) that reflects the different importance of different colors in perceived brightness, the tensor (n, 3) is converted to a one-dimensional representation (n, 1). The noise obtained by sampling is: Normalize z0 to get the noise sample z, 3. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: The camera pose in step 2 is composed of an external parameter matrix that describes the position and direction of the camera coordinate system relative to the world coordinate system and an internal parameter matrix that describes the properties of the camera itself, including focal length, principal point position and pixel size. The external parameter matrix is represented as a 4x4 matrix, and the internal parameter matrix is represented as a 3x3 matrix, with a total of 25 parameters. The pose parameters are simplified to two parameters: horizontal angle and vertical angle. in is the vertical angle, θ is the horizontal angle, and x, y, z are the coordinates of the sampling point in the world coordinate system.
4. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: In step 3, the 3D generator renders the image through the following steps: Step 31: Input the input noise or latent code to the StyleGAN two-dimensional generator, and use the hidden layer in the convolutional layer of the generator as the feature map; Step 32: The feature map is reshaped into three multi-channel planes to form a three-plane feature representation, and then the camera coordinate p in step 2 is projected onto the three feature planes; Step 33: Retrieve the corresponding eigenvector F using trilinear interpolation xy (p), F xz (p), F yz (p), and a lightweight neural decoder converts it into the estimated density σ and color c at position p: Step 34: Passing the Volume Rendering Equation Rendering is performed in the camera direction d in step 2 to obtain the final color of the image surface, where is the exponent of the integral of the negative volume density along the ray to distance t, indicating the attenuation experienced by the ray before reaching distance t, r represents the ray, t n and t f Represents the near and far boundaries of the ray respectively.
5. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: The processing steps of the semantic mask module in step 4 include: Step 41: Binarize the input image; Step 42: Setting a target separation threshold to separate the foreground and background; Step 43: Use contour detection technology to find the longest contour to ensure that the mask of the main target area is extracted. in is the image after shape constraint, I gen is the image generated by the generator, m I is the mask corresponding to the input image.
6. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: The loss of the generator in step 7 is defined as: Where D j Refers to the part of the discriminator that predicts whether it is true or false, D p Refers to the prediction part of the discriminator for the posture, f is the softplus function, L recon is the reconstruction loss between the real image I and the reconstructed image G(E(I)); The loss of the discriminator is defined as: The encoder loss is defined as: Where E z 、E p They refer to the latent code and pose prediction parts of the encoder output respectively.
7. The method for 3D reconstruction of a ship from a single view according to claim 1, characterized in that: In step 8, the three-dimensional reconstruction model is subjected to two-stage adversarial training. In the first stage, the discriminator and the encoder are trained synchronously, and the generator is set to a frozen state, that is, in this stage, the discriminator and the encoder adjust their parameters according to the gradient calculated by the loss function to reduce the error between the model output and the true label, and the generator is used for inference without updating its network parameters; In the second stage, the generator enters the training state, and also adjusts the parameters of the generator through error back propagation to optimize the quality of the generated data. The discriminator and encoder are set to a frozen state; these two stages are executed once in each training cycle.
Citation Information
Patent Citations
Single-view three-dimensional reconstruction system and method based on adversarial training prior learning
CN112489197A
Single-view three-dimensional reconstruction system and method based on semi-supervised learning
CN112489218A