Zero-sample generalization three-dimensional object reconstruction method based on semantic guidance

Through multimodal three-dimensional reconstruction network based on semantic guidance, using the prior knowledge of large language models to generate three-dimensional object structures from single images, solving the problem of difficulty in realizing zero-sample generalization of three-dimensional reconstruction in the existing technology, and achieving efficient and accurate three-dimensional reconstruction effect.

CN120047615AActive Publication Date: 2025-05-27SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510090805.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction methods are difficult to solve the problem of generating three-dimensional object structures in a single image with zero-sample generalization, and usually require high hardware costs or complex algorithms.

Method used

A multimodal three-dimensional reconstruction network is constructed through cross-modal information fusion, and a priori knowledge of large language models is used to guide the generation of three-dimensional structures, realizing the reconstruction from single image to three-dimensional object structure.

Benefits of technology

Without a large amount of data training, a three-dimensional structure can be generated by just a single image, which solves the problems of zero-sample generalization and efficient three-dimensional reconstruction, and improves the efficiency and accuracy of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047615A_ABST
    Figure CN120047615A_ABST
Patent Text Reader

Abstract

The invention relates to a zero-sample generalization three-dimensional object reconstruction method based on semantic guidance, and belongs to the technical field of computer vision and image processing. The method comprises the following steps: utilizing the superiority of a fractional distillation sampling strategy in a single image three-dimensional reconstruction process; cue words are designed, the multi-modal large language model is guided to generate description of the image from coarse granularity to fine granularity, and generation of a three-dimensional result is guided; a multi-modal data alignment strategy is adopted, alignment of semantic and visual modals is achieved, and semantic information is fused into a generated three-dimensional structure. According to the method, the problem of generating a three-dimensional object structure by a single image can be solved through zero sample generalization, experimental verification is carried out on the method in a real data set, and the superiority of the method is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer vision and image processing, and relates to a zero-sample generalized three-dimensional object reconstruction method based on semantic guidance. Background Art

[0002] In recent years, with the rapid development of computing devices and computing power, the field of AI-generated content has achieved remarkable results. Generative models have achieved great success in natural language processing and image generation. Recent developments such as ChatGPT, Wenxinyiyan, kimi, and Doubao have revolutionized many academic and industrial fields. In the field of 3D, with the continuous increase in the amount of 3D data and breakthroughs in other generative fields, 3D generation technology has also made great progress.

[0003] Especially driven by the application of intelligent robots, the requirements for environmental perception technology in this field are also increasing. Two-dimensional visual images alone can no longer meet the perception needs in complex environments, so three-dimensional reconstruction technology is particularly critical in improving the environmental perception ability of intelligent agents. By building accurate three-dimensional models, the system can better understand and perceive the surrounding environment. Three-dimensional vision technology can greatly improve autonomous decision-making, spatial perception and interaction capabilities, enabling the system to respond more intelligently in complex scenarios. Especially in fields that require high-precision perception, such as autonomous driving and medical diagnosis, the value of three-dimensional technology is becoming more and more prominent, laying a solid foundation for the future development of intelligence and automation. By generating realistic three-dimensional models, machines can have a deeper understanding of the complex relationship between spatial structures and objects, providing strong support for the decision-making ability of automated systems.

[0004] However, traditional 3D reconstruction methods usually require high hardware costs or complex algorithms to extract the 3D information of objects. These methods rely on precise image matching and geometric calculations, and are easily affected by factors such as lighting changes, occlusion, and noise. At present, deep learning-based methods have greatly reduced the need for precise pairing of input data. At the same time, multimodal large models based on their rich prior knowledge also provide guidance for the generation of unknown parts of the 3D structure. Therefore, deep learning-based methods and multimodal large models based on their rich prior knowledge have become the main methods for 3D reconstruction. However, existing methods cannot solve the problem of generating 3D object structures from a single image with zero-sample generalization. Summary of the invention

[0005] In view of the above technical deficiencies, the purpose of the present invention is to provide a zero-shot generalized 3D object reconstruction method based on semantic guidance. The invention can use semantic information of different granularities to guide the 3D generation process of the model and complete the task of 3D reconstruction of a single image. A large number of experimental tests on commonly used 3D reconstruction datasets have proved the superiority of the technology.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A zero-shot generalized 3D object reconstruction method based on semantic guidance establishes cross-modal information fusion between semantics and 3D vision, and generates an optimal 3D object structure based on semantic guidance, including the following steps:

[0008] Preprocessing steps: Separate the target object from the background in the image, obtain the target object mask image, extract the depth information of the target object and the normal information of the object surface;

[0009] Model construction steps: construct a 3D implicit structure sub-model based on a two-stage fractional distillation sampling from coarse to fine; introduce a multimodal large language model to generate coarse and fine-grained descriptions of the image and calculate fine-grained similarity; construct a rendering network model for 2D and 3D structures; use the CLIP network for coarse-grained semantic alignment and optimize the generated new perspective; construct a multimodal 3D implicit representation network guided by semantic features through a visual text alignment strategy; construct a model total loss function for back-propagation callback of various network parameters of the multimodal 3D reconstruction network during iterative learning;

[0010] Iterative processing steps: Using a single image as input, the preprocessing results are used to iteratively solve the constructed multimodal 3D reconstruction network, and finally the reconstructed image Mesh is output.

[0011] The two-stage fractional distillation sampling based on coarse to fine includes: constructing a three-dimensional implicit structure sub-model based on NeRF network and DMTet network, initializing implicit neural network to represent three-dimensional scene, sampling spatial points and predicting their color and density, combining volume rendering to obtain multi-view image Novel View, and gradually optimizing the generated realistic image by guiding the quality of the generated multi-view image; the rendering method from three-dimensional structure to multi-view is different in the two stages, the first stage adopts NeRF network, and the second stage adopts DMTet network to achieve refinement; the parameters are updated by back propagation, so as to update the generated three-dimensional structure and obtain the multi-view Novel view of the image rendering result; at the same time, the three-dimensional reconstruction result of the first stage is used as the initialization input of the second stage to accelerate the generation speed of the refinement process.

[0012] Constructing a 3D implicit structure sub-model and using it to generate multi-view images includes the following steps:

[0013] NeRF takes the position (x, y, z) and viewing direction (θ, φ) of each point in three-dimensional space as input, and uses a neural network to predict: the color (r, g, b) and volume density (σ) of the point in a specific direction; the formula is expressed as:

[0014] fΘ(x,y,z,θ,φ)→(r,g,b,σ)

[0015] The neural network f is a multi-layer perceptron (MLP) that implicitly represents the entire 3D scene.

[0016] During initialization, NeRF samples a large number of points within a three-dimensional space, and accumulates and calculates the points on the light through a neural network to obtain the color and density of each sampling point, thereby generating an image pixel value C(r);

[0017]

[0018] Where T(t) is the transmittance, which represents the probability that light is not absorbed before depth t; σ(t) is the volume density; and c(t) is the color value.

[0019] The multimodal large language model specifically includes the following steps:

[0020] Set the prompt word for the LLaVA multimodal large language model to generate a specified description of the image and obtain a coarse-grained description Text c and fine-grained description Text inv ;

[0021] For the generated large segment of fine-grained multi-view invariant text description Text inv Segment the sentences and compare them with the coarse-grained feature Text c The splicing is then performed, and then the similarity is calculated with the input image and the generated new view image respectively; the Softmax method is used to normalize the similarity; since the generated multi-view invariant features are of variable length, the two features (Text inv Components and Text after sentence segmentation c ) and a norm to control the sum of similarities of multi-view invariant features to be equal.

[0022] The similarity between the variable-length multi-view invariant feature description and the image after constrained normalization has the same similarity value under different viewpoints; the loss function is calculated by the following formula:

[0023]

[0024] Among them, the Softmax method is used to normalize the similarity, the Sim operation is used to find the similarity between the image and the multi-view invariant text, r is the view angle of the input image, Ir is the input front view, G θ (v i ) is from v i View rendered by perspective, Text inv Generating fine-grained multi-view invariant text information.

[0025] The two-dimensional structure is rendered using the standard diffusion model:

[0026] Add Gaussian noise and coarse-grained text information to the 2D image rendering results of the multi-view Novel view Test c Together they are used as the input of the standard diffusion model Stable Diffusion to learn the rendered two-dimensional image results; the output of this step is the loss function L 2D ;

[0027] The loss function can be calculated using the following formula:

[0028]

[0029] ∈,∈ φ ,φ,θ are the process parameters for adding noise and predicting noise, and the value of parameter θ is updated according to the solution results. For the coarse-grained stage, θ is the parameter value of NeRF's MLPs, and for the fine-grained stage, it is the parameter value of SDF, triangle deformation and color field; ∈,∈ φ , φ are the added noise, predicted noise, and prior parameters of the 2D diffusion model, z is the image with added noise, and z t is the result of adding noise to the image after t steps, w(t) is the weight corresponding to the prediction of the tth step, and e is the result of text encoding.

[0030] The three-dimensional structure is rendered using a multi-view diffusion model:

[0031] Input the input image of the multi-view Novel view and the translation and rotation matrix (R, T) of the new view into the multi-view diffusion model Zero-1-to-3 to obtain the corresponding new view image, use 3D rendering to obtain the rendering result of this stage, use the new view image obtained by the diffusion model to constrain the rendering result, and update the 3D model;

[0032]

[0033] where (R, T) is the camera pose passed to multi-view generation, z t It is a noise feature map obtained by adding random Gaussian noise with a time step of t to the latent space feature map of I.

[0034] The new perspective generated by the coarse-grained semantic alignment optimization using the CLIP network, the loss function is calculated by the following formula:

[0035] L clip =|Adapter(E m (I r ))-E t (Text c )|

[0036] Text c To provide a coarse-grained text description of the image, the Adapter is composed of two linear fully connected layers and a ReLU activation function. m and E t The gates represent the image encoder and text encoder of CLIP.

[0037] The semantic feature guidance is achieved through the visual text alignment strategy, including: strong constraint calculation of the two-dimensional rendering result L corresponding to the viewpoint of the input image pre , L d , L n ;

[0038] L pre The bi-norm loss of the rendered original view image and the preprocessed RGB and mask images of the input image;

[0039]

[0040] ⊙ Element-by-element multiplication between corresponding elements of the two matrices, M is the foreground mask obtained by the ray-integrated volume density along each pixel;

[0041] L d The Pearson negative correlation loss between the rendered original perspective depth image and the input depth map;

[0042]

[0043] Where cov(.) represents the covariance and σ(.) represents the standard deviation. r The depth estimation map of the original perspective rendering result, d is the depth map obtained by preprocessing;

[0044] L n To estimate the normal vector for each point using a finite difference in depth, a 2D normal map n is rendered from the normal vector and the following loss is applied:

[0045] L n =‖n-τ(g(n,k))‖

[0046] Where τ(.) is used to stop the gradient operation, g(.) is Gaussian blur, and the size of the convolution kernel k is set to 9×9.

[0047] The total loss function for model training is:

[0048] L total =λ 1 L 2D +λ 2 L 3D +λ 3 L clip +λ 4 L inv +λ 5 L pre +λ 6 L d +λ 7 L n

[0049] Among them, λ 1 -λ 7 is the weight coefficient of the corresponding loss function.

[0050] The present invention has the following beneficial effects and advantages:

[0051] 1. The present invention does not require a large amount of data to train the model, and only requires inputting a single image to train the model to generate a three-dimensional structure.

[0052] 2. The present invention introduces a large model to guide the generation of three-dimensional structure. The large model's rich prior knowledge and ability to understand images provide guidance for three-dimensional reconstruction.

[0053] 3. The present invention solves the problem of three-dimensional reconstruction without images and paired three-dimensional structures as well as under multi-view images. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a flow chart of the method of the present invention;

[0055] Figure 2 It is the overall structural diagram of the invented network model;

[0056] Figure 3 This is an effect diagram of three-dimensional reconstruction according to the method of the present invention. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation method of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the invention, so the present invention is not limited by the specific implementation disclosed below.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used in the specification of the invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The method steps are described below in conjunction with the accompanying drawings.

[0059] See also Figure 1 and Figure 2 The present invention provides a zero-sample generalized three-dimensional object reconstruction method based on semantic guidance. It includes volume rendering and multi-granularity semantic guidance, and specifically includes the following steps:

[0060] S1: Preprocess the input image: separate the target object mask from the background in the image, and use existing methods to estimate the depth information and surface normal information of the object. It is necessary to separate the foreground target object from the background of the image, and obtain the depth information depth and surface normal information normal of the image based on the separated target image.

[0061] S2: Construct a multimodal 3D reconstruction network: (S21) Construct a 3D implicit structure sub-model based on a two-stage fractional distillation sampling from coarse to fine; (S22) Introduce a multimodal large language model to generate coarse and fine-grained descriptions of the image and calculate fine-grained similarity; (S23) Construct 2D and 3D optimization network models; (S24) Simultaneously use the CLIP network for coarse-grained semantic alignment to optimize the generated new perspective; (S25) Use a visual text alignment strategy to realize the construction of a multimodal 3D reconstruction network guided by semantic features (3D implicit representation); (S26) Construct a model total loss function for back-propagation callback of various network parameters of the multimodal 3D reconstruction network during the model training process.

[0062] S21: The multimodal 3D reconstruction network adopts a two-stage fractional distillation sampling (SDS) guided generation strategy. The two stages only differ in the rendering method of the 3D structure to multiple perspectives. The first stage uses the Neural Radiance Field NeRF network, and the second stage uses the DMTet network for refinement. The parameters are updated by back propagation to update the generated 3D structure and obtain the multi-view Novel view of the image rendering result. At the same time, the 3D reconstruction result of the first stage is used as the initialization input of the second stage to speed up the generation speed of the refinement process.

[0063] S21-1: The first stage builds a 3D implicit structure representation: initialize the implicit neural network to represent the 3D scene, sample spatial points and predict their color and density, combine volume rendering to obtain multi-view images, and gradually optimize the generated realistic images by guiding the quality of the generated multi-view images. Specifically, it includes the following steps:

[0064] S21-1a: NeRF takes the position (x, y, z) and viewing direction (θ, φ) of each point in three-dimensional space as input, and predicts through a neural network: (r, g, h): represents the color of the point in a specific direction. Volume density (σ): represents the scene occupancy of the point, which is usually related to the absorption or scattering ability of light. The formula is expressed as:

[0065] fΘ(x,y,z,θ,φ)→(r,g,b,σ)

[0066] The neural network f is usually a multi-layer perceptron (MLP) that implicitly represents the entire 3D scene.

[0067] S21-1b: During initialization, NeRF samples a large number of points in a three-dimensional space. The sampled points are volume rendered through a neural network, and the color and density of each sampled point are calculated to obtain the final image.

[0068] S21-1c: In the coarse-grained generation stage, the NeRF neural network weight parameters are initialized using random initialization. Volume density σ: The initial value is small, indicating that most of the space is empty. The color values ​​(r, g, b) are randomly distributed, and the initial color prediction is meaningless. This means that in the early stages of training, the network's representation of the scene is completely chaotic, and gradually a reasonable scene representation is obtained through optimization. In the fine-grained generation stage, the coarse-grained stage training results are used as the initialization of the fine-grained rendering network.

[0069] S21-1d: NeRF generates image pixel values ​​using the volume rendering formula by accumulating the points on the ray, where:

[0070]

[0071] T(t) is the transmittance, which indicates the probability that light is not absorbed before depth t; σ(t) is the volume density; and c(t) is the color value.

[0072] S21-1e: The training goal is to make the multi-view images rendered by the network closer to the real objects by optimizing the loss function. The loss function is the constraints of semantics and vision under two-dimensional and three-dimensional conditions.

[0073] S21-2: In the second stage, the DMTet network is used for refinement. The DMTet is initialized using the results obtained in the NeRF stage, and rendered in the same way as NeRF. The loss function is used to constrain the local structure of the three-dimensional structure.

[0074] S22: Constructing multimodal semantic guidance module:

[0075] Using a multimodal large language model to extract coarse-grained semantic features of images c , Generate fine-grained description Text based on the prior of semantic model inv , guiding the generation of three-dimensional structure. Specifically including the following steps:

[0076] S22-1: Multimodal semantic guidance module, design prompt words, guide LLaVA multimodal large language model to generate a specified description of the image, divided into coarse-grained description Text c and fine-grained description Text inv Among them, Text c An example is A photo of a red drum set with a drum, cymbals, and a stand; Text inv An example is The multiviewinvariant features of the object in the image are a drum set with a red drumshell, a snare drum, a bass drum, and a cymbal. The drum set is set against a black background. The material of the object is not specified in the caption.

[0077] S22-2: Generate large segments of fine-grained multi-view invariant text descriptions inv Segment the sentences and compare them with the coarse-grained feature Text c The output of this step is the loss function L inv .

[0078] The Softmax method is used for similarity normalization. Since the generated multi-view invariant features are of variable length, the two features (Text inv Components and Text after sentence segmentation c ) and a norm to control the sum of the similarities of the multi-view invariant features to be equal. The multi-view invariant features with variable length after constraint normalization describe the similarity with the image and have the same similarity value under different viewpoints. The loss function can be calculated by the following formula:

[0079]

[0080] The Softmax method normalizes the similarity. Sim calculates the similarity between the image and the text that is invariant to multiple perspectives. r is the perspective of the input image, I r is the input front view, G θ (v i ) is from v i View rendered by perspective, Text inv Generating fine-grained multi-view invariant text information.

[0081] S23: 2D and 3D reverse optimization guidance;

[0082] S23-1: The two-dimensional structure is guided by the standard diffusion model:

[0083] Add Gaussian noise and coarse-grained text information to the 2D image rendering results of the multi-view Novel view Test c Together they serve as the conditions of the Stable Diffusion model, and are input into the diffusion model to guide the rendered two-dimensional image results. The output of this step is the loss function L 2D .

[0084] The loss function can be calculated using the following formula:

[0085]

[0086] ∈,∈ φ ,φ,θ are the process parameters of adding noise and predicting noise. The value of parameter θ is updated according to the solution result. For the coarse-grained stage, θ is the parameter value of NeRF's MLPs. For the fine-grained stage, it is the parameter value of SDF, triangle deformation and color field. is the added noise, predicted noise, and prior parameters of the 2D diffusion model. z is the image with added noise, z t is the result of adding noise to the image after t steps, w(t) is the weight corresponding to the prediction of the tth step, and e is the result of text encoding.

[0087] S23-2: The three-dimensional structure is constrained by a multi-view diffusion model:

[0088] The input image of the multi-view Novel view and the translation and rotation matrix (R, T) of the new view are input into the multi-view diffusion model Zero-1-to-3 to obtain the corresponding new view image. The rendering result of this stage is obtained by three-dimensional rendering. The rendering result is constrained by the new view image obtained by the diffusion model, and the three-dimensional model is updated. The constraint calculation can be calculated by the following formula:

[0089]

[0090] where (R, T) is the camera pose passed to multi-view generation, i.e., the view-dependent diffusion model. t It is a noise feature map obtained by adding random Gaussian noise with a time step of t to the latent space feature map of I.

[0091] S24: At the same time, the CLIP network is used for coarse-grained semantic Text c Align and optimize the generated new perspective. This step outputs the loss function L clip The loss function can be calculated by the following formula:

[0092] L clip =|Adapter(E m (I r ))-E t (Text c )|

[0093] Text c To provide a coarse-grained text description of the image, the Adapter is composed of two linear fully connected layers and a ReLU activation function. m and E t They represent the image encoder and text encoder of CLIP respectively.

[0094] S25: Perform strong constraint calculation L on the input perspective 2D rendering result pre , L d , L n . r is the viewpoint corresponding to the input image.

[0095] S25-1: L pre It is the bi-norm loss of the rendered original view image and the preprocessed RGB and mask image mask of the input image. The loss function can be calculated by the following formula:

[0096]

[0097] ⊙ Element-by-element multiplication between corresponding elements of the two matrices, M is the foreground mask obtained by the ray-integrated volume density along each pixel.

[0098] S25-2: L d is the Pearson negative correlation loss between the rendered original depth image depth and the input depth map. The loss function can be calculated using the following formula:

[0099]

[0100] Where cov(.) represents the covariance and σ(.) represents the standard deviation. r The depth estimation map of the original perspective rendering result, d is the depth map obtained by preprocessing.

[0101] S25-3: L n To estimate the normal vector of each point using a finite difference of depth, a 2D normal map n is rendered from the normal vector and the loss applied to it is:

[0102] L n =‖n-τ(g(n,k))‖

[0103] Where τ(.) is the stop gradient operation, g(.) is the Gaussian blur, k is the convolution kernel, and the kernel size is set to 9×9. S26: The total loss function of the model training is:

[0104] L total =λ 1 L 2D +λ 2 L 3D +λ 3 L clip +λ 4 L inv +λ 5 L pre +λ 6 L d +λ 7 L n

[0105] L pre is the bi-norm loss of the rendered original view image and the preprocessed RGB and mask images of the input image, L d is the Pearson negative correlation loss between the rendered original depth image and the input depth map, L n To estimate the normal vector of each point using the finite difference of depth, a 2D normal map n is rendered from the normal vector. λ is the weight coefficient of the corresponding loss function.

[0106] In summary, semantic priors are added to the training process of step S2, and constraints are imposed on the two-dimensional diffusion results and the three-dimensional rendering results. The generation of the three-dimensional structure is finally constrained through back propagation. The coarse-grained rendering stage and the fine-grained rendering stage have the same process, only the rendering method is different. The coarse-grained rendering stage uses NeRF, while the fine-grained rendering stage uses the DMTet rendering method.

[0107] S3: Iterative processing step: Using a single image as input, the preprocessing result is used to iteratively solve the multimodal 3D reconstruction network constructed in step S2, and finally the reconstructed image Mesh is output.

[0108] like Figure 3 As shown in Figure 2, the reconstruction results on three datasets. The multi-view images are the visualization results obtained through 3D rendering.

[0109] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should be regarded as the scope of protection of the present invention.

Claims

1. A semantically guided zero-shot generalized 3D object reconstruction method, characterized in that: Establishing cross-modal information fusion between semantics and 3D vision to generate optimal 3D object structure based on semantic guidance includes the following steps: Preprocessing steps: Separate the target object from the background in the image, obtain the target object mask image, extract the depth information of the target object and the normal information of the object surface; Model construction steps: construct a 3D implicit structure sub-model based on a two-stage fractional distillation sampling from coarse to fine; introduce a multimodal large language model to generate coarse and fine-grained descriptions of the image and calculate fine-grained similarity; construct a rendering network model for 2D and 3D structures; use the CLIP network for coarse-grained semantic alignment and optimize the generated new perspective; construct a multimodal 3D implicit representation network guided by semantic features through a visual text alignment strategy; construct a model total loss function for back-propagation callback of various network parameters of the multimodal 3D reconstruction network during iterative learning; Iterative processing steps: Using a single image as input, the preprocessing results are used to iteratively solve the constructed multimodal 3D reconstruction network, and finally the reconstructed image Mesh is output.

2. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The two-stage fractional distillation sampling based on coarse to fine includes: constructing a three-dimensional implicit structure sub-model based on NeRF network and DMTet network, initializing implicit neural network to represent three-dimensional scene, sampling spatial points and predicting their color and density, combining volume rendering to obtain multi-view image Novel View, and gradually optimizing the generated realistic image by guiding the quality of the generated multi-view image; the rendering method from three-dimensional structure to multi-view is different in the two stages, the first stage adopts NeRF network, and the second stage adopts DMTet network to achieve refinement; the parameters are updated by back propagation, so as to update the generated three-dimensional structure and obtain the multi-view Novel view of the image rendering result; at the same time, the three-dimensional reconstruction result of the first stage is used as the initialization input of the second stage to accelerate the generation speed of the refinement process.

3. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 2, characterized in that: Constructing a 3D implicit structure sub-model and using it to generate multi-view images includes the following steps: NeRF takes the position (x, y, z) and viewing direction (θ, φ) of each point in three-dimensional space as input, and uses a neural network to predict: the color (r, g, b) and volume density (σ) of the point in a specific direction; the formula is expressed as: fΘ(x,y,z,θ,φ)→(r,g,b,σ) The neural network f is a multi-layer perceptron (MLP) that implicitly represents the entire 3D scene. During initialization, NeRF samples a large number of points within a three-dimensional space, and accumulates and calculates the points on the light through a neural network to obtain the color and density of each sampling point, thereby generating an image pixel value C(r); Where T(t) is the transmittance, which represents the probability that light is not absorbed before depth t; σ(t) is the volume density; and c(t) is the color value.

4. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The multimodal large language model specifically includes the following steps: Set the prompt word for the LLaVA multimodal large language model to generate a specified description of the image and obtain a coarse-grained description Text c and fine-grained description Text inv ; For the generated large segment of fine-grained multi-view invariant text description Text inv Segment the sentences and compare them with the coarse-grained feature Text c The splicing is then performed, and then the similarity is calculated with the input image and the generated new view image respectively; the Softmax method is used to normalize the similarity; since the generated multi-view invariant features are of variable length, the two features (Text inv Components and Text after sentence segmentation c ) and a norm to control the sum of similarities of multi-view invariant features to be equal.

5. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 4, characterized in that: The similarity between the variable-length multi-view invariant feature description and the image after constrained normalization has the same similarity value under different viewpoints; the loss function is calculated by the following formula: Among them, the Softmax method is used to normalize the similarity, the Sim operation is used to find the similarity between the image and the multi-view invariant text, r is the view angle of the input image, I r is the input front view, G θ (v i ) is from v i View rendered by perspective, Text inv Generating fine-grained multi-view invariant text information.

6. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The two-dimensional structure is rendered using the standard diffusion model: Add Gaussian noise and coarse-grained text information to the 2D image rendering results of the multi-view Novel view Test c Together they are used as the input of the standard diffusion model Stable Diffusion to learn the rendered two-dimensional image results; the output of this step is the loss function L 2D ; The loss function can be calculated using the following formula: ∈,∈ φ , φ, θ are the process parameters of adding noise and predicting noise, and the value of parameter θ is updated according to the solution result. For the coarse-grained stage, θ is the parameter value of NeRF's MLPs, and for the fine-grained stage, it is the parameter value of SDF, triangle deformation and color field; ∈, ∈ φ , φ are the added noise, predicted noise, and prior parameters of the 2D diffusion model, z is the image with added noise, z t is the result of adding noise to the image after t steps, w(t) is the weight corresponding to the prediction of the tth step, and e is the result of text encoding.

7. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The three-dimensional structure is rendered using a multi-view diffusion model: Input the input image of the multi-view Novel view and the translation and rotation matrix (R, T) of the new view into the multi-view diffusion model Zero-1-to-3 to obtain the corresponding new view image, use 3D rendering to obtain the rendering result of this stage, use the new view image obtained by the diffusion model to constrain the rendering result, and update the 3D model; where (R, T) is the camera pose passed to multi-view generation, z t It is a noise feature map obtained by adding random Gaussian noise with a time step of t to the latent space feature map of I.

8. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The new perspective generated by the coarse-grained semantic alignment optimization using the CLIP network, the loss function is calculated by the following formula: L clip =|Adapter(E m (I r ))-E t (Text c )| Text c To provide a coarse-grained text description of the image, the Adapter is composed of two linear fully connected layers and a ReLU activation function. m and E t The gates represent the image encoder and text encoder of CLIP.

9. The semantically guided zero-shot generalized 3D object reconstruction method according to claim 1, characterized in that: The semantic feature guidance is achieved through the visual text pairing strategy, including: strong constraint calculation of the two-dimensional rendering result L corresponding to the viewpoint of the input image pre , L d , L n ; L pre The bi-norm loss of the rendered original view image and the preprocessed RGB and mask images of the input image; ⊙ Element-by-element multiplication between corresponding elements of the two matrices, M is the foreground mask obtained by the ray-integrated volume density along each pixel; L d The Pearson negative correlation loss between the rendered original perspective depth image and the input depth map; Where cov(.) represents the covariance and σ(.) represents the standard deviation. r The depth estimation map of the original perspective rendering result, d is the depth map obtained by preprocessing; L n To estimate the normal vector for each point using a finite difference in depth, a 2D normal map n is rendered from the normal vector and the following loss is applied: L n =||n-τ(g(n,k))|| Where τ(.) is used to stop the gradient operation, g(.) is Gaussian blur, and the size of the convolution kernel k is set to 9×9.

10. The semantically guided zero-shot generalized 3D object reconstruction method according to any one of claims 1 to 9, characterized in that: The total loss function for model training is: L total =λ1L 2D +λ2L 3D +λ3L clip +λ4L inv +λ5L pre +λ6L d +λ7L n Among them, λ1-λ7 are the weight coefficients of the corresponding loss function.

Citation Information

Patent Citations

  • Efficient and accurate stereoscopic image three-dimensional target detection method

    CN116416117A

  • Method and system for generating interactive object viewer

    KR102551914B1

  • Three-dimensional construction network training method and apparatus, and three-dimensional model generation method and apparatus

    WO2024193622A1

  • Three-dimensional model construction method and system, and related device

    WO2024230843A1