Generating three-dimensional models using machine learning models
By adopting standardized camera coordination and multi-level image prompt controllers in the 3D model generation system, the geometric accuracy and detailed texture problems of 3D object generation under image mode are solved, and a more efficient and robust 3D model generation process is achieved.
Patent Information
- Application Number
- CN202411675459.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-27
AI Technical Summary
When using images as additional modalities generated by 3D objects, the prior art faces the problems of complex feature analysis and interpretation, inaccurate view synthesis, and blurred or incomplete objects, resulting in the generated 3D model lacking geometric accuracy and detailed texture.
Standard camera coordination across different object instances is adopted, and layered control is performed through multi-level image prompt controllers, including global controllers, local controllers and pixel controllers, to guide the diffusion model from image input to each architectural block, simplifying the information transmission path.
Improves the geometric accuracy of the generated 3D objects, improves the robustness and user experience of the 3D generation process, and meets a wider range of creative and practical application needs.
Smart Images

Figure CN120047604A_ABST
Abstract
Description
Background Art
[0001] Machine learning models are increasingly being used in various industries to perform a variety of different tasks. These tasks can include content generation. Improved techniques for performing content generation using machine learning models are desired. Brief Description of the Drawings
[0002] The following detailed description can be better understood when read in conjunction with the accompanying drawings. For purposes of illustration, example embodiments of various aspects of the present disclosure are shown in the drawings; however, the invention is not limited to the specific methods and tools disclosed.
[0003] Figure 1 An example system for generating a three-dimensional (3D) model using a machine learning model in accordance with the present disclosure is shown.
[0004] Figure 2 An example system for generating a 3D model using a machine learning model in accordance with the present disclosure is shown.
[0005] Figure 3 An example first sub-model in accordance with the present disclosure is shown.
[0006] Figure 4 An example first sub-model in accordance with the present disclosure is shown.
[0007] Figure 5 An example global controller and an example local controller in accordance with the present disclosure are shown.
[0008] Figure 6 Example diffusion results at different settings of a multi-level controller in accordance with the present disclosure are shown.
[0009] Figure 7 An example system for generating a 3D model using a machine learning model in accordance with the present disclosure is shown.
[0010] Figures 8A to 8B Background alignment and camera alignment in accordance with the present disclosure are shown.
[0011] Figure 9 An example process for generating a 3D model using a machine learning model in accordance with the present disclosure is shown.
[0012] Figure 10 An example process for generating a 3D model using a machine learning model in accordance with the present disclosure is shown.
[0013] Figure 11 An example process for generating a 3D model using a machine learning model in accordance with the present disclosure is shown.
[0014] Figure 12Illustrates an example process for generating a 3D model using a machine learning model according to the present disclosure.
[0015] Figure 13 Illustrates an example process for generating a 3D model using a machine learning model according to the present disclosure.
[0016] Figure 14 Illustrates an example process for generating a 3D model using a machine learning model according to the present disclosure.
[0017] Figure 15 Illustrates an example process for generating a 3D model using a machine learning model according to the present disclosure.
[0018] Figure 16 Illustrates example evaluation results of a machine learning model configured to generate a 3D model according to the present disclosure.
[0019] Figure 17 Illustrates example evaluation results of a machine learning model configured to generate a 3D model according to the present disclosure.
[0020] Figure 18 Illustrates example evaluation results of a machine learning model configured to generate a 3D model according to the present disclosure.
[0021] Figure 19 Illustrates an example computing device that can be used to perform any of the techniques disclosed herein. Detailed Description
[0022] A machine learning model can be used to generate three-dimensional (3D) assets (e.g., animations, content, models, etc.). Such a machine learning model can generate 3D assets based on text prompts (such as text prompts received from a user). A user can input text associated with a desired 3D asset, and the machine learning model can generate a 3D asset corresponding to the input text. The generated 3D asset can be used as material for education, gaming, or toys. The generated 3D asset can be used for secondary processing, editing, rig intervention, animation, etc.
[0023] Providing images as an additional modality for 3D generation (e.g., in addition to text) offers significant advantages. Images convey rich and accurate visual information that text may describe vaguely or omit entirely. For example, subtle details such as texture, color, and spatial relationships can be captured directly and unambiguously in an image, while text descriptions may struggle to convey the same level of detail comprehensively or may require overly long descriptions. This visual specificity helps generate more accurate and detailed 3D models because the system can directly reference actual visual cues rather than interpret text descriptions, which can vary widely in detail and subjectivity. Further, using images allows users to express their desired outcomes in a more intuitive and direct manner, especially for those who may find it difficult to articulate their vision in text. This multimodal approach combines the richness of visual data with the contextual depth of text, making the 3D generation process more robust, user-friendly, and efficient, thus meeting a wider range of creative and practical application needs.
[0024] Employing images as an additional modality for 3D object generation presents several challenges. Unlike text, images contain multiple features such as color, texture, spatial relationships, etc., and the analysis and interpretation of these features are more complex. Additionally, significant variations in light, shape, or self-occlusion within an object can lead to inaccurate and inconsistent view synthesis, resulting in a blurred or incomplete 3D model. The reconstructed objects often lack geometric precision and detailed texture. During reconstruction, mismatched pixels are averaged in the final 3D object, leading to blurred textures and smoothed geometric shapes. Therefore, improved techniques for 3D model generation are needed.
[0025] This paper describes an improved system for 3D model generation (i.e., ImageDream). The improved system involves canonical camera coordination across different object instances and includes a multi-level image prompt controller, which consists of a global controller, a local controller, and a pixel controller. Applying canonical camera coordination across different object instances improves the geometric precision of the generated 3D objects. The multi-level controller provides hierarchical control, guiding the diffusion model from the image input to each architectural block, thus simplifying the information transmission path.
[0026] Figure 1 An example machine learning model 100 for generating 3D models (e.g., 3D objects) using a machine learning model is illustrated. The machine learning model 100 can be configured to generate three-dimensional (3D) models with precise geometric shapes and detailed textures. The machine learning model 100 can include a first sub-model 102 and a second sub-model 104.
[0027] The first sub - model 102 can be configured to generate a set of multi - view images 103 based at least in part on an input two - dimensional (2D) image 101. The first sub - model 102 can be configured to generate a set of multi - view images 103 based on the input 2D image 101 and input text. The input 2D image 101 can be user - input (e.g., received from a user). The input 2D image 101 can indicate an object for which the user desires to generate a 3D model. The input 2D image 101 can include a 2D object. A set of multi - view images 103 can include a set of four images of the same object from four different orthogonal perspectives. For example, the set of multi - view images can include a set of four images of the same object from a front view (e.g., 0 degrees), a first side view (e.g., 90 degrees), a rear view (e.g., 180 degrees), and a second side view (e.g., 270 degrees). The object can be associated with the input 2D image 101 and the input text. For example, a set of multi - view images 103 can include multiple images of the 2D object depicted in the input image from different perspectives.
[0028] In an embodiment, the first sub - model 102 includes a multi - level image prompt controller. The multi - level image prompt controller can be configured to implement hierarchical control over generating a set of multi - view images based at least in part on the input 2D image. The multi - level image prompt controller will be discussed in more detail below with reference to Figures 3 to 6 The second sub - model 104 can receive the set of multi - view images 103 as input. The second sub - model 104 can generate a 3D model (e.g., 3D object 105) based at least in part on the set of multi - view images 103. The second sub - model 104 can be configured to implement background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model (e.g., 3D object 105). The background alignment and camera alignment will be discussed in more detail below with reference to Figures 8A to 8B The 3D object 105 can depict the object associated with the input 2D image 101. The 3D object 105 can be used for secondary processing, editing, rigging, animation, etc.
[0029] Figure 2 A machine - learning model 100 is illustrated. As described above, the first sub - model 102 can be configured to generate a set of multi - view images 103 based at least on the input 2D image 101. In Figure 2In the example, the input 2D image 101 depicts a robot. The first sub-model 102 can be configured to generate a set of multi-view images 103 based on the 2D image of the robot. The input 2D image 101 can be user input (e.g., received from the user). The user may want to generate a 3D model of the robot (e.g., an asset, an object). A set of multi-view images 103 can include a set of four images of the robot from four different orthogonal perspectives. For example, the set of multi-view images can include a set of four images of the same robot from the front view (e.g., 0 degrees), the first side view (e.g., 90 degrees), the rear view (e.g., 180 degrees), and the second side view (e.g., 270 degrees). The second sub-model 104 can receive the set of multi-view images 103 as input. The second sub-model 104 can generate a 3D model of the robot (e.g., 3D object 105) at least partially based on the set of multi-view images 103.
[0030] Figure 3 FIG. 300 illustrates a diagram showing the first sub-model 102 in more detail. The first sub-model 102 includes a multi-level image prompt controller 202 and a diffusion model 204. The multi-level image prompt controller 202 can be configured to implement hierarchical control over generating a set of multi-view images based at least in part on the input 2D image. The multi-level image prompt controller 202 can include, for example, a local controller 202a, a global controller 202b, and a pixel controller 202c.
[0031] The local controller 202a can be configured to enable the first sub-model 102 to capture detailed structural information from the input image 101. The input image 101 can be encoded by a contrastive language-image pre-training encoder (CLIP). The hidden features from the CLIP encoder can be resampled. The hidden features can be resampled before global pooling. The resampling can be performed by the resampling component of the local controller 202a. The local controller 202a can generate balanced local features. The control of the local controller 202a can be implemented by inputting the balanced local features into the cross-attention layer of the diffusion block 204.
[0032] The global controller 202b can be configured to enable the first sub-model 102 to extract global structural information from the input image 101. The input image 101 can be encoded (e.g., CLIP encoded). The global controller 202b can generate a vector based on the encoded image features. The control of the global controller can be implemented by inputting the vector from the global controller 202b into the cross-attention layer of the diffusion block 204. The vector can be adjusted by a multi-layer perceptron. The adjusted vector can be input into the cross-attention layer of the diffusion block 204.
[0033] The pixel controller 202c can be configured such that the first sub-model 102 can optimize the texture of the generated 3D model (e.g., 3D object) based on the appearance of the input image 101. The input image 101 can be encoded (e.g., variational auto-encoding (VAE)). The VAE-encoded image features can be embedded into all attention layers. The pixel controller 202c can be incorporated into the first sub-model 102 to implement a 3D self-attention process between the input image 101 and a set of multi-view images 103.
[0034] Figure 4 FIG. 400 illustrates a diagram showing the first sub-model 102 in more detail. The first sub-model 102 includes a multi-level image prompt controller 202 and a diffusion block 204. The multi-level image prompt controller 202 can be configured to implement hierarchical control over generating a set of multi-view images based at least in part on an input 2D image. The multi-level image prompt controller 202 can include a local controller 202a, a global controller 202b, and a pixel controller 202c.
[0035] The input 2D image 101 can be encoded (e.g., CLIP encoding) by a CLIP encoder 401. The local controller 202a and the global controller 202b obtain the input of the image features after CLIP encoding. The local controller 202a and the global controller 202b can output the adjusted features to the cross-attention layer of the diffusion block 204. The adjusted features can represent image semantic information. The input image 101 can be encoded (e.g., VAE encoding) by a VAE encoder 403. The pixel controller 202c can send the VAE-encoded features to the 3D self-attention layer of the diffusion block 204. The 3D self-attention layer of the diffusion block 204 can perform pixel-level dense self-attention with corresponding hidden features at each layer of the four-view diffusion.
[0036] Figure 5 FIG. 500 illustrates a diagram showing the global controller and the local controller in more detail. The input image (e.g., 2D image 101) can be encoded by a CLIP encoder 401. The global controller 202b can integrate a global embedding 504 into the first sub-model 102. The global embedding 504 can include a 1024-dimensional vector with a symbol (token) length (denoted as f g ) of 4. A multi-layer perceptron (MLP) θg serving as an adapter 506 can further adapt the vector to 1024 as the input to the cross-attention layer of the diffusion block 204. The adapter 506 can align the image features with the text features. The adapter 506 can be an effective and lightweight adapter that implements the image prompt ability of the pre-trained text-to-image diffusion model. The adapter 506 can be a decoupled cross-attention mechanism for separating text features and image features in the cross-attention layer.
[0037] The hidden features 502 from the CLIP encoder can be resampled by the resampling component 508 of the local controller 202a. The hidden features can be resampled before global pooling. The local controller 202a can generate balanced local features. The balanced local features from the local controller 202a can be adjusted by the adapter 510. The adapter 510 can be a multi-layer perceptron. The adapter 510 can be the same as or different from the adapter 506. The adjusted balanced local features can be input into the cross-attention layer of the diffusion block 204. The adjusted balanced local features from the local controller enable the first sub-model 102 to capture detailed structural information from the input image.
[0038] Figure 6 Examples of diffusion results under different settings of the multi-level controller in the machine learning model 100 are shown. As described above, a multi-layer perceptron (MLP)(θ g ) can be inserted after the CLIP image global embedding to be used as an adapter. This can align the image features with the text features. Specifically, the CLIP image encoding can encode the image features into a 1024-dimensional vector (f g ) with a symbol length of 4. The MLP can further adjust the image features into the input of the cross-attention in the diffusion block 204. In the diffusion block 204, inside the attention layer l, a new set of MLPs θ kg,l and θ vg,l receive the adjusted features as input and output their attention key matrix and value matrix. Then, based on the query feature matrix q l , the attention key matrix and value matrix are aggregated to produce the corresponding image cross-attention feature h g,l . Here, a weight λ = 1.0 can be introduced to balance the hidden from the text and the image, and the final output of layer l is h l = h t,l + λh g,l .
[0039] To train such a model, the diffusion model can be frozen, and only {θ g , θ kg,l , θ vg,l} l is fine-tuned. After adjusting the model, the model can extract some information from the input image, such as the structure of the object. As shown in example 602 of Figure 6 , when compared with the input 2D image, the diffusion output can put a pirate hat similar to the image on the bulldog's head but lose some detailed pose and appearance information.
[0040] To enhance control, the hidden features from the CLIP encoder can be utilized before global pooling, as this hidden feature may contain more detailed structural information. The hidden feature (denoted as f h ) has a symbol length of 257 and a feature dimension of 1280. An MLP adapter θ h can be introduced to feed f h into the cross-attention module of the diffusion network, where θ kh,l and θ vh,l constitute the key matrix and the value matrix. These parameters {θ kh,l , θ vh,l , θ h} l can be jointly trained as learnable elements similar to the global controller 202b. The trained result is overly sensitive to the image symbols, leading to overexposed and unrealistic images, especially in the case where the classifier-free guidance (CFG) setting is high, as shown in Figure 6 's example 604. To alleviate this situation, a resampling module θ r is used to reduce the hidden symbol count from 257 to 16, resulting in more balanced local image features f r . The corresponding local controller parameters are {θ r , θ kr,l , θ vr,l} l . As shown in Figure 6 's example 606, after this resampling, even with a high CFG level, the diffusion images look more realistic. It can be clearly seen from the generated images that the model captures the overall layout and object shape, but it is still difficult to capture more fine-grained feature details, such as the skin texture of the object.
[0041] The pixel controller 202c can be used to optimally integrate the appearance texture of the object. To optimally integrate the appearance texture of the object, the image prompt pixel latent variable x can be embedded across all attention layers of the machine learning model 100. Specifically, the machine learning model 100 employs a 3D dense self-attention mechanism with a shape of (bz, 4, c, h l , w l ) across four views within the transformer layer. The machine learning model 100 further adopts additional frames by concatenating the input images, resulting in a feature shape of (bz, 5, c, h l , w l ). This enables a similar 3D self-attention process between the four-view image and the input image.
[0042] During the training of the diffusion block 204, noise can be not added to the latent variables of the input image prompt, thereby ensuring that the network clearly captures the image information. Additionally, to distinguish the input image features, a zero vector can be assigned to the camera embedding of the input image. Assuming that the pixel controller 202c is integrated into the multi-view diffusion without the need for additional parameters, all feature parameters can be uniformly fine-tuned, adopting the same training mechanism as the local controller 202a and the global controller 202b but with a learning rate reduced by a factor of ten. This method can more effectively preserve the representation of the original features. As Figure 6 depicted in example 608 of
[0043] Figure 7 FIG. illustrates a system 700 that more particularly shows the machine learning model 100. The first sub-model 102 can be configured to generate a set of multi-view images 103 based at least on an input 2D image 101. For example, the first sub-model 102 can be configured to generate a set of multi-view images 103 based on the input 2D image 101 and a text input. The input 2D image 101 and the text can be user input (e.g., received from a user). The input 2D image 101 and the text can indicate the 3D model that the user desires to generate. The input 2D image 101 can include a 2D object. A set of multi-view images 103 can include a set of four images of the same object from four different orthogonal perspectives. For example, the set of multi-view images can include a set of four images of the same object from a front view (e.g., 0 degrees), a first side view (e.g., 90 degrees), a rear view (e.g., 180 degrees), and a second side view (e.g., 270 degrees). The object can be associated with the input 2D image 101 and the text. For example, a set of multi-view images 103 can include multiple images of the 2D object depicted in the 2D image 101 and / or described by the text from different perspectives.
[0044] As described above, the multi-level image prompt controller 202 can be configured to implement hierarchical control over generating a set of multi-view images 103 based at least in part on an input image. The diffusion block 204 can receive the output of the multi-level image prompt controller 202 and the embedding of the input text. Based on the output of the multi-level image prompt controller 202 and the text embedding, the diffusion block 204 can generate a set of multi-view images 103. The second sub-model 104 can receive the set of multi-view images 103 as input. The second sub-model 104 can generate a 3D model (e.g., a 3D object) based at least in part on the set of multi-view images 103.
[0045] During the training of the first sub-model 102, multiple views can be rendered based on canonical camera coordination, and at least one other image prompt, the front-view image, can be rendered with a random setting. The multi-view images can be fed as training targets of the multi-view diffusion network, and the image prompts can be encoded with a multi-level controller as the input to the diffusion. During the training of the second sub-model 104, the trained diffusion is used for image prompt score distillation.
[0046] To train the first sub-model 102, assume a text-image dataset X = {x, y}, and a multi-view dataset X mv ={x mv , y, c mv}, where x is the latent image embedding from the VAE, y is the text embedding from CLIP, and c is its self-designed camera embedding. The multi-view (MV) diffusion loss can be formulated as:
[0047] where,
[0048]
[0049] where x is the noisy latent image generated according to the random noise ∈ and the image latent variable, and ∈ θ is the multi-view diffusion model parameterized by θ.
[0050] After the first sub-model 102 is trained, the first sub-model 102 can be inserted into the pipeline, where score distillation sampling (SDS) is performed based on four generated views. Specifically, at each iteration step, four random orthogonal views are rendered with four random views, the extrinsic and intrinsic camera parameters. Then, the four random orthogonal views can be encoded as the latent variable x mv and inserted into the multi-view diffusion network to calculate the diffusion loss in the image space, which is backpropagated to optimize the parameters of the second sub-model 104. Its form is
[0051]
[0052] where, is the denoised MV image at time step 0 from the diffusion block.
[0053] The second sub-model 104 can be configured to perform background alignment to improve the quality of the generated 3D model. During the score distillation sampling (SDS) optimization, the 3D model generated by the second sub-model 104 includes a randomly colored background to distinguish the inside and outside of the 3D object. When input into the diffusion network together with the object, this random background may conflict with the background from the image prompt, resulting in floating artifacts in the 3D model generated by the second sub-model 104, such asFigure 8A as shown in Example 800. To solve this problem, the image prompt background can be adjusted to match the rendered background color from the second sub-model 104, thereby successfully eliminating these artifacts.
[0054] The second sub-model 104 can be configured to implement camera alignment for the geometric accuracy of the generated 3D model. The first sub-model 102 can generate multi-view images that reflect (mirror) the camera parameters (e.g., elevation angle, field of view (FoV)) of the input image prompt, which remain unknown to the second sub-model 104 during 3D model rendering. Random sampling parameters for rendering may cause the images to be inconsistent with the rendering settings of the image prompt, thereby affecting the geometry of the detailed image structure. To mitigate this situation, the parameter sampling ranges of the camera FoV and elevation angle ([15,60], [0,30]) can be narrowed to [45,50] and [0,5] respectively, which is a more typical range for generating photos. As Figure 8B shown in Example 801, such adjustment can significantly improve the geometric accuracy of the 3D object. In an embodiment, a camera parameter estimation module can be used to determine the parameter sampling range. In other embodiments, during diffusion training, increased randomness can be used in image prompt rendering to better synchronize the settings between 3D model rendering and the first sub-model 102.
[0055] Figure 9 FIG. illustrates an example process 900 for generating a 3D model (e.g., an object) using a machine learning model. Although depicted as a series of operations in Figure 9 those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.
[0056] At 902, a two-dimensional (2D) image can be input into a machine learning model (e.g., machine learning model 100). The machine learning model is configured to generate a three-dimensional (3D) model with precise geometry and detailed texture. The input 2D image can include a 2D object. The generated 3D model can include a 3D object corresponding to the 2D object in the input image.
[0057] At 904, a set of multi-view images can be generated. The set of multi-view images can be generated at least in part based on the input 2D image. The set of multi-view images can be generated by a first sub-model of a machine learning model. The set of multi-view images can include a set of four images of the same object from four different orthogonal perspectives. For example, the set of multi-view images can include a set of four images of the same object from a front view (e.g., 0 degrees), a first side view (e.g., 90 degrees), a rear view (e.g., 180 degrees), and a second side view (e.g., 270 degrees). The object can be associated with the input 2D image. The first sub-model can include a multi-level image prompt controller configured to perform hierarchical control over generating the multi-view images by the first sub-model based on the input image. The multi-level image prompt controller can include, for example, a local controller, a global controller, and a pixel controller.
[0058] At 906, a 3D model can be generated. The 3D model can be generated at least in part based on the set of multi-view images. The 3D model can be generated by a second sub-model of a machine learning model. The second sub-model can be configured to perform background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model. The 3D model can depict the object associated with the input 2D image. The 3D object can be used for secondary processing, editing, skeletal rigging, animation, etc.
[0059] Figure 10 An example process 1000 for generating a 3D model using a machine learning model is illustrated. Although depicted as a series of operations in Figure 10 those of ordinary skill in the art will understand that various embodiments can add, remove, reorder, or modify the depicted operations.
[0060] The first sub-model (e.g., first sub-model 102) can include a multi-level image prompt controller (e.g., multi-level prompt controller 202) configured to perform hierarchical control over generating the multi-view images by the first sub-model based on the input image. The multi-level image prompt controller can include, for example, a local controller, a global controller, and a pixel controller. The global controller can be configured to enable the first sub-model to extract global structural information from the input image.
[0061] At 1002, encoding can be performed on a 2D image input to a machine learning model (e.g., machine learning model 100). The encoding can be performed by a contrastive language-image pre-training (CLIP) encoder. At 1004, a vector can be generated. The vector can be generated by a global controller of a multi-level image prompt controller. The vector can be generated based on the encoded image features. At 1006, a first control for generating a set of multi-view images can be implemented. The first control for generating a set of multi-view images can be implemented by inputting the vector from the global controller into a cross-attention layer of a first sub-model. The cross-attention layer can be in a diffusion block of the first sub-model. At 1008, the vector from the global controller can be adjusted. The vector can be adjusted by a multi-layer perceptron. At 1010, the adjusted vector can be input into the cross-attention layer. The adjusted vector can be input into the cross-attention layer of the diffusion block.
[0062] Figure 11 FIG. illustrates an example process 1100 for generating a 3D model using a machine learning model. Although depicted as a series of operations in Figure 11 those skilled in the art will understand that various embodiments can add, remove, reorder, or modify the depicted operations.
[0063] The first sub-model (e.g., first sub-model 102) can include a multi-level image prompt controller (e.g., multi-level prompt controller 202) configured to implement hierarchical control over the generation of multi-view images by the first sub-model based on an input image. The multi-level image prompt controller can include, for example, a local controller, a global controller, and a pixel controller. The local controller can be configured to enable the first sub-model to capture detailed structural information from the input image.
[0064] At 1102, encoding can be performed on a 2D image input to a machine learning model (e.g., machine learning model 100). The encoding can be performed by a Contrastive Language-Image Pretraining (CLIP) encoder. At 1104, the hidden features from the CLIP encoder can be resampled. The hidden features can be resampled before global pooling. The hidden features can be resampled by a resampling component of a local controller of a multi-level image prompt controller. Balanced local features can be generated. At 1106, a second control for generating the set of multi-view images can be implemented. The second control for generating the set of multi-view images can be implemented by inputting the balanced local features into a cross-attention layer of a first sub-model. The cross-attention layer can be in a diffusion block of the first sub-model. At 1108, the balanced local features from the local controller can be adjusted. The balanced local features from the local controller can be adjusted by a multi-layer perceptron. At 1110, the adjusted balanced local features can be input into the cross-attention layer. The adjusted balanced local features can be input into the cross-attention layer of the diffusion block.
[0065] Figure 12 FIG. illustrates an example process 1200 for generating a 3D model using a machine learning model. Although depicted as a series of operations in Figure 12 those skilled in the art will understand that various embodiments can add, remove, reorder, or modify the depicted operations.
[0066] The first sub-model (e.g., first sub-model 102) can include a multi-level image prompt controller (e.g., multi-level prompt controller 202) configured to implement hierarchical control over generating multi-view images by the first sub-model based on an input image. The multi-level image prompt controller can include, for example, a local controller, a global controller, and a pixel controller. The pixel controller can be configured to enable the first sub-model to optimize the texture of the generated 3D model based on the appearance of the 2D image input to the machine learning model (e.g., machine learning model 100).
[0067] The input 2D image can be encoded by a VAE encoder (e.g., VAE encoding) to generate VAE-encoded features. At 1202, the 2D image input to the machine learning model can be VAE-encoded. At 1204, the VAE-encoded image features can be embedded. The VAE-encoded image features can be embedded by integrating the pixel controller of the multi-level image prompt controller into the first sub-model, so as to enable a 3D self-attention process between the 2D image and a set of multi-view images. The pixel controller can receive the VAE-encoded features as input. The pixel controller can send the VAE-encoded features to the 3D self-attention layer of the diffusion block. The 3D self-attention layer of the diffusion block can perform pixel-level dense self-attention with corresponding hidden features at each layer of the four-view diffusion.
[0068] Figure 13 FIG. illustrates an example process 1300 for generating a 3D model using a machine learning model. Although depicted as a series of operations in Figure 13 , those of ordinary skill in the art will understand that various embodiments can add, remove, reorder, or modify the depicted operations.
[0069] At 1302, a two-dimensional (2D) image can be input into a machine learning model (e.g., machine learning model 100). The machine learning model is configured to generate a three-dimensional (3D) model with precise geometry and detailed texture. The 2D image can include or depict a 2D object. The generated 3D model can include a 3D object corresponding to the 2D object in the input image.
[0070] At 1304, a set of multi-view images can be generated. The set of multi-view images can be generated at least in part based on the 2D image. The set of multi-view images can be generated by a first sub-model (e.g., first sub-model 102) of the machine learning model. The set of multi-view images can include a set of four images of the 2D object from four different orthogonal perspectives. For example, the set of multi-view images can include a set of four images of the 2D object from the front view (e.g., 0 degrees), the first side view (e.g., 90 degrees), the rear view (e.g., 180 degrees), and the second side view (e.g., 270 degrees).
[0071] At 1306, a 3D model can be generated. The 3D model can be generated at least in part based on the set of multi-view images. The 3D model can be generated by a second sub-model (e.g., second sub-model 104) of the machine learning model. The second sub-model can be configured to implement background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model. The 3D model can include a 3D object corresponding to the 2D object in the input image. The 3D object can be used for secondary processing, editing, action skeleton intervention, animation, etc.
[0072] During score distillation sampling (SDS) optimization, the generated 3D model includes a randomly colored background to distinguish the inside and outside of the 3D object. When input into the diffusion network together with the object, this random background may conflict with the background of the image prompt, resulting in floating artifacts in the 3D model generated by the second sub-model 104. During the optimization of the 3D model, the background of the image prompt can be adjusted to match the randomly colored background of the 3D model, which successfully eliminates the floating artifacts in the generated 3D model. At 1308, the background of the 2D image prompt can be adjusted. During the optimization of the 3D model, the background of the image prompt can be adjusted to match the randomly colored background of the 3D model, thereby eliminating the floating artifacts in the generated 3D model.
[0073] Figure 14 An example process 1400 for generating a 3D model using a machine learning model is illustrated. Although depicted as a series of operations in Figure 14 , those of ordinary skill in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.
[0074] At 1402, a two-dimensional (2D) image can be input into a machine learning model (e.g., machine learning model 100). The machine learning model is configured to generate a three-dimensional (3D) model with precise geometry and detailed textures. The input 2D image can indicate the 3D model that the user wants to generate. The 2D image can include or depict 2D objects.
[0075] At 1404, a set of multi-view images can be generated. The set of multi-view images can be generated at least in part based on the 2D image. The set of multi-view images can be generated by a first sub-model (e.g., first sub-model 102) of the machine learning model. The set of multi-view images can include a set of four images of the 2D object from four different orthogonal perspectives. For example, the set of multi-view images can include a set of four images of the 2D object from a front view (e.g., 0 degrees), a first side view (e.g., 90 degrees), a rear view (e.g., 180 degrees), and a second side view (e.g., 270 degrees).
[0076] At 1406, a 3D model can be generated. The 3D model can be generated at least in part based on the set of multi-view images. The 3D model can be generated by a second sub-model (e.g., second sub-model 104) of the machine learning model. The second sub-model can be configured to implement background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model. The 3D model can include a 3D object corresponding to the 2D object. The 3D object can be used for secondary processing, editing, action skeleton intervention, animation, etc.
[0077] At 1408, the geometric accuracy of a 3D object can be improved by narrowing the parameter sampling range of camera parameters. The first sub-model can generate multi-view images reflecting the camera parameters (e.g., elevation angle, FoV) prompted by the input image, and the parameters remain unknown to the second sub-model during the rendering of the 3D model. Randomly sampled parameters for rendering may cause the images to be inconsistent with the rendering settings of the image prompts, thereby affecting the geometry of the detailed image structure. To alleviate this situation, the parameter sampling ranges of the camera FoV and elevation angle ([15, 60], [0, 30]) can be narrowed to [45, 50] and [0, 5] respectively. This adjustment can significantly improve the geometric accuracy of the 3D object.
[0078] Figure 15 An example process 1500 for generating a 3D model using a machine learning model is illustrated. Although depicted as a series of operations in Figure 15 those skilled in the art will understand that various embodiments can add, remove, reorder, or modify the depicted operations.
[0079] A two-dimensional (2D) image can be input into a machine learning model (e.g., machine learning model 100). The machine learning model is configured to generate a three-dimensional (3D) model with precise geometry and detailed texture. The input 2D image can include or depict a 2D object. At 1502, a 3D model can be generated. The 3D model can be generated at least in part based on the 2D image and a set of multi-view images. The set of multi-view images can depict the 2D object. The set of multi-view images can include multiple images of the 2D object from different perspectives. The set of multi-view images can be generated by a first sub-model (e.g., first sub-model 102) of the machine learning model. The 3D model can be generated by a second sub-model (e.g., second sub-model 104) of the machine learning model. The second sub-model can be configured to implement background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model. The 3D model can include a 3D object corresponding to the 2D object included in the input image prompt. The 3D object can be used for secondary processing, editing, rig intervention, animation, etc.
[0080] At 1504, the geometric accuracy of the 3D object can be improved. The geometric accuracy of the 3D object can be improved by applying canonical camera coordination across different object instances. Canonical camera coordination can be adopted in the machine learning model. The first sub-model aims to regress towards a canonical multi-view image of the 2D object depicted in the input image. Canonical camera coordination stipulates that, under default camera settings (i.e., identity rotation and zero translation), the rendered image represents a front view of the object centered. This significantly simplifies the task of mapping the variations in the input image to the 3D object. The application of canonical camera coordination results in a 3D object with higher geometric accuracy.
[0081] Evaluate the performance of the machine learning model 100. The performance of the machine learning model 100 is evaluated using a combined dataset for 3D multi-view rendering and a 2D image dataset for training the controller. For the image prompts in the 3D dataset, one of the 16 front views is randomly selected from a total of 32 circular views, where the azimuth range is [-90, 90] degrees. For the 2D dataset, the input image is used as the image prompt. During training, the random dropout rate of the image prompts is set to 0.1, and the dropped prompts are replaced with random monochromatic images. For all experiments (i.e., using the global controller, local controller, and local controller plus pixel controller), the machine learning model 100 is trained for 60,000 steps, where the batch size is 256 and the gradient accumulation is two. The learning rate is set to 1e-4, except for the model with the pixel controller, where its learning rate is reduced to 1e-5. The size of the test image prompts is resized to 256×256, and the diffusion CFG is set to 5.0.
[0082] To evaluate the performance of the machine learning model 100, carefully curated prompts covering a variety of objects with relatively complex geometries and appearances are selected. Multiple images are generated according to each prompt, and the aesthetic object images are selected. Then the backgrounds of these images are removed, and the objects are recentered. The comparison criteria are geometric quality and similarity to the image prompt (IP). "Geometric quality" refers to the consistency of the generated 3D assets with common sense in terms of shape and minimal artifacts, while "similarity to IP" evaluates the similarity of the results to the input image.
[0083] Due to the lack of ground truth for the test image prompts, a real-user study is conducted to evaluate the quality of the generated 3D models. The participants are briefed on the evaluation criteria and are asked to select their favorite models based on these criteria. The experiment is a double-blind experiment, where the 3D assets generated by different methods are presented to the participants without identification labels. Figure 16 The comparison results 1600 depicted in show that the machine learning model 100 (e.g., ImageDream) (whether it is a machine learning model with a full pixel controller (e.g., ImageDream-P) or a machine learning model without a pixel controller (e.g., ImageDream-G)) is significantly better than other baselines. ImageDream-P is particularly favored, while ImageDream-G also obtains a positive preference rate. SyncDreamer is omitted from the figure because the preference rate of its NeuS results is 0%.
[0084] Figure 17A representative case 1700 presenting the comparison results between the diffusion model and the final second sub-model 104 (e.g., the Neural Radiance Field (NeRF) model) is shown. Systems like Magic123 and Zero123 rely on single-view diffusion with relative camera embeddings, which typically produce correct geometries. As shown, they are unable to accurately represent the length of the horse's body. In contrast, the machine learning model 100 (e.g., Image-Dream) effectively solves this problem through its unique design, resulting in a more satisfactory model.
[0085] To comprehensively evaluate the image quality of the machine learning model 100 at various stages, the Inception Score (IS) and the CLIP score are calculated using text prompts and image prompts respectively. The IS evaluates the image quality, while the CLIP score evaluates the text-image and image-image alignment. However, since the IS traditionally evaluates both the image quality and diversity within the set, and the number of prompts used for evaluation is limited, the reliability of the diversity aspect score is relatively low. Therefore, the IS is modified by omitting its diversity evaluation and replacing the average distribution with a uniform distribution. Specifically, we set q i in the IS to 1 / N, such that the IS of the image is i p i log N p i , where N is the inception class count, and p i is the predicted probability of the i-th class. This modified metric is denoted as the Quality-only IS (QIS). For the CLIP score, the average score between each generated view and the provided text prompt or image prompt is calculated.
[0086] Figure 18Table 1800 showing the comparison results is presented. SD-XL, which reflects the scores of the test images, achieved the highest QIS and CLIP scores. MVDream was listed as the benchmark for the final 3D model quality, and due to the multi-view consistency, it showed an improvement in the synthetic image quality after 3D fusion. In contrast, due to diffusion inconsistency, the image quality of both Zero123 and Zero123-XL decreased after 3D fusion. Magic123 had a higher CLIP score than Zero123 by integrating a joint diffusion model. The quality of SyncDreamer decreased because it only diffused 16 fixed views, making the reconstruction complex. For the machine learning model 100 (e.g., ImageDream), ablation evaluations were conducted on three models: one with a global controller, another with a local controller, and the last one integrating both the local controller and the pixel controller. The machine learning model 100 maintained high image quality in both the diffusion stage and the post-3D fusion stage. In particular, due to richer image feature representations, the local controller provided a higher image CLIP score after fusion. During both stages, the pixel controller model performed well in terms of the image CLIP score.
[0087] Figure 19 Illustrated is a computing device that can be used for various aspects, such as Figure 1 the models, components, and / or devices depicted in any one of FIGS. 6 Figure 1 to 8. With respect to Figure 19 FIGS. 6 to 8, any or all of the components can each be implemented by Figure 19 one or more instances of the computing device 1900.
[0088] The computing device 1900 can include a substrate or "motherboard", which is a printed circuit board to which a plurality of components or devices can be connected via a system bus or other electrical communication paths. One or more central processing units (CPUs) 1904 can operate in conjunction with a chipset 1906. The (multiple) CPUs 1904 can be standard programmable processors that perform the arithmetic and logical operations required for the operation of the computing device 1900.
[0089] (Multiple) CPUs 1904 can perform necessary operations by transitioning from one discrete physical state to the next by manipulating switching elements that can distinguish and change these states. The switching elements can typically include electronic circuits (such as flip-flops) that hold one of two binary states, and electronic circuits (such as logic gates) that provide an output state based on a logical combination of the states of one or more other switching elements. These basic switching elements can be combined to create more complex logic circuits, including registers, adder-subtractors, arithmetic logic units, floating-point units, and the like.
[0090] (Multiple) CPUs 1904 can be augmented or replaced with other processing units, such as (multiple) GPUs 1905. (Multiple) GPUs 1905 can include processing units dedicated to, but not necessarily limited to, highly parallel computing, such as graphics and other visualization-related processing.
[0091] The chipset 1906 can provide an interface between (multiple) CPUs 1904 and the remaining components and devices on the substrate. The chipset 1906 can provide an interface to random access memory (RAM) 1908, which is used as the main memory in the computing device 1900. The chipset 1906 can further provide an interface to a computer-readable storage medium, such as read-only memory (ROM) 1920 or non-volatile RAM (NVRAM) (not shown), for storing basic routines that can assist in booting the computing device 1900 and transferring information between the various components and devices. ROM 1920 or NVRAM can also store other software components required for the operation of the computing device 1900 in accordance with the various aspects described herein.
[0092] The computing device 1900 can operate in a networked environment using a logical connection to remote computing nodes and computer systems via a local area network (LAN). The chipset 1906 can include functionality for providing network connectivity via a network interface controller (NIC) 1922, such as a gigabit Ethernet adapter. The NIC 1922 can be capable of connecting the computing device 1900 to other computing nodes via the network 1916. It should be understood that multiple NICs 1922 can be present in the computing device 1900 to connect the computing device to other types of networks and remote computer systems.
[0093] The computing device 1900 can be connected to a mass storage device 1928 that provides non-volatile storage for the computer. The mass storage device 1928 can store system programs, application programs, other program modules, and data, which have been described in more detail herein. The mass storage device 1928 can be connected to the computing device 1900 through a storage controller 1924 connected to the chipset 1906. The mass storage device 1928 can be composed of one or more physical storage units. The mass storage device 1928 can include a management component 1910. The storage controller 1924 can interface with the physical storage units through a Serial Attached SCSI (SAS) interface, a Serial Advanced Technology Attachment (SATA) interface, a Fibre Channel (FC) interface, or other types of interfaces for physically connecting and transferring data between the computer and the physical storage units.
[0094] The computing device 1900 can store data on the mass storage device 1928 by changing the physical state of the physical storage units to reflect the stored information. The specific transformation of the physical state can depend on various factors and different embodiments of this specification. Examples of such factors can include, but are not limited to, the technology used to implement the physical storage units and whether the mass storage device 1928 is characterized as a main storage device or an auxiliary storage device, etc.
[0095] For example, the computing device 1900 can issue instructions through the storage controller 1924 to change the magnetic characteristics of a specific location within a disk drive unit, the reflection or refraction characteristics of a specific location within an optical storage unit, or the electrical characteristics of specific capacitors, transistors, or other discrete components within a solid-state storage unit, so as to store information on the mass storage device 1928. Other transformations of the physical medium are possible without departing from the scope and spirit of this specification, and the foregoing examples are provided only to assist this specification. The computing device 1900 can further read information from the mass storage device 1928 by detecting the physical state or characteristics of one or more specific locations within the physical storage units.
[0096] In addition to the mass storage device 1928 described above, the computing device 1900 can also access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. Those skilled in the art should understand that a computer-readable storage medium can be any available medium that provides non-transitory data storage and can be accessed by the computing device 1900.
[0097] By way of example and not limitation, a computer-readable storage medium can include volatile and non-volatile computer-readable storage media, transient computer-readable storage media and non-transient computer-readable storage media, and removable and non-removable media implemented in any method or technology. Computer-readable storage media includes, but is not limited to, RAM, ROM, erasable programmable ROM (“EPROM”), electrically erasable programmable ROM (“EEPROM”), flash memory or other solid state memory technologies, compact disc ROM (“CD-ROM”), digital versatile disc (“DVD”), high definition DVD (“HD-DVD”), BLU-RAY, or other optical storage devices, magnetic tape cartridges, tapes, magnetic disk storage devices, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transient manner.
[0098] A mass storage device (such as Figure 19 the mass storage device 1928 depicted in
[0099] can store an operating system for controlling the operation of the computing device 1900. The operating system can include a version of the LINUX operating system. The operating system can include a version of the WINDOWS SERVER operating system from Microsoft Corporation. According to other aspects, the operating system can include a version of the UNIX operating system. Various mobile phone operating systems, such as IOS and Android, can also be utilized. It should be understood that other operating systems can also be utilized. The mass storage device 1928 can store other systems or applications and data used by the computing device 1900.
[0100] A computing device (such as Figure 19 the computing device 1900 depicted inFigure 19 All components in the components shown in Figure 19 may include other components not explicitly shown in Figure 19 or may utilize a completely different architecture from that shown in
[0101] As described herein, a computing device may be a physical computing device, such as Figure 19 the computing device 1900. The computing node may also include a virtual machine host process and one or more virtual machine instances. Computer-executable instructions may be indirectly executed by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed in the context of the virtual machine.
[0102] It should be understood that the method and system are not limited to a particular method, particular components, or specific embodiments. It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.
[0103] Unless the context clearly dictates otherwise, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when values are expressed as approximations by use of the antecedent "about", it will be understood that the particular value forms another embodiment. It will be further understood that each end point of each range is significant with respect to the other end point and independent of the other end point.
[0104] "Optional" or "optionally" means that the subsequent described event or circumstance may or may not occur, and the description includes instances where the event or circumstance occurs and does not occur.
[0105] Throughout this specification and the claims, the word "comprise" and variations of the word (such as "comprising" and "comprises") mean "including but not limited to" and are not intended to exclude, for example, other components, integers, or steps. "Exemplary" means "an example of" and is not intended to convey an indication of a preferred or ideal embodiment. "Such as" is not used in a limiting sense, but for explanatory purposes.
[0106] Components that can be used to perform the described methods and systems are described. When describing combinations, subsets, interactions, groups, etc. of these components, it should be understood that although specific references to each of the various individual and collective combinations and permutations of these components may not be explicitly described, each component is specifically contemplated and described herein for all methods and systems. This applies to all aspects of the present application, including but not limited to operations in the described methods. Thus, if there are various additional operations that can be performed, it should be understood that each of these additional operations can be performed with any specific embodiment or combination of embodiments of the described methods.
[0107] The present methods and systems can be more readily understood by reference to the following detailed description of preferred embodiments and examples included therein, as well as the accompanying drawings and their description.
[0108] As will be understood by those skilled in the art, the present methods and systems can take the form of a full hardware embodiment, a full software embodiment, or an embodiment combining software and hardware aspects. In addition, the present methods and systems can take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied in the storage medium. More specifically, the present methods and systems can take the form of web-implemented computer software. Any suitable computer-readable storage medium can be utilized, including a hard disk, a CD-ROM, an optical storage device, or a magnetic storage device.
[0109] Embodiments of the present methods and systems are described below with reference to block diagrams and flowchart illustrations of methods, systems, apparatuses, and computer program products. It should be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, can be implemented respectively by computer program instructions. These computer program instructions can be loaded onto a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed on the computer or other programmable data processing apparatus create means for implementing the functions specified in one or more blocks of the flowchart.
[0110] These computer program instructions can also be stored in a computer-readable memory, and the computer program instructions can direct a computer or other programmable data processing apparatus to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture that includes computer-readable instructions for implementing the functions specified in one or more blocks of the flowchart. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more blocks of the flowchart.
[0111] The various features and processes described above can be used independently of one another or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Additionally, in some embodiments, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated with the methods and processes can be performed in other suitable orders. For example, the described blocks or states can be performed in an order different from that specifically described, or multiple blocks or states can be combined in a single block or state. The example blocks or states can be performed serially, in parallel, or in some other manner. Blocks or states can be added to or removed from the described example embodiments. The example systems and components described herein can be configured differently than described. For example, elements can be added, removed, or rearranged compared to the described example embodiments.
[0112] It should also be understood that the various items are illustrated as being stored in memory or a storage device during use, and for purposes of memory management and data integrity, these items or portions thereof may be transferred between memory and other storage devices. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Additionally, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least partially in firmware and / or hardware, including but not limited to one or more application specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions and including microcontrollers and / or embedded controllers), field programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all of the modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on a computer-readable medium, such as a hard disk, memory, network, or portable media article to be read by an appropriate device or via an appropriate connection. The systems, modules, and data structures may also be transmitted on various computer-readable transmission media (including wireless and wire / cable-based media) as generated data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) and may take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). In other embodiments, such computer program products may take other forms. Accordingly, the invention may be practiced with other computer system configurations.
[0113] Although the methods and systems have been described in connection with preferred embodiments and specific examples, this does not mean that the scope is limited to the specific embodiments set forth, since the embodiments herein are meant to be illustrative rather than restrictive in all respects.
[0114] Unless otherwise expressly stated, in no way is any method set forth herein to be construed as requiring that its operations be performed in a particular order. Accordingly, where a method claim does not actually recite an order of its operations or where no order is otherwise specifically stated in the claim or in the specification, no order is, in any way, intended to be inferred, regardless of any non-explicitly stated basis for interpretation, including: logical issues regarding step arrangement or operational flow; the plain meaning derived from grammatical organization or punctuation; and the number or type of embodiments described in the specification.
[0115] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of the disclosure. Considering the specification and practice described herein, other embodiments will be apparent to those skilled in the art. This specification and the example figures are only considered to be exemplary, and the true scope and spirit are indicated by the appended claims.
Claims
1. A method for generating a three-dimensional model using a machine learning model, the method comprising: Inputting the two-dimensional (2D) image into a machine learning model, wherein the machine learning model is configured to generate a three-dimensional (3D) model with accurate geometry and detailed texture; generating, by a first sub-model of the machine learning model, a set of multi-view images based at least in part on the 2D image, wherein the first sub-model includes a multi-level image hint controller configured to implement hierarchical control over the generation of the multi-view images by the first sub-model based at least in part on the input image; and A 3D model is generated by a second sub-model of the machine learning model based at least in part on the set of multi-view images, wherein the second sub-model is configured to perform background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model.
2. The method according to claim 1, wherein: The multi-level image hinting controller includes a global controller configured to enable the first sub-model to extract global structural information from the input image.
3. The method according to claim 2, further comprising: performing encoding on the 2D image by a contrastive language-image pre-trained CLIP encoder; generating, by the global controller, a vector based on the encoded image features; as well as A first control over generating the set of multi-view images is implemented by inputting the vector from the global controller into a criss-cross attention layer of the first sub-model.
4. The method according to claim 3, further comprising: adjusting the vector from the global controller by a multilayer perceptron; as well as The adjusted vector is input into the criss-cross attention layer.
5. The method according to claim 1, wherein: The multi-level image hinting controller further includes a local controller configured to enable the first sub-model to capture detailed structural information from the input image.
6. The method according to claim 5, further comprising: performing encoding on the 2D image by a contrastive language-image pre-trained CLIP encoder; Resampling the hidden features from the CLIP encoder by a resampling component of the local controller before global pooling and generating balanced local features; as well as A second control on generating the set of multi-view images is implemented by inputting the balanced local features into a cross-attention layer of the first sub-model.
7. The method according to claim 1, wherein: The multi-level image hinting controller further includes a pixel controller configured to enable the first sub-model to optimize a texture of the generated 3D model based on an appearance of the 2D image.
8. The method according to claim 7, further comprising: Performing variational autoencoder (VAE) encoding on the 2D image; as well as The VAE-encoded image features are embedded by integrating the pixel controller into the first sub-model so as to enable a 3D self-attention process between the 2D image and the set of multi-view images.
9. The method according to claim 1, wherein: The 2D image comprises a 2D object, wherein the set of multi-view images comprises a plurality of images of the 2D object from different viewing angles, and wherein the 3D model comprises a 3D object corresponding to the 2D object.
10. The method according to claim 9, further comprising: During optimization of the 3D model, the background of the 2D image is adjusted to match a randomly colored background used to distinguish the interior and exterior of the 3D object, thereby eliminating floating artifacts in the generated 3D model.
11. The method according to claim 9, further comprising: The geometric accuracy of the 3D object is improved by reducing the parameter sampling range of the camera parameters during the process of generating the 3D model.
12. The method according to claim 9, further comprising: The geometric accuracy of the 3D model is improved by applying canonical camera coordination across different object instances.
13. A system for generating a three-dimensional model using a machine learning model, the system comprising: at least one processor; as well as at least one memory, the at least one memory being communicatively coupled to the at least one processor and comprising computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: Inputting the two-dimensional (2D) image into a machine learning model, wherein the machine learning model is configured to generate a three-dimensional (3D) model with accurate geometry and detailed texture; generating, by a first sub-model of the machine learning model, a set of multi-view images based at least in part on the 2D image, wherein the first sub-model includes a multi-level image hint controller configured to implement hierarchical control over the generation of the multi-view images by the first sub-model based at least in part on the input image; and A 3D model is generated by a second sub-model of the machine learning model based at least in part on the set of multi-view images, wherein the second sub-model is configured to perform background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model.
14. The system according to claim 13, wherein: The multi-level image hinting controller includes a global controller configured to enable the first sub-model to extract global structural information from the input image, and wherein the operation further includes: performing encoding on the 2D image by a contrastive language-image pre-trained CLIP encoder; generating, by the global controller, a vector based on the encoded image features; and A first control over generating the set of multi-view images is implemented by inputting the vector from the global controller into a criss-cross attention layer of the first sub-model.
15. The system of claim 13, wherein: The multi-level image hinting controller further includes a local controller configured to enable the first sub-model to capture detailed structural information from the input image, and wherein the operations further include: performing encoding on the 2D image by a contrastive language-image pre-trained CLIP encoder; resampling the latent features from the CLIP encoder by a resampling component of the local controller before global pooling and generating balanced local features; and A second control on generating the set of multi-view images is implemented by inputting the balanced local features into a cross-attention layer of the first sub-model.
16. The system of claim 13, wherein: The multi-level image hinting controller further comprises a pixel controller configured to enable the first sub-model to optimize the texture of the generated 3D model based on the appearance of the 2D image, and wherein the operation further comprises: encoding the 2D image using a variational autoencoder (VAE); and The VAE-encoded image features are embedded by integrating the pixel controller into the first sub-model so as to enable a 3D self-attention process between the 2D image and the set of multi-view images.
17. The system of claim 13, the operations further comprising: During optimization of the 3D model, adjusting the background of the 2D image to match a randomly colored background used to distinguish between the interior and exterior of the 3D object; as well as During the process of generating the 3D model, a parameter sampling range of camera parameters is reduced.
18. A non-transitory computer-readable storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the processor performs operations comprising: Inputting the two-dimensional (2D) image into a machine learning model, wherein the machine learning model is configured to generate a three-dimensional (3D) model with accurate geometry and detailed texture; generating, by a first sub-model of the machine learning model, a set of multi-view images based at least in part on the 2D image, wherein the first sub-model includes a multi-level image hint controller configured to implement hierarchical control over the generation of the multi-view images by the first sub-model based at least in part on the input image; and A 3D model is generated by a second sub-model of the machine learning model based at least in part on the set of multi-view images, wherein the second sub-model is configured to perform background alignment and camera alignment to improve the quality and geometric accuracy of the generated 3D model.
19. The non-transitory computer readable storage medium of claim 18, in, The multi-level image hinting controller includes a global controller configured to enable the first sub-model to extract global structural information from the input image; Wherein, the multi-level image prompting controller further comprises a local controller, wherein the local controller is configured to enable the first sub-model to capture detailed structural information from the input image; and The multi-level image hinting controller further comprises a pixel controller, and the pixel controller is configured to enable the first sub-model to optimize the texture of the generated 3D model based on the appearance of the 2D image.
20. The non-transitory computer readable storage medium of claim 18, the operations further comprising: During optimization of the 3D model, adjusting the background of the 2D image to match a randomly colored background used to distinguish between the interior and exterior of the 3D object; as well as During the process of generating the 3D model, a parameter sampling range of camera parameters is reduced.