3D model generation method and device, computer storage medium and processor

By obtaining the prompt word set and using the multimodal diffusion model and Wonder3D model to convert 2D images into 3D mesh grids, the problem of 2D image conversion 3D models in the prior art is solved, and efficient and low-cost 3D model generation is achieved.

CN120047612APending Publication Date: 2025-05-27中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017968.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing 2D image conversion 3D model is based on parallax, resulting in complex implementation processes and inability to cope with large-scale production.

Method used

By obtaining the prompt word set, the initial 2D image is generated using the multimodal diffusion model, and then the 2D image is converted into a 3D mesh grid using the Wonder3D model, and on this basis, the 3D model is generated using 3D modeling software.

Benefits of technology

The process from 2D images to 3D models is simplified, manual intervention is reduced, technical difficulty and production costs are reduced, and production efficiency and model quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047612A_ABST
    Figure CN120047612A_ABST
Patent Text Reader

Abstract

The invention provides a 3D model generation method and device, a computer storage medium and a processor. The method comprises the steps of obtaining a current prompt word set; processing the current cue word set by adopting a multi-modal diffusion model to generate an initial 2D image; a Wonder3D model is adopted to process the initial 2D image, and a 3D mesh is generated; and generating a 3D model based on the 3D mesh by adopting 3D modeling software. Through the combination of the prompt word set, the diffusion model, the multi-modal generation technology and the Wonder3D model, the whole process from the 2D image to the 3D model is simplified, manual intervention is reduced, and the technical difficulty and the production cost are reduced. A user can generate a 3D model with a complex geometric structure and details only through short text description, so that the production efficiency and the model quality are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of 3D models. Specifically, it relates to a method for generating a 3D model, a device for generating a 3D model, a computer storage medium, and a processor. Background Technique

[0002] Many current methods for converting 2D images into 3D models are based on parallax. 3D modeling techniques mainly focus on 3D models generated by using 3D modeling software. This method has a relatively high technical complexity, high cost, complex processes, and low efficiency, and requires steps such as 3D modeling, UV unwrapping (texture mapping coordinates), texture production, bone binding, and rendering.

[0003] Many current methods for converting 2D images into 3D models are based on parallax. 3D modeling techniques mainly focus on 3D models generated by using 3D modeling software. This method has a relatively high technical complexity, high cost, complex processes, and low efficiency.

[0004] In the existing solutions, the conversion of 2D images into 3D models is based on parallax, which makes the implementation process relatively complex and unable to handle mass production. Summary of the Invention

[0005] The main purpose of this application is to provide a method for generating a 3D model, a device for generating a 3D model, a computer storage medium, and a processor, so as to at least solve the problem that in the existing solutions, the conversion of 2D images into 3D models is based on parallax, which makes the implementation process relatively complex and unable to handle mass production.

[0006] To achieve the above object, according to one aspect of this application, a method for generating a 3D model is provided, including: obtaining a current prompt set, where the current prompt set includes a plurality of current prompts, and the current prompts represent prompts for describing the 3D model to be generated; processing the current prompt set by using a multi-modal diffusion model to generate an initial 2D image; processing the initial 2D image by using a Wonder3D model to generate a 3D mesh; and generating a 3D model based on the 3D mesh by using 3D modeling software.

[0007] Optionally, after generating the initial 2D image, the method further includes: processing the initial 2D image by using a Controlnet model to obtain a final 2D image to enhance the details of the initial 2D image; processing the initial 2D image by using a Wonder3D model to generate a 3D mesh, including: processing the final 2D image by using a Wonder3D model to generate the 3D mesh.

[0008] Optionally, the initial 2D image is processed using the Controlnet model to obtain a final 2D image, including: performing zero convolution processing on the initial 2D image to obtain a first zero convolution result, and storing the first zero convolution result in an initial trainable copy of the Controlnet model to obtain a final trainable copy; performing zero convolution processing on the final trainable copy to obtain a second zero convolution result; processing the initial 2D image using a neural network model to obtain a neural network output result; and processing the neural network output result and the second zero convolution result using the Controlnet model to obtain the final 2D image.

[0009] Optionally, the Wonder3D model is used to process the final 2D image to generate the 3D mesh grid, including: processing the final 2D image using a cross-domain diffusion model to generate a consistent multi-view normal map and a corresponding color image; and extracting the 3D mesh grid from the consistent multi-view normal map and the corresponding color image using a normal fusion method.

[0010] Optionally, in the process of generating a 3D model based on the 3D mesh grid using 3D modeling software, the method further includes: extracting the skeleton features of the human pose using the controllable conditions of Openpose to make the 3D model have actions.

[0011] Optionally, the multi-modal diffusion model includes the Stable Diffusion SDXL model and the GPT4o model.

[0012] Optionally, using 3D modeling software to generate a 3D model based on the 3D mesh grid includes: processing the 3D mesh grid using a ray tracing image method and a gradient-based optimization algorithm to obtain an OBJ format file; and performing preset adjustments on the OBJ format file using the 3D modeling software to generate the 3D model, where the preset adjustments include adjustments to rendering colors and texture maps.

[0013] According to another aspect of the present application, a 3D model generation device is provided, including: an acquisition unit for acquiring a current prompt word set, the current prompt word set including a plurality of current prompt words, and the current prompt words representing prompt words for describing a 3D model to be generated; a first generation unit for processing the current prompt word set using a multi-modal diffusion model to generate an initial 2D image; a second generation unit for processing the initial 2D image using the Wonder3D model to generate a 3D mesh grid; and a third generation unit for generating a 3D model based on the 3D mesh grid using 3D modeling software.

[0014] According to another aspect of the present application, there is provided a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods for generating a 3D model.

[0015] According to yet another aspect of the present application, there is provided a processor for running a program. When the program runs, it executes any one of the methods for generating a 3D model.

[0016] Applying the technical solution of the present application, first, obtain the current set of prompt words. The current set of prompt words includes multiple current prompt words, and the current prompt words represent the prompt words used to describe the 3D model to be generated. Then, use a multi-modal diffusion model to process the current set of prompt words to generate an initial 2D image. Next, use the Wonder3D model to process the initial 2D image to generate a 3D mesh grid. Finally, use 3D modeling software to generate a 3D model based on the 3D mesh grid. By combining the set of prompt words, the diffusion model, the multi-modal generation technology, and the Wonder3D model, the entire process from 2D image to 3D model is simplified, manual intervention is reduced, and the technical difficulty and production cost are lowered. Users can generate 3D models with complex geometric structures and details only through a short text description, greatly improving the production efficiency and model quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0018] Figure 1 Shows a hardware structure block diagram of a mobile terminal for executing a method for generating a 3D model provided in an embodiment of the present application;

[0019] Figure 2 Shows a schematic flowchart of a method for generating a 3D model provided in an embodiment of the present application;

[0020] Figure 3 Shows a structure block diagram of a device for generating a 3D model provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0022] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of this application here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] For the convenience of description, some nouns or terms related to the embodiments of this application are explained below:

[0025] Generative Adversarial Networks (GANs): GANs consist of two parts, a Generator and a Discriminator. The task of the Generator is to create realistic images, while the task of the Discriminator is to distinguish whether the image is created by the Generator or is real. The two compete with each other and continuously improve the authenticity of the generated images.

[0026] Deep learning algorithms: Through a large amount of data training, deep learning models can learn the features of images and generate new images. These models can generate digital collectibles based on existing art styles or brand new ideas.

[0027] 3D model: It refers to a data format in computer graphics used to represent the geometric shape, appearance attributes, and layout structure of an object or scene in three-dimensional space. It simulates entities in the real world or a fictional world through mathematical algorithms and polygon meshes, thus providing three-dimensional information for visual presentation.

[0028] 3D Character Design: 3D character design refers to the process of creating and shaping three-dimensional, vivid digital characters through computer software technology in a three-dimensional space. Designers use 3D modeling software to construct the geometric shapes of objects, endowing them with material properties, colors, textures, and lighting effects, thereby creating three-dimensional models with a sense of realism or artistic style. These models can be static, such as physical models in product design, or dynamic, such as characters in animated movies and video games.

[0029] 3D Modeling Software: Refers to software that can perform various operations such as 3D modeling, animation, rendering, materials, lighting, etc., including Autodesk 3ds Max, Maya, Blender, SketchUp, Cinema 4D, SolidWorks, Rhino, etc.

[0030] Wonder3D Model: Wonder3D is a lightweight 3D engine based on WebGL and Three.js that can generate high-fidelity texture meshes from single-view images, a method for generating high-quality 3D models from a single image.

[0031] X-Dreamer Model: A high-quality 3D generation model based on diffusion models, which creates high-quality 3D assets by bridging the gap between Text-to-2D and Text-to-3D generation fields.

[0032] Stable Diffusion Model: A generative artificial intelligence (generative AI) model that can generate unique and realistic images based on text and image prompts.

[0033] LoRA Model (Low-Rank Adaptation of Large Language Models): A lightweight model fine-tuning training method that fine-tunes the model on the basis of the original image model, enabling it to generate specific characters, items, or painting styles.

[0034] As introduced in the background art, in the prior art, the conversion of 2D images to 3D models is completed based on parallax, which makes the implementation process relatively complex and unable to cope with the problem of mass production. To solve the problems existing in the prior art, an embodiment of the present application provides a method for generating 3D models.

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0036] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1It is a hardware block diagram of a mobile terminal for a method of generating a 3D model according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0037] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method of generating a 3D model in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0038] In this embodiment, a method for generating a 3D model running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0039] Figure 2 It is a schematic flowchart of a method for generating a 3D model provided according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0040] Step S201, obtain the current prompt set. The current prompt set includes multiple current prompts, and the current prompts represent prompts for describing the 3D model to be generated;

[0041] Specifically, by obtaining the current prompt set, multiple keywords or phrases for describing the target 3D model are defined. Each prompt represents the characteristics of the 3D model in different aspects, such as shape, texture, style, color, etc. The selection of the prompt set provides descriptive information and directional guidance for the subsequent generation process. The current prompt set can be manually input by the user, generated based on certain rules, or extracted from text descriptions through natural language processing algorithms. By using the prompt set, precise control of the 3D model generation process can be achieved, ensuring that the generated 3D model meets the user's requirements. Compared with traditional modeling methods, this step simplifies the input method. Users do not need to master complex modeling skills, and only need to provide a set of simple descriptive keywords to start the model generation process.

[0042] Step S202, process the current prompt set using a multimodal diffusion model to generate an initial 2D image;

[0043] The multimodal diffusion model is a model used to describe the diffusion and propagation of information in different modalities (such as sound, vision, touch, etc.) in a network. This model considers the interaction and influence between different modalities, and how they propagate and affect other nodes in the network. Through the multimodal diffusion model, researchers can better understand and predict the propagation and influence process of different modality information in various networks such as social networks and sensor networks.

[0044] Specifically, use a multimodal diffusion model (such as the CLIP model combined with diffusion generation technology) to process the prompt set. This model can not only understand text descriptions, but also combine image generation capabilities to generate an initial 2D image that conforms to the description based on the input prompt set. The core advantage of the multimodal diffusion model is that it can bridge the gap between different modalities (text and image), understand complex semantic information, and convert it into visual content. Using the multimodal diffusion model to generate the initial 2D image can accelerate the conversion process from text to image, reducing the steps that need to be manually drawn and designed in traditional 3D modeling. At the same time, the diffusion model can flexibly generate high-quality images according to the prompts, making the 3D modeling process more efficient and creative.

[0045] Step S203: Process the initial 2D image using the Wonder3D model to generate a 3D mesh grid;

[0046] The Wonder3D model is a 3D modeling technology. By using the Wonder3D software, 3D models with strong realism and rich details can be created. Such models can be used in fields such as game development, film and television production, virtual reality, etc., and can provide more accurate and realistic visual effects. The Wonder3D model can help designers and artists better express their creative ideas, improve work efficiency and save time costs.

[0047] A 3D mesh grid is a structure used to represent the surface of a 3D object. It consists of many small triangles or quadrilaterals, and each triangle or quadrilateral is called a face. These faces are connected together through their vertices to form a continuous surface. 3D mesh grids are usually used in computer graphics for creating and rendering 3D models. It can describe attributes such as the shape, curvature, and texture of an object, and can be processed and manipulated through different algorithms.

[0048] Specifically, the Wonder3D model is used in this step to convert the initial 2D image into a 3D mesh. This model can analyze the visual information in the 2D image, predict and generate mesh data with a 3D structure. The 3D mesh is the basis of a 3D model, which can provide geometric information such as point clouds, patches, and boundaries, and support subsequent 3D modeling and rendering. The application of the Wonder3D model effectively simplifies the conversion process from 2D to 3D. By automatically generating the mesh through deep learning and computer vision technologies, it avoids the cumbersome manual drawing and modeling steps in traditional 3D modeling, greatly improving the conversion efficiency and accuracy.

[0049] Step S204: Use 3D modeling software to generate a 3D model based on the 3D mesh grid.

[0050] Specifically, use professional 3D modeling software (such as Blender, Maya, etc.) to further refine and optimize the 3D mesh, and finally generate a complete 3D model. During this process, operations such as mesh repair, topology optimization, texture mapping, and bone binding can be performed to ensure that the generated 3D model meets the technical requirements and has high artistic expressiveness.

[0051] As can be seen, the embodiment of the present application provides a method for generating a 3D model. First, obtain the current prompt set, where the current prompt set includes multiple current prompts, and the current prompts represent the prompts used to describe the 3D model to be generated; then, use a multimodal diffusion model to process the current prompt set to generate an initial 2D image; then, use the Wonder3D model to process the initial 2D image to generate a 3D mesh grid; finally, use 3D modeling software to generate a 3D model based on the 3D mesh grid. By combining the prompt set, diffusion model, multimodal generation technology, and Wonder3D model, the entire process from 2D image to 3D model is simplified, manual intervention is reduced, and the technical difficulty and production cost are lowered. Users can generate 3D models with complex geometric structures and details through a short text description, greatly improving production efficiency and model quality.

[0052] As a possible implementation, after generating the initial 2D image, the method further includes the following steps:

[0053] Step S301, use the Controlnet model to process the initial 2D image to obtain a final 2D image to enhance the details of the initial 2D image;

[0054] The Controlnet model is a mathematical model used to describe and analyze control systems. It is usually composed of difference equations or state-space equations and is used to describe the relationship between the input, output, and state of the system. The Controlnet model is usually used for designing, analyzing, and optimizing control systems to achieve system stability, performance, and robustness. This model can be continuous-time or discrete-time, depending on the characteristics and requirements of the system.

[0055] Specifically, after generating the initial 2D image, use the Controlnet model to refine it. Controlnet is an image processing model based on deep learning that can enhance the details in the initial 2D image and improve the clarity and fineness of the image. By learning the structural, textural, and other detailed information in the image, Controlnet can generate a more detailed and realistic final 2D image. This step aims to enhance the quality of the initial image and provide richer visual information for subsequent 3D modeling.

[0056] Step S302, use the Wonder3D model to process the initial 2D image to generate a 3D mesh grid, including: using the Wonder3D model to process the final 2D image to generate a 3D mesh grid.

[0057] Specifically, the Wonder3D model processes the final 2D image (the image optimized by the Controlnet model) to generate a 3D mesh. The Wonder3D model can extract three-dimensional geometric information from the 2D image and construct a 3D mesh that conforms to visual representation. By providing more detailed information, the final 2D image enables the Wonder3D model to generate 3D mesh data with rich details more precisely. This ensures that the geometric structure of the 3D mesh is more in line with the physical properties of the real world, enhances the expressiveness and applicability of the model, reduces the need for manual repair and adjustment, and improves the modeling efficiency.

[0058] Thus, after generating the initial 2D image, further using the Controlnet model for detail enhancement processing improves the clarity and detail performance of the image, providing more accurate and rich visual information for subsequent 3D mesh generation. Processing the finally optimized 2D image with the Wonder3D model generates a 3D mesh that is not only more refined but also has a high-quality geometric structure, significantly improving the realism and accuracy of the generated 3D model. Compared with traditional methods, this method effectively improves the overall quality and efficiency of 3D model generation by introducing multiple optimization and refinement steps, can meet the requirements of more complex application scenarios, reduces the workload of subsequent manual correction, and enhances the reliability and usability of the automated modeling process.

[0059] As a possible implementation, using the Controlnet model to process the initial 2D image to obtain the final 2D image includes the following steps:

[0060] Step S401: Perform zero convolution processing on the initial 2D image to obtain the first zero convolution result, and store the first zero convolution result in the initial trainable copy of the Controlnet model to obtain the final trainable copy;

[0061] Zero convolution refers to a situation in convolution operation where there is no overlapping part between the convolution kernel and the input data, resulting in all values in the output result being zero. This usually occurs when the size of the convolution kernel is larger than the size of the input data, or when there is no overlap between the boundary part of the input data and the convolution kernel. Zero convolution may cause information loss or a decline in model performance in deep learning, so attention needs to be paid to avoiding the occurrence of zero convolution when designing convolutional neural networks.

[0062] A trainable copy of the ControlNet model refers to a model that uses the same architecture and parameters as the original ControlNet model but is trained on a different dataset. This copy model can improve the generalization ability and performance of the model by training on different datasets, thus adapting to a wider range of application scenarios. By training the trainable copy, the model can better adapt to different input data and improve the effectiveness and accuracy of the model in practical applications.

[0063] Specifically, first perform zero convolution on the initial 2D image. Zero convolution is a specific type of convolution operation mainly used to preserve the original structure of the image without introducing additional changes when processing images, and is usually used for image smoothing and detail enhancement. After the zero convolution process, the resulting first zero convolution result will be used as input and stored in the initial trainable copy of the ControlNet model, and then the final trainable copy of the model will be generated. After such processing, the ControlNet model can further optimize its learning process to adapt to more complex image features. Using the initial trainable copy for training enables the model to be more effectively optimized for specific tasks and reduces interference factors during the training process.

[0064] Step S402: Perform zero convolution on the final trainable copy to obtain a second zero convolution result;

[0065] Specifically, perform zero convolution on the final trainable copy to further optimize the performance of the model. Similar to the first zero convolution process, this step generates a second zero convolution result by applying the zero convolution operation to the model copy. Through two zero convolution processes, the training process of the model can be gradually refined, ultimately improving the quality of the image. Through two zero convolution operations, the training process of the model can be adjusted more meticulously, further enhancing the details and texture performance of the image. This method of multiple convolution processes can reduce information loss and make the finally output image more delicate and clear.

[0066] Step S403: Process the initial 2D image using a neural network model to obtain a neural network output result;

[0067] Specifically, use a neural network model to process the initial 2D image. Through its complex structure and hierarchical processing, the neural network can extract high-level features in the image, such as texture, shape, lighting changes, etc., and then generate the output result of the neural network. This result is a high-level representation of the image content and provides important visual information for subsequent processing. The neural network model can extract deeper and multi-level features from the initial 2D image, enhancing the image understanding ability. This process can not only improve the performance of image details but also provide richer semantic information for subsequent image generation.

[0068] Step S404: Use the Controlnet model to process the neural network output result and the second zero convolution result to obtain the final 2D image.

[0069] Specifically, the Controlnet model combines the neural network output result and the second zero convolution result and further processes them to generate the final 2D image. This processing step optimizes the details and accuracy of the image by integrating the deep features of the neural network and the smoothing effect of zero convolution, and finally obtains a 2D image with better detail performance. By combining the output of the neural network with the zero convolution result, the Controlnet model can better coordinate the detail performance and overall structure of the image and generate the final optimized 2D image. The fineness and accuracy of the image are significantly improved, reducing unnecessary noise or distortion in the image and making the image clearer and more natural.

[0070] It can be seen that by introducing multiple zero convolution processes and the deep feature extraction of the neural network model, the accuracy and detail performance of image processing are significantly improved. Through the multiple processes of zero convolution on the initial 2D image and the final trainable copy, information loss is reduced and the smoothness of the image is optimized; at the same time, the high-level features extracted by the neural network model provide a deep understanding for enhancing image details. With the combination of the two, the finally generated 2D image is not only richer and more realistic in details, but also has higher image quality, providing more accurate input data for subsequent 3D modeling and generation. This method improves the accuracy of the model, reduces the need for manual intervention, enhances the effect and efficiency of automatic image generation, and adapts to more complex image generation scenarios through the optimization and processing of multiple steps.

[0071] As a possible implementation, use the Wonder3D model to process the final 2D image to generate a 3D mesh, including the following steps:

[0072] Step S501: Use the cross-domain diffusion model to process the final 2D image to generate a consistent multi-view normal map and the corresponding color image;

[0073] The cross-domain diffusion model of the Wonder3D model is a model used to transfer knowledge and information between different domains, which can help the model perform effective transfer learning and knowledge sharing between different domains. The normal fusion method of the Wonder3D model is a method used to merge the normal information of different surfaces in order to better present the surface details and textures of the 3D model. This method can help improve the realism and visual effect of the model.

[0074] Specifically, the cross - domain diffusion model is applied to the final 2D image. The cross - domain diffusion model can map image data from one domain to another, and is typically used to handle complex features such as perspective and depth information in images. Through this method, the model generates consistent multi - perspective normal maps and corresponding color images. The multi - perspective normal map is an image representing the surface normals (i.e., surface orientations) of an object, presenting the 3D structure of the object from different perspectives. The corresponding color image helps to retain the texture information of the object, thus providing more detailed visual information for subsequent 3D mesh generation. By generating consistent multi - perspective normal maps, the 3D surface details of the object can be accurately represented, and consistent normal data can be obtained when observing the same object from different perspectives. This makes the subsequent 3D mesh generation process more stable and efficient. In addition, the generation of the color image retains rich texture information, making the final 3D model have higher visual quality and detail performance, thereby improving the accuracy of the image generation and modeling process.

[0075] Step S502: Use the normal fusion method to extract a 3D mesh from the consistent multi - perspective normal map and the corresponding color image.

[0076] Specifically, use the normal fusion method to extract a 3D mesh from the consistent multi - perspective normal map and the corresponding color image. The normal fusion method combines the normal information from multiple perspectives to ensure that the generated 3D mesh performs consistently under different perspectives and can accurately represent the geometric shape of the object surface. The color image provides texture information, making the 3D mesh not only accurate in shape but also more realistic in color and texture. By fusing the normal map and the color image, the generated 3D mesh will have a complete geometric structure and texture details. The normal fusion method can ensure that the generated 3D mesh performs consistently under different perspectives, thus avoiding geometric inconsistencies between different perspectives. This method improves the accuracy and consistency of the generated 3D mesh, reducing errors or distortions caused by perspective changes. At the same time, fusing the texture information of the color image makes the 3D mesh more realistic, capable of showing higher details and quality in rendering or subsequent applications.

[0077] It can be seen that by adopting a cross-domain diffusion model to generate consistent multi-view normal maps and corresponding color images, the 3D mesh generated from the final 2D image can have higher accuracy and consistency. The application of the cross-domain diffusion model ensures the consistency of normal maps from multiple views and provides detailed texture information for subsequent 3D mesh generation. Through the normal fusion method, an accurate 3D mesh can be extracted from the normal maps and color images of multiple views, ensuring that the generated mesh is both accurate in shape and rich in texture details. This method significantly improves the efficiency and accuracy of 3D mesh generation, reduces the manual adjustment and correction work in the traditional modeling process, and provides an efficient, automated, and high-quality 3D modeling solution. Finally, the generated 3D mesh not only has high-quality geometric details but also can more realistically reflect the texture and appearance of the object, meeting the requirements of more complex application scenarios.

[0078] As a possible implementation method, in the process of using 3D modeling software to generate a 3D model based on a 3D mesh grid, the method further includes the following steps:

[0079] Adopt the controllable conditions of Openpose to extract the skeleton features of the human body posture, so that the 3D model has actions.

[0080] The controllable conditions of Openpose refer to that users can control the behavior and output results of Openpose by adjusting parameters according to their own needs and application scenarios. These controllable conditions include but are not limited to: 1. The number and types of human key points detected: Users can choose the number and types of human key points to be detected, such as only detecting the key points of the upper body or the whole body. 2. Confidence threshold: Users can set the confidence threshold to only retain the key points with a confidence higher than the threshold, thereby reducing the situation of false detection. 3. Output result format: Users can choose the output result format, such as JSON format, XML format, etc. 4. Input data format: Users can choose the input data format, such as pictures, videos, camera live streams, etc. 5. Running mode: Users can choose the running mode, including single-person detection, multi-person detection, real-time detection, etc. By adjusting these controllable conditions, users can optimize the performance and output results of Openpose according to specific needs and scenarios. For example, controllable conditions such as Canny, Depth, HED, MLSD, Normal, Openpose, Scribble, Seg, etc. can be selected. If users want the finally generated 3D model to have actions, they can introduce the controllable conditions of Openpose to extract the skeleton features of the human body posture.

[0081] Specifically, the Openpose model is used to extract the skeletal features of human postures. Openpose is a deep learning-based keypoint detection technology that can detect and extract the skeletal structure features of the human body in real time in images or videos, such as the positions and angles of each joint. The Openpose model extracts the skeletal features of the human body according to controllable conditions (such as set actions or postures) and combines these features with the generated 3D mesh, thus endowing the 3D model with certain actions or dynamic performances. This method can not only capture the static form of the human body, but also introduce action information in the 3D modeling process, making the generated 3D model have higher realism and action performance capabilities. By introducing the Openpose model, the human postures and skeletal features can be efficiently and accurately extracted and combined with the 3D mesh, enabling the 3D model to show dynamic effects and posture changes. This technology adds action elements to the traditional static 3D modeling process, enhances the expressiveness of the 3D model, provides richer action capture data for subsequent applications such as animation and games, and improves the practicality and visual expressiveness of the model.

[0082] As a possible implementation, the multimodal diffusion model includes the Stable Diffusion SDXL model and the GPT4o model.

[0083] Specifically, the Stable Diffusion SDXL model: Stable Diffusion is a deep generative model based on the diffusion process and is widely used in image generation tasks. SDXL (Stable Diffusion eXtra Large) is an improved version of this model, with stronger generative capabilities and higher image quality, and can generate more detailed and high-resolution 2D images. In this application, the Stable Diffusion SDXL model is used to generate an initial 2D image according to a given set of prompt words, ensuring that the generated image meets the expected visual effects. The GPT4o model: The GPT4o is an optimized version based on the GPT-4 architecture and is specifically used for the processing of multimodal tasks, with powerful text understanding and generation capabilities. This model can understand and generate descriptions and prompts related to the image content, providing higher-level semantic information for the diffusion process. The GPT4o model optimizes the generation process through the fusion of text and images, ensuring that the 2D image and the subsequent 3D model are highly consistent with the semantic content in the given set of prompt words.

[0084] It can be seen that by combining the Stable Diffusion SDXL model and the GPT4o model, this application can leverage the advantages of both models to improve the quality and accuracy of image generation. The Stable Diffusion SDXL model can generate high-quality 2D images, while the GPT4o model can provide strong semantic support and intelligent reasoning capabilities to ensure that the generated images are more in line with the content of the text prompt. This multimodal combination method improves the diversity, accuracy, and flexibility of the generation process, greatly enhancing the effect of the entire modeling process.

[0085] As a possible implementation method, a 3D modeling software is used to generate a 3D model based on a 3D mesh grid, including the following steps:

[0086] Step S601: Process the 3D mesh grid using the method of ray tracing images and a gradient-based optimization algorithm to obtain an OBJ format file;

[0087] The method of ray tracing images (nerf) and the gradient-based optimization algorithm (instant-nsr-plx) are both methods used for neural network training.

[0088] Nerf (Neural Radiance Fields) is a method for generating high-quality ray tracing images. It trains a neural network to predict the view and color of each point in a scene and uses this information to generate realistic images. The nerf method performs well in rendering complex lighting and material effects and is widely used in the fields of computer graphics and computer vision. Instant-nsr-plx is a method for neural network training and is a gradient-based optimization algorithm. This method updates the weights and parameters of the network by using the backpropagation algorithm of the neural network to minimize the loss function and improve the performance of the network. The instant-nsr-plx method has good convergence and efficiency when training deep neural networks and is widely used in various deep learning tasks.

[0089] Specifically, first, the 3D mesh is processed using ray tracing technology. Ray tracing is a computational method for simulating the propagation of light. It calculates the effects of light reflection, refraction, shadow, etc. by tracing the process of light intersecting with the object surface, thereby generating highly realistic images. In 3D modeling, ray tracing can be used to render the lighting, shadow, and reflection effects of objects, making the generated 3D images more realistic. At the same time, a gradient-based optimization algorithm is applied to the optimization process of the 3D mesh. The gradient optimization algorithm adjusts the positions of the mesh points iteratively to minimize the error and optimize the shape of the mesh, enabling the mesh to present a smoother and more realistic three-dimensional shape while maintaining details. The processed mesh information is finally saved as an OBJ format file. OBJ is a common 3D file format that can contain information such as meshes, vertices, and texture coordinates, facilitating subsequent processing and applications.

[0090] Step S602, use 3D modeling software to perform preset adjustments on the OBJ format file to generate a 3D model. The preset adjustments include adjustments to the rendering color and texture mapping.

[0091] Specifically, the obtained OBJ format file will be preset adjusted through 3D modeling software, mainly including adjustments to the rendering color and texture mapping. The rendering color refers to adjusting the color of the 3D model surface to make it more in line with the expected visual effect. The adjustment of texture mapping involves applying texture images (such as surface materials, patterns, etc.) to different areas of the 3D mesh to increase the details and realism of the model. These adjustments will make the 3D model more delicate and rich in visual performance, especially in terms of surface details and material effects. By adjusting the rendering color and texture mapping of the OBJ format file, the appearance quality of the 3D model can be greatly improved, making it more in line with the physical properties of the real world. For example, through texture mapping, the surface of the model can present different material effects, such as wood grain, metallic texture, cloth texture, etc. These details make the final 3D model more realistic. In addition, the color rendering adjustment ensures that the appearance color of the model meets the design requirements, enhancing the visual effect and expressiveness. This preset adjustment process effectively improves the visual realism and artistry of the 3D model, ensuring its effect in subsequent applications.

[0092] It can be seen that by using the ray tracing image method and the gradient-based optimization algorithm to process the 3D mesh grid, it is ensured that the generated 3D grid has higher geometric accuracy and lighting expressiveness. Ray tracing enhances the realism and details of the image, while the gradient optimization algorithm optimizes the mesh structure to make it smoother and more accurate, thus generating a high-quality OBJ format file. On this basis, through preset adjustments in 3D modeling software, including adjustments to rendering colors and texture maps, the appearance and detail performance of the 3D model are further enhanced, making it more in line with actual needs and having stronger visual impact and realism. This method not only improves the accuracy and quality of 3D model generation, but also greatly enhances the visual effect of the model, meeting the higher requirements of modeling and rendering needs.

[0093] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0094] The embodiment of the present application also provides a 3D model generation device. It should be noted that the 3D model generation device of the embodiment of the present application can be used to execute the 3D model generation method provided by the embodiment of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0095] The following introduces the 3D model generation device provided by the embodiment of the present application.

[0096] Figure 3 is a structural block diagram of a 3D model generation device provided according to an embodiment of the present application. As Figure 3 shown, the device includes: an acquisition unit 10, a first generation unit 20, a second generation unit 30, and a third generation unit 40.

[0097] The acquisition unit 10 is used to acquire the current prompt word set, the current prompt word set includes a plurality of current prompt words, and the current prompt words represent the prompt words used to describe the 3D model to be generated;

[0098] The first generation unit 20 is used to process the current prompt word set by using a multimodal diffusion model to generate an initial 2D image;

[0099] The second generation unit 30 is used to process the initial 2D image by using the Wonder3D model to generate a 3D mesh grid;

[0100] A third generation unit 40, configured to generate a 3D model based on a 3D mesh using 3D modeling software.

[0101] As a possible implementation, the 3D model generation device further includes:

[0102] An enhancement unit, configured to process an initial 2D image using a Controlnet model to obtain a final 2D image, so as to enhance the details of the initial 2D image;

[0103] The second generation unit is specifically configured to process the final 2D image using a Wonder3D model to generate a 3D mesh.

[0104] As a possible implementation, the enhancement unit includes:

[0105] A first zero convolution module, configured to perform zero convolution processing on the initial 2D image to obtain a first zero convolution result, and store the first zero convolution result in an initial trainable copy of the Controlnet model to obtain a final trainable copy;

[0106] A second zero convolution module, configured to perform zero convolution processing on the final trainable copy to obtain a second zero convolution result;

[0107] A result output module, configured to process the initial 2D image using a neural network model to obtain a neural network output result;

[0108] An acquisition module, configured to process the neural network output result and the second zero convolution result using the Controlnet model to obtain a final 2D image.

[0109] As a possible implementation, the enhancement unit includes:

[0110] A processing module, configured to process the final 2D image using a cross-domain diffusion model to generate a consistent multi-view normal map and a corresponding color image;

[0111] An extraction module, configured to extract a 3D mesh from the consistent multi-view normal map and the corresponding color image using a normal fusion method.

[0112] As a possible implementation, the 3D model generation device further includes:

[0113] A skeleton feature extraction unit, configured to extract the skeleton features of a human pose using the controllable conditions of Openpose, so that the 3D model has actions.

[0114] As a possible implementation, the third generation unit includes:

[0115] A mesh processing module, which is used to process a 3D mesh using a ray tracing image method and a gradient-based optimization algorithm to obtain an OBJ format file;

[0116] An adjustment module, which is used to perform preset adjustments on the OBJ format file using 3D modeling software to generate a 3D model. The preset adjustments include adjustments of rendering colors and texture mapping.

[0117] The 3D model generation device includes a processor and a memory. The above-mentioned acquisition unit, first generation unit, second generation unit, third generation unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; alternatively, the above modules are separately located in different processors in any combination form.

[0118] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set. By adjusting the kernel parameters, the problem that the existing solution uses parallax to complete the conversion of 2D images to 3D models, making the implementation process relatively complex and unable to handle mass production, can be solved.

[0119] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0120] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the 3D model generation method.

[0121] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, it executes the 3D model generation method.

[0122] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. The device herein can be a server, PC, PAD, mobile phone, etc. When the processor executes the program, it implements at least the steps of the 3D model generation method.

[0123] This application also provides a computer program product, which is suitable for executing a program initialized with at least the steps of the 3D model generation method when executed on a data processing device.

[0124] Obviously, those skilled in the art should understand that each module or step of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0125] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0126] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions in the flow Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0129] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0130] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0131] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0132] It should also be noted that the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "includes a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements. From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects:

[0133] 1), The embodiments of the present application provide a method for generating a 3D model. By combining a prompt set, a diffusion model, multi-modal generation technology, and the Wonder3D model, the entire process from 2D images to 3D models is simplified, manual intervention is reduced, and the technical difficulty and production cost are lowered. Users can generate 3D models with complex geometric structures and details through a short text description, greatly improving production efficiency and model quality.

[0134] 2), The embodiments of the present application provide a device for generating a 3D model. The device includes: an acquisition unit 10, a first generation unit 20, a second generation unit 30, and a third generation unit 40. By combining a prompt set, a diffusion model, multi-modal generation technology, and the Wonder3D model, the entire process from 2D images to 3D models is simplified, manual intervention is reduced, and the technical difficulty and production cost are lowered. Users can generate 3D models with complex geometric structures and details through a short text description, greatly improving production efficiency and model quality.

[0135] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for generating a 3D model, characterized in that: include: Acquire a current prompt word set, wherein the current prompt word set includes a plurality of current prompt words, and the current prompt words represent prompt words used to describe a 3D model to be generated; Processing the current prompt word set using a multimodal diffusion model to generate an initial 2D image; The initial 2D image is processed using the Wonder3D model to generate a 3D mesh; A 3D modeling software is used to generate a 3D model based on the 3D mesh.

2. The method according to claim 1, characterized in that After generating the initial 2D image, the method further includes: processing the initial 2D image using a Controlnet model to obtain a final 2D image to enhance the details of the initial 2D image; The method of processing the initial 2D image by using the Wonder3D model to generate a 3D mesh includes: processing the final 2D image by using the Wonder3D model to generate the 3D mesh.

3. The method according to claim 2, characterized in that The initial 2D image is processed using the Controlnet model to obtain a final 2D image, including: Performing zero convolution processing on the initial 2D image to obtain a first zero convolution result, and storing the first zero convolution result in an initial trainable copy of the Controlnet model to obtain a final trainable copy; Performing zero convolution processing on the final trainable copy to obtain a second zero convolution result; Processing the initial 2D image using a neural network model to obtain a neural network output result; The Controlnet model is used to process the neural network output result and the second zero convolution result to obtain the final 2D image.

4. The method according to claim 2, characterized in that: The final 2D image is processed using the Wonder3D model to generate the 3D mesh, including: Processing the final 2D image using a cross-domain diffusion model to generate a consistent multi-view normal map and a corresponding color image; The 3D mesh is extracted from the consistent multi-view normal map and the corresponding color image using a normal fusion method.

5. The method according to claim 1, characterized in that In the process of using 3D modeling software to generate a 3D model based on the 3D mesh, the method further includes: The controllable conditions of Openpose are used to extract the skeleton features of the human body posture so that the 3D model can have movements.

6. The method according to claim 1, characterized in that The multimodal diffusion model includes a StableDiffusion SDXL model and a GPT4o model.

7. The method according to any one of claims 1 to 6, characterized in that Using 3D modeling software, based on the 3D mesh, a 3D model is generated, including: The 3D mesh is processed by using a ray tracing image method and a gradient-based optimization algorithm to obtain an OBJ format file; The 3D modeling software is used to perform preset adjustments on the OBJ format file to generate the 3D model, wherein the preset adjustments include adjustments to rendering colors and texture maps.

8. A 3D model generation device, characterized in that: include: An acquisition unit, configured to acquire a current prompt word set, wherein the current prompt word set includes a plurality of current prompt words, and the current prompt words represent prompt words used to describe a 3D model to be generated; A first generating unit, configured to process the current prompt word set using a multimodal diffusion model to generate an initial 2D image; A second generating unit is used to process the initial 2D image using a Wonder3D model to generate a 3D mesh; The third generating unit is used to generate a 3D model based on the 3D mesh by using 3D modeling software.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for generating a 3D model according to any one of claims 1 to 7.

10. A 3D model generation system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the 3D model generation method described in any one of claims 1 to 7.