Three-dimensional model generation method and apparatus, and device, medium and program product

By using multi-view image generation models and normal map optimization techniques, the problem of time-consuming and labor-intensive generation of 3D models in existing technologies has been solved, achieving efficient and low-cost 3D model generation.

WO2026081558A1PCT designated stage Publication Date: 2026-04-23BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-07-01
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

In existing technologies, generating 3D models requires the participation of professionals, which is time-consuming and labor-intensive, increasing the generation cost.

Method used

By acquiring the first image and texture description information of the target object, a second image and normal map from different perspectives are generated using a multi-view image generation model. 3D reconstruction and optimization are then performed, and finally, texture mapping is applied to generate a high-quality 3D model.

Benefits of technology

It enables the rapid generation of high-quality 3D models without the need for professional personnel, reducing generation costs and improving convenience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106415_23042026_PF_FP_ABST
    Figure CN2025106415_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a three-dimensional model generation method and apparatus, and a device, a medium and a program product. The three-dimensional model generation method comprises: acquiring a first image and texture description information of a target object; on the basis of a multi-view image generation model and the first image, generating a second image of the target object at each preset view and a normal map corresponding to each second image; on the basis of a plurality of second images, performing three-dimensional reconstruction, so as to obtain a first three-dimensional model of the target object; on the basis of the normal maps, optimizing the first three-dimensional model, so as to obtain an optimized second three-dimensional model; and on the basis of the first image and the texture description information, performing texture mapping on the second three-dimensional model, so as to obtain a target three-dimensional model of the target object. By means of the technical solution in the embodiments of the present disclosure, the convenience of generating a three-dimensional model can be improved, thereby reducing the generation cost of the three-dimensional model, and also ensuring the generation quality of the three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

3D model generation methods, apparatus, equipment, media and program products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411441264.0, filed on October 15, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to a method, apparatus, device, medium, and program product for generating three-dimensional models. Background Technology

[0004] With the rapid development of computer technology, it is often necessary to reconstruct 3D objects to obtain their 3D models. Some methods involve using specialized imaging equipment to scan and photograph the object in 3D, generating a 3D model based on multiple images. However, this method requires professional personnel, is time-consuming and labor-intensive, and increases the cost of generating the 3D model. Summary of the Invention

[0005] This disclosure provides a method, apparatus, device, medium, and program product for generating three-dimensional models, thereby improving the convenience of generating three-dimensional models, reducing the cost of generating three-dimensional models, and ensuring the quality of generated three-dimensional models.

[0006] In a first aspect, embodiments of this disclosure provide a method for generating a three-dimensional model, including:

[0007] Obtain the first image and texture description information of the target object;

[0008] Based on the multi-view image generation model and the first image, a second image of the target object under each preset view and the normal map corresponding to the second image are generated.

[0009] Based on multiple second images, a three-dimensional reconstruction is performed to obtain a first three-dimensional model of the target object;

[0010] Based on the normal map, the first three-dimensional model is optimized to obtain the optimized second three-dimensional model;

[0011] Based on the first image and the texture description information, texture mapping is performed on the second 3D model to obtain the target 3D model of the target object.

[0012] Secondly, embodiments of this disclosure also provide a three-dimensional model generation apparatus, comprising:

[0013] The information acquisition module is used to acquire the first image and texture description information of the target object;

[0014] A multi-view image generation module is used to generate a second image of the target object under each preset view and a normal map corresponding to the second image, based on a multi-view image generation model and the first image;

[0015] A 3D reconstruction module is used to perform 3D reconstruction based on multiple second images to obtain a first 3D model of the target object;

[0016] The model optimization module is used to optimize the first three-dimensional model based on the normal map to obtain an optimized second three-dimensional model.

[0017] The texture mapping module is used to perform texture mapping on the second three-dimensional model based on the first image and the texture description information to obtain the target three-dimensional model of the target object.

[0018] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0019] One or more processors;

[0020] Storage device for storing one or more programs.

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional model generation method as described in any embodiment of this disclosure.

[0022] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the three-dimensional model generation method as described in any of the embodiments of this disclosure.

[0023] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the three-dimensional model generation method as described in any of the embodiments of this disclosure. Attached Figure Description

[0024] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0025] Figure 1 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure;

[0026] Figure 2 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure;

[0027] Figure 3 is an example diagram of the network architecture of a multi-view image generation model according to an embodiment of this disclosure;

[0028] Figure 4 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure;

[0029] Figure 5 is an example diagram of a first three-dimensional model generation process according to an embodiment of this disclosure;

[0030] Figure 6 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure;

[0031] Figure 7 is a schematic diagram of the structure of a three-dimensional model generation device provided in an embodiment of this disclosure; and

[0032] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0035] The term "comprising" and its variations as used in this disclosure are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions for other terms will be given in the description below.

[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0037] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0038] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0039] It is understood that the data involved in this disclosure (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and provisions.

[0040] Figure 1 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure. This embodiment is applicable to generating three-dimensional models of any object such as items. The method can be executed by a three-dimensional model generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0041] As shown in Figure 1, the three-dimensional model generation method specifically includes the following steps:

[0042] S110. Obtain the first image and texture description information of the target object.

[0043] The target object can refer to a single object for which a 3D model needs to be generated. For example, the target object can be any object such as an item or a person. The first image can refer to a color image of the target object taken from any viewpoint. The viewpoint of the first image is not limited. Texture description information can refer to text information used to describe the texture of the target object. The first image and texture description information describe the presentation style of the target object using image and text methods, respectively.

[0044] Specifically, the system acquires any color image containing the target object taken by the user using a regular camera, along with the input texture description information of the target object. Background removal can be performed on the acquired color image to obtain a first image containing only the target object, thus avoiding the influence of background objects. A higher-quality 3D model can then be reconstructed based on this first image containing only the target object.

[0045] S120. Based on the multi-view image generation model and the first image, generate a second image of the target object under each preset view and a normal map corresponding to the second image.

[0046] The preset viewpoints are pre-set perspectives, and there can be multiple preset viewpoints. For example, preset viewpoints include perspectives taken from six directions: top, bottom, left, right, front, and back of the target object. The second image refers to the generated color image of the target object from the preset viewpoints. There is a one-to-one correspondence between the second image and the preset viewpoint. There is also a one-to-one correspondence between the second image and the normal map. The normal map can include the normal information corresponding to each pixel in the second image. The normal map can be used to represent the geometric shape information of the target object from the preset viewpoints. The multi-view image generation model can be a deep learning network model used to generate multiple new viewpoint images and normal maps based on a single image. The multi-view image generation model can be obtained by fine-tuning a generative pre-trained model based on a sample dataset. The sample dataset can include sample images of the sample object from any viewpoint, as well as standard images and standard normal maps from each preset viewpoint. The generative pre-trained model can be any pre-trained model used to generate images. For example, the generative pre-trained model can refer to a diffusion model, such as the Stable Diffusion latent diffusion model.

[0047] Specifically, a first image of the target object is input into a multi-view image generation model to generate multiple images from preset viewpoints. Based on the output of the multi-view image generation model, a second image of the target object under each preset viewpoint and a corresponding normal map for each second image are obtained. By utilizing the multi-view image generation model, multiple second images and normal maps can be generated simultaneously and quickly, thereby obtaining multiple images from different viewpoints and richer geometric shape information. Furthermore, the normal map can effectively improve the precision of geometric modeling.

[0048] S130. Perform three-dimensional reconstruction based on multiple second images to obtain the first three-dimensional model of the target object.

[0049] The first 3D model can refer to the initially reconstructed 3D mesh model. The first 3D model is a relatively coarse 3D mesh model.

[0050] Specifically, after generating multiple second images from different perspectives, the initial reconstruction of the target object's 3D model can be performed using these images to obtain the target object's first 3D model. For example, by using a hybrid 3D representation of triplanes and Gaussians, the first 3D model of the target object can be reconstructed from multiple second images with high quality and speed.

[0051] S140. Based on the normal map, optimize the first three-dimensional model to obtain the optimized second three-dimensional model.

[0052] The second 3D model can refer to the optimized 3D mesh model. The second 3D model is a more refined 3D mesh model, with richer details and higher precision than the first 3D model.

[0053] Specifically, based on the normal map corresponding to each preset viewpoint, the first 3D model lacking details is optimized to recover the geometric details on the normal map, thereby obtaining a second 3D model with more refined and richer details. By using the normal map to optimize the details of the coarse first 3D model, the problem of difficulty in recovering details under sparse viewpoint reconstruction is solved.

[0054] For example, step S140 may include: using a continuous reconstructed mesh method, performing mesh smoothing optimization on the first three-dimensional model based on the normal map to obtain an optimized second three-dimensional model.

[0055] Specifically, by using the continuous-remeshing method to reconstruct the mesh, the mesh of the first 3D model is simplified based on the normal map, degenerate triangles on the mesh are removed, and the neighborhood area of ​​the initial mesh vertices is increased, thereby enhancing the normal fitting ability of the 3D model, restoring the geometric details on the normal map, and improving the detail refinement of the second 3D model.

[0056] S150. Based on the first image and texture description information, perform texture mapping on the second three-dimensional model to obtain the target three-dimensional model of the target object.

[0057] Specifically, based on the first image and texture description information, a target texture image of the target object can be generated more accurately. This target texture image is then applied to the second 3D model, giving the second 3D model its presentation style and obtaining the final target 3D model of the target object. Since a 3D model of the target object can be generated efficiently and with high quality based on only one image, eliminating the need to scan and capture multiple images, it simplifies user operation, improves the convenience of 3D model generation, and enhances the geometric rationality, fidelity, and detail of the 3D model. Compared to the traditional modeling time measured in days, this embodiment can reduce modeling time to minutes, significantly improving the efficiency of 3D model generation.

[0058] The technical solution of this disclosure generates a first 3D model of the target object from a multi-view image generation model and a first image of the target object. This generates a second image of the target object from each preset viewpoint and a normal map corresponding to the second image. Based on multiple second images, 3D reconstruction is performed to obtain a first 3D model of the target object. The first 3D model is then optimized based on the normal map corresponding to the second image to obtain a higher-quality second 3D model. Finally, based on the first image and texture description information, a more realistic texture mapping is applied to the second 3D model to obtain a higher-quality target 3D model of the target object, thus ensuring the quality of the generated 3D model. By requiring only a single first image from any viewpoint, without the need for 3D scanning, and eliminating the need for professional personnel, the convenience of 3D model generation is improved, and the cost of 3D model generation is reduced.

[0059] Based on the above technical solution, the multi-view image generation model in step S120 is generated by fine-tuning the generative pre-trained model in advance based on the sample dataset and the target loss function. The target loss function includes a perceptual loss function and a noise loss function. The perceptual loss function is used to characterize the loss value between the sample denoised image and the predicted denoised image. The sample denoised image is determined based on the sample denoised image and the sample noise, and the predicted denoised image is determined based on the sample denoised image and the predicted noise. The noise loss function is used to characterize the loss value between the sample noise and the predicted noise.

[0060] The sample dataset can include sample images of the sample object from any viewpoint, as well as standard images and standard normal maps from each preset viewpoint. The generative pre-trained model can be any pre-trained model used to generate images. For example, a generative pre-trained model can refer to a diffusion model, such as the Stable Diffusion latent diffusion model.

[0061] Here, sample noise refers to randomly generated image noise at the current moment. Prediction noise refers to the noise predicted by the generative pre-trained model. A sample-denoised image is obtained by adding sample noise to a sample image. A sample-denoised image is obtained by removing sample noise from the sample-denoised image. A prediction-denoised image is obtained by removing prediction noise from the sample-denoised image. A one-step denoising method allows for faster acquisition of both sample-denoised and prediction-denoised images, reducing computational cost. The perceptual loss function can be used to characterize prediction errors in viewpoint perception. The noise loss function can be used to characterize prediction errors in noise.

[0062] Specifically, the generative pre-trained model can simultaneously output color images and normal maps from a preset viewpoint. Since both color images and normal maps are in image format, their output training processes are identical, allowing for independent training. For example, the fine-tuning training process of the generative pre-trained model when outputting either a color image or a normal map is as follows: A perceptual loss value is determined based on the perceptual loss function, the denoised sample image, and the predicted denoised image; a noise loss value is determined based on the noise loss function, the sample noise, and the predicted noise; the perceptual loss value and the noise loss value are weighted and summed to obtain the target loss value; this target loss value is then backpropagated to the generative pre-trained model, adjusting the network parameters until a preset convergence condition is met, such as the number of iterations equaling a preset number, or the loss value becoming stable. At this point, the fine-tuning of the generative pre-trained model is considered complete, resulting in a multi-view image generation model. By using the perceptual loss function to determine the perceptual loss value for training, perceptual constraints between multiple viewpoints can be introduced, leading to higher quality and content consistency in the final generated multi-view images.

[0063] It should be noted that generative pre-trained models generate new images by using the input image as a conditional image and performing multiple denoising operations on randomly generated noisy images. When the generative pre-trained model is a latent diffusion model, the fine-tuning training process requires encoding the sample images to obtain sample images in the latent space, and then performing image training in the latent space. That is, the sample noisy image, the sample denoised image, and the predicted denoised image are all image data represented in the latent space.

[0064] Figure 2 is a flowchart illustrating a three-dimensional model generation method provided in this embodiment. Based on the above-described embodiments, this embodiment optimizes the step of "generating a second image of the target object under each preset viewpoint and a normal map corresponding to the second image based on a multi-view image generation model and a first image." Explanations of terms identical or corresponding to those in the above-described embodiments are not repeated here.

[0065] As shown in Figure 2, the three-dimensional model generation method specifically includes the following steps:

[0066] S210. Obtain the first image and texture description information of the target object.

[0067] S220. Input the first image into the image encoder in the multi-view image generation model for feature extraction to obtain the feature information of the first image.

[0068] The image encoder can be any encoder used to extract image features. Specifically, referring to Figure 3, the feature information of the first image can be extracted by utilizing the image encoder in the multi-view image generation model.

[0069] S230. Input the feature information of the first image and the initial image and initial normal map corresponding to each preset viewpoint into multiple denoising networks cascaded in the multi-view image generation model for multiple iterations of denoising, and obtain the second image of the target object under each preset viewpoint and the normal map corresponding to the second image after the last iteration of denoising.

[0070] The initial image can be a noisy color image generated for each preset viewpoint, which can be initialized using random Gaussian noise. The initial normal map can be a noise normal map generated for each preset viewpoint. Multiple denoising networks are used, so that each network performs denoising processing once. Multiple denoising networks are cascaded, with the output of the current denoising network used as the input of the next denoising network, as shown in Figure 3.

[0071] Specifically, by utilizing multiple cascaded denoising networks, the feature information of the first image can be used as conditional information to perform multiple iterations of denoising on the initial image and initial normal map corresponding to each preset viewpoint. The result of the last iteration of denoising is then used as the final generated second image and normal map for each preset viewpoint. Multiple iterations of denoising improve the generation quality of the second image and normal map, thereby improving the generation quality of the 3D model.

[0072] It should be noted that when the generative pre-trained model is a latent diffusion model, the input initial image and initial normal map are both encoded representations of the initial image and initial normal map in the latent space. The final result obtained by the iteration is also a second image and normal map represented in the latent space. Therefore, it is necessary to decode the second image and normal map represented in the latent space to obtain the final second image and normal map.

[0073] For example, referring to Figure 3, each denoising network may include: a residual module, a first attention module, a second attention module, a third attention module, and a fourth attention module. Each denoising network corresponds to one denoising process. Each denoising process is the same, and each denoising process may include the following steps S231-235:

[0074] S231. Input the current image and current normal map corresponding to each preset viewpoint into the residual module for convolution processing.

[0075] The residual module can be a network formed by stacking convolutional modules with skip connections. Specifically, for each denoising network, i.e., for each denoising process, the current image and current normal map to be denoised are input into the residual module for convolution processing to obtain the processed current image and current normal map, which are then output. Referring to Figure 3, during the first denoising, the initial image and initial normal map corresponding to each preset viewpoint are used as the current image and current normal map, respectively, and input into the residual module of the first denoising network for convolution processing. In subsequent denoising, the current image and current normal map output from the previous denoising network are input into the residual module of the current denoising network for convolution processing.

[0076] S232. Input the feature information of the first image and the current image and current normal map corresponding to each preset viewpoint output by the residual module into the first attention module to perform the embedding processing of the feature information of the first image.

[0077] The first attention module can be a multi-head attention module, used to embed feature information of the first image. Specifically, the first attention module adds the current noise level to the feature information of the input first image and embeds it into the key and value information of each current image and current normal map for attention processing, so that the feature information of the first image is present in each preset viewpoint corresponding to the current image and current normal map after embedding.

[0078] S233. Input the current image and current normal map corresponding to each preset viewpoint output by the first attention module into the second attention module to perform data fusion between multiple viewpoints.

[0079] Specifically, the second attention module performs multi-view data fusion on the current image corresponding to each preset viewpoint, obtaining and outputting the fused current image corresponding to each preset viewpoint. It also performs multi-view data fusion on the current normal map corresponding to each preset viewpoint, obtaining and outputting the fused current normal map corresponding to each preset viewpoint. Multi-view data fusion improves the correlation between different viewpoints, thereby enhancing content consistency across multiple viewpoints.

[0080] In the data fusion process across multiple perspectives, for each preset perspective, based on the projection relationships between different preset perspectives, related perspectives that overlap with that preset perspective can be identified. The current image corresponding to the preset perspective is then fused with the current images corresponding to the related perspectives, such as by multiplication or addition, to obtain the fused current image corresponding to that preset perspective. Similarly, the current normal map corresponding to the preset perspective is fused with the current normal map corresponding to the related perspectives, also by multiplication or addition, to obtain the fused current normal map corresponding to that preset perspective. By fusing only the data from related perspectives, fusion efficiency can be improved while maintaining the consistency of content across multiple perspectives.

[0081] S234. Input the current image and current normal map corresponding to each preset viewpoint output by the second attention module into the third attention module for joint optimization of the image and normal map.

[0082] Specifically, the third attention module performs two optimizations on the current image and current normal map corresponding to each preset viewpoint, so that the optimized current image and current normal map have a higher degree of alignment.

[0083] S235. Perform cross-attention processing on the feature information of the first image, the view information of each preset view, and the current image and current normal map corresponding to each preset view output by the third attention module to obtain the current image and current normal map corresponding to each preset view after the current denoising.

[0084] Specifically, the fourth attention module performs cross-attention processing on the input current image and current normal map based on the feature information of the first input image and the view information of each preset viewpoint, so that the processed current image and current normal map can retain the feature information of the first image. The current image and current normal map corresponding to each preset viewpoint output by the fourth attention module in the current denoising network are input into the residual module in the next denoising network for the next denoising process, and the current image and current normal map corresponding to each preset viewpoint output by the fourth attention module in the last denoising network are used as the target object for the final second image and normal map under each preset viewpoint.

[0085] S240. Perform three-dimensional reconstruction based on multiple second images to obtain the first three-dimensional model of the target object.

[0086] S250. Based on the normal map, the first three-dimensional model is optimized to obtain the optimized second three-dimensional model.

[0087] S260. Based on the first image and texture description information, perform texture mapping on the second three-dimensional model to obtain the target three-dimensional model of the target object.

[0088] The technical solution of this disclosure utilizes multiple denoising networks cascaded in a multi-view image generation model to perform iterative denoising processing on the initial image and initial normal map corresponding to each preset viewpoint based on the first image, thereby generating a second image and normal map of higher quality, further improving the generation quality of the 3D model.

[0089] Figure 4 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure. Based on the aforementioned embodiments, this disclosure optimizes the step of "performing three-dimensional reconstruction based on multiple second images to obtain a first three-dimensional model of the target object." Explanations of terms identical or corresponding to those in the aforementioned embodiments are not repeated here.

[0090] As shown in Figure 4, the three-dimensional model generation method specifically includes the following steps:

[0091] S410. Obtain the first image and texture description information of the target object.

[0092] S420. Based on the multi-view image generation model and the first image, generate a second image of the target object under each preset view and a normal map corresponding to the second image.

[0093] S430. Encode each second image to obtain the image feature information corresponding to each second image.

[0094] Specifically, referring to Figure 5, each second image is input into an image encoder for encoding processing to obtain the image feature information corresponding to each second image output by the image encoder. The image encoder can be, but is not limited to, a pre-trained ViT (Vision Transformer) model. A pre-trained ViT model is a deep learning model based on a visual attention mechanism, obtained through pre-training using a large-scale image dataset.

[0095] S440. Based on the initial three-plane model, the initial offset grid, and image feature information, synchronous iterative optimization is performed on denoising and anchor point position updates to obtain the target three-plane model and the target offset grid.

[0096] The initial three-plane model can be a Triplane model generated using random Gaussian noise initialization. A Triplane model represents a 3D object using three planes. The initial offset mesh can be a randomly generated offset network. The initial offset mesh can include the initial offset corresponding to each anchor point in 3D space. The target three-plane model can be the final three-plane model generated after denoising. The target offset mesh can include the target position offset corresponding to each anchor point in 3D space. An anchor point can refer to the center point of each voxel in 3D space.

[0097] Specifically, efficient and high-quality 3D reconstruction can be achieved by employing a hybrid triplane model and a Gaussian 3D representation. For example, an initial triplane model and an initial offset mesh are generated using random Gaussian noise initialization. Based on the image feature information of all second images, the initial triplane model and the initial offset mesh are iteratively optimized by synchronously denoising and updating anchor points to obtain a more accurate target triplane model and target offset mesh.

[0098] For example, step S440 may include the following steps S441-S446:

[0099] S441. The initial three-plane model and the initial offset mesh are respectively used as the current three-plane model and the current offset mesh in the first iteration, wherein the current offset mesh includes the current position deviation of each anchor point in the three-dimensional space.

[0100] S442. Input the current three-plane model and image feature information into the transform network model for denoising processing to obtain the current three-plane model after the current iteration.

[0101] Specifically, in each iteration, the current three-plane model is first denoised before the anchor point positions are updated. By utilizing a transform network model, such as the Transformer model, the current three-plane model is denoised based on the image feature information of all second images to obtain the denoised current three-plane model.

[0102] S443. Determine the current position of each anchor point based on the current three-plane model and the current offset mesh after the current iteration.

[0103] Specifically, based on the current position of each anchor point in the current three-plane model after denoising in the current iteration and the current position offset of each anchor point in the current offset grid, the current position and current position offset of the same anchor point in the current three-plane model after denoising in the current iteration and the current position offset are added together, and the result of the addition is used as the updated current position to offset each anchor point and update the current position corresponding to each anchor point.

[0104] S444. Based on the image feature information set and the current position corresponding to each anchor point, determine the current image feature information corresponding to each anchor point, and based on the current three-plane model after the current iteration and the current position corresponding to each anchor point, determine the current spatial feature information corresponding to each anchor point.

[0105] Specifically, a projection mapping is performed on the current position corresponding to each anchor point to obtain the projection point of each anchor point in the second image. Based on the image feature information of each second image, the current image feature information corresponding to each anchor point is obtained. For example, the image feature information corresponding to the projection point of each anchor point in the second image is used as the current image feature information corresponding to that anchor point. Based on the current position corresponding to each anchor point, the current spatial feature information corresponding to each anchor point is determined from the current three-plane model after the current iteration. For example, the spatial feature information corresponding to the current position of a certain anchor point in the current three-plane model after the current iteration is used as the current spatial feature information corresponding to that anchor point.

[0106] S445. The current image feature information and the current spatial feature information corresponding to each anchor point are fused to predict the offset grid generated in the current iteration, and the offset grid generated in the current iteration is added to the current offset grid to update the current offset grid.

[0107] Specifically, the current image feature information and current spatial feature information corresponding to each anchor point are input into the offset prediction model (such as the ConvGRU convolutional gated recurrent model) to predict the positional offset of each anchor point during the current iteration, thus obtaining the offset grid generated in the current iteration. The offset grid generated in the current iteration and the current offset grid are added together by the positional offset of the same anchor point, and the resulting offset grid is used as the updated current offset grid.

[0108] S446. Based on the updated current offset mesh and the current three-plane model after the current iteration, the next iteration is performed, and the current three-plane model and the current offset mesh obtained in the last iteration are used as the target three-plane model and the target offset mesh, respectively.

[0109] Specifically, based on the current offset mesh updated in the current iteration and the current three-plane model after the current iteration, the next iteration optimization is performed by returning to execute the above steps S442-S445, and the current three-plane model and the current offset mesh obtained in the last iteration are used as the final target three-plane model and the target offset mesh, respectively.

[0110] S450. Based on the target three-plane model and the target offset mesh, determine the target spatial feature information corresponding to each anchor point, and perform Gaussian decoding on the target spatial feature information corresponding to each anchor point to obtain the Gaussian volume corresponding to each anchor point.

[0111] Specifically, referring to Figure 5, the target position of each anchor point is determined based on the current position of each anchor point in the target three-plane model and the target position offset of each anchor point in the target offset grid. For example, the current position and target offset of the same anchor point are added together to obtain the target position of that anchor point. Spatial feature information at the target position is obtained from the target three-plane model and used as the target spatial feature information corresponding to that anchor point. A Gaussian decoder is then used to perform Gaussian decoding on the target spatial feature information corresponding to each anchor point, thereby transforming each anchor point into an independent Gaussian volume.

[0112] S460. Synthesize the Gaussian volume of all anchor points to obtain the first three-dimensional model of the target object.

[0113] Specifically, based on the target position of each anchor point, the Gaussian volumes of all anchor points are synthesized to obtain the complete first 3D model. Utilizing a hybrid three-plane model and a Gaussian 3D representation allows for faster generation of the 3D model, and the generation of a Gaussian volume on each voxel improves the geometric detail depiction of the 3D model. Using Gaussian volume representation also reduces memory overhead, and the output Gaussian volume position is not fixed at the anchor point position in the three-plane model but is obtained through iterative optimization, resulting in higher accuracy and flexibility.

[0114] S470. Based on the normal map, the first three-dimensional model is optimized to obtain the optimized second three-dimensional model.

[0115] S480. Based on the first image and texture description information, perform texture mapping on the second three-dimensional model to obtain the target three-dimensional model of the target object.

[0116] The technical solution of this disclosure, through synchronous iterative optimization of denoising and anchor point position updates based on the initial three-plane model, the initial offset grid, and the image feature information of each second image, can quickly obtain a more accurate target three-plane model and target offset grid; based on the target three-plane model and target offset grid, the target spatial feature information corresponding to each anchor point is determined, and Gaussian decoding is performed on the target spatial feature information corresponding to each anchor point to obtain the Gaussian volume corresponding to each anchor point. The Gaussian volumes of all anchor points are synthesized to obtain the first three-dimensional model of the target object. Thus, by using a hybrid three-plane model and Gaussian three-dimensional representation, efficient and high-quality three-dimensional reconstruction can be achieved.

[0117] Figure 6 is a flowchart illustrating a three-dimensional model generation method provided in an embodiment of this disclosure. Based on the aforementioned embodiments, this disclosure optimizes the step of "texturing the second three-dimensional model based on the first image and texture description information to obtain the target three-dimensional model of the target object." Explanations of terms identical or corresponding to those in the aforementioned embodiments are not repeated here.

[0118] As shown in Figure 6, the 3D model generation method specifically includes the following steps:

[0119] S610. Obtain the first image and texture description information of the target object.

[0120] S620. Based on the multi-view image generation model and the first image, generate a second image of the target object under each preset view and a normal map corresponding to the second image.

[0121] S630. Perform three-dimensional reconstruction based on multiple second images to obtain the first three-dimensional model of the target object.

[0122] S640. Based on the normal map, optimize the first three-dimensional model to obtain the optimized second three-dimensional model.

[0123] S650. Unfold the texture of the second three-dimensional model, determine the texture coordinates corresponding to the surface position points in the second three-dimensional model, initialize the texture values ​​corresponding to the texture coordinates, obtain the initial texture image corresponding to the second three-dimensional model, encode the initial texture image, and obtain the texture feature information of the initial model.

[0124] The initial texture image is the noisy texture image of the entire model. The texture feature information of the initial model can be the texture feature information of the entire model.

[0125] Specifically, the second 3D model is unfolded in a texture coordinate system to obtain the texture coordinates corresponding to the surface position points in the second 3D model, thereby establishing a mapping relationship between surface position points and texture coordinates. The texture value corresponding to each texture coordinate can be randomly initialized to obtain a complete initial texture image. The initial texture image is then input into an image encoder for encoding processing to obtain the initial model texture feature information in the latent space.

[0126] S660. Based on the first image and texture description information, the initial model texture feature information is denoised multiple times to obtain the target model texture feature information corresponding to the second three-dimensional model.

[0127] Specifically, the first image, texture description information, and initial model texture feature information can be input into the latent diffusion model for multiple iterations of denoising. The latent diffusion model uses the input first image and texture description information as conditional information to perform multiple iterations of denoising on the initial model texture feature information to obtain the final target model texture feature information in the latent space.

[0128] For example, step S660 may include the following steps S661-S664:

[0129] S661. Use the initial model texture feature information as the current model texture feature information during the first iteration.

[0130] S662. Perform multi-view rendering on the current model's texture feature information to obtain the first texture feature information corresponding to each view.

[0131] Specifically, the current model's texture feature information is rendered and projected onto multiple viewpoints to obtain the first texture feature information for each viewpoint. The first texture feature information refers to the texture feature information of a single viewpoint. By sharing the current model's texture feature information across multiple viewpoints, content consistency between different viewpoints can be guaranteed.

[0132] S663. Based on the first image and texture description information, perform noise reduction processing on the first texture feature information to determine the second texture feature information corresponding to each viewpoint.

[0133] Specifically, the first image and texture description information are used as conditional information to denoise the first texture feature information, and the second texture feature information corresponding to each viewpoint after denoising is determined.

[0134] For example, step S663 may include: inputting the first image and texture description information into a latent diffusion model introduced by the image cue adapter to perform denoising processing on the first texture feature information, and obtaining the second texture feature information corresponding to each viewpoint.

[0135] The image cue adapter (IP adapter) is a method for simultaneously introducing images and text into the latent diffusion model, enabling the model to use both images and text as cues for denoising. The image cue adapter can consist of two parts: an image encoder that extracts image features from the image cue, and a decoupled cross-attention adaptation module that embeds the image features into the latent diffusion model.

[0136] Specifically, by utilizing a latent diffusion model with an image cue adapter, a connection between images and text can be established. This allows for more accurate and efficient denoising based on the first image and texture description information, directly generating second texture feature information corresponding to each viewpoint, while ensuring content consistency across different viewpoints. For example, the internal processing of the latent diffusion model with an image cue adapter involves: extracting features from the first image and texture description information to obtain image feature information and text feature information; performing cross-attention processing on the image feature information and text feature information to obtain image attention information and text attention information; and inputting the image attention information, text attention information, and first texture feature information into the latent diffusion model for denoising processing to obtain second texture feature information corresponding to each viewpoint.

[0137] S664. Merge the second texture feature information corresponding to all viewpoints to obtain the current model texture feature information after the current iteration. Perform the next iteration based on the current model texture feature information after the current iteration, and use the current model texture feature information obtained in the last iteration as the target model texture feature information corresponding to the second three-dimensional model.

[0138] Specifically, based on the weights corresponding to each viewpoint, the second texture feature information corresponding to all viewpoints is weighted and summed to obtain the current model texture feature information of the entire model. The weight for each viewpoint can be the cosine value formed by the viewpoint and the normal. Based on the current model texture feature information after the current iteration, denoising is performed in the next iteration by returning to and executing steps S662-S664 above. This process is repeated multiple times, and the current model texture feature information obtained in the last iteration is used as the final target model texture feature information.

[0139] S670. Decode the texture feature information of the target model to obtain the target texture image, and paste the target texture image onto the second three-dimensional model to obtain the target three-dimensional model of the target object.

[0140] Specifically, the texture feature information of the target model in the latent space is decoded to obtain the target texture image in the explicit space. Based on the mapping relationship between the surface position point and the texture coordinate, the target texture image is pasted onto the surface of the second three-dimensional model, thereby completing the texture mapping of the second three-dimensional model and obtaining a more realistic target three-dimensional model.

[0141] The technical solution of this disclosure involves encoding an initial texture image to obtain initial model texture feature information; performing multiple iterations of denoising on the initial model texture feature information based on a first image and texture description information to obtain target model texture feature information corresponding to a second three-dimensional model; decoding the target model texture feature information to obtain a higher quality target texture image; and pasting the target texture image onto the second three-dimensional model, thereby improving the realism and reconstruction quality of the target three-dimensional model.

[0142] Figure 7 is a schematic diagram of the structure of a three-dimensional model generation device provided in an embodiment of the present disclosure. As shown in Figure 7, the device specifically includes: an information acquisition module 710, a multi-view image generation module 720, a three-dimensional reconstruction module 730, a model optimization module 740, and a texture mapping module 750.

[0143] The system includes: an information acquisition module 710 for acquiring a first image and texture description information of the target object; a multi-view image generation module 720 for generating a second image of the target object from each preset viewpoint and a normal map corresponding to the second image based on a multi-view image generation model and the first image; a 3D reconstruction module 730 for performing 3D reconstruction based on multiple second images to obtain a first 3D model of the target object; a model optimization module 740 for optimizing the first 3D model based on the normal map to obtain an optimized second 3D model; and a texture mapping module 750 for performing texture mapping on the second 3D model based on the first image and the texture description information to obtain a target 3D model of the target object.

[0144] The technical solution provided in this disclosure generates a first 3D model of the target object from a multi-view image generation model and a first image of the target object. This model generates a second image of the target object from each preset viewpoint and a normal map corresponding to the second image. Based on multiple second images, 3D reconstruction is performed to obtain a first 3D model of the target object. The first 3D model is then optimized based on the normal map corresponding to the second image to obtain a higher-quality second 3D model. Finally, based on the first image and texture description information, a more realistic texture mapping is applied to the second 3D model, resulting in a higher-quality target 3D model of the target object. This ensures the quality of the generated 3D model. By requiring only one first image from any viewpoint, without the need for 3D scanning, and eliminating the need for professional personnel, the convenience of 3D model generation is improved, and the cost of 3D model generation is reduced.

[0145] Based on the above technical solution, the multi-view image generation module 720 includes:

[0146] An image feature extraction unit is used to input the first image into the image encoder in the multi-view image generation model for feature extraction to obtain the feature information of the first image.

[0147] The multi-view image generation unit is used to input the feature information of the first image and the initial image and initial normal map corresponding to each preset viewpoint into multiple denoising networks cascaded in the multi-view image generation model for multiple iterations of denoising, so as to obtain the second image of the target object under each preset viewpoint and the normal map corresponding to the second image after the last iteration of denoising.

[0148] Based on the above technical solutions, each denoising network includes: a residual module, a first attention module, a second attention module, a third attention module, and a fourth attention module; wherein, the multi-view image generation unit performs denoising each time by executing the following steps:

[0149] The current image and current normal map corresponding to each preset viewpoint are input into the residual module for convolution processing;

[0150] The feature information of the first image and the current image and current normal map corresponding to each preset viewpoint output by the residual module are input into the first attention module to perform the embedding processing of the feature information of the first image;

[0151] The current image and current normal map corresponding to each preset viewpoint output by the first attention module are input into the second attention module for data fusion between multiple viewpoints;

[0152] The current image and current normal map corresponding to each preset viewpoint output by the second attention module are input into the third attention module for joint optimization of the image and normal map;

[0153] The feature information of the first image, the view information of each preset viewpoint, and the current image and current normal map corresponding to each preset viewpoint output by the third attention module are subjected to cross-attention processing to obtain the current image and current normal map corresponding to each preset viewpoint after the current denoising.

[0154] Based on the above technical solutions, the multi-view image generation model is generated by fine-tuning a generative pre-trained model based on a sample dataset and a target loss function.

[0155] The target loss function includes a perceptual loss function and a noise loss function. The perceptual loss function is used to characterize the loss value between the sample denoised image and the predicted denoised image. The sample denoised image is determined based on the sample denoised image and the sample noise. The predicted denoised image is determined based on the sample denoised image and the predicted noise.

[0156] The noise loss function is used to characterize the loss value between sample noise and predicted noise.

[0157] Based on the above technical solutions, the three-dimensional reconstruction module 730 includes:

[0158] An image encoding unit is used to encode each of the second images to obtain image feature information corresponding to each of the second images;

[0159] The iterative optimization unit is used to synchronously iteratively optimize denoising and anchor point position updates based on the initial three-plane model, the initial offset grid and the image feature information to obtain the target three-plane model and the target offset grid, wherein the target offset grid includes the target position offset corresponding to each anchor point in the three-dimensional space.

[0160] The Gaussian volume generation unit is used to determine the target spatial feature information corresponding to each anchor point based on the target three-plane model and the target offset mesh, and to perform Gaussian decoding on the target spatial feature information corresponding to each anchor point to obtain the Gaussian volume corresponding to each anchor point.

[0161] The first 3D model generation unit is used to synthesize the Gaussian volume of all anchor points to obtain the first 3D model of the target object.

[0162] Based on the above technical solutions, the iterative optimization unit is specifically used for:

[0163] The initial three-plane model and the initial offset mesh are used as the current three-plane model and the current offset mesh in the first iteration, respectively. The current offset mesh includes the current position deviation of each anchor point in 3D space. The current three-plane model and the image feature information are input into the transform network model for denoising to obtain the current three-plane model after the current iteration. The current position of each anchor point is determined based on the current three-plane model and the current offset mesh after the current iteration. Based on the image feature information set and the current position of each anchor point, the current image feature information of each anchor point is determined, and the current spatial feature information of each anchor point is determined based on the current three-plane model and the current position of each anchor point after the current iteration. The current image feature information and the current spatial feature information of each anchor point are fused to predict the offset mesh generated in the current iteration, and the offset mesh generated in the current iteration is added to the current offset mesh to update the current offset mesh. The next iteration is based on the updated current offset mesh and the current three-plane model after the current iteration, and the current three-plane model and the current offset mesh obtained in the last iteration are used as the target three-plane model and the target offset mesh, respectively.

[0164] Based on the above technical solutions, the model optimization module 740 is specifically used for:

[0165] By using a continuous mesh reconstruction method, the first 3D model is optimized by mesh smoothing based on the normal map to obtain an optimized second 3D model.

[0166] Based on the above technical solutions, the texture mapping module 750 includes:

[0167] The initial model texture feature information determination unit is used to perform texture unrolling on the second three-dimensional model, determine the texture coordinates corresponding to the surface position points in the second three-dimensional model, initialize the texture values ​​corresponding to the texture coordinates, obtain the initial texture image corresponding to the second three-dimensional model, encode the initial texture image, and obtain the initial model texture feature information.

[0168] The target model texture feature information determination unit is used to perform multiple iterations of denoising on the initial model texture feature information based on the first image and the texture description information to obtain the target model texture feature information corresponding to the second three-dimensional model;

[0169] The target 3D model generation unit is used to decode the texture feature information of the target model to obtain a target texture image, and then paste the target texture image onto the second 3D model to obtain the target 3D model of the target object.

[0170] Based on the above technical solutions, the target model texture feature information determination unit is specifically used for:

[0171] The initial model texture feature information is used as the current model texture feature information in the first iteration; the current model texture feature information is rendered from multiple perspectives to obtain the first texture feature information corresponding to each perspective; based on the first image and the texture description information, the first texture feature information is denoised to determine the second texture feature information corresponding to each perspective; the second texture feature information corresponding to all perspectives is merged to obtain the current model texture feature information after the current iteration; the next iteration is performed based on the current model texture feature information after the current iteration, and the current model texture feature information obtained in the last iteration is used as the target model texture feature information corresponding to the second three-dimensional model.

[0172] Based on the above technical solutions, the target model texture feature information determination unit is specifically used for:

[0173] The first image and the texture description information are input into the latent diffusion model of the image cue adapter to perform denoising processing on the first texture feature information, thereby obtaining the second texture feature information corresponding to each viewpoint.

[0174] The three-dimensional model generation apparatus provided in this disclosure can execute the three-dimensional model generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.

[0175] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0176] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Referring now to Figure 8, a schematic diagram of the structure of an electronic device (e.g., the terminal device or server in Figure 8) 500 suitable for implementing embodiments of this disclosure is shown. The terminal device in embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 8 is merely an example and should not impose any limitations on the functionality and scope of use of embodiments of this disclosure.

[0177] As shown in Figure 8, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0178] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG8 shows an electronic device 500 with various devices, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0179] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0180] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0181] The electronic device provided in this embodiment and the three-dimensional model generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0182] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the three-dimensional model generation method provided in the above embodiments.

[0183] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0184] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0185] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0186] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a first image and texture description information of a target object; generate a second image of the target object at each preset viewpoint and a normal map corresponding to the second image based on a multi-view image generation model and the first image; perform three-dimensional reconstruction based on multiple second images to obtain a first three-dimensional model of the target object; optimize the first three-dimensional model based on the normal map to obtain an optimized second three-dimensional model; and perform texture mapping on the second three-dimensional model based on the first image and the texture description information to obtain a target three-dimensional model of the target object.

[0187] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0188] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the three-dimensional model generation method provided in the above embodiments.

[0189] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0191] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0192] The functions described above in this disclosure can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0193] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0194] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0195] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0196] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for generating a three-dimensional model, comprising: Obtain the first image and texture description information of the target object; Based on the multi-view image generation model and the first image, a second image of the target object under each preset view and the normal map corresponding to the second image are generated. Based on multiple second images, a three-dimensional reconstruction is performed to obtain a first three-dimensional model of the target object; Based on the normal map, the first three-dimensional model is optimized to obtain the optimized second three-dimensional model; Based on the first image and the texture description information, texture mapping is performed on the second 3D model to obtain the target 3D model of the target object.

2. The three-dimensional model generation method according to claim 1, wherein Based on a multi-view image generation model and the first image, a second image of the target object under each preset viewpoint and a normal map corresponding to the second image are generated, including: The first image is input into the image encoder in the multi-view image generation model for feature extraction to obtain the feature information of the first image. The feature information of the first image and the initial image and initial normal map corresponding to each preset viewpoint are input into multiple denoising networks cascaded in the multi-view image generation model for multiple iterations of denoising, so as to obtain the second image of the target object under each preset viewpoint and the normal map corresponding to the second image after the last iteration of denoising.

3. The three-dimensional model generation method according to claim 2, wherein Each denoising network includes: a residual module, a first attention module, a second attention module, a third attention module, and a fourth attention module; wherein each denoising process includes: The current image and current normal map corresponding to each preset viewpoint are input into the residual module for convolution processing; The feature information of the first image and the current image and current normal map corresponding to each preset viewpoint output by the residual module are input into the first attention module to perform the embedding processing of the feature information of the first image; The current image and current normal map corresponding to each preset viewpoint output by the first attention module are input into the second attention module for data fusion between multiple viewpoints; The current image and current normal map corresponding to each preset viewpoint output by the second attention module are input into the third attention module for joint optimization of the image and normal map; The feature information of the first image, the view information of each preset view, and the current image and current normal map corresponding to each preset view output by the third attention module are subjected to cross-attention processing to obtain the current image and current normal map corresponding to each preset view after the current denoising.

4. The three-dimensional model generation method according to any one of claims 1 to 3, wherein The multi-view image generation model is generated by fine-tuning a generative pre-trained model based on a sample dataset and a target loss function. The target loss function includes a perceptual loss function and a noise loss function. The perceptual loss function is used to characterize the loss value between the sample denoised image and the predicted denoised image. The sample denoised image is determined based on the sample denoised image and the sample noise. The predicted denoised image is determined based on the sample denoised image and the predicted noise. The noise loss function is used to characterize the loss value between sample noise and predicted noise.

5. The three-dimensional model generation method according to any one of claims 1 to 4, wherein Based on multiple second images, a three-dimensional reconstruction is performed to obtain a first three-dimensional model of the target object, including: Each of the second images is encoded to obtain the image feature information corresponding to each of the second images; Based on the initial three-plane model, the initial offset grid, and the image feature information, synchronous iterative optimization of denoising and anchor point position update is performed to obtain the target three-plane model and the target offset grid, wherein the target offset grid includes the target position offset corresponding to each anchor point in three-dimensional space; Based on the target three-plane model and the target offset mesh, the target spatial feature information corresponding to each anchor point is determined, and Gaussian decoding is performed on the target spatial feature information corresponding to each anchor point to obtain the Gaussian volume corresponding to each anchor point. The Gaussian volumes of all anchor points are synthesized to obtain the first three-dimensional model of the target object.

6. The three-dimensional model generation method according to claim 5, wherein Synchronous iterative optimization of denoising and anchor point position updating based on the initial three-plane model, initial offset grid, and the image feature information yields the target three-plane model and target offset grid, including: The initial three-plane model and the initial offset mesh are used as the current three-plane model and the current offset mesh during the first iteration, respectively. The current offset mesh includes the current position offset of each anchor point in the three-dimensional space. The current three-plane model and the image feature information are input into the transform network model for denoising to obtain the current three-plane model after the current iteration. The current position of each anchor point is determined based on the current three-plane model and the current offset mesh after the current iteration. Based on the image feature information set and the current position corresponding to each anchor point, the current image feature information corresponding to each anchor point is determined, and based on the current three-plane model after the current iteration and the current position corresponding to each anchor point, the current spatial feature information corresponding to each anchor point is determined. The current image feature information and current spatial feature information corresponding to each anchor point are fused to predict the offset grid generated in the current iteration, and the offset grid generated in the current iteration is added to the current offset grid to update the current offset grid; The next iteration is based on the updated current offset mesh and the current three-plane model after the current iteration, and the current three-plane model and current offset mesh obtained in the last iteration are used as the target three-plane model and target offset mesh, respectively.

7. The three-dimensional model generation method according to any one of claims 1 to 6, wherein, Based on the normal map, the first 3D model is optimized to obtain an optimized second 3D model, including: By using a continuous mesh reconstruction method, the first 3D model is optimized by mesh smoothing based on the normal map to obtain an optimized second 3D model.

8. The three-dimensional model generation method according to any one of claims 1 to 7, wherein, Based on the first image and the texture description information, texture mapping is performed on the second 3D model to obtain the target 3D model of the target object, including: The second three-dimensional model is textured and unfolded to determine the texture coordinates corresponding to the surface position points in the second three-dimensional model. The texture values ​​corresponding to the texture coordinates are initialized to obtain the initial texture image corresponding to the second three-dimensional model. The initial texture image is encoded to obtain the initial model texture feature information. Based on the first image and the texture description information, the initial model texture feature information is denoised multiple times to obtain the target model texture feature information corresponding to the second three-dimensional model. The target model texture feature information is decoded to obtain a target texture image, and the target texture image is pasted onto the second three-dimensional model to obtain the target three-dimensional model of the target object.

9. The three-dimensional model generation method of claim 8, wherein, Based on the first image and the texture description information, the initial model texture feature information is iteratively optimized multiple times to obtain the target model texture feature information corresponding to the second 3D model, including: The initial model texture feature information is used as the current model texture feature information during the first iteration; Perform multi-view rendering on the current model's texture feature information to obtain the first texture feature information corresponding to each view; Based on the first image and the texture description information, the first texture feature information is denoised to determine the second texture feature information corresponding to each viewpoint. The second texture feature information corresponding to all viewpoints is merged to obtain the current model texture feature information after the current iteration. The next iteration is based on the current model texture feature information after the current iteration, and the current model texture feature information obtained in the last iteration is used as the target model texture feature information corresponding to the second three-dimensional model.

10. The three-dimensional model generation method of claim 9, wherein, Based on the first image and the texture description information, the first texture feature information is denoised to determine the second texture feature information corresponding to each viewpoint, including: The first image and the texture description information are input into the latent diffusion model of the image cue adapter to perform denoising processing on the first texture feature information, thereby obtaining the second texture feature information corresponding to each viewpoint.

11. A three-dimensional model generation device, comprising: The information acquisition module is configured to acquire the first image and texture description information of the target object; The multi-view image generation module is configured to generate a second image of the target object under each preset view and a normal map corresponding to the second image based on the multi-view image generation model and the first image; The 3D reconstruction module is configured to perform 3D reconstruction based on multiple second images to obtain a first 3D model of the target object; The model optimization module is configured to optimize the first three-dimensional model based on the normal map to obtain an optimized second three-dimensional model; The texture mapping module is configured to perform texture mapping on the second 3D model based on the first image and the texture description information to obtain the target 3D model of the target object.

12. An electronic device, comprising: One or more processors; Storage device, storing one or more programs When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional model generation method as described in any one of claims 1-10.

13. A storage medium containing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the three-dimensional model generation method as described in any one of claims 1-10.

14. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the three-dimensional model generation method as claimed in any one of claims 1-10.