Method and device for generating three-dimensional model

Through the two-stage three-dimensional reconstruction method, the object generation model is used to predict and optimize images of other perspectives based on single-view image images, and the problem of inaccurate three-dimensional reconstruction of single-view image is solved, achieving more accurate and finer three-dimensional model generation.

CN120014133APending Publication Date: 2025-05-16BEIJING DAJIA INTERNET INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065501.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The results of the three-dimensional reconstruction based on single-view images are inaccurate, resulting in the chaotic structure of the three-dimensional model and lack of realism, making it difficult to directly apply it.

Method used

Two-stage three-dimensional reconstruction method is adopted: the first stage uses the first object generation model to predict images from other perspectives based on a single-view image, and generates an initial three-dimensional model; the second stage uses the second object generation model to optimize images from other perspectives to further optimize the three-dimensional model.

Benefits of technology

The accuracy and refinement of the three-dimensional model is improved, the authenticity and application value of the three-dimensional reconstruction results are enhanced, and the conditions are provided for the implementation of the three-dimensional reconstruction solution of single-view image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014133A_ABST
    Figure CN120014133A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method and device for generating a three-dimensional model. The main technical scheme comprises the steps of obtaining a single-view image of a target object; utilizing the first object generation model to predict and obtain images of other visual angles based on the single-visual-angle image of the target object; obtaining a three-dimensional model of the target object by using the images of other visual angles; utilizing a second object generation model to obtain optimized images of other visual angles based on the images of other visual angles; optimizing the three-dimensional model of the target object by taking the optimized images of other visual angles as true value images; wherein the first object generation model and the second object generation model are obtained by pre-training based on a diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular to a method and device for generating a three-dimensional model. Background Art

[0002] 3D reconstruction plays a key role in computer graphics and virtual reality, and is widely used in the creation of virtual characters, film and television animation production, virtual fitting and other fields. Realizing 3D reconstruction based on single-view images has very important theoretical significance and application value. For example, in virtual reality and augmented reality applications, users often need to generate a complete 3D model through a single-view image. In order to achieve a highly realistic effect, accurate reconstruction of the single-view image of the target object is crucial. However, due to the limitations of related technologies, the 3D reconstruction results based on single-view images are not very ideal. The structure of the 3D model is often disordered, lacking in realism and other quality issues, making it difficult to directly implement it. Summary of the invention

[0003] In view of this, the present application provides a method and device for generating a three-dimensional model, so as to solve the problem of inaccuracy in reconstructing a three-dimensional model based on a single-view image.

[0004] This application provides the following solutions:

[0005] In a first aspect, a method for generating a three-dimensional model is provided, the method for generating a three-dimensional model comprising:

[0006] Acquire a single-view image of the target object;

[0007] Generate images of other perspectives based on the single-perspective image of the target object using the first object generation model;

[0008] Using images from other perspectives to obtain a three-dimensional model of the target object;

[0009] Generate images based on other perspectives using the second object model to obtain optimized images of other perspectives;

[0010] Using optimized images from other perspectives as true value images to optimize the three-dimensional model of the target object;

[0011] The first object generation model and the second object generation model are both pre-trained based on the diffusion model.

[0012] Optionally, the images from other perspectives include images from multiple perspectives;

[0013] Using images from other perspectives to obtain a three-dimensional model of the target object includes: using images from multiple perspectives to perform three-dimensional reconstruction on the target object to obtain a three-dimensional model of the target object.

[0014] Optionally, obtaining a three-dimensional model of the target object using images from other perspectives includes:

[0015] Initialize the three-dimensional model of the target object;

[0016] Images from other perspectives are used as true images to optimize the 3D model of the target object.

[0017] Optionally, optimizing the three-dimensional model of the target object includes:

[0018] Determine the viewing angle corresponding to the true image;

[0019] Obtain an image of the target object's three-dimensional model rendered according to the viewing angle corresponding to the true value image;

[0020] The difference between the rendered image and the ground-truth image is used to optimize the 3D model.

[0021] Optionally, using the first object generation model to predict images of other perspectives based on the single-perspective image of the target object includes:

[0022] The single-view image is denoised, and the noisy single-view image and information of other viewpoints are input into a first object generation model. The first object generation model uses the information of other viewpoints as a guiding condition to denoise the noisy single-view image to obtain images of other viewpoints.

[0023] Optionally, the first object generation model is trained in the following manner:

[0024] Acquire an image sample of a first perspective and an image sample of a second perspective;

[0025] After adding noise to the image samples of the first perspective, the noisy image samples of the first perspective and the information of the second perspective are input into the first diffusion model, and the first diffusion model is used to denoise the noisy image samples of the first perspective based on the information of the second perspective to obtain a predicted image of the second perspective, and the model parameters of the first diffusion model are updated using the loss function corresponding to the first training objective, wherein the first training objective includes minimizing the difference between the predicted image and the image samples of the second perspective, and the distance between the feature representation of the image samples of the first perspective and the feature representation of the predicted image.

[0026] Optionally, using the second object to generate the model based on images from other perspectives to obtain optimized images from other perspectives includes:

[0027] The images from other perspectives and the noise image are input into the second object generation model, and the second object generation model uses the images from other perspectives as guidance conditions to perform denoising on the noise image to obtain optimized images from other perspectives.

[0028] Optionally, the second object generation model is trained in the following manner:

[0029] Acquire an image sample whose texture clarity meets a preset condition of the target object, and perform noise processing on the image sample;

[0030] The denoised image samples and the noisy image are input into a second diffusion model, a predicted image is obtained by denoising the noisy image based on the denoised image samples by the second diffusion model, and the model parameters of the second diffusion model are updated using a loss function corresponding to a second training objective, wherein the second training objective includes minimizing the difference between the predicted image and the image samples before denoising, and the distance between the feature representation of the denoised image samples and the feature representation of the predicted image.

[0031] Optionally, the target object is hair;

[0032] Acquiring a single-view image of the target object includes: acquiring a single-view hair image and a body template image, aligning and synthesizing a hair part in the hair image and a body part in the body template image to obtain a single-view image.

[0033] Optionally, the three-dimensional model adopts three-dimensional Gaussian representation, implicit field representation or voxel representation.

[0034] In a second aspect, a device for generating a three-dimensional model is provided, the device for generating a three-dimensional model comprising:

[0035] An image acquisition unit, configured to acquire a single-view image of a target object;

[0036] A first optimization unit is configured to use the first object generation model to predict images of other perspectives based on a single-perspective image of the target object; and use the images of other perspectives to obtain a three-dimensional model of the target object;

[0037] A second optimization unit is configured to generate images of other perspectives based on the second object generation model to obtain optimized images of other perspectives; and optimize the three-dimensional model of the target object using the optimized images of other perspectives as true value images;

[0038] The first object generation model and the second object generation model are both pre-trained based on the diffusion model.

[0039] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.

[0040] In a fourth aspect, an electronic device is provided, including:

[0041] one or more processors; and

[0042] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first aspects above.

[0043] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.

[0044] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0045] 1) This application adopts a two-stage 3D reconstruction method: in the first stage, the first object generation model is used to predict and generate images of other perspectives based on the single-view image of the target object, and the images of other perspectives are used to obtain the 3D model of the target object; in the second stage, the second object generation model is used to optimize the images of other perspectives to obtain the optimized images of other perspectives, and the optimized images of other perspectives are used as the true value images to optimize the 3D model of the target object. This method can further optimize the 3D model of the target object in the second stage after reconstructing the 3D model of the target object in the first stage, so that the 3D model reconstructed based on the single-view image is more accurate and refined, which provides conditions for the implementation of the 3D reconstruction solution based on the single-view image.

[0046] 2) This application uses images from multiple perspectives to perform three-dimensional reconstruction on a target object to obtain a three-dimensional model of the target object. It provides a method for generating a three-dimensional model based on a single-perspective image, and generates multi-perspective images based on a single-perspective image, thereby achieving preliminary three-dimensional reconstruction.

[0047] 3) The present application also provides a method for generating a three-dimensional model based on a single-view image. After initializing the three-dimensional model of the target object, images from other viewpoints are used as true-value images to optimize the three-dimensional model of the target object.

[0048] 4) The present application first determines the viewing angle corresponding to the true image. After obtaining the image of the target object's three-dimensional model rendered according to the viewing angle corresponding to the true image, by comparing the difference between the rendered image and the true image, the deficiencies of the three-dimensional model in terms of shape, texture, etc. can be identified, and these deficiencies can be optimized to make the three-dimensional model closer to the target object in the real world, thereby improving its realism.

[0049] 5) The present application adds noise to a single-view image, and inputs the noisy single-view image and information of other viewpoints into a first object generation model. The first object generation model uses the information of other viewpoints as a guiding condition to denoise the noisy single-view image to obtain images of other viewpoints. This enables the noisy single-view image to be denoised at the image level, allowing the first object generation model to output high-quality images of other viewpoints.

[0050] 6) The present application uses image samples of the first perspective and image samples of the second perspective as samples for training the first object generation model, and updates the model parameters of the first diffusion model based on the loss function corresponding to the first training objective. The first training objective includes minimizing the difference between the predicted image and the image samples of the second perspective, and the distance between the feature representation of the image samples of the first perspective and the feature representation of the predicted image, so as to train a high-precision first object generation model and further improve the quality of images of other perspectives predicted and output by the first object generation model.

[0051] 7) The present application inputs images from other perspectives and noise images into a second object generation model, and the second object generation model uses the images from other perspectives as guidance conditions to denoise the noise images to obtain optimized images from other perspectives, thereby optimizing each pixel in the images from other perspectives to obtain optimized images from other perspectives.

[0052] 8) The present application uses the image samples after noise processing and the noisy image as training samples for training the second object generation model, and uses the loss function corresponding to the second training objective to update the model parameters of the second diffusion model. The second training objective includes minimizing the difference between the predicted image and the image samples before noise processing, and the distance between the feature representation of the image samples after noise processing and the feature representation of the predicted image, so as to train a high-precision second object generation model and further improve the quality of optimized images of other perspectives output by the second object generation model.

[0053] 9) In the present application, the target object is hair; after obtaining a single-view hair image and a body template image, the hair part in the hair image and the body part in the body template image are aligned and synthesized to obtain a single-view image. In this scenario, accurate three-dimensional reconstruction can be performed based on the single-view image containing the hair to obtain a three-dimensional model of the hair.

[0054] 10) This application can choose to use three-dimensional Gaussian representation, implicit field representation or voxel representation for the three-dimensional model according to the specific usage requirements of the three-dimensional model, so as to further improve the flexibility and adaptability of three-dimensional modeling.

[0055] Of course, any product implementing the present application does not necessarily need to achieve all of the above advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;

[0058] Figure 2 A flowchart of a method for generating a three-dimensional model provided in an embodiment of the present application;

[0059] Figure 3 A schematic diagram of using the first object generation model to predict images from other perspectives provided in an embodiment of the present application;

[0060] Figure 4 A schematic diagram of obtaining hair images at different viewing angles provided in an embodiment of the present application;

[0061] Figure 5 A schematic diagram of training a first object generation model provided in an embodiment of the present application;

[0062] Figure 6 A schematic diagram of obtaining optimized images of other viewing angles by using a second object generation model provided in an embodiment of the present application;

[0063] Figure 7 A schematic diagram of training a second object generation model provided in an embodiment of the present application;

[0064] Figure 8 A schematic diagram of a device for generating a three-dimensional model provided in an embodiment of the present application;

[0065] Fig. 9 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0067] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0068] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0069] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0070] 3D reconstruction is an important research direction in the field of computer vision, and the 3D reconstruction of the target part to be reconstructed in the image based on the monocular image has very important theoretical significance and application value. In the related technology, the 3D reconstruction results based on the single-view image are not very ideal. The structure of the 3D model is often disordered, lacking in realism and other quality problems, making it difficult to directly implement it.

[0071] In view of this, the present application provides a new idea. In order to facilitate the understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: a user end, a terminal device and a server end.

[0072] Among them, a user terminal for displaying the target object is running in the terminal device. The terminal device can exchange data with the server terminal through the network. As one of the achievable methods, the server terminal can use the method provided in the embodiment of the present application to perform three-dimensional reconstruction based on a single-view image from the user terminal or a single-view image from a database, etc., to generate a three-dimensional model of the target object; in response to a request from the user terminal, the request carries information about the target view, uses the three-dimensional model to render an image of the target view, and returns the rendered image to the user terminal for display.

[0073] As another feasible method, when a 3D model of a target object is generated on the server side, the 3D model of the target object can be directly sent to a terminal device, which then renders an image of the target perspective locally and displays the target object.

[0074] Among them, the user end involved in the embodiments of the present application can be a client running on a terminal device, a small program, or a Web application running through a browser, etc.

[0075] Terminal devices may include, but are not limited to, smart mobile terminals, wearable devices, PCs (Personal Computers), smart home devices, etc. Smart mobile devices may include, for example, mobile phones, tablet computers, laptops, PDAs (Personal Digital Assistants), Internet car terminals, etc. Wearable devices may include, for example, smart watches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality), mixed reality devices (i.e., devices that can support both virtual reality and augmented reality), etc. Smart home devices may include, for example, smart TVs, smart refrigerators with display screens, etc.

[0076] The server side can be a platform responsible for 3D reconstruction, a single server, a server group consisting of multiple servers, or a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a host product in the cloud computing service system to solve the defects of difficult management and weak service scalability in traditional physical hosts and virtual private servers (VPS) services.

[0077] The above-mentioned network may include but is not limited to: wired network, wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication.

[0078] It should be understood that Figure 1 The number of user terminals, terminal devices and server terminals in the embodiment is only for illustration. Any number of user terminals, terminal devices and server terminals may be provided according to implementation requirements.

[0079] Figure 2 A flow chart of a method for generating a three-dimensional model provided in an embodiment of the present application. The method for generating a three-dimensional model can be performed by Figure 1 The server side execution in the system shown. Figure 2 As shown in , the method for generating a three-dimensional model may include the following steps:

[0080] Step 201: Acquire a single-view image of a target object;

[0081] Step 203: using the first object generation model to predict images of other perspectives based on the single-perspective image of the target object;

[0082] Step 205: Obtain a three-dimensional model of the target object using images from other perspectives;

[0083] Step 207: using the second object to generate the model based on images of other perspectives, to obtain optimized images of other perspectives;

[0084] Step 209: using the optimized images from other perspectives as true value images to optimize the three-dimensional model of the target object;

[0085] The first object generation model and the second object generation model are both pre-trained based on the diffusion model.

[0086] It can be seen that the present application adopts a two-stage 3D reconstruction method: in the first stage, the first object generation model is used to predict and generate images of other perspectives based on the single-view image of the target object, and the images of other perspectives are used to obtain the 3D model of the target object; in the second stage, the second object generation model is used to optimize the images of other perspectives, and the optimized images of other perspectives are obtained, and the optimized images of other perspectives are used as the true value images to optimize the 3D model of the target object. This method can further optimize the 3D model of the target object in the second stage after reconstructing the 3D model of the target object in the first stage, so that the 3D model reconstructed based on the single-view image is more accurate and refined, which provides conditions for the implementation of the 3D reconstruction solution based on the single-view image.

[0087] The following is a detailed description of each step in the above process and the effects that can be further produced in conjunction with the embodiments. It should be noted that the "first" and "second" limitations involved in the present disclosure do not have limitations in terms of size, order, and quantity, and are only used to distinguish them in name. For example, "first object generation model" and "second object generation model" are used to distinguish two models in name.

[0088] First, the above step 201, namely “obtaining a single-view image of the target object”, is described in detail in conjunction with the embodiment.

[0089] In an embodiment of the present application, a three-dimensional reconstruction of the target object is performed based on a single-view image of the target object, that is, only one image of the target object needs to be captured or collected. In an embodiment of the present application, the image is referred to as a single-view image, and the information about the viewing angle when the image is captured is not limited; that is, the single viewing angle can be any viewing angle. Among them, the server side responsible for three-dimensional reconstruction can obtain a single-view image of the target object. Optionally, the target object can be captured by an independent camera device and the captured single-view image of the target object can be uploaded to the server side. Optionally, the camera device can be a camera on a terminal device.

[0090] The target object in the embodiment of the present application can be any object, such as clothes, shoes, decorations, etc., or it can be a component of the human body, such as hair, eyelashes, etc.

[0091] In one example, the target object is hair. The following takes the reconstruction of the three-dimensional hair model of hair as an example. For other objects (such as clothes), the hair can be replaced with other objects. For the reconstruction of other corresponding three-dimensional models, the process of reconstructing the three-dimensional hair model of hair can be referred to, and no further details are given here.

[0092] The 3D hair reconstruction involved in the embodiments of the present application may refer to the reconstruction of a 3D hair model based on a single-view image, which achieves preliminary 3D reconstruction. Reconstructing a 3D hair model plays a key role in computer graphics and virtual reality, and is widely used in the creation of virtual characters, film and television animation production, and virtual fitting. For example, in virtual reality and augmented reality applications, users often need to generate a complete 3D character model from a single-view image.

[0093] In one example, obtaining a monoscopic image of a target object may include: obtaining a monoscopic hair image and a body template image, aligning and synthesizing a hair portion in the hair image and a body portion in the body template image to obtain a monoscopic image. The hair image may intuitively display the appearance characteristics of the hair, including color, texture, length, hairstyle, etc. The body template image generally refers to an image used to display the morphology, structure, and proportion of the human body, and its purpose is to provide a standard or reference for alignment with the hair portion in the hair image.

[0094] The above step 203, namely "using the first object generation model to predict images of other perspectives based on the single-perspective image of the target object" is described in detail below in conjunction with the embodiments.

[0095] The first object generation model in the embodiment of the present application is pre-trained based on the diffusion model. That is, using a single-view image of a given target object at any perspective, the first object generation model is used to predict the corresponding images of other perspectives, so that images of other perspectives can be generated based on the first object generation model. The single-view image at any perspective and the images of other perspectives correspond to the images of the target object at different perspectives.

[0096] In one example, using the first object generation model based on a single-view image of the target object, such as an image of view θ0, to predict images of other viewpoints includes: Figure 3 As shown in , a single-view image is denoised, and the noisy single-view image and information of other viewpoints (e.g., a specified viewpoint θ1) are input into a first object generation model, and the first object generation model uses the information of other viewpoints as a guide condition to denoise the noisy single-view image to obtain images of other viewpoints, i.e., images based on the viewpoint θ1 of the target object. Optionally, the above-mentioned denoising may refer to randomly adding noise.

[0097] In the embodiment of the present application, when the first object generation model is used to predict images of other perspectives based on the single-perspective image of the target object, information of other perspectives can be specified. The single-perspective image is denoised, and the noisy single-perspective image and information of other perspectives are input into the first object generation model. The first object generation model uses the information of other perspectives as a guiding condition to denoise the noisy single-perspective image to obtain images of other perspectives.

[0098] Here, the guiding condition may refer to using the information of other perspectives as a constraint condition for the first object generation model to perform denoising on the noisy single-perspective image, so as to predict images of other perspectives.

[0099] By specifying information of multiple other perspectives in the above manner, the embodiment of the present application can generate images of multiple perspectives through the first object generation model to provide data support for reconstructing the three-dimensional model.

[0100] In one example, the first object generation model is trained in the following manner:

[0101] Acquire an image sample of a first perspective and an image sample of a second perspective;

[0102] After adding noise to the image samples of the first perspective, the noisy image samples of the first perspective and the information of the second perspective are input into the first diffusion model, and the first diffusion model is used to denoise the noisy image samples of the first perspective based on the information of the second perspective to obtain a predicted image of the second perspective, and the model parameters of the first diffusion model are updated using the loss function corresponding to the first training objective, wherein the first training objective includes minimizing the difference between the predicted image and the image samples of the second perspective, and the distance between the feature representation of the image samples of the first perspective and the feature representation of the predicted image.

[0103] Here, the image samples of the first perspective and the image samples of the second perspective are used as samples for training the first object generation model, and the image samples of the first perspective and the image samples of the second perspective are respectively image samples of the same target object with different perspectives.

[0104] When constructing the above-mentioned image samples of the first perspective and the second perspective, a combination of multiple different perspectives can be used as much as possible to cover as many perspective ranges as possible, so as to train the first object generation model to be able to generate clear images at different perspectives.

[0105] Next, after denoising the image samples of the first perspective, the noisy image samples of the first perspective and the information of the second perspective are input into the first diffusion model, and the first diffusion model denoises the noisy image samples of the first perspective based on the information of the second perspective to obtain a predicted image of the second perspective; next, the model parameters of the first diffusion model are updated using the loss function corresponding to the first training objective.

[0106] It should be noted that the above-mentioned noise addition process is the forward diffusion process in the first diffusion model, which can be regarded as a series of processes of gradually adding random noise. Specifically, the first diffusion model defines a diffusion Markov chain to slowly add random noise to the image samples of the first perspective. As the time step increases, the proportion of noise in the image samples of the first perspective becomes higher and higher.

[0107] The image samples of the first perspective and the image samples of the second perspective can be obtained by rendering an existing three-dimensional model for hair.

[0108] like Figure 4As shown in , the existing 3D model of hair can be obtained, and the 3D model is fused with UV texture maps (generally multiple UV texture maps, such as 5 to 188), and further aligned and spliced ​​with the 3D body template to obtain a 3D human body model, which includes the aligned human body part and the hair part. The 3D human body model is rendered to obtain hair images at different viewing angles, so that the image samples of the first viewing angle and the image samples of the second viewing angle can be formed in pairs.

[0109] It should also be noted that the existing 3D hair model can cover as many hairstyles and types as possible, especially braided hairstyles, providing complex hairstyles such as single ponytail, double ponytail, French braid, French twist, and bun. The first generative model trained in this way is very powerful in processing complex hairstyles and multi-view Figure 1 It shows significant advantages in terms of consistency, data dependence and computational efficiency, which can greatly improve the effect and application value of reconstructing three-dimensional hair models.

[0110] The above denoising process is the reverse diffusion process in the first diffusion model. The reverse diffusion process is to learn how to recover the predicted image of the second perspective from the image sample of the first perspective after adding noise in the forward diffusion process, so as to have the ability to generate a new image.

[0111] Furthermore, if Figure 5 As shown in , the first training objective can be used to adjust the model parameters of the first object generation model to minimize the difference between the predicted image of the second perspective and the image sample of the second perspective (the image loss function, i.e., L1 loss, can be used), and the distance between the feature representation of the image sample of the first perspective and the feature representation of the predicted image of the second perspective (the perceptual loss function, i.e., Perceptual loss, can be used). The distance between the feature representation of the image sample of the first perspective and the feature representation of the predicted image of the second perspective can be determined by the cosine formula.

[0112] Among them, the image loss function, namely L1 loss, ensures the closeness between the predicted image and the true image, and the perceptual loss function, namely Perceptual loss, ensures the semantic similarity between the predicted image of the second perspective and the image sample of the first time, that is, the basic semantics of the image sample of the first perspective is retained as much as possible when predicting the second perspective.

[0113] In the embodiment of the present application, by training on hair images at multiple perspectives, the first diffusion model can learn the performance of hair at different perspectives, thereby generating predicted images of other perspectives from image sample input from the first perspective.

[0114] The above step 205, namely "obtaining a three-dimensional model of the target object using images from other perspectives", is described in detail below in conjunction with an embodiment.

[0115] In various application scenarios such as panoramic display, game modeling, 3D product display, and virtual character display, a 3D model of the target object is required. To this end, the target object needs to be reconstructed in 3D. When images from other perspectives are obtained, the initial 3D model of the target object can be obtained based on the images from other perspectives.

[0116] As one of the achievable methods, when obtaining the initial 3D model of the target object based on images from other perspectives, the single-perspective image may be used as a basis, and step 203, i.e., "using the first object generation model to predict and obtain images from other perspectives based on the single-perspective image of the target object", may be performed multiple times to obtain images from other perspectives, wherein the images from other perspectives are images from multiple perspectives. Then, the images from multiple perspectives are used to perform 3D reconstruction on the target object to obtain a 3D model of the target object.

[0117] Among them, when using images from multiple perspectives to perform three-dimensional reconstruction on the target object, feature points can be extracted from the images from multiple perspectives and matched, and the three-dimensional coordinates of the feature points can be calculated using the matching of feature points and camera parameters to generate a three-dimensional point cloud or mesh model. The feature representation of each point on the surface of the target object is determined using the three-dimensional point cloud or mesh model to obtain a three-dimensional model of the target object. This part of the reconstruction process is an existing technology and will not be described in detail here.

[0118] The feature representation of each point on the surface of the target object can be represented by a three-dimensional Gaussian representation, an implicit field representation or a voxel representation.

[0119] 3D Gaussian representation is a method of modeling a point or region in 3D space using a Gaussian function. 3D Gaussian representation uses Gaussian distribution, color, and opacity to represent a point in space. In order to improve the flexibility and computational efficiency of the 3D model, the embodiment of the present application uses a 3D Gaussian representation of the 3D model, which can achieve fast rendering while maintaining high details.

[0120] Implicit Field Representation describes the shape and properties of a three-dimensional model by defining a field, such as a signed distance function field (SDF), an occupancy field (Occupancy Field), or a neural radiation field (NeRF). In an embodiment of the present application, the implicit field representation represents the shape and structure of a target object (such as hair) by a continuous implicit function, which can capture complex geometric forms more naturally; for example, when processing high-frequency details such as hair strands and braids, the implicit field representation can better express the continuity and details of the hair. It should also be noted that the implicit field representation can be generated directly through the network, which is suitable for combination with the deep learning framework of the neural network, and may bring higher generation accuracy.

[0121] Voxel representation refers to discretizing three-dimensional space into regular voxel (Volume Pixels) grids, which can represent the details of target objects (such as hair) at multiple resolutions, especially multi-level processing of the rough shape and local details of hair. Each voxel can store attribute data such as color and density. The advantage of using voxel representation in the embodiment of the present application is its simple and intuitive structure, which can be easily combined with existing three-dimensional processing tools and algorithms. In addition, the voxel grid is suitable for spatial subdivision and local adjustment, and can obtain better control over detail processing.

[0122] As another feasible method, when obtaining the initial three-dimensional model of the target object based on images from other perspectives, the three-dimensional model of the target object can be initialized first, for example, the model parameters of the three-dimensional model of the target object can be initialized; then, the images from other perspectives are used as true images to optimize the initialized three-dimensional model of the target object.

[0123] The true value image in the embodiment of the present application generally refers to an image with high precision and high authenticity in three-dimensional reconstruction, which is used as a reference standard to evaluate and optimize the accuracy of the three-dimensional model. Here, images from other perspectives are used as true value images to optimize the initialized three-dimensional model. Optionally, part or all of the images from multiple perspectives are used as true value images to optimize the initialized three-dimensional model of the target object.

[0124] Correspondingly, in this example, optimizing the three-dimensional model of the target object includes: determining the viewing angle corresponding to the true image; obtaining an image of the three-dimensional model of the target object rendered according to the viewing angle corresponding to the true image; and optimizing the three-dimensional model using the difference between the rendered image and the true image.

[0125] Specifically, after determining the perspective corresponding to the images of other perspectives (i.e., the true image), the three-dimensional model of the target object is obtained by rendering the image according to the perspective corresponding to the images of other perspectives, and the difference between the rendered image and the true image is used to optimize the three-dimensional model.

[0126] Here, the difference between the rendered image and the image from other perspectives may include color, brightness, contrast, texture, etc. Next, based on the difference information and the optimization goal, the 3D model is optimized. The optimization includes adjusting the feature representation of the points on the surface of the target object in the 3D space. Taking the 3D Gaussian representation as an example, the position, color, opacity, etc. of the Gaussian distribution of each point on the surface of the target object in the 3D space are continuously optimized. The optimization goal is to minimize the difference between the rendered image and the image from other perspectives.

[0127] After obtaining an image of the three-dimensional model of the target object rendered according to the perspective corresponding to images from other perspectives, the embodiment of the present application can identify the deficiencies of the three-dimensional model in terms of shape, texture, etc. by comparing the differences between the rendered image and the true image, and optimize these deficiencies, so that the three-dimensional model can be made closer to the target object in the real world, thereby improving its realism.

[0128] The above step 207, namely "using the second object generation model to generate images based on other perspectives to obtain optimized images of other perspectives", is described in detail below in conjunction with the embodiments.

[0129] The second object generation model in the embodiment of the present application is pre-trained based on the diffusion model. The second object generation model is used to further optimize the images of other perspectives obtained by the first object generation model to obtain optimized images of other perspectives. The optimized images of other perspectives have improved image quality compared to images of other perspectives.

[0130] Continuing from the above, the second object generation model refines each pixel in images from other perspectives on the basis of optimizing the three-dimensional hair model by the first object generation model, thereby enhancing the clarity of the hair texture and ensuring the consistency of the three-dimensional hair model in images from multiple perspectives.

[0131] In one example, using the second object to generate the model based on images from other perspectives to obtain optimized images from other perspectives includes:

[0132] like Figure 6 As shown in , images from other perspectives and noise images are input into the second object generation model, and the second object generation model uses the images from other perspectives as guidance conditions to denoise the noise images to obtain optimized images from other perspectives.

[0133] Here, the noise image may be a randomly generated noise image, such as sampling a random vector from a certain data distribution and generating a noise image based on the sampled random vector.

[0134] In one example, a noisy image refers to an image in which random interference or random signals are added. Noise causes the image to present some random, undesirable visual changes. Noise images can be obtained by using additive noise, multiplicative noise, uniform noise, etc.

[0135] The guiding condition may refer to a constraint condition for using images from other perspectives as a second object generation model to perform denoising on the noisy image to obtain optimized images from other perspectives. The images from other perspectives may be images predicted and outputted based on the single-perspective image of the target object using the first object generation model.

[0136] In one example, the second object generation model can be trained in the following manner:

[0137] Acquire an image sample whose texture clarity meets a preset condition of the target object, and perform noise processing on the image sample;

[0138] like Figure 7 As shown in, the denoised image samples and the noisy image are input into the second diffusion model, and the predicted image is obtained by denoising the noisy image based on the denoised image samples by the second diffusion model, and the model parameters of the second diffusion model are updated using the loss function corresponding to the second training objective, wherein the second training objective includes minimizing the difference between the predicted image and the image samples before denoising, and the distance between the feature representation of the denoised image samples and the feature representation of the predicted image.

[0139] Here, the image samples and noise images whose texture clarity of the target object meets the element conditions are used as samples for training the second object generation model. The texture clarity of the target object meets the preset conditions, which can be set according to the accuracy of the three-dimensional model and / or prior knowledge.

[0140] Next, after the image samples whose texture clarity meets the preset conditions are processed by denoising, the image samples after denoising are used as the images before optimization, in order to obtain an image with lower quality than the image samples by appropriate denoising as the input of the second diffusion model. Then, the image samples after denoising and the noise image are input into the second diffusion model, which obtains the predicted image by denoising the noise image based on the image samples after denoising; next, the model parameters of the second diffusion model are updated using the loss function corresponding to the second training target.

[0141] Furthermore, the model parameters of the second object generation model are adjusted using the second training objective to minimize the difference between the predicted image and the image sample before noise addition (the image loss function, i.e., L1 loss, can be used), and the distance between the feature representation of the image sample after noise addition and the feature representation of the predicted image (the perceptual loss function, i.e., Perceptual loss, can be used). The distance between the feature representation of the image sample after noise addition and the feature representation of the predicted image can be determined by the cosine formula.

[0142] Among them, the image loss function, namely L1 loss, ensures the closeness between the predicted image and the image samples before noise addition (that is, the image samples whose texture clarity of the target object meets the preset conditions), and the perceptual loss function, namely Perceptual loss, ensures the semantic similarity between the predicted image and the image samples before noise addition, that is, the basic semantics of the input image samples is retained as much as possible when predicting the optimized image.

[0143] It should be noted that the above-mentioned first object generation model and second object generation model are both implemented based on the diffusion model. The denoising network in the diffusion model (usually using UNet) predicts noise based on the guidance conditions and the feature representation corresponding to the input noise image, removes the predicted noise from the feature representation corresponding to the input noise image, and obtains the denoised feature representation after denoising for multiple time steps. The predicted image is then decoded based on the denoised feature representation through the decoding network. Since the diffusion model is the existing basic model, it will not be described in detail here.

[0144] The above step 209, namely "using optimized images from other viewing angles as the true value image to optimize the three-dimensional model of the target object", is described in detail below in conjunction with the embodiments.

[0145] In the embodiment of the present application, the optimized image obtained in step 207 is used as a true value image to further optimize the feature representation of each point in the three-dimensional model, so as to achieve the purpose of optimizing the three-dimensional model.

[0146] In one example, optimizing the three-dimensional model of the target object includes: determining the viewing angle corresponding to the true image; obtaining an image of the three-dimensional model of the target object rendered according to the viewing angle corresponding to the true image; and optimizing the three-dimensional model using the difference between the rendered image and the true image.

[0147] Specifically, after determining the viewing angle corresponding to the true image, an image of the three-dimensional model of the target object rendered according to the viewing angle is obtained, and the three-dimensional model is optimized using the difference between the rendered image and the optimized image of the corresponding viewing angle.

[0148] Here, the difference between the rendered image and the optimized image from other perspectives may include color, brightness, contrast, texture, etc. Next, the 3D model is optimized based on the difference information and the optimization target.

[0149] After obtaining an image rendered according to the perspective corresponding to the optimized images from other perspectives, the embodiment of the present application can identify the deficiencies of the three-dimensional model in terms of shape, texture, etc. by comparing the differences between the rendered image and the optimized images from other perspectives, and optimize these deficiencies, thereby making the three-dimensional model closer to the target object in the real world and improving its realism.

[0150] By comparing the method provided by this application with the existing single-view based reconstruction method for hair 3D model, the following conclusions are drawn:

[0151] 1) The three-dimensional hair model obtained by the existing methods has serious artifacts and blurred texture. The three-dimensional hair model generated by the method provided in the embodiment of the present application is more visually realistic, rich in geometric details, and can accurately present the structure and texture of complex hairstyles.

[0152] 2) Compared with existing methods, the method provided by this application improves the consistency of texture and shape presented at different viewing angles.

[0153] 3) Compared with the existing methods, the method provided by this application reduces the dependence on large-scale real data. The method provided by this application can still generate high-quality three-dimensional hair models without real data, especially in the reconstruction of complex hairstyles.

[0154] 4) Compared with the existing methods, the method provided by this application greatly improves the computing efficiency. When processing complex hairstyles, the reconstruction speed is increased by about 70%.

[0155] 5) Compared with the existing methods, the method provided by this application is superior in terms of indicators such as L1 loss, perceptual loss and PSNR (peak signal-to-noise ratio).

[0156] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0157] According to an embodiment of another aspect, an apparatus for generating a three-dimensional model is provided. Figure 8 A schematic block diagram of a device for generating a three-dimensional model according to an embodiment is shown. The device for generating a three-dimensional model is arranged at Figure 1 The server side of the architecture shown in Figure 1. Figure 8 As shown, the device for generating a three-dimensional model includes: an image acquisition unit 801, a first optimization unit 802, and a second optimization unit 803. The main functions of each component unit are as follows:

[0158] An image acquisition unit 801 is configured to acquire a single-view image of a target object;

[0159] The first optimization unit 802 is configured to use the first object generation model to predict images of other perspectives based on the single-perspective image of the target object; and use the images of other perspectives to obtain a three-dimensional model of the target object;

[0160] The second optimization unit 803 is configured to use the second object generation model to generate images based on other perspectives to obtain optimized images of other perspectives; and use the optimized images of other perspectives as true value images to optimize the three-dimensional model of the target object;

[0161] The first object generation model and the second object generation model are both pre-trained based on the diffusion model.

[0162] As one of the possible implementations, the images from other perspectives include images from multiple perspectives;

[0163] The first optimization unit 802 is further configured to: perform three-dimensional reconstruction on the target object using images from multiple perspectives to obtain a three-dimensional model of the target object.

[0164] As one possible implementation method, the first optimization unit 802 is further configured:

[0165] Initialize the three-dimensional model of the target object;

[0166] Images from other perspectives are used as true images to optimize the 3D model of the target object.

[0167] As one possible implementation method, the second optimization unit 803 is further configured:

[0168] Determine the viewing angle corresponding to the true image;

[0169] Obtain an image of the target object's three-dimensional model rendered according to the viewing angle corresponding to the true value image;

[0170] The difference between the rendered image and the ground-truth image is used to optimize the 3D model.

[0171] As one achievable manner, the first optimization unit 802 is further configured as follows:

[0172] The single-view image is denoised, and the noisy single-view image and information of other viewpoints are input into a first object generation model. The first object generation model uses the information of other viewpoints as a guiding condition to denoise the noisy single-view image to obtain images of other viewpoints.

[0173] As one of the achievable methods, the first object generation model is trained using the following method:

[0174] Acquire an image sample of a first perspective and an image sample of a second perspective;

[0175] After adding noise to the image samples of the first perspective, the noisy image samples of the first perspective and the information of the second perspective are input into the first diffusion model, and the first diffusion model is used to denoise the noisy image samples of the first perspective based on the information of the second perspective to obtain a predicted image of the second perspective, and the model parameters of the first diffusion model are updated using the loss function corresponding to the first training objective, wherein the first training objective includes minimizing the difference between the predicted image and the image samples of the second perspective, and the distance between the feature representation of the image samples of the first perspective and the feature representation of the predicted image.

[0176] As one achievable manner, the second optimization unit 803 is further configured as follows:

[0177] The images from other perspectives and the noise image are input into the second object generation model, and the second object generation model uses the images from other perspectives as guidance conditions to perform denoising on the noise image to obtain optimized images from other perspectives.

[0178] As one of the achievable methods, the second object generation model is trained in the following manner:

[0179] Acquire an image sample whose texture clarity meets a preset condition of the target object, and perform noise processing on the image sample;

[0180] The denoised image samples and the noisy image are input into a second diffusion model, a predicted image is obtained by denoising the noisy image based on the denoised image samples by the second diffusion model, and the model parameters of the second diffusion model are updated using a loss function corresponding to a second training objective, wherein the second training objective includes minimizing the difference between the predicted image and the image samples before denoising, and the distance between the feature representation of the denoised image samples and the feature representation of the predicted image.

[0181] As one of the possible ways to achieve this, the target object is hair;

[0182] The image acquisition unit 801 is further configured to: acquire a single-view hair image and a body template image, align and synthesize the hair part in the hair image and the body part in the body template image to obtain a single-view image.

[0183] As one of the possible implementations, the three-dimensional model adopts a three-dimensional Gaussian representation, an implicit field representation or a voxel representation.

[0184] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system and device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0185] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0186] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0187] And an electronic device, comprising:

[0188] one or more processors; and

[0189] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0190] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the methods in the aforementioned method embodiments when executed by a processor.

[0191] in, Fig. 9 The architecture of the electronic device is shown as an example, which may include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920 may be communicatively connected via a communication bus 930.

[0192] Among them, the processor 910 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solution provided in this application.

[0193] The memory 920 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device 900, and a basic input and output system (BIOS) 922 for controlling the low-level operation of the electronic device 900. In addition, a web browser 923, a data storage management system 924, and a device 925 for generating a three-dimensional model, etc. can also be stored. The above-mentioned device 925 for generating a three-dimensional model can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 920 and is called and executed by the processor 910.

[0194] The input / output interface 913 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0195] The network interface 914 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0196] The bus 930 comprises a pathway for transmitting information between the various components of the device (eg, the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920).

[0197] It should be noted that, although the above device only shows a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, a memory 920, a bus 930, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.

[0198] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0199] The technical solution provided by the present application is described in detail above. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.

Claims

1. A method for generating a three-dimensional model, characterized in that: The method comprises: Acquire a single-view image of the target object; Generate images of other perspectives based on the single-perspective image of the target object using the first object generation model; Obtaining a three-dimensional model of the target object using the images from other perspectives; Generate images of the other perspectives based on the second object generation model to obtain optimized images of the other perspectives; Using the optimized images from other viewing angles as true images to optimize the three-dimensional model of the target object; The first object generation model and the second object generation model are both pre-trained based on a diffusion model.

2. The method according to claim 1, characterized in that The images from other perspectives include images from multiple perspectives; Obtaining the three-dimensional model of the target object by using the images from other perspectives includes: performing three-dimensional reconstruction on the target object by using the images from multiple perspectives to obtain the three-dimensional model of the target object.

3. The method according to claim 1, characterized in that Obtaining the three-dimensional model of the target object by using the images from other perspectives includes: Initializing a three-dimensional model of the target object; The images from other perspectives are used as true-value images to optimize the three-dimensional model of the target object.

4. The method according to claim 1 or 3, characterized in that: Optimizing the three-dimensional model of the target object includes: Determine the viewing angle corresponding to the true value image; Obtain an image of the three-dimensional model of the target object rendered according to the viewing angle corresponding to the true value image; The three-dimensional model is optimized using the difference between the rendered image and the true image.

5. The method according to claim 1, characterized in that The method of using the first object generation model to predict images of other perspectives based on the single-perspective image of the target object comprises: The single-view image is denoised, and the noisy single-view image and the information of the other viewpoints are input into the first object generation model. The first object generation model uses the information of the other viewpoints as a guiding condition to denoise the single-view image after noisy, so as to obtain images of the other viewpoints.

6. The method according to claim 5, characterized in that The first object generation model is trained in the following manner: Acquire an image sample of a first perspective and an image sample of a second perspective; After adding noise to the image samples of the first perspective, the image samples of the first perspective after adding noise and information of the second perspective are input into a first diffusion model, the first diffusion model is obtained to denoise the image samples of the first perspective after adding noise based on the information of the second perspective to obtain a predicted image of the second perspective, and the model parameters of the first diffusion model are updated using a loss function corresponding to a first training objective, wherein the first training objective includes minimizing the difference between the predicted image and the image samples of the second perspective, and the distance between the feature representation of the image samples of the first perspective and the feature representation of the predicted image.

7. The method according to claim 1, characterized in that The step of using the second object generation model to generate the image from the other perspectives to obtain the optimized image from the other perspectives includes: The images from other perspectives and the noise image are input into a second object generation model, and the second object generation model uses the images from other perspectives as guidance conditions to perform denoising on the noise image to obtain optimized images from other perspectives.

8. The method according to claim 7, characterized in that The second object generation model is trained in the following manner: Acquire an image sample whose texture clarity of the target object meets a preset condition, and perform noise processing on the image sample; The denoised image samples and the noisy image are input into a second diffusion model, a predicted image is obtained by denoising the noisy image based on the denoised image samples by the second diffusion model, and the model parameters of the second diffusion model are updated using a loss function corresponding to a second training objective, wherein the second training objective includes minimizing the difference between the predicted image and the image samples before denoising, and the distance between the feature representation of the denoised image samples and the feature representation of the predicted image.

9. The method according to any one of claims 1 to 3 and 5 to 8, characterized in that: The target object is hair; Acquiring a single-view image of the target object includes: acquiring a single-view hair image and a body template image, aligning and synthesizing a hair part in the hair image and a body part in the body template image to obtain the single-view image.

10. The method according to any one of claims 1 to 3 and 5 to 8, characterized in that: The three-dimensional model adopts three-dimensional Gaussian representation, implicit field representation or voxel representation.

11. A device for generating a three-dimensional model, characterized in that: The device comprises: An image acquisition unit, configured to acquire a single-view image of a target object; A first optimization unit is configured to use a first object generation model to predict images of other perspectives based on a single-perspective image of the target object; and use the images of other perspectives to obtain a three-dimensional model of the target object; A second optimization unit is configured to generate an optimized image of the other perspective based on the image of the other perspective using a second object generation model; and optimize the three-dimensional model of the target object using the optimized image of the other perspective as a true value image; The first object generation model and the second object generation model are both pre-trained based on a diffusion model.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.