Method and apparatus for generating three-dimensional avatar, device, and storage medium
By acquiring images from multiple reference perspectives and optimizing the initial three-dimensional image information using multi-perspective feature fusion and diffusion models, the difficulty of generating high-precision and high-quality 3D images in existing technologies is solved, and efficient generation of high-quality 3D images is achieved.
Patent Information
- Application Number
- PCT/CN2024/140040
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-07
- Filing Date
- 2024-12-17
- Publication Date
- 2025-10-16
AI Technical Summary
Existing technologies have difficulty generating high-precision and high-quality 3D images, and machine learning-based generation methods are time-consuming and consume excessive computing resources.
By obtaining the initial three-dimensional image information of the target object and images under multiple reference perspectives, the initial three-dimensional image information is optimized using multi-perspective feature fusion and diffusion model to generate the target three-dimensional image information.
It improves the geometry and texture details of 3D images, reduces computing resource consumption and generation time, and achieves efficient generation of high-quality 3D images.
Smart Images

Figure CN2024140040_16102025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and storage medium for three-dimensional avatar generation
[0001] This application claims priority to the Chinese patent application for "Method, apparatus, device and storage medium for three-dimensional avatar generation", filed on April 7, 2024, with the title of "Method, apparatus, device and storage medium for three-dimensional avatar generation", and the application number of 202410410581.X, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device and computer-readable storage medium for three-dimensional avatar generation. BACKGROUND
[0003] It is desirable to generate and use diversified three-dimensional (3D) avatars in many application scenarios, including people's daily socializing, gaming, video, etc. A 3D avatar, also referred to as a 3D digital avatar, refers to a three-dimensional body that can reflect the visual features of an object figuratively. With the advancement of computer vision technology based on machine learning, it has become possible to generate 3D data using machine learning models. It is desirable to be able to generate 3D avatars with rich details. SUMMARY
[0004] In a first aspect of the present disclosure, a method for three-dimensional avatar generation is provided. The method comprises: obtaining initial three-dimensional avatar information of a target object and a plurality of reference images of the target object respectively under a plurality of reference view angles, each reference image corresponding to one of the plurality of reference view angles; generating a target image of the target object under a target view angle based on the plurality of reference images; and updating the initial three-dimensional avatar information based on the target image to obtain target three-dimensional avatar information of the target object.
[0005] In a second aspect of the present disclosure, an apparatus for three-dimensional avatar generation is provided. The apparatus comprises: an information obtaining module configured to obtain initial three-dimensional avatar information of a target object and a plurality of reference images of the target object respectively under a plurality of reference view angles, each reference image corresponding to one of the plurality of reference view angles; an image generating module configured to generate a target image of the target object under a target view angle based on the plurality of reference images; and an optimization module configured to update the initial three-dimensional avatar information based on the target image to obtain target three-dimensional avatar information of the target object.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, which is executable by a processor to implement the method of the first aspect.
[0008] It should be understood that the content described in this part of the content is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:
[0010] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0011] FIG. 2 shows a schematic block diagram of a system for 3D avatar generation according to some embodiments of the present disclosure;
[0012] FIG. 3 shows a schematic block diagram of an optimization subsystem according to some embodiments of the present disclosure;
[0013] FIG. 4 shows a flowchart of a process of 3D avatar generation according to some embodiments of the present disclosure;
[0014] FIG. 5 shows a block diagram of an apparatus for 3D avatar generation according to some embodiments of the present disclosure; and
[0015] FIG. 6 shows a block diagram of a device capable of implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] It can be understood that, before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means in accordance with relevant laws and regulations.
[0017] For example, in response to receiving a user's active request, a prompt message is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware, such as electronic devices, application programs, servers, or storage media, etc. that perform the operation of the technical solutions of the present disclosure according to the prompt message.
[0018] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0019] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0020] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0022] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.
[0023] In this document, unless explicitly stated otherwise, performing a step "in response to A" does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0024] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are to be understood as open-ended, i.e., "including but not limited to". The term "based on" is to be understood as "based at least in part on". The term "one embodiment" or "the embodiment" is to be understood as "at least one embodiment". The term "some embodiments" is to be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.
[0025] As used herein, the term“machine learning model” can learn the association between respective inputs and outputs from training data, such that after training is completed, the corresponding output can be generated for a given input. Deep learning is a type of machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. In this document, a“machine learning model” can also be referred to as a“machine learning network” or a“network,” which terms are used interchangeably herein.
[0026] As used herein, the term“3D avatar,” also referred to as a 3D digital avatar or a 3D object, refers to a three-dimensional entity that can visually reflect the visual features of an object. The 3D avatar information of a target object refers to the digital representation of the 3D avatar, which can represent the target object in any suitable form.
[0027] Example Environment
[0028] FIG. 1 illustrates a block diagram of an example environment 100 in which multiple implementations of the disclosure can be implemented. In the environment of FIG. 1, a system 105 is configured to generate a 3D avatar 132. The system 105 can be implemented at a user device 110 and / or a remote 3D avatar generation device 120. A 3D avatar, also referred to as a 3D digital avatar, refers to a three-dimensional entity, a three-dimensional model, or a three-dimensional asset that can visually reflect the visual features of an object. The objects that can be modeled in three dimensions by the system 105 can include, but are not limited to, a person (e.g., a head, a half-body, or a full-body avatar of a person), an animal, a plant, or other static and dynamic objects, and can even include composite objects or scenes, etc. It should be understood that the target objects described in some embodiments below are merely exemplary and are not intended to be limiting in any way. The embodiments described with reference to these example objects are applicable to other types of objects.
[0029] In some embodiments, the system 105 can generate the 3D avatar 132 based on a user request of the user device 110. In some embodiments, as shown in FIG. 1, the user request can include description information 115 for a target object. The description information 115 can include, for example, an image of the target object and / or text describing the target object, etc. In some embodiments, if the system 105 is running at the remote 3D avatar generation device 120 instead of locally at the user device 110, the user request can be sent to the 3D avatar generation device 120 via a network 130. The generated 3D avatar 132 can be sent to the user device 110 via the network 130. In some embodiments, the 3D avatar generation device 120 can generate 3D avatars based on other triggering events. In some embodiments, the 3D avatar generation device 120 can provide 3D avatar generation services in response to requests from multiple user devices.
[0030] User devices 110 can include any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia player, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. Servers include, but are not limited to, mainframe computers, edge computing nodes, computing devices in a cloud environment, and the like. 3D avatar generation devices 120 can be any electronic device with computing capabilities, including servers, blade servers, mainframes, edge devices, and other suitable computing devices.
[0031] It should be appreciated that the components and arrangements of the environment 100 illustrated in FIG. 1 are merely examples and that a computing system suitable for implementing the example implementations described in this disclosure can include one or more different components, other components, and / or different arrangements. Some specific examples of avatars and images are illustrated in FIG. 1 and other figures, but this is for illustrative purposes only and does not impose any limitation on the specific implementations of this disclosure.
[0032] As briefly mentioned above, machine learning models have been used for 3D avatar generation. In some approaches, a 3D avatar can be directly generated using a feed-forward generative network. However, such approaches are limited by the generation resolution of the generative network and the size of the training dataset, making it difficult to generate high-precision and high-quality 3D avatars. In other approaches, a way of score distillation sampling (SDS) is used to optimize the representation of an object in three-dimensional space to generate a 3D avatar. However, these approaches require a large amount of time and computing resources to be consumed, making it difficult to be used in large quantities in practical applications.
[0033] Embodiments of the present disclosure propose an approach for 3D avatar generation. According to various embodiments of the present disclosure, initial three-dimensional avatar information of a target object and a plurality of reference images of the target object at a plurality of reference perspectives are obtained. Each reference image corresponds to one of the plurality of reference perspectives. Based on the plurality of reference images, a target image of the target object at a target perspective is generated. Based on the target image, the initial three-dimensional avatar information is updated to obtain target three-dimensional avatar information of the target object.
[0034] In embodiments of the present disclosure, images at a plurality of different perspectives are used in optimizing the initial three-dimensional avatar information. These images at different perspectives can provide more information about the geometry and texture of the target object. In this way, by using images at different perspectives, a three-dimensional avatar with geometric and texture details can be generated.
[0035] Some example implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0036] Example system for 3D avatar generation
[0037] FIG. 2 illustrates a schematic block diagram of a system for 3D avatar generation, according to some embodiments of the present disclosure. The system of FIG. 2 may, for example, be implemented as the system 105 of FIG. 1. As shown in FIG. 2, the system 105 can include an initial generation subsystem 210, which is also referred to simply as subsystem 210. The initial generation subsystem 210 is configured to obtain initial three-dimensional avatar information 202 of a target object and a plurality of reference images 203 of the target object at a plurality of viewing angles, respectively. Each reference image corresponds to a viewing angle. In the following, these reference images can also be referred to as reference views. The initial three-dimensional avatar information 202 can represent the target object in three-dimensional space in various suitable manners. In some embodiments, the initial three-dimensional avatar information 202 can be represented by a mesh. In some embodiments, the initial three-dimensional avatar information 202 can be represented by a neural radiance field (NeRF).
[0038] These viewing angles can be any different angles. In particular, in the case of a certain number, these viewing angles can be uniformly distributed within a range of 360 degrees, to reflect as many appearances, e.g., geometry and texture, of the target object as possible. As an example and without any limitation, the plurality of viewing angles can include a 0° viewing angle, a 90° viewing angle, a 180° viewing angle, and a 270° viewing angle.
[0039] In some embodiments, the initial generation subsystem 210 can generate the plurality of reference images 203 and the initial three-dimensional avatar information 202 based on description information 201 about the target object. The description information 201 can include any suitable form of information to describe the desired target object. For example, as shown in FIG. 2, the description information 201 can include an image of the target object and / or text describing the target object.
[0040] The initial generation subsystem 210 can employ any suitable operation or algorithm to generate the reference images 203 and the initial avatar information 202. Illustratively, the initial generation subsystem 210 can generate the reference images at the plurality of viewing angles based on the description information 201 using an image generation model, and then perform three-dimensional reconstruction on these reference images using a three-dimensional reconstruction model to obtain the initial three-dimensional avatar information 202. It should be understood that the above description of the initial generation subsystem 210 is merely illustrative and is not intended to be any limitation.
[0041] In some embodiments, a plurality of candidate images of the target object can be obtained, each candidate image corresponding to a viewing angle. These candidate images can correspond to different viewing angles, or multiple images can be obtained for the same viewing angle. These candidate images can be presented to a user, and the plurality of reference images for subsequent optimization can be determined based on the received user selection. In such embodiments, the user can be allowed to select a view with higher quality for subsequent optimization according to the quality of each view. In this way, the quality of the optimization can be improved.
[0042] It should be appreciated that the above description of the initial generation subsystem 210 is merely exemplary and is not intended to be limiting in any way. In some embodiments, the system 105 can not include the initial generation subsystem 210, but can receive pre-acquired initial appearance information and reference images from an external source (e.g., by user input).
[0043] As shown in FIG. 1, the system 105 also includes an optimization subsystem 220. The optimization subsystem 220 is configured to update the initial three-dimensional appearance information 202 based on the reference images 203 under different viewpoints to obtain target three-dimensional appearance information 205 of the target object. Similar to the initial three-dimensional appearance information 202, the target three-dimensional appearance information 205 can represent the target object in a three-dimensional space in various suitable manners. In some embodiments, the target three-dimensional appearance information 205 can be represented by patches. In some embodiments, the target three-dimensional appearance information 205 can be represented by neural radiance fields.
[0044] To optimize the initial three-dimensional appearance information 202, the optimization subsystem 220 can generate a target image of the target object under a target viewpoint based on the plurality of reference images 203. To this end, the optimization subsystem 220 can include a machine learning model (e.g., a diffusion model) to generate the target image, as will be described below with reference to FIG. 3. In the following, the target image under the target viewpoint is also referred to as a target view. Then, the optimization subsystem 220 can update the initial three-dimensional appearance information 202 based on the target image to obtain the target three-dimensional appearance information 205. In other words, with the target image, the initial 3D appearance is optimized to a target 3D appearance.
[0045] Exemplarily, the operation of the optimization subsystem 220 can be described by the following expression. Given n+1 views as input, where each view captures the target object from a different viewpoint n . One goal of the optimization subsystem 220 is to synthesize a target view x under a target viewpoint x . To this end, the optimization subsystem 220 can compute the camera rotation i and translation relative to each viewpoint , i e (0, 1,..., n, x). The camera poses relative to each other can be represented as c i = (R i , T i ), i e (0, 1,..., n, x). Accordingly, the process of generating the target view can be represented as follows:
[0046] where representing a camera pose c x of a target view.
[0047] Although some embodiments above and below are described with reference to a single target view, it should be understood that this is merely exemplary and not intended to be limiting. In some embodiments, respective target images of a target object at multiple different target view angles can be generated separately. These target view angles can include the same view angles as the reference images, or can include different view angles from the reference images. These target images can then be utilized to iteratively update the initial three-dimensional appearance information 202. In this way, the generated 3D appearance can be optimized at various different view angles.
[0048] An example of the system 105 for 3D appearance generation is described above with reference to FIG. 2. In general, the system 105 utilizes multiple views of a target object to optimize an imperfect initial three-dimensional appearance information to obtain an optimized 3D appearance. As a result, the 3D appearance can have more geometric and textural details.
[0049] Example optimization of 3D appearance information
[0050] FIG. 3 illustrates a schematic block diagram of the optimization subsystem 220, according to some embodiments of the present disclosure. The architecture illustrated in FIG. 3 can be considered as an example implementation of the optimization subsystem 220, without intending to be limiting. As shown in FIG. 3, the optimization subsystem 220 can include a reference image processing branch 301 and a target image generation branch 302, which can be implemented by any suitable machine learning model, e.g., a diffusion model. In addition, as shown in FIG. 3, in some embodiments, the optimization subsystem 220 can further include an encoder 303 for encoding an image at a certain view (i.e., a conditional view 351) to extract global features. The conditional view 351 can be one of the reference views, or can be a different view from the reference views. In some embodiments, the conditional view is the best quality view among the multiple reference views, e.g., the frontal view.
[0051] The reference image processing branch 301 is configured to determine fused features of a target object at multiple reference view angles based on multiple reference images. The fused features can characterize the appearance features of the target object from multiple view angles. That is, the reference image processing branch 301 can be configured to implement multi-view feature fusion.
[0052] In some embodiments, the weights of the reference images under different reference view angles in the multi-view feature fusion, i.e., the conditional strengths of the reference images in the multi-view feature fusion, can be controlled. To this end, a plurality of control factors 333 corresponding to the plurality of reference view angles can be obtained, each control factor indicating the weight of the reference image under the corresponding reference view angle in the feature fusion. For example, the greater the control factor (which can also be referred to as a conditional label), the greater the proportion or degree of influence of the features of the reference image under the corresponding reference view angle in the fused features.
[0053] In some embodiments, the control factors 333 can be provided by a user, such as input via the user device 110. For example, the user can be presented with the reference images and prompt information, such as by the user device 110. The prompt information indicates that the user controls the weights of the different reference images in the multi-view feature fusion. In turn, the control factors 333 can be determined based on the received user input. In such embodiments, the user can be allowed to determine the size of the control factor according to the quality of the reference image or the view angle that the user is interested in, thereby adjusting the influence strength of different views on the fused features. For example, the influence of high-quality views can be enhanced, while the influence of low-quality views can be weakened.
[0054] In some embodiments, the control factors 333 can also be determined in other suitable manners. For example, the control factors 333 can be predetermined or determined in the training of a machine learning model. For another example, evaluation information (e.g., quality scores) of the quality of the reference images can be obtained, and the size of the control factor of each reference image can be determined according to the evaluation information.
[0055] In such embodiments, the reference image processing branch 301 can generate the fused features based on the plurality of reference images and the corresponding control factors. In addition, the reference view information 332 can also be introduced in the generation of the fused features to identify the view angles to which the reference images belong. In some embodiments, the generation of the fused features can also be based on time information, such as time 0 shown in FIG. 3. For example, if a diffusion model is used to implement the target image generation branch 302. The reference image processing branch 301 can only be executed at the beginning of the denoising process (i.e., at time 0). In this way, the computational resource consumption and time can be reduced, thereby improving the processing efficiency.
[0056] An example of the reference image processing branch 301 is described with reference to FIG. 3. The reference image processing branch 301 can include a plurality of processing blocks 310. Each processing block 310 can in turn include a feature fusion block 311, a self-attention block 312, and a cross-attention block 313.
[0057] The feature fusion block 311 can include a plurality of residual blocks, for example. The feature fusion block 311 in the first processing block 310 can receive the reference image 331, the reference view information 332, and optionally the control factor 333 and the temporal information, and generate the fused features in the first processing block 310. The feature fusion block 311 in the subsequent processing block 310 can receive the features output by the previous processing block, and generate the fused features in the processing block 310.
[0058] The self-attention block 312 can determine the value (V) features, the key (K) features, and the query (Q) features in the attention mechanism based on the fused features output by the feature fusion block 311, and apply the attention mechanism to the value features, the key features, and the query features. In this way, the output features of the self-attention block 312 can be obtained. The cross-attention block 313 can apply the cross-attention mechanism to the output features of the self-attention block 312 and the image features of the conditional view output by the encoder 303, thereby obtaining the output of the processing block 310.
[0059] The above describes the acquisition of the fused features. For the target image generation branch 302, an initial image 340 of the target object in the target view can be generated based on the initial three-dimensional avatar information. For example, the patches can be converted to neural radiance fields, and the image in the target view can be rendered. In this way, the target image generation branch 302 can generate the target image 355 in the target view from the initial image 340 based on the fused features from the reference image processing branch 301.
[0060] The generation of the target image 355 from the initial image 340 can utilize any suitable image generation, image optimization, image reconstruction technique. In some embodiments, the target image generation branch 302 can utilize or be implemented as a diffusion model. Accordingly, a predetermined noise can be added to the initial image 340 to obtain a noisy image 341. Then, the target image 355 can be generated from the noisy image 341 by a plurality of denoising steps. The fused features from the reference image processing branch 301 are used in each denoising step.
[0061] FIG. 3 illustrates an example implementation of a given denoising step (denoted by time t). The target image generation branch 302 can include a plurality of processing blocks 320. Each processing block 320 can in turn include a feature extraction block 321, a self-attention block 322, and a cross-attention block 323.
[0062] The feature extraction block 321 can include a plurality of residual blocks, for example. The feature extraction block 321 in the first processing block 320 can receive the noisy image 341, the target view information 342 and optionally the control factor 343 (note that the control factor can be fixed for the target view) and the temporal information (e.g. time t as shown in FIG. 3), and generate the image features in the first processing block 320, i.e. the features of the input image for the given denoising step. The feature extraction block 321 in the subsequent processing block 320 can receive the features output by the previous processing block, and generate the image features in the processing block 320.
[0063] The self-attention block 322 can determine the value (V) features and the key (K) features in the attention mechanism based on the fused features from the reference image processing branch 301 and the image features output by the feature extraction block 321. For example, as shown by arrows 361 and 362 in FIG. 3, the value features and the key features determined in the self-attention block 312 from the fused features can be concatenated to the value features and the key features determined from the image features. Thereby, the final value features and the key features of the self-attention block 322 can be obtained. The self-attention block 322 can determine the query (Q) features based on the image features output by the feature extraction block 321, and apply the attention mechanism to the value features, the key features and the query features. Thereby, the output features of the self-attention block 322 can be obtained. The cross-attention block 323 can apply the cross-attention mechanism to the output features of the self-attention block 322 and the image features of the conditional view output by the encoder 303, and thereby obtain the output of the processing block 320.
[0064] After the processing by the plurality of processing blocks 320, the output image for the given denoising step can be obtained. By a plurality of denoising steps, the final target image 355 can be obtained from the noisy image 341.
[0065] The above describes an example process of generating the target image 355 from the initial image 340. The following describes an example of updating the initial three-dimensional avatar information based on the target image 355. The goal of updating the initial three-dimensional avatar information based on the target image 355 is to feedback the geometric and texture details in the target image 355 to the initial three-dimensional avatar information to obtain the final three-dimensional avatar information.
[0066] In some embodiments, a loss in the process of generating the target image 355 from the initial image 340 under the target view can be determined, and the initial three-dimensional avatar information can be updated based on the loss. The initial image 340 can be rendered from the initial three-dimensional avatar information can be differentiable, so that the loss described above can be back propagated to the initial three-dimensional avatar information to optimize it. For example, if the initial three-dimensional avatar information is a mesh representation that is not differentiable, the mesh representation can be converted to a differentiable neural radiance field representation first, and the initial image 340 can be rendered according to the neural radiance field representation. In this way, the loss described above can be back propagated to the three-dimensional avatar information in the neural radiance field representation and updated.
[0067] In some embodiments, the loss described above can be an SDS loss. Specifically, a predetermined noise is added to the initial image 340 to obtain a noisy image 341. The target image 355 is generated based on the noisy image 341, as described above. In generating the target image 355, the target image generation branch 302 can generate a predicted noise. The difference between the predicted noise and the added predetermined noise can be used to determine the loss described above, which is back propagated to update the initial three-dimensional avatar information.
[0068] In such embodiments, the multi-view based SDS loss provides precise guidance for the optimization of the three-dimensional avatar information. While keeping the features of the initial generation result consistent with the multi-view images, the local details of the geometry and texture of the initial relatively rough 3D avatar are enriched. In addition, the use of the geometry and texture information provided by the multi-view also helps to shorten the optimization time.
[0069] Example training
[0070] The above describes an example of using multi-view to optimize an initial three-dimensional avatar. The training of the optimization subsystem 220, i.e., the training of the machine learning models that constitute the optimization subsystem 220, is described below. In some embodiments, in the training of the optimization subsystem 220, additional processing can be performed on the training samples to ensure the training effect.
[0071] In the inference process such as described above, the reference images can not be perfect, e.g., have artifacts, due to the performance of the initial generation subsystem 210, etc. In contrast, in training, the images in the training dataset as ground truth are typically perfect. To this end, in some embodiments, each reference view in the training sample can be degraded to make it more consistent with the actual inference situation. For example, Gaussian noise can be added to the reference view in the training sample as a perturbation. As another example, a portion of the reference views can be randomly scaled (e.g., down-sampled and then up-sampled) to produce blurry training inputs. In this way, the training samples can be made more consistent with the reference views in the actual inference, thereby ensuring that the trained optimization subsystem 220 is able to handle imperfect reference views.
[0072] Further, when trained with images under multiple views, the optimization subsystem 220 can be biased towards a certain view and ignore information from other views, thereby possibly leading to failure to maintain 3D consistency. To this end, in some embodiments, in training, one or more views can be randomly removed from the multiple reference views, thereby determining one or more remaining views for training. The optimization subsystem 220 can then be trained using the corresponding images under the one or more remaining views from the reference pairs in the training. For example, the training of the optimization subsystem 220 can include multiple training steps. In one or more of the steps, a reference image under one or more views can be randomly removed from the reference images under the multiple views. That is, the removed reference images do not participate in the training in that step. In this way, the network can be forced to synthesize information from all available views to generate more consistent results.
[0073] Example processes, apparatuses, and devices
[0074] FIG. 4 illustrates a flowchart of a process 400 for generation of a 3D avatar, in accordance with some embodiments of the present disclosure. The process 400 can be implemented at the system 105 of FIG. 1.
[0075] At block 410, the system 105 obtains initial 3D avatar information of a target object and multiple reference images of the target object under multiple reference views, respectively. Each reference image corresponds to one of the multiple reference views. At block 420, the system 105 generates a target image of the target object under a target view based on the multiple reference images. At block 430, the system 105 updates the initial 3D avatar information based on the target image to obtain target 3D avatar information of the target object.
[0076] In some embodiments, generating the target image of the target object under the target view includes: generating an initial image of the target object under the target view according to the initial three-dimensional appearance information; determining a fusion feature of the target object under the plurality of reference views based on the plurality of reference images; and obtaining the target image from the initial image based on the fusion feature.
[0077] In some embodiments, determining the fusion feature of the target object under the plurality of reference views includes: obtaining a plurality of control factors respectively corresponding to the plurality of reference views, each control factor indicating a weight of a reference image under a corresponding reference view in multi-view feature fusion; and generating the fusion feature based on the plurality of reference images and the plurality of control factors.
[0078] In some embodiments, obtaining the plurality of control factors respectively corresponding to the plurality of views includes: presenting the plurality of reference images and prompt information indicating that a user controls the weight of the plurality of reference images in multi-view feature fusion; and determining the plurality of control factors based on received user input.
[0079] In some embodiments, obtaining the target image from the initial image includes: obtaining the target image from the initial image with added noise through a plurality of denoising steps, wherein the fusion feature is used in the plurality of denoising steps.
[0080] In some embodiments, a given denoising step in the plurality of denoising steps includes: determining key features and value features in an attention mechanism based on features of an input image of the given denoising step and the fusion feature; determining query features in the attention mechanism based on the features of the input image; and determining an output image of the given denoising step based on the query features, the key features, and the value features.
[0081] In some embodiments, updating the initial three-dimensional appearance information includes: generating an initial image of the target object under the target view according to the initial three-dimensional appearance information; determining a loss in a process of generating the target image from the initial image; and updating the initial three-dimensional appearance information based on the loss.
[0082] In some embodiments, determining the loss includes: adding a predetermined noise to the initial image to obtain a noise image, wherein the target image is generated based on the noise image; and obtaining the loss based on a difference between the predicted noise in the generation of the target image and the predetermined noise.
[0083] In some embodiments, obtaining the plurality of reference images includes: presenting a plurality of candidate images of the target object, each candidate image corresponding to a view; and determining the plurality of reference images based on user selection of the plurality of candidate images.
[0084] In some embodiments, the target image is generated by using a machine learning model, and training of the machine learning model comprises: removing one or more of the plurality of reference perspectives randomly to determine one or more remaining perspectives; and training the machine learning model based on respective images of the reference object under the one or more remaining perspectives.
[0085] FIG. 5 illustrates a schematic structural block diagram of an apparatus 500 for 3D avatar generation, according to certain embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the system 105, the user device 110, or the 3D avatar generation device 120. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0086] As shown, the apparatus 500 includes an information obtaining module 510 configured to obtain initial three-dimensional avatar information of a target object and a plurality of reference images of the target object respectively under a plurality of reference perspectives, each reference image corresponding to one of the plurality of reference perspectives. The apparatus 500 includes an image generating module 520 configured to generate a target image of the target object under a target perspective based on the plurality of reference images. The apparatus 500 includes an optimization module 530 configured to update the initial three-dimensional avatar information based on the target image to obtain target three-dimensional avatar information of the target object.
[0087] In some embodiments, the image generating module 520 includes: an initial image generating module configured to generate an initial image of the target object under the target perspective according to the initial three-dimensional avatar information; a feature fusion module configured to determine a fusion feature of the target object under the plurality of reference perspectives based on the plurality of reference images; and an initial image converting module configured to obtain the target image from the initial image based on the fusion feature.
[0088] In some embodiments, the feature fusion module is further configured to: obtain a plurality of control factors respectively corresponding to the plurality of reference perspectives, each control factor indicating a weight of a reference image under a corresponding reference perspective in multi-perspective feature fusion; and generate the fusion feature based on the plurality of reference images and the plurality of control factors.
[0089] In some embodiments, the feature fusion module is further configured to: present the plurality of reference images and prompt information indicating that a user controls weights of the plurality of reference images in multi-perspective feature fusion; and determine the plurality of control factors based on received user input.
[0090] In some embodiments, the initial image converting module is further configured to: obtain the target image from the initial image with added noise through a plurality of denoising steps, wherein the fusion feature is used in the plurality of denoising steps.
[0091] In some embodiments, the initial image conversion module is further configured to: determine the key feature and the value feature in the attention mechanism based on the feature of the input image of the given denoising step and the fusion feature; determine the query feature in the attention mechanism based on the feature of the input image; and determine the output image of the given denoising step based on the query feature, the key feature and the value feature.
[0092] In some embodiments, the optimization module 530 includes: a loss determination module configured to generate an initial image of the target object at the target view angle according to the initial three-dimensional figure information; determine a loss in the process of generating the target image from the initial image; and a loss propagation module configured to update the initial three-dimensional figure information based on the loss.
[0093] In some embodiments, the loss determination module is further configured to: add a predetermined noise to the initial image to obtain a noise image, wherein the target image is generated based on the noise image; and obtain the loss based on a difference between the predicted noise in the generation of the target image and the predetermined noise.
[0094] In some embodiments, the information acquisition module 510 is further configured to: present a plurality of candidate images of the target object, each candidate image corresponding to a view angle; and determine the plurality of reference images based on user selection for the plurality of candidate images.
[0095] In some embodiments, the target image is generated using a machine learning model, and the training of the machine learning model includes: removing one or more view angles from the plurality of reference view angles randomly to determine one or more remaining view angles; and training the machine learning model based on the corresponding images of the reference object at the one or more remaining view angles.
[0096] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 illustrated in FIG. 6 is merely exemplary and should not be construed as any limitation to the functionality and scope of the embodiments described herein. The electronic device 600 illustrated in FIG. 6 can be used to implement the system 105, the user device 110, or the 3D figure generation device 120 of FIG. 1.
[0097] As shown in FIG. 6, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 can include, but are not limited to, one or more processors or processing units 610, memory 620, storage 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit(s) 610 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 620. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 600.
[0098] Electronic device 600 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 620 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 630 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 600.
[0099] Electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0100] Communication unit(s) 640 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines capable of communication through a communication connection. Thus, electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.
[0101] The input device 650 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., one or more devices that enable a user to interact with the electronic device 600, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 600 to communicate with one or more other electronic devices, as desired, via the communication unit 640. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0102] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0103] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0104] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The computer program product can also be an article of manufacture that comprises a computer readable medium. The medium can comprise a hard disk drive, a memory, or a floppy diskette, which can be accessed using a drive unit. Additionally, the medium can comprise a storage device that can store program codes. The storage device can include, but is not limited to, devices needing a platter and a read / write head, optical disk drives such as CD-ROM, DVD, Blu-ray Disc™ drives, memory devices such as flash drives, memory sticks, or any device that stores digital information. Additionally, the medium can include a single storage device or a plurality of storage devices.
[0107] Various implementations of the disclosure have been described in detail above. The foregoing description is exemplary and explanatory only, and not exhaustive, of the disclosed implementations. Many modifications and variations of the implementations described herein are possible and will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the implementations described. The description of the term used herein is chosen for the purpose of explaining the principles of the implementations, practical application, or improvement over technology in the field, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating a three-dimensional image, comprising: Acquire initial three-dimensional image information of a target object and a plurality of reference images of the target object at a plurality of reference viewing angles, each reference image corresponding to one of the plurality of reference viewing angles; generating a target image of the target object at a target perspective based on the multiple reference images; as well as Based on the target image, the initial three-dimensional image information is updated to obtain target three-dimensional image information of the target object.
2. The method according to claim 1, wherein generating a target image of the target object at a target perspective comprises: generating an initial image of the target object at the target viewing angle according to the initial three-dimensional image information; determining, based on the multiple reference images, fusion features of the target object under the multiple reference perspectives; as well as The target image is obtained from the initial image based on the fused features.
3. The method according to claim 2, wherein determining the fusion features of the target object under the multiple reference perspectives comprises: Acquire a plurality of control factors corresponding to the plurality of reference perspectives, each control factor indicating a weight of a reference image at a corresponding reference perspective in multi-perspective feature fusion; as well as The fusion feature is generated based on the multiple reference images and the multiple control factors.
4. The method according to claim 3, wherein obtaining a plurality of control factors respectively corresponding to the plurality of viewing angles comprises: Presenting the multiple reference images and prompt information, wherein the prompt information instructs a user to control weights of the multiple reference images in multi-view feature fusion; as well as Based on the received user input, the plurality of control factors are determined.
5. The method of claim 2, wherein obtaining the target image from the initial image comprises: The target image is obtained from the initial image with noise added thereto through a plurality of denoising steps, wherein the fused features are used in the plurality of denoising steps.
6. The method of claim 5, wherein a given denoising step in the plurality of denoising steps comprises: Determining key features and value features in an attention mechanism based on features of the input image of the given denoising step and the fused features; Determining query features in an attention mechanism based on features of the input image; as well as An output image of the given denoising step is determined based on the query feature, the key feature, and the value feature.
7. The method according to claim 1, wherein updating the initial 3D image information comprises: generating an initial image of the target object at the target viewing angle according to the initial three-dimensional image information; determining a loss in generating the target image from the initial image; as well as Based on the loss, the initial three-dimensional image information is updated.
8. The method of claim 7, wherein determining the loss comprises: adding predetermined noise to the initial image to obtain a noise image, wherein the target image is generated based on the noise image; as well as The loss is obtained based on a difference between noise predicted in generating the target image and the predetermined noise.
9. The method according to claim 1 , wherein acquiring the plurality of reference images comprises: presenting a plurality of candidate images of the target object, each candidate image corresponding to a perspective; as well as The plurality of reference images are determined based on user selections of the plurality of candidate images.
10. The method of claim 1 , wherein the target image is generated using a machine learning model, and the training of the machine learning model comprises: randomly removing one or more perspectives from the plurality of reference perspectives to determine one or more retained perspectives; as well as The machine learning model is trained based on corresponding images of a reference object at the one or more retained viewpoints.
11. A device for generating a three-dimensional image, comprising: an information acquisition module configured to acquire initial three-dimensional image information of a target object and a plurality of reference images of the target object at a plurality of reference viewing angles, each reference image corresponding to one of the plurality of reference viewing angles; an image generation module, configured to generate a target image of the target object at a target perspective based on the multiple reference images; as well as The optimization module is configured to update the initial three-dimensional image information based on the target image to obtain target three-dimensional image information of the target object.
12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Visual angle adjustment method and device for three-dimensional model
CN109242978A
Three-dimensional display method and device of virtual object, electronic equipment and storage medium
CN114399614A
Method and device for content generation, equipment and storage medium
CN116012561A
Three-dimensional virtual image generation method and device, equipment and storage medium
CN116468849A
Image generation from 3D model using neural network
US11403800B1