Three-dimensional model generation method and device, medium, equipment and product

Through technical means of receiving target images and generating three-dimensional models, the role customization problem in UGC products is solved, and the diversity and user experience of three-dimensional models are improved.

CN120088393APending Publication Date: 2025-06-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411379800.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve high-quality character configuration and creation in UGC products, and it is difficult for players to customize the three-dimensional model of standardized characters.

Method used

By receiving the target image, extracting image features, and generating the target features of the three-dimensional model based on the image generation model, a three-dimensional model representing the three-dimensional surface information of the target object is finally generated.

Benefits of technology

It realizes the generation of two-dimensional images of the target object to three-dimensional models, supports users to create new roles outside of standardized roles during the interaction process, and improves the diversity of role configurations and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088393A_ABST
    Figure CN120088393A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional model generation method and device, a medium, equipment and a product, and the method comprises the steps: receiving a target image used for generating a three-dimensional model, and the target image comprises an image corresponding to a target object; performing feature extraction on the target image to obtain image features; based on the image features and an image generation model, obtaining target features of a three-dimensional model corresponding to the target object; based on the target features, a target three-dimensional model corresponding to the target object is generated, and the target three-dimensional model is used for representing three-dimensional surface information corresponding to the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a medium, a device, and a product for generating a three-dimensional model. Background Art

[0002] In UGC (User Generated Content) products, it is usually necessary to support players to configure and create characters based on their needs. However, high-quality art production often requires high technical requirements, making it difficult for players to configure. In related technologies, it is usually possible to support users to perform personalized configuration on the models of existing characters in the database, such as configuring the skin of a character. In this process, the user operates based on the standardized characters and elements preset in the character library. Summary of the Invention

[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a method for generating a three-dimensional model. The method includes: Receiving a target image for generating a three-dimensional model, where the target image includes an image corresponding to a target object; Performing feature extraction on the target image to obtain image features; Based on the image features and an image generation model, obtaining target features of the three-dimensional model corresponding to the target object; Based on the target features, generating a target three-dimensional model corresponding to the target object, where the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

[0005] In a second aspect, the present disclosure provides an apparatus for generating a three-dimensional model. The apparatus includes: A receiving module, configured to receive a target image for generating a three-dimensional model, where the target image includes an image corresponding to a target object; A first processing module, configured to perform feature extraction on the target image to obtain image features; A second processing module, configured to obtain target features of the three-dimensional model corresponding to the target object based on the image features and an image generation model; A first generation module, configured to generate a target three-dimensional model corresponding to the target object based on the target features, where the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the method described in the first aspect are implemented.

[0007] In a fourth aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0008] In a fifth aspect, the present disclosure provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] Through the above technical solutions, the target image provided by the user can be processed to obtain the three-dimensional surface information corresponding to the target object, and the generation of the three-dimensional model corresponding to the target object from the two-dimensional image of the target object can be realized. Thus, it is possible to support the user to trigger the generation of new characters other than the standardized characters during the interaction process, improve the diversity of character configuration in the UGC product, and at the same time effectively simplify the configuration operation process of the three-dimensional model, reduce the technical requirements for players, and improve the user experience and interaction activity.

[0010] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In combination with the accompanying drawings and with reference to the following specific implementation manners, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a method for generating a three-dimensional model provided according to an embodiment of the present disclosure.

[0012] Figure 2 is a schematic diagram of a training process of an image generation model provided according to an embodiment of the present disclosure.

[0013] Figure 3 is a schematic diagram of a training process of a feature processing model provided according to an embodiment of the present disclosure.

[0014] Figure 4 is a block diagram of a device for generating a three-dimensional model provided according to an embodiment of the present disclosure.

[0015] Figure 5The structural schematic diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. Detailed implementation manners

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0017] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0018] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.

[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0026] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0027] Figure 1 As shown, it is a flowchart of a method for generating a three-dimensional model provided according to an embodiment of the present disclosure. As Figure 1 shown, the method may include: In step 11, a target image for generating a three-dimensional model is received, and the target image contains an image corresponding to a target object.

[0028] Among them, in this embodiment, the user can be supported to upload an image of the target object to generate a three-dimensional model corresponding to the target object. As an example, the three-dimensional model can be represented by 3D mesh. For example, when a user wants to add a character to a character library, the user can upload an image of the character, and then a corresponding three-dimensional model can be generated based on the image, that is, converting a two-dimensional character image into a three-dimensional model corresponding to the character. After that, the user can perform subsequent personalized configuration based on the three-dimensional model, such as skin texture map generation, etc. As an example, the target object may be a character object to be added by the user, and the image of the target object can be shown in the target image.

[0029] In step 12, feature extraction is performed on the target image to obtain image features.

[0030] As an example, the feature extraction of the target image can be based on the CLIP (Contrastive Language-Image Pre-Training) model. The CLIP model is a multi-modal pre-training model based on contrastive learning, which can be pre-trained with a large amount of paired data of images and texts to learn the alignment relationship between images and texts. In this disclosure, the encoder in the pre-trained CLIP model can be used to extract and encode the features of the target image to obtain the image features.

[0031] As another example, this step may include: Perform local image feature extraction on the target image to obtain the local features corresponding to the target image. Among them, the local features obtained by local image feature extraction can be used to represent the features of pixel units. For example, based on the feature representation accurate to pixels, it can be used to represent what content corresponds to certain pixels in the target image. As an example, local image feature extraction can be based on DINOv2 (Dual-Stage Implicit Object-Oriented Network), and DINOv2 is a vision model based on Transformer.

[0032] Perform global image feature extraction on the target image to obtain the global features corresponding to the target image.

[0033] Among them, the global features can be used to represent the element features included in the target image, such as globally representing what content is included in the target image. As an example, the global image feature extraction of the target image can be performed through the CLIP model. It should be noted that the acquisition and application of the above CLIP model, DINOv2 model, etc. can be realized based on the general methods in the art, and this disclosure does not limit this.

[0034] After that, based on the local features and the global features, generate the image features.

[0035] As an example, the local features and the global features can be feature-stitched to obtain the image features. Thus, the expression accuracy of the image features for the target object can be improved, so as to provide accurate and effective data support for the subsequent generation of the 3D model.

[0036] In step 13, based on the image features and the image generation model, obtain the target features of the 3D model corresponding to the target object.

[0037] Among them, the image generation model can be implemented based on a diffusion model, such as the Diffusion model, etc. For example, image features and Gaussian noise can be input into the image generation model, so that the Gaussian noise can be gradually denoised based on the image features and the image generation model to obtain target features.

[0038] In step 14, based on the target features, a target 3D model corresponding to the target object is generated, where the target 3D model is used to represent the 3D surface information corresponding to the target object. For example, the 3D surface information includes vertex information and triangular facet information of the 3D model.

[0039] Among them, the target 3D model corresponding to the target object can be a 3D mesh corresponding to the target object. The vertex information is used to represent points on the surface of the 3D model, and the triangular facet information is used to represent the faces on the surface of the 3D model, which can be represented by an array. Each triangular facet is composed of three vertex indices to indicate which vertices the triangular facet is composed of, so that the 3D model corresponding to the target object can be obtained. Subsequently, the user can perform texture mapping configuration based on the target 3D model to achieve the configuration of the character object.

[0040] Through the above technical solution, the target image provided by the user can be processed to obtain the 3D surface information corresponding to the target object, realizing the generation of the 3D model corresponding to the target object from the 2D image of the target object. Thus, it can support the user to trigger the generation of new characters other than the standardized characters during the interaction process, improve the diversity of character configuration in UGC products, and at the same time effectively simplify the configuration operation process of the 3D model, reduce the technical requirements for players, and improve the user experience and interaction activity.

[0041] In a possible embodiment, the image generation model can be determined in the following manner: Obtain a training 3D model. As an example, a 3D mesh corresponding to a character in the character library can be obtained as the training 3D model.

[0042] Based on a feature processing model, perform feature extraction on the training 3D model to obtain the 3D surface features corresponding to the training 3D model, and perform rendering processing on the training 3D model to obtain the rendered image corresponding to the training 3D model.

[0043] Among them, the number of vertices and triangular facets in different training 3D models may be different, and the input length corresponding to the feature extraction by the feature processing model is usually fixed. Therefore, performing feature extraction on the training 3D model based on the feature processing model to obtain the 3D surface features corresponding to the training 3D model may include: Perform point cloud extraction on the training 3D model to obtain the point cloud information corresponding to the training 3D model.

[0044] As an example, the number of point clouds can be preset in advance, so that different 3D models can be represented based on the same number of point clouds. For example, the number of point clouds can be set to 4096. Then, point cloud sampling can be performed on the surface of the training 3D model to sample out 4096 point clouds, and each training 3D model can be converted into a representation corresponding to 4096 point clouds.

[0045] Perform feature encoding on the point cloud information based on the feature processing model to obtain the 3D surface features corresponding to the training 3D model.

[0046] As an example, the point cloud information can be feature-encoded based on the encoder of the feature processing model, so that encoded features can be obtained, which are used as the 3D surface features.

[0047] Thus, through the above technical solutions, 3D models containing different vertices and triangular patches can be converted into point cloud feature representations, so that different 3D models can be represented by a fixed number of point clouds, ensuring the uniformity of the input data of the feature processing model, improving the accuracy and efficiency of feature extraction by the feature processing model, and providing accurate data support for the subsequent generation of 3D models based on 3D surface features.

[0048] As an example, it can be rendered through the general rendering method of 3D models in the field to obtain its corresponding image, so as to ensure the matching degree between the rendered image corresponding to the training 3D model and the 3D surface features.

[0049] After that, generate training samples based on the training 3D model, the 3D surface features, and the rendered image, and train a preset generation model based on the training samples to obtain the image generation model.

[0050] Among them, one training 3D model and the 3D surface features and rendered image corresponding to this training 3D model can be used as a set of training samples, so that training samples can be quickly constructed based on the training 3D model, without the need for manual annotation and verification, and at the same time, the accuracy of the training samples can be ensured, thereby ensuring the prediction accuracy of the image generation model obtained by training based on this training sample.

[0051] As an example, an exemplary implementation manner of training a preset generation model based on the training samples to obtain the image generation model is as follows, as Figure 2 shown: Extract features from the rendered image A1 to obtain rendered image features. Among them, the features of the rendered image can be extracted in the same way as the features of the target image mentioned above, which will not be elaborated here.

[0052] Add target noise to the three-dimensional surface feature B1 to obtain a noisy three-dimensional surface feature.

[0053] Among them, the target noise can be pre-set Gaussian noise. The pre-set generation model can be a diffusion model, such as the Diffusion model. The process of adding noise can be a step-by-step process. For example, it can start from the real data distribution (i.e., the three-dimensional surface feature), and gradually convert the data distribution into a standard Gaussian distribution by adding Gaussian noise in each step. The target noise in each step can be pre-configured and the target noise in each step is different.

[0054] In this step, when adding target noise to the three-dimensional surface feature, the time step corresponding to the target noise can be marked at the same time. Then, some details in the three-dimensional surface feature can be masked by adding target noise to the three-dimensional surface feature to obtain the noisy three-dimensional surface feature.

[0055] After that, based on the noisy three-dimensional surface feature, the rendered image feature, and the pre-set generation model, obtain the predicted noise.

[0056] As an example, in the process of predicting noise based on the generation model, the prediction can be combined with the rendered image feature. For example, the noisy three-dimensional surface feature and the rendered image feature can be fused through the cross-attention mechanism Cross Attention. In this process, a query vector Query can be generated based on the noisy three-dimensional surface feature, and a key vector Key and a value vector Value can be generated based on the rendered image feature. Then, a dot product operation can be performed using the query vector and the key vector, and the attention weight can be obtained through the softmax function. Then, the attention weight is multiplied by the value vector and the results are summed to obtain the fused feature. Among them, the methods of generating the query vector, the key vector, and the value vector can be based on the general calculation methods under the attention mechanism in the art, and the present disclosure does not limit this. After obtaining the fused feature, the fused feature can be input into the pre-set generation model so that the generation model can predict the added noise based on the fused feature to obtain the predicted noise. In this process, by fusing the noisy three-dimensional surface feature and the rendered image feature, the attention of the noisy three-dimensional surface feature can be adjusted based on the rendered image feature.

[0057] Based on the target noise and the predicted noise, determine the generation loss of the pre-set generation model.

[0058] Training the preset generation model based on the generated loss to obtain the image generation model. Among them, the training process is as follows Figure 2 shown.

[0059] Among them, the loss function can be calculated based on the loss function commonly used in diffusion models in this field. For example, it can be a distance function that determines the distribution between the target noise and the predicted noise, such as the KL (Kullback-Leibler Divergence) divergence, etc. Then, the parameters in the generation model can be adjusted based on the direction of minimizing the generated loss, and then the image generation model can be obtained. The adjustment method of the model parameters can be adjusted based on the parameter adjustment method of the diffusion model in this field, which is not limited here.

[0060] Thus, through the above technical solution, the model can be trained to obtain an image generation model, so that the image generation model can learn the feature distribution of the three-dimensional model, and thus the three-dimensional surface features of the corresponding three-dimensional model can be generated based on the image generation model, realizing the conversion from a two-dimensional image to a three-dimensional model.

[0061] As an example, the feature processing model includes an encoder and a decoder. For example, the feature processing model can be implemented based on a VAE (Variational Auto-Encoder) model. The feature processing model is determined in the following manner. Such as Figure 3 shown: Obtain a training three-dimensional model, the implementation method of which has been described in detail above.

[0062] After that, encode the training three-dimensional model A2 based on the encoder in the preset processing model to obtain the latent space feature B2 corresponding to the training three-dimensional model. As an example, the target feature of the three-dimensional model generated based on the image generation model is also represented as the latent space feature latent.

[0063] As an example, in this step, it can also be to first extract the point cloud of the training three-dimensional model, obtain the point cloud information corresponding to the training three-dimensional model, and then perform feature encoding on the point cloud information based on the encoder to obtain the latent space feature. Among them, the encoder Encoder is used to map the input data to a latent space to determine the feature distribution in the latent space and perform feature sampling to obtain the latent space feature latent.

[0064] Decode the latent space feature based on the decoder in the preset processing model to obtain the predicted voxel feature, where the voxel feature contains the directed distance field information of multiple voxels.

[0065] Among them, the signed distance field (SDF) information is used to represent the shortest distance from a point to the object surface. If the point is outside the object, its value is positive; if the point is inside the object, its value is negative; if the point is on the object surface, its value is zero. In this embodiment, the decoder Decoder can be used to decode the latent space features latent to reconstruct the input data.

[0066] After that, based on the latent space features and the predicted voxel features, the prediction loss of the preset processing model is determined, and the preset processing model is trained based on the prediction loss to obtain the feature processing model.

[0067] As an example, the determining the prediction loss of the preset processing model based on the latent space features and the predicted voxel features may include: Determine a first loss based on the distribution corresponding to the latent space features.

[0068] Among them, the VAE model performs feature encoding to make the variable distribution in the latent space as close as possible to the standard normal distribution. Therefore, in this step, the distance between the distribution corresponding to the latent space features and the standard normal distribution can be used as the first loss. For example, the KL divergence between the two distributions can be used as the first loss.

[0069] Perform voxel conversion on the training three-dimensional model to obtain the target voxel features corresponding to the training three-dimensional model, and determine a second loss based on the target voxel features and the predicted voxel features.

[0070] In the feature processing model, it is necessary to ensure the accuracy of reconstructing the original input. In this step, the second loss can be determined based on the predicted voxel features and the true voxel features corresponding to the training three-dimensional model. In this process, the training three-dimensional model can be first converted into voxel features based on the dimension of the voxel features to represent it in the corresponding voxel feature form. For example, if the voxel features are represented by voxels of 256×256×256, and the value of each voxel is the signed distance field information of the point represented by the voxel, the training three-dimensional model can be converted into a voxel representation of 256×256×256, and the value of each voxel can be determined based on the three-dimensional model to obtain the target voxel features.

[0071] Furthermore, the space corresponding to the voxel features is a sparse space. To avoid the influence of voxels in the background on model training and data calculation, the surface voxels corresponding to the surface can be determined from the target voxel features.

[0072] As an example, the determining the second loss based on the target voxel features and the predicted voxel features may include: Voxels with the absolute value of the signed distance field information in the target voxel features less than the target threshold are used as the surface voxels; obtain the predicted SDF corresponding to the surface voxels in the predicted voxel features, and determine the second loss based on the target SDF of the surface voxels in the target voxel features and the predicted SDF in the predicted voxel features. For example, the second loss can be determined based on the mean-square error (MSE).

[0073] Based on the first loss and the second loss, determine the predicted loss.

[0074] As an example, the weighted sum of the first loss and the second loss can be used as the predicted loss. The weights corresponding to different losses can be set in advance.

[0075] After obtaining the predicted loss, the processing model can be trained based on the predicted loss according to the model parameter adjustment method of the VAE model in the art, so as to obtain the feature processing model. The training process of the feature processing model is as Figure 3 shown.

[0076] Thus, through the above technical solution, the feature processing model can be obtained. Through the feature processing model, the training three-dimensional model can be encoded into the low-dimensional space to obtain the latent space features, so as to reduce the redundancy in the three-dimensional space and reduce the generation difficulty of the three-dimensional model. And in this process, the predicted loss is determined based on the distribution of the latent space features and the voxel features to train the model, so that better surface detail quality can be obtained while ensuring the accuracy of feature encoding.

[0077] In a possible embodiment, the generating the target three-dimensional model corresponding to the target object based on the target features may include: Decode the target features to obtain the voxel features corresponding to the target features, where the voxel features contain the signed distance field information of multiple voxels. Among them, the target features can be input into the decoder of the feature processing model to obtain the voxel features corresponding to the target features.

[0078] After that, generate the target three-dimensional model based on the target voxels in the voxel features, where the target voxels are voxels with the absolute value of the signed distance field information less than the target threshold.

[0079] Among them, the target threshold can be set according to the actual application scenario. Voxels with the absolute value of the signed distance field information less than the target threshold can be considered as voxels near the surface of the three-dimensional object. Then, the above target voxels can be used as the surface voxels of the three-dimensional object, and further construct the target three-dimensional model of its surface features.

[0080] As an example, based on the target voxels in the voxel features, the target 3D model can be obtained through the Marching Cube algorithm, which can take the isosurface in space and represent it in the triangular facet mesh data format.

[0081] Thus, through the above technical solution, the target features can be decoded to obtain voxel features and further a target 3D model can be obtained based on the voxel features, so as to realize the conversion from the target image to the 3D model and improve the diversity of user role configuration.

[0082] Based on the same inventive concept, the present disclosure also provides a device for generating a 3D model, as Figure 4 shown, the device 10 includes: A receiving module 100, configured to receive a target image for generating a 3D model, where the target image includes an image corresponding to a target object; A first processing module 200, configured to perform feature extraction on the target image to obtain image features; A second processing module 300, configured to obtain target features of a 3D model corresponding to the target object based on the image features and an image generation model; A first generation module 400, configured to generate a target 3D model corresponding to the target object based on the target features, where the target 3D model is used to represent three-dimensional surface information corresponding to the target object.

[0083] Optionally, the image generation model is obtained through a first determination module, and the first determination module includes: An acquisition sub-module, configured to acquire a training 3D model; A first processing sub-module, configured to perform feature extraction on the training 3D model based on a feature processing model to obtain three-dimensional surface features corresponding to the training 3D model, and perform rendering processing on the training 3D model to obtain a rendered image corresponding to the training 3D model; A first training sub-module, configured to generate training samples according to the training 3D model, the three-dimensional surface features, and the rendered image, and train a preset generation model based on the training samples to obtain the image generation model.

[0084] Optionally, the first processing sub-module includes: A first extraction sub-module, configured to perform point cloud extraction on the training 3D model to obtain point cloud information corresponding to the training 3D model; A first encoding sub-module, configured to perform feature encoding on the point cloud information based on the feature processing model to obtain three-dimensional surface features corresponding to the training 3D model.

[0085] Optionally, the first training sub-module includes: A second processing sub-module, configured to extract features from the rendered image to obtain rendered image features; An adding sub-module, configured to add target noise to the three-dimensional surface features to obtain noise-added three-dimensional surface features; A third processing sub-module, configured to obtain predicted noise based on the noise-added three-dimensional surface features, the rendered image features, and the preset generation model; A fourth processing sub-module, configured to determine the generation loss of the preset generation model based on the target noise and the predicted noise; A second training sub-module, configured to train the preset generation model based on the generation loss to obtain the image generation model.

[0086] Optionally, the first generation module includes: A first decoding sub-module, configured to decode the target features to obtain voxel features corresponding to the target features, where the voxel features include directed distance field information of multiple voxels; A first generation sub-module, configured to generate the target three-dimensional model based on target voxels in the voxel features, where the target voxels are voxels with an absolute value of the directed distance field information less than a target threshold.

[0087] Optionally, the feature processing model includes an encoder and a decoder, and the feature processing model is obtained through a second determination module, where the second determination module includes: An acquisition sub-module, configured to acquire a training three-dimensional model; A second encoding sub-module, configured to encode the training three-dimensional model based on the encoder in the preset processing model to obtain latent space features corresponding to the training three-dimensional model; A second decoding sub-module, configured to decode the latent space features based on the decoder in the preset processing model to obtain predicted voxel features, where the voxel features include directed distance field information of multiple voxels; A third training sub-module, configured to determine the prediction loss of the preset processing model based on the latent space features and the predicted voxel features, and train the preset processing model based on the prediction loss to obtain the feature processing model.

[0088] Optionally, the third training sub-module includes: A first determination sub-module, configured to determine a first loss based on the distribution corresponding to the latent space features; A second determination sub-module, configured to perform voxel conversion on the training three-dimensional model to obtain target voxel features corresponding to the training three-dimensional model, and determine a second loss based on the target voxel features and the predicted voxel features; A third determination sub-module, configured to determine the prediction loss based on the first loss and the second loss.

[0089] Optionally, the first processing module includes: A second extraction sub-module, configured to perform local image feature extraction on the target image to obtain local features corresponding to the target image; A third extraction sub-module, configured to perform global image feature extraction on the target image to obtain global features corresponding to the target image; A second generation sub-module, configured to generate the image features based on the local features and the global features.

[0090] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0091] As Figure 5 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0092] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5An electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or have all the devices shown. Instead, more or fewer devices may be implemented or had.

[0093] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, the above functions defined in the methods of the embodiments of the present disclosure are performed.

[0094] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0096] The above computer-readable medium can be included in the above electronic device; it can also exist separately and not be assembled into the electronic device.

[0097] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: receive a target image for generating a three-dimensional model, where the target image contains an image corresponding to a target object; extract features from the target image to obtain image features; based on the image features and an image generation model, obtain target features of the three-dimensional model corresponding to the target object; and based on the target features, generate a target three-dimensional model corresponding to the target object, where the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

[0098] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0100] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the receiving module can also be described as "the module for receiving the target image for generating a three-dimensional model".

[0101] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0103] According to one or more embodiments of the present disclosure, Example 1 provides a method for generating a three-dimensional model, wherein the method includes: Receiving a target image for generating a three-dimensional model, where the target image contains an image corresponding to a target object; Performing feature extraction on the target image to obtain image features; Based on the image features and an image generation model, obtaining target features of the three-dimensional model corresponding to the target object; Based on the target features, generating a target three-dimensional model corresponding to the target object, where the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

[0104] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the image generation model is determined by the following method: Obtaining a training three-dimensional model; Performing feature extraction on the training three-dimensional model based on a feature processing model to obtain three-dimensional surface features corresponding to the training three-dimensional model, and performing rendering processing on the training three-dimensional model to obtain a rendered image corresponding to the training three-dimensional model; Generating a training sample according to the training three-dimensional model, the three-dimensional surface features, and the rendered image, and training a preset generation model based on the training sample to obtain the image generation model.

[0105] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein the performing feature extraction on the training three-dimensional model based on a feature processing model to obtain three-dimensional surface features corresponding to the training three-dimensional model includes: Performing point cloud extraction on the training three-dimensional model to obtain point cloud information corresponding to the training three-dimensional model; Performing feature encoding on the point cloud information based on the feature processing model to obtain three-dimensional surface features corresponding to the training three-dimensional model.

[0106] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2, wherein the training a preset generation model based on the training sample to obtain the image generation model includes: Performing feature extraction on the rendered image to obtain rendered image features; Adding target noise to the three-dimensional surface features to obtain noise-added three-dimensional surface features; Based on the noise-added three-dimensional surface features, the rendered image features, and the preset generation model, obtaining predicted noise; Based on the target noise and the predicted noise, determining the generation loss of the preset generation model; Training the preset generation model based on the generated loss to obtain the image generation model.

[0107] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein generating the target three-dimensional model corresponding to the target object based on the target feature includes: Decoding the target feature to obtain the voxel feature corresponding to the target feature, where the voxel feature contains the signed distance field information of multiple voxels; Generating the target three-dimensional model based on the target voxels in the voxel feature, where the target voxels are the voxels whose absolute value of the signed distance field information is less than the target threshold.

[0108] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 2, wherein the feature processing model includes an encoder and a decoder, and the feature processing model is determined by the following method: Obtaining a training three-dimensional model; Encoding the training three-dimensional model based on the encoder in the preset processing model to obtain the latent space feature corresponding to the training three-dimensional model; Decoding the latent space feature based on the decoder in the preset processing model to obtain the predicted voxel feature, where the voxel feature contains the signed distance field information of multiple voxels; Determining the prediction loss of the preset processing model based on the latent space feature and the predicted voxel feature, and training the preset processing model based on the prediction loss to obtain the feature processing model.

[0109] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6, wherein determining the prediction loss of the preset processing model based on the latent space feature and the predicted voxel feature includes: Determining a first loss based on the distribution corresponding to the latent space feature; Performing voxel conversion on the training three-dimensional model to obtain the target voxel feature corresponding to the training three-dimensional model, and determining a second loss based on the target voxel feature and the predicted voxel feature; Determining the prediction loss based on the first loss and the second loss.

[0110] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 1, wherein extracting the image feature from the target image includes: Extracting local image features from the target image to obtain the local features corresponding to the target image; Extract global image features from the target image to obtain the global features corresponding to the target image; Generate the image features based on the local features and the global features.

[0111] According to one or more embodiments of the present disclosure, Example 9 provides a device for generating a three-dimensional model, the device including: A receiving module, configured to receive a target image for generating a three-dimensional model, where the target image includes an image corresponding to a target object; A first processing module, configured to perform feature extraction on the target image to obtain image features; A second processing module, configured to obtain target features of a three-dimensional model corresponding to the target object based on the image features and an image generation model; A first generation module, configured to generate a target three-dimensional model corresponding to the target object based on the target features, where the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

[0112] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method described in any one of Examples 1-8 are implemented.

[0113] According to one or more embodiments of the present disclosure, Example 11 provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-8.

[0114] According to one or more embodiments of the present disclosure, Example 12 provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of Examples 1-8 are implemented.

[0115] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0116] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0117] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the foregoing embodiments, the specific manner in which each module performs an operation has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A method for generating a three-dimensional model, characterized in that: The method comprises: Receiving a target image for generating a three-dimensional model, wherein the target image includes an image corresponding to a target object; Performing feature extraction on the target image to obtain image features; Based on the image features and the image generation model, obtaining target features of the three-dimensional model corresponding to the target object; Based on the target features, a target three-dimensional model corresponding to the target object is generated, wherein the target three-dimensional model is used to represent three-dimensional surface information corresponding to the target object.

2. The method according to claim 1, characterized in that The image generation model is determined by: Obtain a training 3D model; Performing feature extraction on the training three-dimensional model based on the feature processing model to obtain three-dimensional surface features corresponding to the training three-dimensional model, and performing rendering processing on the training three-dimensional model to obtain a rendered image corresponding to the training three-dimensional model; A training sample is generated according to the training three-dimensional model, the three-dimensional surface features and the rendered image, and a preset generation model is trained based on the training sample to obtain the image generation model.

3. The method according to claim 2, characterized in that The extracting features of the training three-dimensional model based on the feature processing model to obtain the three-dimensional surface features corresponding to the training three-dimensional model includes: Performing point cloud extraction on the training three-dimensional model to obtain point cloud information corresponding to the training three-dimensional model; The point cloud information is feature encoded based on the feature processing model to obtain three-dimensional surface features corresponding to the training three-dimensional model.

4. The method according to claim 2, characterized in that: The step of training a preset generation model based on the training sample to obtain the image generation model includes: Performing feature extraction on the rendered image to obtain rendered image features; Adding target noise to the three-dimensional surface feature to obtain a noisy three-dimensional surface feature; Obtaining predicted noise based on the noisy three-dimensional surface features, the rendered image features and the preset generation model; Determining a generation loss of the preset generation model based on the target noise and the predicted noise; The preset generation model is trained based on the generation loss to obtain the image generation model.

5. The method according to claim 1, characterized in that The step of generating a target three-dimensional model corresponding to the target object based on the target feature includes: Decoding the target feature to obtain a voxel feature corresponding to the target feature, wherein the voxel feature includes directed distance field information of a plurality of voxels; The target three-dimensional model is generated based on the target voxel in the voxel feature, wherein the target voxel is a voxel whose absolute value of the signed distance field information is less than a target threshold.

6. The method according to claim 2, characterized in that The feature processing model includes an encoder and a decoder, and the feature processing model is determined in the following manner: Obtain a training 3D model; Encoding the training three-dimensional model based on an encoder in a preset processing model to obtain latent space features corresponding to the training three-dimensional model; Decoding the latent space feature based on the decoder in the preset processing model to obtain a predicted voxel feature, wherein the voxel feature includes the directed distance field information of a plurality of voxels; Based on the latent space features and the predicted voxel features, the prediction loss of the preset processing model is determined, and the preset processing model is trained based on the prediction loss to obtain the feature processing model.

7. The method according to claim 6, characterized in that The step of determining the prediction loss of the preset processing model based on the latent space feature and the predicted voxel feature comprises: Determining a first loss based on a distribution corresponding to the latent space feature; Performing voxel conversion on the training three-dimensional model to obtain target voxel features corresponding to the training three-dimensional model, and determining a second loss based on the target voxel features and the predicted voxel features; The predicted loss is determined based on the first loss and the second loss.

8. The method according to claim 1, characterized in that: The step of extracting features from the target image to obtain image features includes: Performing local image feature extraction on the target image to obtain local features corresponding to the target image; Performing global image feature extraction on the target image to obtain global features corresponding to the target image; The image feature is generated based on the local feature and the global feature.

9. A three-dimensional model generation device, characterized in that: The device comprises: A receiving module, used for receiving a target image for generating a three-dimensional model, wherein the target image includes an image corresponding to a target object; A first processing module is used to extract features from the target image to obtain image features; A second processing module, configured to obtain target features of a three-dimensional model corresponding to the target object based on the image features and the image generation model; The first generating module is used to generate a target three-dimensional model corresponding to the target object based on the target feature, wherein the target three-dimensional model is used to represent the three-dimensional surface information corresponding to the target object.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 8 are implemented.

11. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.