A method, apparatus and system for object generation

By combining a 2D image generation model and a 3D object generation model, and by using the Transformer structure and NeRF network to optimize the loss function, the problem of slow speed and poor quality of 3D object generation in existing technologies is solved, and high-quality and fast 3D object generation is achieved.

CN116228959BActive Publication Date: 2026-01-23HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211515098.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2026-01-23
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

Existing technologies generate 3D objects that are not realistic enough, lack texture, are slow to generate, consume a lot of resources, and require a long optimization process.

Method used

This paper combines a 2D image generation model and a 3D object generation model. By generating 2D images from multiple perspectives and calculating their similarity, the paper optimizes the 3D object generation process. It uses a Transformer structure and a NeRF network for image rendering and optimizes the loss function and perspective contrast loss function to improve generation quality and speed.

Benefits of technology

It improves the quality of generated 3D objects, enhances viewpoint consistency and texture realism, significantly speeds up the generation process, and reduces resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116228959B_ABST
    Figure CN116228959B_ABST
Patent Text Reader

Abstract

The application provides a method for generating an object, comprising: inputting text into a two-dimensional picture generation model to output two-dimensional pictures of multiple perspectives of an object; the text is used to describe characteristics of the object, the characteristics including an object category, a color and a shape; the two-dimensional picture generation model is used to generate the two-dimensional pictures of the multiple perspectives according to the text; the perspective is a spatial angle at which the object is presented; a similarity value of the two-dimensional pictures of the multiple perspectives and the text is calculated, and two-dimensional pictures of multiple perspective-enhanced perspectives are obtained according to the similarity value; and the two-dimensional pictures of the multiple perspective-enhanced perspectives are input into a three-dimensional object generation model, the three-dimensional object generation model renders two-dimensional pictures of other perspectives based on the two-dimensional pictures of the multiple perspective-enhanced perspectives, and outputs a three-dimensional object that meets the text description. The application generates two-dimensional pictures of multiple perspectives of a corresponding object according to text by using a two-dimensional picture generation model, and generates a corresponding 3D object according to the two-dimensional pictures of the multiple perspectives by using a three-dimensional object generation model, so that the quality of the generated 3D object can be improved and the speed of generating the 3D object can be accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device and system for object generation. BACKGROUND

[0002] The goal of a text-guided 3D object generation task is to generate a 3D image of an object that matches the description in a sentence given the sentence, where the sentence describes the characteristics of the object, such as the class, color, shape, etc. of the object. Figure 1 As shown in the example, the left is the text "charcoal leather loveseat", and the right is the generated 3D object, a 2D picture rendered.

[0003] Text-guided 3D object generation technology can be used in (1) data augmentation: such as generating training data for autonomous driving scenarios; (2) entertainment, content creation: users generate corresponding 3D objects through text, which improves user experience and increases efficiency in practical application scenarios.

[0004] The 3D objects generated by existing technologies are not realistic and lack real textures; optimization is required for each text, which requires a large amount of resources and a long time, and the 3D object generation speed is very slow.

[0005] How to improve the quality of the generated 3D objects and speed up the 3D object generation speed is a problem that needs to be solved. SUMMARY

[0006] The embodiments of the present application provide a method, device, system and computer and equipment for object generation, which can improve the quality of the generated 3D objects and speed up the 3D object generation speed.

[0007] In a first aspect, the embodiments of the present application provide a method for generating an object, comprising: inputting a text into a two-dimensional picture generation model to output a plurality of two-dimensional pictures of different perspectives of the object; the text is used to describe features of the object, the features including a category, a color and a shape of the object; the two-dimensional picture generation model is used to generate the plurality of two-dimensional pictures of different perspectives according to the text; the perspective is a spatial angle at which the object is presented; calculating a similarity value between the plurality of two-dimensional pictures of different perspectives and the text, and obtaining a plurality of two-dimensional pictures of enhanced perspectives according to the similarity value; inputting the plurality of two-dimensional pictures of enhanced perspectives into a three-dimensional object generation model, and rendering a two-dimensional picture of another perspective based on the plurality of two-dimensional pictures of enhanced perspectives, and outputting a three-dimensional object that meets the description of the text. In this way, the two-dimensional picture generation model can be used to generate a plurality of two-dimensional pictures of different perspectives of the corresponding object according to the text, and the three-dimensional object generation model can be used to generate the corresponding 3D object according to the plurality of two-dimensional pictures of different perspectives, thereby improving the quality of the generated 3D object and speeding up the generation of the 3D object.

[0008] In some implementable embodiments, the two-dimensional picture generation model comprises a text marker, a Transformer structure and a first image marker, the text is cut into a first set of markers by inputting the text into the text marker; the first set of markers is used to indicate the category, color and shape features of the object; the number of markers in the first set of markers is a plurality; the first set of markers is input into the Transformer structure to obtain a second set of markers; the number of markers in the second set of markers is n; the second set of markers is input into the first image marker to output a plurality of two-dimensional pictures of different perspectives of the object. In this way, the text marker, the Transformer structure and the first image marker can be used to generate a plurality of two-dimensional pictures of different perspectives of the corresponding object according to the text, thereby improving the quality of the generated 3D object and speeding up the generation of the 3D object.

[0009] In some implementable embodiments, the Transformer structure adopts an autoregressive generation mode, and the first n-1 markers in the second set of markers are used as input to generate the n-th marker in the second set of markers. In this way, the quality of the plurality of two-dimensional pictures of different perspectives can be improved.

[0010] In some implementable embodiments, the input of the two-dimensional picture generation model further comprises camera parameters, the camera parameters are used to indicate the perspective of the generated two-dimensional picture; the camera parameters are mapped to a corresponding third set of markers; the number of markers in the third set of markers is k; the third set of markers is sequentially input into the Transformer structure to obtain the second set of markers. In this way, the camera parameters and the text can be used to control the generation of a plurality of two-dimensional pictures of different perspectives of the corresponding object, thereby improving the quality of the generated 3D object and speeding up the generation of the 3D object.

[0011] In some implementable embodiments, the camera parameter indicates that the perspective of generating the two-dimensional picture further includes a pre-perspective, the two-dimensional picture generation model further includes a second image marker, a two-dimensional picture of the pre-perspective is obtained, and the two-dimensional picture of the pre-perspective is converted into a fourth set of markers through the second image marker; the number of markers in the fourth set of markers is multiple; and the fourth set of markers is input into the Transformer structure for decoding to obtain the second set of markers. In this way, the consistency between different perspectives of the same text can be improved through the pre-perspective guidance.

[0012] In some implementable embodiments, the two-dimensional picture of each perspective in the multiple perspectives is m; the semantic similarity between the m two-dimensional pictures of each perspective in the multiple perspectives and the text is calculated to obtain m values of the similarity; the m values of the similarity are sorted to determine s two-dimensional pictures whose values of the similarity meet the threshold requirement as the enhanced two-dimensional picture of each perspective, where m > s. In this way, the semantic consistency between the generated perspective and the text can be enhanced, and the generation quality of the 3D object can be improved, where m and s are natural numbers.

[0013] In some implementable embodiments, the text is input into the text encoder to output a first feature vector; the m two-dimensional pictures of each perspective are input into the image encoder to output m second feature vectors; and the inner product between the first feature vector and the m second feature vectors is calculated to obtain the value of the similarity between the two-dimensional picture of each perspective and the text. In this way, the value of the similarity between the two-dimensional picture of each perspective and the text can be obtained.

[0014] In some implementable embodiments, the perspective other than the multiple perspectives is a second perspective, the three-dimensional object generation model is a pixelNeRF network, the multiple enhanced two-dimensional pictures of the perspectives are input into the pixelNeRF network, the corresponding image features are extracted from each enhanced two-dimensional picture of the perspectives along the query point x of the target ray d of each perspective in the multiple perspectives through projection and interpolation, each image feature is input into the NeRF network together with the spatial coordinates, the output RGB and density values are volume rendered to obtain the NERF model of the object; the NERF model of the object is a hidden model; and based on the NERF model of the object, the images of other perspectives other than the multiple perspectives are rendered to obtain the 3D object model conforming to the text description; and the 3D object model is a mesh model of the object. In this way, the NERF can be used as the 3D object representation method, the generation quality of the 3D object is improved, time-consuming optimization operations are not required in the inference stage, and thus the generation speed of the 3D object is improved.

[0015] In some implementable embodiments, the training step of the two-dimensional picture generation model comprises: inputting the text and the training set GT picture corresponding to the text into the two-dimensional picture generation model; the picture of the training set can be collected by a device or obtained by rendering a 3D model; and an optimization loss function is used to make the similarity between the generated two-dimensional pictures of multiple perspectives and the training set GT picture converge, so as to obtain a trained two-dimensional picture generation model. In this way, the quality of the 2D picture generated by the two-dimensional picture generation model can be improved by training and optimizing the loss function.

[0016] In some implementable embodiments, one perspective in the camera parameters and the previous perspective picture of the perspective are input into the two-dimensional picture generation model; an optimization loss function is used to make the similarity between the generated two-dimensional pictures of multiple perspectives and the multiple perspective pictures of the training set corresponding to the text converge, so as to obtain a trained two-dimensional picture generation model. In this way, the consistency between different perspectives generated for the same text can be improved by training and optimizing the loss function according to the previous perspective, the two-dimensional picture generation model can be optimized, and the quality of the generated 2D picture can be improved.

[0017] In some implementable embodiments, the semantic similarity between the two-dimensional pictures of multiple perspectives and the text is calculated, and in the case of semantic similarity convergence, an optimized semantic loss function L is obtained:

[0018] L = -I · T

[0019] where I is the feature vector of the two-dimensional picture, and T is the feature vector of the text. In this way, the semantic consistency between the generated perspective and the text can be enhanced by training and optimizing the semantic loss function, the two-dimensional picture generation model can be optimized, and the quality of the generated 2D picture can be improved.

[0020] In some implementable embodiments, the input of the two-dimensional picture generation model further comprises a current perspective image and its previous perspective image, the two-dimensional picture generation model further comprises an image analyzer, a cross-entropy loss function is calculated for the multiple tokens output by the Transformer structure decoding, and the Transformer model is trained at the token level. In this way, the two-dimensional picture generation model can be optimized, and the quality of the generated 2D picture can be improved.

[0021] In some implementable embodiments, based on the cross-entropy loss function, an L1 loss function is calculated between the output second 2D picture and the training set picture GT at the pixel level, and an optimized detail loss function is obtained. In this way, the two-dimensional picture generation model can be optimized, the details of the generated multiple perspective 2D pictures can be enhanced, and the quality of the generated 3D object can be better.

[0022] In some feasible implementations, optimizing the viewpoint contrast loss function can reduce the distance between 2D images generated from different viewpoints of the same text, while increasing the distance between 2D images generated from different texts and from different viewpoints. This can improve the consistency between different viewpoints of the same text generation, resulting in better quality 3D object generation.

[0023] In some feasible implementations, the viewpoint contrast loss function is L. contrastive :

[0024]

[0025] In the formula, and These are 2D images generated from the same text but from different perspectives. Is with 2D images generated from different texts from different perspectives; sim() is a similarity function used to calculate and The similarity is obtained by the inner product of the feature vectors; f enc () is the feature extraction function, used to extract... and The eigenvectors; τ is the temperature coefficient, the smaller the value of τ, the better. and The greater the distance, the greater the value of τ. and The smaller the distance, the better; exp() is used when the value of τ is minimized, making one value close to 1 and the others close to 0. In this way, the consistency between different perspectives of the same text generation can be improved, resulting in better quality 3D object generation.

[0026] Secondly, embodiments of this application provide an object generation apparatus for performing the method as described in any of the first aspects, comprising at least: a two-dimensional image generation model for taking text as input and outputting two-dimensional images of an object from multiple perspectives; the text is used to describe the features of the object, including object category, color, and shape; the two-dimensional image generation model is used to generate two-dimensional images from multiple perspectives based on the text; the perspective is the spatial angle presented by the object; an enhancement module is used to improve the similarity between the two-dimensional images from multiple perspectives and the text, thereby obtaining two-dimensional images enhanced from multiple perspectives; and a three-dimensional object generation model is used to take the two-dimensional images enhanced from multiple perspectives as input, render two-dimensional images from other angles based on the two-dimensional images enhanced from multiple perspectives, and output a three-dimensional object that conforms to the text description.

[0027] Thirdly, embodiments of this application provide a system for generating objects, the system including an apparatus for generating objects according to the second aspect, wherein the apparatus is used to perform the method as described in any of the first aspects.

[0028] In a fourth aspect, an embodiment of the present application provides a computer device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein the processor is configured to execute the method according to any one of the first aspect when the program stored in the memory is executed.

[0029] In a fifth aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores an instruction, and the instruction, when executed on a computer, causes the computer to execute the method according to any one of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments disclosed in the specification, the following will briefly introduce the drawings needed in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the specification, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0031] The following will briefly introduce the drawings needed in the embodiments or prior art description.

[0032] Figure 1 is a flowchart of a three-dimensional object generation method provided by the background technology;

[0033] Figure 2a is a technical solution schematic diagram of CLIP-Forge provided by the first scheme;

[0034] Figure 2b is a texture schematic diagram shown by the first scheme;

[0035] Figure 3 is a technical solution schematic diagram of DreamFields provided by the second scheme;

[0036] Figure 4 is a method flow of object generation provided by the embodiments of the present application;

[0037] Figure 5 is a two-dimensional picture generation model schematic diagram provided by the embodiments of the present application;

[0038] Figure 6 is a semantic enhancement schematic diagram in the method of object generation provided by the embodiments of the present application;

[0039] Figure 7 is a flowchart of two-dimensional picture generation model training in the method of object generation provided by the embodiments of the present application;

[0040] Figure 8 is a view angle comparison schematic diagram in the method of object generation provided by the embodiments of the present application;

[0041] Figure 9 A schematic diagram for comparing the performance of the object generation method provided in the embodiments of the present application with existing algorithms;

[0042] Figure 10 A schematic diagram for controlling the visual effect of the color of the generated 3D object by the object generation method provided in the embodiments of the present application;

[0043] Figure 11 A schematic diagram for controlling the visual effect of the color of the generated 3D object by the object generation method provided in the embodiments of the present application;

[0044] Figure 12 A schematic diagram for controlling the visual effect of the shape of the generated 3D object by the object generation method provided in the embodiments of the present application;

[0045] Figure 13 A schematic diagram of the object generation device provided in the embodiments of the present application;

[0046] Figure 14 A schematic diagram of the object generation system provided in the embodiments of the present application;

[0047] Figure 15 A schematic diagram of the computing device provided in the embodiments of the present application;

[0048] Figure 16 A schematic diagram of the computing device cluster provided in the embodiments of the present application;

[0049] Figure 17 A schematic diagram of a possible implementation of the computing device cluster provided in the embodiments of the present application. DETAILED DESCRIPTION

[0050] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0051] In the description of the embodiments of the present application, the words “exemplary”, “for example”, or “for instance” are used to mean serving as an example, instance or illustration. Any embodiment or design solution described as “exemplary”, “for example” or “for instance” in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of the words “exemplary”, “for example” or “for instance” is intended to present the relevant concept in a specific way.

[0052] In the description of the embodiments of the present application, the term "and / or", merely describes an association relationship for associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, B alone, and A and B together. In addition, unless otherwise specified, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of terminals means two or more terminals.

[0053] In addition, the terms "first", "second", "third" and the like are used only for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined as "first", "second" can explicitly or implicitly include one or more of the features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.

[0054] In the description of the embodiments of the present application, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.

[0055] In the description of the embodiments of the present application, the terms "first", "second", "third", etc. or module A, module B, module C, etc. are used only to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that the specific order or sequence can be interchanged as permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0056] In the description of the embodiments of the present application, the steps represented by the labels such as S110, S120, etc. do not necessarily mean that the steps are executed in this order, and the order of the steps can be interchanged or executed simultaneously as permitted.

[0057] Token means token in lexical analysis. The tokenizer is a process in computer science that converts a sequence of characters into a sequence of tokens. The process of generating tokens from text is called tokenization, in which tokenizer also classifies tokens.

[0058] PixelNeRF is a learning framework that predicts a neural radiance field (NeRF) representation from a single or multiple images. PixelNeRF can be trained on a set of multiple-view 2D pictures, allowing it to generate trustworthy new view synthesis from very few input images without test-time optimization.

[0059] PixelNeRF takes 2D images with known multiple views as input for neural rendering. Specifically, for a query point x along the target camera ray with direction d, the corresponding image feature is extracted from the feature volume W (object map) by projection and interpolation, and then the feature is input into the NeRF network together with the spatial coordinates, and the output RGB and density values are volume rendered and compared with the target pixel values. The coordinates x and d are in the camera coordinate system of the view.

[0060] 2D images as used herein is a short form of two-dimensional images, and 3D objects is a short form of three-dimensional objects.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.

[0062] Figure 2a A technical scheme diagram of CLIP-Forge is provided for the first scheme. As shown in FIG. 2, the training stage is mainly divided into the following steps:

[0063] S11, training an auto-encoder 11 to encode 3D shapes 12 into 3D embedding 13, and the auto-encoder 11 can also be used to decode 3D embedding into 3D shapes;

[0064] S12, training a normalizing flow 14 to generate 3D embedding based on the clip-image embedding 16 of the 2D images 15 rendered by the 3D shapes 12, wherein the images are images extracted by the CLIP model.

[0065] The inference stage is mainly divided into the following steps:

[0066] S13, given a text, using the CLIP model to extract a clip-text embedding.

[0067] S14, using the normalizing flow to generate 3D embedding based on the clip-text embedding in S13. Since the CLIP model is trained to align the image and text embedding, the text embedding is used instead of the image embedding used in the training stage in the inference stage.

[0068] S15, using the auto-encoder 11 to decode the 3D embedding to generate the corresponding 3D shapes.

[0069] As Figure 2bAs shown, in order to encode and decode the 3D object 12 using the autoencoder 11, the first scheme uses a display 3D representation to represent the 3D object, which is shown as a voxel in the figure. This results in the generated 3D object not being realistic enough and lacking realistic textures.

[0070] Figure 3 A schematic diagram of the technical solution for DreamFields provided for the second option. (See diagram below.) Figure 3 As shown, the Neural Radiance Fields (NERF) model is used as the representation of 3D objects. For each text, Dreamfields is optimized to obtain a NERF model. The CLIP model is used to calculate the semantic similarity between the 2D images rendered by the NERF model from different perspectives and the text as the loss function, while constraints are placed on the sparsity of the NERF model. Experimental results are shown below. Figure 3 The example on the right is shown in the middle.

[0071] The drawback of the second approach is that it requires optimization of each text element to obtain the final 3D object. Experiments using publicly available code show that, with eight TPUs used for optimization, each optimization takes over 70 minutes, resulting in a very slow generation speed.

[0072] Figure 4 The method flow for generating objects provided in the embodiments of this application is shown in Figure 4, which includes the following steps S21-S24.

[0073] S21, Define the text. The text describes the characteristics of the object, including its category, color, and shape. The text can be in Chinese or English.

[0074] For example, the text is determined to be the Chinese phrase "charcoal-colored leather double sofa".

[0075] For example, the text is determined to be the English phrase "charcoal leather loveseat".

[0076] S22: Input text into a 2D image to generate a model, outputting 2D images of the object from multiple perspectives; the text describes the object's features, including its category, color, and shape. The perspective is the angle from which the object is viewed. The 2D images from multiple perspectives can be denoted as the first 2D image.

[0077] For example, the image generation model takes "charcoal leather loveseat" as input and outputs images of the sofa from multiple perspectives, including front, back, side, top, bottom, 10 degrees to the left, and 20 degrees to the right.

[0078] In some embodiments, the 2D picture of each view of the plurality of views of 2D pictures comprises a plurality of 2D pictures of the view, and the plurality of 2D pictures have randomness.

[0079] Exemplarily, the 2D picture generation model outputs 20 2D pictures of each view of the sofa.

[0080] In some embodiments, the input of the 2D picture generation model further comprises camera parameters, a current view and a previous view, wherein the current view is only used for training the 2D picture generation model. The camera parameters are used to control the generation of the plurality of views of 2D pictures.

[0081] In some embodiments, the camera parameters can be set to control the views of the generated 2D pictures, and the output order of the plurality of views of 2D pictures is determined according to a plurality of camera parameter orders.

[0082] Exemplarily, the views of the generated 2D pictures are set to 9, and the 9 views are determined according to 9 camera parameter orders. The 2D picture generation model can sequentially generate 9 2D pictures of different views for each object described by the text. According to the order of the 9 camera parameter orders, each view has a previous view, i.e. the previous view of the current view. When the model generates a 2D picture of a view, it inputs the 2D picture of the previous view as input. The input of the 2D picture of the previous view is to improve the consistency of the 2D pictures of different views.

[0083] Figure 5 A schematic diagram of the 2D picture generation model provided in the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the 2D picture generation model comprises a Transformer structure 51, a text Tokenizer 52, an image Tokenizer 53 and an image Tokenizer 54. The image Tokenizer 54 can be referred to as a first image marker, the image Tokenizer 53 can be referred to as a second image marker, and the text Tokenizer 52 can be referred to as a text marker. Figure 5 The text Tokenizer 52 converts the input text into an output text marker token1, which is used to indicate the category, color and shape features of the object. The token1 can be referred to as a first group of markers. The number of markers in the first group of markers is multiple.

[0084] The Transformer structure 51 is used to sequentially decode and generate token4 according to the input token1. The token4 can be referred to as a second group of markers. The number of markers in the second group of markers is multiple. The markers in the second group of markers are different from the markers in the first group of markers.

[0085] The Transformer structure 51 is used to sequentially decode and generate token4 according to the input token1. The token4 can be referred to as a second group of markers. The number of markers in the second group of markers is multiple. The markers in the second group of markers are different from the markers in the first group of markers.

[0086] The Transformer structure 51 adopts an autoregressive generation manner, and generates an nth token in the token 4 by taking the first n-1 tokens in the token 4 as input, where n is a natural number greater than 0.

[0087] The image Tokenizer 54 takes the token 4 as input and outputs a two-dimensional picture of a certain view angle.

[0088] The previous view angle of a certain view angle is referred to as a previous view angle.

[0089] The image Tokenizer 53 takes a two-dimensional picture of the previous view angle as input and outputs the token 3. The token 3 can be referred to as the fourth group of tokens. The number of tokens in the fourth group of tokens is a plurality.

[0090] In some embodiments, the step S22 includes the following steps S221-S225 to determine the text.

[0091] S221, obtaining the text, and converting the text into text tokens token 1 through the text tokenizer 52. The token 1 is used to indicate the category, color and shape features of the object; the number of tokens in the token 1 is a plurality.

[0092] S222, obtaining the camera parameters, a plurality of camera parameters can be set, and each of the camera parameters is mapped to a corresponding parameter token token 2 in a specified order. The token 2 is referred to as the third group of tokens. The number of camera parameters is k, and the number of tokens in the third group of tokens is k.

[0093] Exemplarily, the number of camera parameters can be 9, and the 9 set camera parameters are used to control 9 view angles of the generated two-dimensional picture. The token 2 can be referred to as the second token.

[0094] S223, obtaining the previous view angle picture, and converting the previous view angle picture into the token 3 through the image Tokenizer 53. The token 3 can be referred to as the fourth token. The number of tokens in the fourth group of tokens is a plurality.

[0095] The 2D picture of the i-1th view angle that has been generated can be referred to as the previous view angle picture, and the 2D picture of the ith view angle that is being generated can be referred to as the current view angle picture, where 1 < i ≤ k, and i is a natural number.

[0096] In some embodiments, the previous view angle picture can be obtained when i is a natural number greater than 1, and the token 3 of the previous view angle is converted through the image Tokenizer 53.

[0097] S224, input token1, token2 and token3 into the Transformer structure 51 to decode token4; token4 has n tokens; n is a natural number greater than 0.

[0098] In some embodiments, the Transformer structure 51 adopts an autoregressive generation manner, taking the first n-1 tokens in token4 as input to generate the nth token in token4.

[0099] In some embodiments, when i = 1, input token1 and token2 into the Transformer structure 51 for autoregressive decoding, and add the first n-1 tokens in token4 as input during the decoding process to generate the nth token in token4, and output token4.

[0100] In some embodiments, when i > 1, input token1, token2 and token3 into the Transformer structure 51 for autoregressive decoding, and add the first n-1 tokens in token4 as input during the decoding process to generate the nth token in token4, and output token4.

[0101] S225, input token4 into the image Tokenizer 54 to obtain the 2D picture of the ith view.

[0102] S226, when i ≤ k, repeat steps S221-S225 to obtain 2D pictures of k views in sequence according to the order of views.

[0103] For example, when i = 1, there is no previous view picture, and the 2D picture generation model generates the 2D picture of the first view according to the first camera parameter and the text.

[0104] For example, when i = 2, the previous view picture is the 2D picture of the first view, and the 2D picture generation model generates the 2D picture of the second view according to the second camera parameter, the previous view picture and the text.

[0105] For example, when i = 3, the previous view picture is the 2D picture of the second view, and the 2D picture generation model generates the 2D picture of the third view according to the third camera parameter, the previous view picture and the text. The subsequent is similar until i = k.

[0106] In some embodiments, the generation of text into 2D pictures of multiple views can adopt an image generation model based on a convolutional neural network structure.

[0107] In addition to the image generation model based on the Transformer structure used in this embodiment, the image generation model based on the convolutional neural network structure, and other image generation models that can be easily thought of by those skilled in the art are within the protection scope of this application.

[0108] For the same set of inputs, due to the randomness in the 2D picture generation model, each 2D picture of multiple views in the multiple view 2D pictures includes multiple random 2D pictures of the view, and thus multiple sets of different multiple view 2D pictures can be obtained by inputting the same camera parameters and text multiple times.

[0109] S23, determine the semantic similarity between the multiple view 2D pictures and the text, and obtain the multiple view enhanced 2D pictures according to the value of the semantic similarity. The enhanced 2D picture is the 2D picture whose semantic similarity with the text meets the threshold requirement.

[0110] Figure 6 The semantic enhancement schematic diagram in the object generation method provided by the embodiment of the application is shown in FIG. 6. Figure 6 As shown in FIG. 6, the semantic similarity between each of the multiple random 2D pictures of each view in the multiple view 2D pictures and the text can be calculated to obtain multiple similarity values of each view, the multiple similarity values of each view are sorted, and the 2D picture with the highest value in the multiple similarity values is determined as the enhanced 2D picture of each view, so as to enhance the consistency of the semantics between the multiple view 2D pictures and the text.

[0111] In some embodiments, each 2D picture of the multiple view 2D pictures is m; the semantic similarity between each 2D picture and the text is calculated, including steps S231-S232.

[0112] S231, calculate the semantic similarity between the m 2D pictures of each view and the text to obtain m similarity values.

[0113] In some embodiments, step S231 includes steps S2311-S2313.

[0114] S2311, input the text into the text encoder 61 to output a first feature vector.

[0115] S2312, input the m 2D pictures of each view into the image encoder 62 to output m second feature vectors.

[0116] For example, the text passes through the text encoder 61 to obtain a first feature The 2D picture passes through the image encoder 62 to obtain a second feature

[0117] S2313, calculate the inner product between the first feature vector and each second feature vector, to obtain the value of the semantic similarity of each view.

[0118] Exemplarily, the inner product S between the first feature F t and the second feature F i may be calculated to obtain the value of the semantic similarity of the text and the 2D picture. t i

[0119] In some embodiments, the cosine similarity between the first feature F t and the second feature F i may be calculated to obtain the value of the semantic similarity of the text and the 2D picture.

[0120] S232, sort the values of the m semantic similarities, and determine the 2D pictures whose similarity values meet the threshold requirement of n as the enhanced 2D pictures of each view; wherein, m>n.

[0121] In some embodiments, the n 2D pictures with the highest semantic similarity values can be selected as the enhanced 2D pictures according to the sorting of the semantic similarity values.

[0122] S24, input the plurality of view-enhanced 2D pictures into a three-dimensional (3D) object generation model, and the 3D object generation model renders based on the plurality of view-enhanced 2D pictures, and outputs a 3D object meeting the text description.

[0123] In some embodiments, the three-dimensional object generation model is a pixelNeRF network, and step S24 includes the following steps:

[0124] S241, taking the enhanced 2D pictures of the plurality of views of the object as the input of the pixelNeRF network, and outputting the NERF model of the object through internal parameter learning.

[0125] The pixelNeRF network is used for neural rendering with known 2D pictures of multiple views as input. The principle is that for a query point x along the target camera ray in the observation direction d, the corresponding image feature is extracted from the feature volume W (object map) through projection and interpolation, and then the feature is input into the neural radiance field (NERF) network together with the spatial coordinates. The output RGB and density values are volume rendered and compared with the target pixel values. The coordinates x and d are in the camera coordinate system of the view. Wherein, rendering is the process of integrating the colors on the picture to visualize the complete picture.

[0126] ​​In some embodiments, the pixelNeRF network can be input with multiple perspective-enhanced 2D pictures, and at a query point x along a target ray d of each perspective in the multiple perspectives, corresponding image features can be extracted from each perspective-enhanced 2D picture through projection and interpolation, and then each image feature can be input into the NeRF network together with the spatial coordinates, and the RGB and density values of the output image can be volume rendered to obtain the NERF model of the object.

[0127] The NERF model of the object is a parameterized object model optimized by learning parameters inside the NERF network, and is an implicit model.

[0128] S242, based on the NERF model of the object, rendering images of other perspectives other than the multiple perspectives to obtain a 3D object consistent with the text description.

[0129] In some embodiments, the 3D object is a mesh model. The mesh is a collection of points (also referred to as vertices), normal vectors, and faces, which defines the 3D shape of an object.

[0130] Figure 7 A flowchart of a 2D picture generation model training process in a method of object generation provided by embodiments of the present application is shown in FIG. 2. Figure 7 As shown in FIG. 2, in the training phase, the 2D picture generation model is input with the input text t, the camera parameters P, the 2D picture of the current perspective, and the 2D picture of the previous perspective.

[0131] The input text t is the text of the training set, the 2D picture of the current perspective is the current perspective picture corresponding to the training set, and the 2D picture of the previous perspective is the previous perspective picture corresponding to the current perspective picture of the training set. The current perspective picture of the training set can be captured by a camera device or rendered by a 3D model. The camera parameters P are the set camera parameters, which are the same as the camera parameters input in the inference phase.

[0132] The training phase includes the following steps:

[0133] S31, input the camera parameters, the text of the training set, and the corresponding current perspective picture of the training set into the 2D picture generation model.

[0134] In some embodiments, S31 includes the following steps S311-S313.

[0135] S311, map the set camera parameters to the corresponding token2; and input the token2 into the Transformer structure 51 in sequence to decode and output the parameter token5.

[0136] S312, the training set text is converted by the text tokenizer 52 to obtain token1, and the token1 is input into the transformer structure 51 to decode and output token4.

[0137] S313, the current view picture of the training set is converted by the image tokenizer 53 to obtain token7, and the token7 is input into the transformer structure 51 to decode and output image token8.

[0138] In some embodiments, in order to enhance the continuity between the 2D pictures of different views generated by the same training text, the 2D pictures of two continuous views in the training set can be input into the 2D picture generation model together. The one with earlier view order in the two continuous view 2D pictures is recorded as the previous view picture, and the one with later view order is recorded as the current view picture. The following step S314 is included.

[0139] S314, the previous view picture of the training set is converted by the image tokenizer 53 to obtain token3, and the token3 is input into the transformer structure 51 to decode and output token6.

[0140] S32, token4, token5, token6 and token8 are synthesized by the image tokenizer 54 to obtain the trained view 2D picture. The trained view 2D picture can be recorded as the second 2D picture.

[0141] S33, the semantic loss function is optimized. The semantic consistency between the trained view 2D picture and the text can be enhanced by optimizing the semantic loss function.

[0142] In some embodiments, the semantic similarity between the plurality of second 2D pictures and the training set text is calculated, and the optimized semantic loss function L is used to make the value of the semantic similarity converge. Wherein, the semantic loss function L is:

[0143] L=-I·T

[0144] Wherein I is the feature vector of the second 2D picture, and T is the feature vector of the training set text.

[0145] S34, the detail loss function is optimized.

[0146] The details of the generated results can be enhanced by optimizing the detail loss function. The following steps are included:

[0147] S341, calculate the cross-entropy loss function for the tokens token4, token5, token6 and token8 output by the Transformer structure, thereby training the Transformer model at the token level.

[0148] Experiments show that this token-level loss function is insufficient to guide the network model to generate enough detail.

[0149] S342 computes a detail loss function at the pixel level between multiple second 2D images and the training set image GT.

[0150] In some embodiments, a token-level cross-entropy loss function is used to compute a detail loss function at the pixel level between the output second 2D image and the training set image GT, thereby providing the Transformer model with a more fine-grained training signal.

[0151] Those skilled in the art will readily recognize that other image generation modules can be used as replacements, such as perceptual loss, which is also a viable solution.

[0152] S35, optimize the viewpoint contrast loss function.

[0153] In some embodiments, by optimizing the perspective contrast loss function, the distance between 2D images generated from different perspectives of the same text can be reduced, while the distance between 2D images generated from different texts and different perspectives can be increased.

[0154] Figure 8 This is a schematic diagram illustrating the perspective comparison in the object generation method provided in this application embodiment. For example... Figure 8 As shown, and These are 2D images generated from the same text but from different perspectives. Is with Different texts generate 2D images from different perspectives. Optimizing the perspective contrast loss function includes the following steps:

[0155] S351, using the feature extraction network FINC to extract... eigenvectors and eigenvectors; The eigenvector of is denoted as the third eigenvector. The eigenvector of is denoted as the fourth eigenvector.

[0156] S352, calculate the inner product of the third and fourth feature vectors to obtain the similarity value sim(), and calculate the perspective contrast loss function L based on the similarity value. contrastive :

[0157]

[0158] where, sim(,) is a similarity function that computes the inner product of the feature vectors of f enc () is a feature extraction function that extracts the feature vector of ; and are 2D pictures of different perspectives generated by the same text, are 2D pictures of different perspectives generated by different texts; τ is a temperature coefficient, the smaller the value of τ, the greater the distance between and ; the greater the value of τ, the smaller the distance between and ; exp() is used to make a value close to 1 and the other close to 0 when the value of τ is the smallest.

[0159] By optimizing the perspective contrast loss function, the value of the similarity of the 2D pictures of different perspectives generated by the same text is made larger, so as to realize the narrowing of the distance between and , and the widening of the distance between and .

[0160] The PixelNeRF network can be trained on a set of multiple perspective 2D pictures, allowing it to generate credible new view synthesis from very few input images without the need for test-time optimization.

[0161] In order to compare the performance with existing algorithms, the embodiments of the present application compare the baseline model Text-NeRF, DreamField on the public dataset (amazon-berkeley objects, ABO).

[0162] Figure 9 The schematic diagram of the object generation method provided by the embodiments of the present application compared with the performance of the existing algorithms. As Figure 9 shown, the embodiments of the present application (Ours) are compared in terms of 3D object generation quality and semantic consistency between 3D objects and texts. Figure 9

[0163] The embodiments of the present application use a total of 6 numerical indicators. Among them, PSNR, SSIM, LPIPS are automatic indicators for measuring the quality of 3D object generation, and CLIP-Score is an automatic indicator for measuring the semantic consistency between texts.

[0164] In addition, the embodiments of the present application also use artificial evaluation methods to compare the pros and cons between the results of different methods, among which Object Fidelity is an artificial evaluation indicator for measuring the quality of 3D object generation, and Caption Similarity is an artificial evaluation indicator for measuring the semantic consistency between texts.

[0165] By Figure 9 It can be seen that the embodiments of the present application are superior to the existing text-NeRF algorithm and Dreamfields in all indicators.

[0166] In addition, in terms of generation speed, compared with Dreamfields which can also generate better textures, the embodiments of the present application do not need to optimize each text, but only need to perform inference after training the neural network model in the training stage to generate 3D objects in the inference stage. The embodiments of the present application only need to use one V100, and the time consumption is only 6 minutes.

[0167] Figure 10 The visual effect schematic diagram of the object generation method provided by the embodiments of the present application controls the visual effect of the category of the generated 3D object. As shown in Figure 10 , the input text is English, and the embodiments of the present application can control the category attribute of the generated 3D object through the input English text.

[0168] Figure 11 The visual effect schematic diagram of the object generation method provided by the embodiments of the present application controls the color of the generated 3D object. As shown in Figure 11 , the input text is English, and the embodiments of the present application can control the color attribute of the generated 3D object through the input English text.

[0169] Figure 12 The schematic diagram of the object generation method provided by the embodiments of the present application controls the shape of the generated 3D object. As shown in Figure 12 , the input text is English, and the embodiments of the present application can control the shape attribute of the generated 3D object through the input English text.

[0170] The object generation method proposed by the embodiments of the present application first generates a plurality of 2D pictures of different perspectives of the corresponding object according to the text by using a text-to-2D perspective generation module, and then generates a new method of generating a 3D object according to the plurality of perspectives by using a perspective-to-3D object generation module.

[0171] The view angle to 3D object generation module proposed in the embodiments of the present application supports using a NERF as a representation manner of a 3D object, and can generate a 3D object with higher quality. Time-consuming optimization is not required in the inference stage, thereby realizing faster generation speed.

[0172] The object generation method proposed in the embodiments of the present application can be applied in different scenarios, and a corresponding 3D object is generated according to a text.

[0173] In some embodiments, the object generation method proposed in the embodiments of the present application can be applied in a data enhancement scenario.

[0174] Exemplarily, a text meeting task requirements can be designed, a corresponding 3D object is generated, the 3D object is directly used or the generated 3D object is fused with other data to obtain new training data, thereby saving data acquisition cost, improving the performance of a machine learning model, and realizing data enhancement.

[0175] In machine learning, more training data is generally beneficial to the training of a machine learning model.

[0176] In some embodiments, the object generation method proposed in the embodiments of the present application can be applied in an entertainment scenario.

[0177] In some embodiments, the object generation method proposed in the embodiments of the present application can be deployed on different intelligent devices to interact with users to achieve the effect of mass entertainment. The intelligent devices include a mobile phone, a PAD, a PC, and the like.

[0178] Exemplarily, a user inputs a text on a mobile phone, the mobile phone processes the text to obtain a generation result, the mobile phone performs secondary editing on the generation result to achieve the effect of mass entertainment.

[0179] In some embodiments, the text-guided 3D object generation method proposed in the embodiments of the present application can be applied in content creation.

[0180] Exemplarily, in 3D object design, a corresponding 3D object is generated based on a text, and secondary editing design is performed according to actual requirements, thereby improving the creation efficiency and also widening the creation ideas of creators.

[0181] Exemplarily, in the design of a commodity, a corresponding 3D object is generated based on a text, and secondary editing design is performed according to actual requirements, thereby improving the creation efficiency and also widening the creation ideas of creators.

[0182] The object generation method provided in the embodiments of the present application can be applied in cloud products, terminal devices, and the like.

[0183] The method for generating an object provided in the embodiment of the present application is deployed on a computing node of a related device, and can improve the generation quality and generation speed of a text-oriented 3D object generation through software modification.

[0184] Figure 13 The device for generating an object provided in the embodiment of the present application is shown in a schematic diagram. The device for generating an object provided in the embodiment of the present application executes any one of the methods described above, and the device comprises a 2D picture generation model 131, an enhancement module 132, and a three-dimensional object generation model 133. The 2D picture generation model 131 takes text as input and outputs 2D pictures of multiple perspectives of an object. The text is used to describe the features of the object, and the features include the object category, color, and shape. The 2D picture generation model is used to generate 2D pictures of multiple perspectives according to the text. The perspective is a spatial angle at which the object is presented. The enhancement module 132 improves the similarity between the 2D pictures of multiple perspectives and the text to obtain 2D pictures of multiple enhanced perspectives. The three-dimensional object generation model 133 takes the 2D pictures of multiple enhanced perspectives as input, renders 2D pictures of other angles based on the 2D pictures of multiple enhanced perspectives, and outputs a three-dimensional object that conforms to the text description.

[0185] Figure 14 The system for generating an object provided in the embodiment of the present application, as shown in Figure 14 , comprises:

[0186] The device for generating an object 141 is configured to input text into a 2D picture generation model and output 2D pictures of multiple perspectives of an object. The text is used to describe the features of the object, and the features include the object category, color, and shape. The 2D picture generation model is used to generate 2D pictures of multiple perspectives according to the text. The perspective is a spatial angle at which the object is presented. The similarity between the 2D pictures of multiple perspectives and the text is calculated, and 2D pictures of multiple enhanced perspectives are obtained according to the similarity value. The 2D pictures of multiple enhanced perspectives are input into a three-dimensional object generation model. The three-dimensional object generation model renders 2D pictures of other angles based on the 2D pictures of multiple enhanced perspectives and outputs a three-dimensional object that conforms to the text description. In this way, the 2D picture generation model can be used to generate 2D pictures of multiple perspectives of a corresponding object according to the text, and the three-dimensional object generation model can be used to generate a corresponding 3D object according to the 2D pictures of multiple perspectives, thereby improving the quality of the generated 3D object and speeding up the generation of the 3D object.

[0187] The training device 142 is configured to input the text of a training set and the corresponding training set GT pictures into the 2D picture generation model. The pictures of the training set can be collected by a device or obtained by rendering a 3D model. An optimization loss function is used to make the similarity between the generated 2D pictures of multiple perspectives and the training set GT pictures converge, thereby obtaining a trained 2D picture generation model. In this way, the quality of the 2D pictures generated by the 2D picture generation model can be improved through training and optimization of the loss function.

[0188] The object generation apparatus 141 and the training apparatus 142 can be implemented by software or by hardware. As an example, the implementation of the object generation apparatus 141 is described below. Similarly, the implementation of the training apparatus 142 can refer to the implementation of the object generation apparatus 141.

[0189] As an example of a software functional unit, the object generation apparatus 141 can include code running on a computing instance. The computing instance can be at least one of a physical host (computing device), a virtual machine, a container, and the like. Further, the computing device can be one or more. For example, the object generation apparatus 141 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers for running the code can be distributed in the same AZ or in different AZs, each of which includes a data center or multiple data centers in a similar geographical location. Generally, one region can include multiple AZs.

[0190] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same VPC or in multiple VPCs. Generally, one VPC is set in one region. Communication between two VPCs in the same region or between VPCs in different regions requires a communication gateway in each VPC to realize the interconnection between VPCs.

[0191] As an example of a hardware functional unit, the object generation apparatus 141 can include at least one computing device, such as a server or the like. Alternatively, the object generation apparatus 141 can be a device implemented by ASIC or PLD, and the like. The PLD can be CPLD, FPGA, GAL, or any combination thereof.

[0192] The multiple computing devices included in the object generation apparatus 141 can be distributed in the same region or in different regions. The multiple computing devices included in the object generation apparatus 141 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the object generation apparatus 141 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs, and the like.

[0193] The present application also provides a computing device 100. As shown in Figure 15 The computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100.

[0194] The bus 102 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 15 only one line is used, but it does not mean that there is only one bus or only one type of bus. The bus 102 can include a path for transmitting information between various components (e.g., the memory 106, the processor 104, the communication interface 108) of the computing device 100.

[0195] The processor 104 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.

[0196] The memory 106 can include a volatile memory, such as a random access memory (RAM). The processor 104 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), or a solid state drive (SSD).

[0197] The memory 106 stores executable program code, and the processor 104 executes the executable program code to respectively implement the functions of the aforementioned 2D picture generation model, the enhancement module, and the three-dimensional object generation model, thereby implementing the object generation method. That is, the memory 106 has instructions for executing the object generation method.

[0198] Alternatively, the memory 106 stores executable code, and the processor 104 executes the executable code to implement the functions of the aforementioned object generation apparatus 141 and the training apparatus 142 respectively, so as to implement the object generation method. That is, the memory 106 stores instructions for executing the object generation method.

[0199] The communication interface 108 uses a transceiving module such as, but not limited to, a network interface card and a transceiver, to implement communication between the computing device 100 and other devices or communication networks.

[0200] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0201] As shown in Figure 16 The computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster can store the same instructions for executing the object generation method.

[0202] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster can also respectively store partial instructions for executing the object generation method. In other words, the combination of one or more computing devices 100 can collectively execute the instructions for executing the object generation method.

[0203] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster can store different instructions for respectively executing partial functions of the object generation apparatus. That is, the instructions stored in the memories 106 in different computing devices 100 can implement the functions of one or more of the 2D picture generation model, the enhancement module, and the three-dimensional object generation model.

[0204] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc.

[0205] Figure 17 A possible implementation is shown. As shown in Figure 17As shown, two computing devices 100A and 100B are connected through a network. Specifically, the computing devices are connected to the network through communication interfaces in the computing devices. In this kind of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the 2D picture generation model and the enhancement module. Meanwhile, the memory 106 in the computing device 100B stores instructions for executing the functions of the three-dimensional object generation model.

[0206] Figure 17 The connection between the computing device cluster shown can be that the object generation method provided in the present application needs to store a large amount of data and perform calculations, and therefore the functions of the 2D picture generation model and the enhancement module are implemented by the computing device 100B.

[0207] It should be understood that Figure 17 The functions of the computing device 100A shown in the above embodiment can also be completed by a plurality of computing devices 100. Similarly, the functions of the computing device 100B can also be completed by a plurality of computing devices 100.

[0208] The present embodiment also provides another computing device cluster. The connection between the computing devices in the computing device cluster can be similar to the connection between the computing devices in the computing device cluster shown in the above embodiment. Figure 16 and Figure 17 The connection between the computing devices in the computing device cluster. The difference is that the memory 106 in one or more computing devices 100 in the computing device cluster can store the same instructions for executing the object generation method.

[0209] In some possible implementations, the memory 106 in one or more computing devices 100 in the computing device cluster can also respectively store partial instructions for executing the object generation method. In other words, the combination of one or more computing devices 100 can collectively execute the instructions for executing the object generation method.

[0210] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for executing part of the functions of the object generation system. That is, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more of the devices 141, 142 in the object generation system.

[0211] The present embodiment also provides a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the object generation method or the training method.

[0212] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage that can be used to store data and that can be accessed by a computing device. The computer readable storage medium can be a magnetic-based, (e.g., a floppy disk, a hard disk drive, a magnetic tape, or the like), an optical-based, (e.g., a compact disc, a DVD, or the like) or a semiconductor-based, (e.g., a solid-state drive, a RAM, or the like) or any other available medium that can be used to store data and that can be accessed by a computing device. The computer readable storage medium includes instructions that instruct the computing device to perform the object generation method, or instruct the computing device to perform the object generation method.

[0213] It can be understood that the processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0214] The method steps in the embodiments of the present application can be implemented by means of hardware, or by means of a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.

[0215] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in or transmitted by a computer readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.

[0216] It can be understood that various numerical numbers involved in the embodiments of the present application are only distinguished for convenience of description, and are not used to limit the scope of the embodiments of the present application.

Claims

1. A method for generating an object, characterized in that, The method includes: The model takes text as input and outputs 2D images of an object from multiple perspectives. The text describes the object's features, including its category, color, and shape. The 2D image generation model generates 2D images from multiple perspectives based on the text. The model includes a text marker, a Transformer structure, and a first image marker. The input to the 2D image generation model also includes camera parameters, which indicate the perspectives from which the 2D images are generated. The perspective is the spatial angle presented by the object. Calculate the similarity values ​​between the 2D images from multiple viewpoints and the text, and obtain 2D images with enhanced viewpoints based on the similarity values, including: Calculate the semantic similarity between m two-dimensional images from each of the multiple viewpoints and the text to obtain m similarity values; sort the m similarity values ​​and determine s two-dimensional images whose similarity values ​​meet the threshold requirements as the enhanced two-dimensional images for each viewpoint, where m>s; The multiple enhanced 2D images are input into a 3D object generation model. The 3D object generation model renders 2D images from other angles based on the enhanced 2D images and outputs a 3D object that conforms to the text description, including: The enhanced 2D images from multiple viewpoints are input into a pixelNeRF network, which is a 3D object generation model. Along the query point x of the target ray d in each of the multiple viewpoints, corresponding image features are extracted from the enhanced 2D images from each viewpoint through projection and interpolation. Each image feature, along with its spatial coordinates, is then input into the NeRF network. Volume rendering is performed on the output RGB and density values ​​to obtain the NERF model of the object; the NERF model of the object is a latent model. Based on the NERF model of the object, images from other perspectives besides the multiple viewpoints are rendered to obtain a three-dimensional object model that conforms to the text description; the three-dimensional object model is a mesh model of the object.

2. The method according to claim 1, characterized in that, The method of inputting text into a 2D image generation model and outputting 2D images of an object from multiple perspectives includes: The text input text tagger converts and outputs a first set of tags; the first set of tags is used to indicate the category, color, and shape characteristics of the object; the number of tags in the first set of tags is multiple; The first set of tags is input into the Transformer structure and decoded to obtain the second set of tags; the second set of tags has n tags; n is a natural number greater than 0; The second set of markers is input into the first image marker, and two-dimensional images of the object from multiple perspectives are output.

3. The method according to claim 2, characterized in that, The Transformer structure adopts an autoregressive generation method, using the first n-1 tags in the second group of tags as input to generate the nth tag in the second group of tags.

4. The method according to claim 2, characterized in that, The method of inputting text into a 2D image generation model and outputting 2D images of an object from multiple perspectives includes: The camera parameters are mapped to the corresponding third set of markers; the third set of markers contains k markers. The third set of tags is sequentially input into the Transformer structure for decoding to obtain the second set of tags.

5. The method according to claim 2, characterized in that, The camera parameters indicating the viewpoints for generating 2D images also include preceding viewpoints. The 2D image generation model also includes a second image marker. The process of inputting text into the 2D image generation model and outputting 2D images of objects from multiple viewpoints includes: A two-dimensional image from a preceding viewpoint is obtained, and the two-dimensional image from the preceding viewpoint is converted into a fourth set of markers by a second image marker; the fourth set of markers contains multiple markers. The fourth set of tags is input into the Transformer structure for decoding to obtain the second set of tags.

6. The method according to claim 1, characterized in that, The semantic similarity between the text and m two-dimensional images from each viewpoint in the multi-view two-dimensional images is calculated, resulting in m similarity values, including: The text is input into the text encoder, which outputs the first feature vector. The m two-dimensional images from each viewpoint are input into an image encoder, which outputs m second feature vectors. Calculate the inner product between the first feature vector and the m second feature vectors to obtain the m similarity values ​​for each viewpoint.

7. The method according to claim 1 or 2, characterized in that, The method further includes a training step for the two-dimensional image generation model, including: The text and the corresponding training set GT images are input into a 2D image generation model; the training set images are acquired by a device or obtained by rendering a 3D model. The loss function is optimized to make the similarity between the generated 2D images from multiple perspectives and the ground truth images in the training set converge, thus obtaining the trained 2D image generation model.

8. The method according to claim 7, characterized in that, The training steps of the two-dimensional image generation model also include: Input one of the camera parameters and the preceding view image of the view into the two-dimensional image generation model; The loss function is optimized to make the similarity between the generated two-dimensional images from multiple perspectives and the images from multiple perspectives in the training set corresponding to the text converge, thereby obtaining the trained two-dimensional image generation model.

9. The method according to claim 7, characterized in that, The optimized loss function includes: Calculate the semantic similarity between the 2D images from multiple perspectives and the text. If the semantic similarity converges, obtain the optimized semantic loss function L. in .

10. The method according to claim 7, characterized in that, The input to the 2D image generation model also includes the current view image and its preceding view image. The 2D image generation model also includes an image analyzer. The optimization loss function includes: The cross-entropy loss function is calculated for the multiple tokens output by the decoded Transformer structure, and the Transformer model is trained at the token level.

11. The method according to claim 7, characterized in that, The optimized loss function includes: The L1 loss function between the second 2D image and the training set image GT is calculated at the pixel level to obtain the optimized detail loss function; the second 2D image is a two-dimensional image from the training viewpoint.

12. The method according to claim 7, characterized in that, The optimized loss function includes: optimizing the viewpoint contrast loss function, reducing the distance between 2D images generated from different viewpoints of the same text, and increasing the distance between 2D images generated from different viewpoints of different texts.

13. The method according to claim 12, characterized in that, The viewpoint contrast loss function is: : In the formula, and These are 2D images generated from the same text but from different perspectives. Is with 2D images generated from different texts from different perspectives; X= ; This is a similarity function used to calculate... and The similarity is obtained by the inner product of the feature vectors; This is a feature extraction function used to extract... and eigenvectors; For temperature coefficient, The smaller the value, and The greater the distance; The larger the value, and The smaller the distance; Used to make one value close to 1 and the others close to 0.

14. An apparatus for generating an object, used to perform the method as described in any one of claims 1-13, characterized in that, At least including: A 2D image generation model that takes text as input and outputs 2D images of objects from multiple perspectives. The text is used to describe the features of the object, including the object's category, color, and shape; the 2D image generation model is used to generate 2D images from multiple perspectives based on the text; the perspective is the spatial angle presented by the object. An enhancement module is used to improve the similarity between the two-dimensional images from multiple perspectives and the text, thereby obtaining two-dimensional images with enhanced multiple perspectives. A 3D object generation model is used to take the multiple viewpoint-enhanced 2D images as input, render 2D images from other angles based on the multiple viewpoint-enhanced 2D images, and output a 3D object that conforms to the text description.

15. A system for generating objects, characterized in that, The system includes means for generating the object, wherein the means is used to perform the method as described in any one of claims 1-13.

16. A computer device, characterized in that, include: At least one memory for storing programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-13.

17. A computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Object surface material analysis method and device

    CN113920433A

  • Image rendering method and device of three-dimensional object and electronic equipment

    CN114863007A