Three-dimensional model attribute binding method, system, equipment, medium and product

By performing deep extraction of the target three-dimensional model and extracting reference image features, combined with the text cross attention mechanism of the IP-Adapter model, the problem of difficulty in binding the properties of different parts of the same object in the prior art is solved, and efficient attribute binding and calculation reduction is achieved.

CN120220134APending Publication Date: 2025-06-27GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358132.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing attribute binding technology is prone to difficult binding when dealing with attribute binding between different parts of the same object, and has a large computing power, resulting in an increase in the time to generate pictures.

Method used

By obtaining the target three-dimensional model and its corresponding attribute prompt words and reference pictures, the faces in the target three-dimensional model are extracted in depth, the image features of the reference picture are extracted, the attribute prompt words, depth maps and image features are input into the IP-Adapter model, and the modifiers and nouns are semantically bound based on text cross attention, and iterative training is combined with the depth maps and image features to output the image of the binding attribute prompt words.

Benefits of technology

It realizes attribute binding between different parts of the same object, reduces the calculation amount, avoids attribute leakage problems, and improves the accuracy and efficiency of attribute binding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220134A_ABST
    Figure CN120220134A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of three-dimensional models, and discloses a three-dimensional model attribute binding method, system and device, a medium and a product, and the method comprises the steps: carrying out the depth extraction of each surface in a target three-dimensional model, extracting the image features of a reference picture through a pre-trained CLIP image encoder, and carrying out the recognition of the target three-dimensional model; inputting the attribute cues, the depth map and the image features into an IP-Adapter model, performing semantic binding on the modifiers and the nouns based on text cross attention, forcing the modifiers to be associated with the nouns, promoting the associated attributes to be associated with each other semantically, and improving the user experience. Therefore, the problem of attribute leakage occurring during attribute binding of different parts of the same object is avoided, and attribute binding of different parts of the same object is achieved. Meanwhile, loss calculation is not needed, and back propagation is not needed to update variables of a potential space, so that the calculation amount is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 3D models, and in particular, to a method, system, device, medium and product for binding 3D model attributes. Background Art

[0002] Research on attribute binding has attracted increasing attention in the field of text-to-image generation based on diffusion models. Attribute binding studies the problem that the model cannot correctly bind the attributes (such as color) in the text prompt to the corresponding theme.

[0003] Incorrect attribute binding is a very common failure phenomenon caused by incorrect attribute binding. Among them, the modifiers in the text prompt cannot affect the visual attributes of the entity nouns that are grammatically related to them. When multiple objects are involved, the model is prone to confusing the modifiers of the two entity nouns.

[0004] Existing attribute binding technologies are all applied to the attribute binding of two or more different objects. It is easy to have problems with difficult binding when dealing with the attribute binding between different parts of the same object. It is difficult to handle the attribute binding between different parts of the same object. Moreover, most steps in the denoising process of existing attribute binding technologies require calculating losses and updating the variables in the latent space to control the direction of attention, which will undoubtedly bring a greater amount of calculation, and the time to generate an image will also increase accordingly. Summary of the Invention

[0005] In view of this, the present invention provides a method, system, device, medium and product for binding 3D model attributes, which solves the technical problems that existing attribute binding technologies are prone to difficult binding problems when dealing with the attribute binding between different parts of the same object, are difficult to handle the attribute binding between different parts of the same object, and have a large computational power.

[0006] The first aspect of the present invention provides a method for binding 3D model attributes, including:

[0007] Obtain a target 3D model, as well as an attribute prompt word and a reference picture corresponding to the target 3D model; wherein, the attribute prompt word includes a modifier and a noun;

[0008] Perform depth extraction on each face of the target 3D model to obtain a depth map corresponding to the target 3D model;

[0009] Extract the image features of the reference picture based on a pre-trained CLIP image encoder;

[0010] Input the attribute prompt words, the depth map, and the image features into the IP-Adapter model. Based on text cross-attention, semantically bind the modifier and the noun, and perform iterative training in combination with the depth map and the image features to output an image bound with the attribute prompt words.

[0011] Preferably, the method further includes:

[0012] Use the spaCY library to extract the modifier and the noun in the attribute prompt words.

[0013] Preferably, the depth extraction of each face in the target 3D model to obtain the depth map corresponding to the target 3D model includes:

[0014] Perform depth extraction on each face in the target 3D model to obtain the depth map of each face in the 3D model;

[0015] For each face in the 3D model, perform normalization processing on the depth map to obtain the depth map corresponding to the target 3D model.

[0016] Preferably, the extraction of the image features of the reference picture based on the pre-trained CLIP image encoder includes:

[0017] Input the reference picture and the attribute prompt words corresponding to the reference picture into the pre-trained CLIP image encoder to obtain a picture embedding and a prompt word embedding;

[0018] Divide the cross-attention of the Unet network of the pre-trained CLIP image encoder into a picture cross-attention mechanism and a text cross-attention mechanism;

[0019] Based on the picture cross-attention mechanism and the text cross-attention mechanism, perform cross-attention distribution on the picture embedding and the prompt word embedding respectively with the same latent variable to obtain the image features of the reference picture.

[0020] Preferably, the inputting the attribute prompt words, the depth map, and the image features into the IP-Adapter model, semantically binding the modifier and the noun based on text cross-attention, and performing iterative training in combination with the depth map and the image features to output an image bound with the attribute prompt words includes:

[0021] Input the attribute prompt words, the depth map, and the image features into the IP-Adapter model. Based on the text cross-attention in the denoising network Unet in the IP-Adapter model, perform attention distribution on the modifier and the noun, so that the modifier and the noun are semantically bound;

[0022] Under the constraints of the attribute prompt words after attention distribution, the depth map, and the image features, the noise at the current step is predicted through the denoising network Unet;

[0023] Update the current latent variable based on the noise at the current step, predict the noise at the next step according to the current latent variable, and repeat this step until the convergence condition is reached, and then generate the image under the current latent variable;

[0024] Decode the image under the current latent variable to obtain the image bound with the attribute prompt words.

[0025] Preferably, the attention calculation formula of the text cross-attention is:

[0026]

[0027] In the formula, D is the attention score, is the Softmax function, is the query vector, is the key vector, T is the matrix transpose, is the value vector, d is the dimension of the key vector, i is the noun index, n is the number of nouns, K i is the key vector of the i-th noun, is the value vector of the modifier corresponding to the i-th noun.

[0028] In a second aspect, the present invention also provides a three-dimensional model attribute binding system, including:

[0029] A data acquisition module for acquiring a target three-dimensional model, as well as the attribute prompt words and reference pictures corresponding to the target three-dimensional model; wherein, the attribute prompt words include modifiers and nouns;

[0030] A depth extraction module for extracting the depth of each face in the target three-dimensional model to obtain the depth map corresponding to the target three-dimensional model;

[0031] A feature extraction module for extracting the image features of the reference picture based on a pre-trained CLIP image encoder;

[0032] An attribute binding module for inputting the attribute prompt words, the depth map, and the image features into the IP-Adapter model, semantically binding the modifier and the noun based on text cross-attention, and performing iterative training in combination with the depth map and the image features, and outputting the image bound with the attribute prompt words.

[0033] In a third aspect, the present invention further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the three-dimensional model attribute binding method as described in the first aspect.

[0034] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the three-dimensional model attribute binding method as described in the first aspect are implemented.

[0035] In a fifth aspect, the present invention further provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the steps of the three-dimensional model attribute binding method as described in the first aspect.

[0036] As can be seen from the above technical solutions, the present invention deeply extracts each face in the target three-dimensional model, extracts the image features of the reference picture through a pre-trained CLIP image encoder, inputs the attribute prompt words, depth map, and image features into the IP-Adapter model, and binds the modifier and the noun semantically based on text cross-attention, forcing the modifier and the noun to be associated, promoting the semantic association of related attributes, so as to avoid the attribute leakage problem that occurs when binding the attributes of different parts of the same object, and realizing the attribute binding between different parts of the same object. At the same time, since there is no need to calculate the loss and there is no need to backpropagate to update the variables in the latent space, the computational amount is significantly reduced. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is an application environment of a three-dimensional model attribute binding method provided by an embodiment of the present invention;

[0039] Figure 2 It is a flowchart of a three-dimensional model attribute binding method provided by an embodiment of the present invention;

[0040] Figure 3 It is a reference picture provided by an embodiment of the present invention;

[0041] Figure 4Depth map of the target 3D model provided by the embodiments of the present invention;

[0042] Figure 5 Image generated before modification provided by the embodiments of the present invention;

[0043] Figure 6 Image generated after modification provided by the embodiments of the present invention;

[0044] Figure 7 Attention map with the token "orange" at the tenth step of denoising before modification provided by the embodiments of the present invention;

[0045] Figure 8 Attention map with the token "legs" at the tenth step of denoising before modification provided by the embodiments of the present invention;

[0046] Figure 9 Attention map with the token "orange" at the tenth step of denoising after modification provided by the embodiments of the present invention;

[0047] Figure 10 Attention map with the token "legs" at the tenth step of denoising after modification provided by the embodiments of the present invention;

[0048] Figure 11 Schematic structural diagram of a 3D model attribute binding system provided by the embodiments of the present invention;

[0049] Figure 12 Schematic structural diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] The 3D model attribute binding method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed on the cloud or other network servers. The terminal 101 or the server 102 obtains a target 3D model, as well as an attribute prompt word and a reference picture corresponding to the target 3D model; among them, the attribute prompt word includes a modifier and a noun; the depth of each face in the target 3D model is extracted to obtain a depth map corresponding to the target 3D model; based on the pre-trained CLIP image encoder, the image features of the reference picture are extracted; the attribute prompt word, the depth map and the image features are input into the IP-Adapter model, and the modifier and the noun are semantically bound based on text cross-attention, and iterative training is performed in combination with the depth map and the image features to output an image with the bound attribute prompt word.

[0052] The terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc.

[0053] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0054] Such as Figure 2 As shown, an embodiment of the present application provides a 3D model attribute binding method. Taking the method applied to Figure 1 the terminal 101 or the server 102 in it as an example for description, it includes the following steps S1 to step S4. Among them:

[0055] Step S1, obtain a target 3D model, as well as an attribute prompt word and a reference picture corresponding to the target 3D model; among them, the attribute prompt word includes a modifier and a noun.

[0056] Among them, by obtaining the target 3D model, at the same time, obtain the attribute prompt word and the reference picture related to the target 3D model. The attribute prompt word is used to describe the specific attributes of the target 3D model, such as color, material or texture, etc. These attribute prompt words are composed of a modifier and a noun. The modifier is used to limit the specific manifestation of the noun. For example, "red" in "red chair" is the modifier, and "chair" is the noun. The reference picture provides the appearance reference of the target 3D model in the real world or virtual environment, helping the model to understand and bind attributes more accurately.

[0057] Step S2, extract the depth of each face in the target 3D model to obtain a depth map corresponding to the target 3D model.

[0058] Step S3, extract the image features of the reference picture based on the pre-trained CLIP image encoder.

[0059] The CLIP image encoder is the Contrastive Language–Image Pre-training model. This model is jointly trained with a large number of text-image pairs, learning rich visual and semantic feature representations, enabling effective association and matching of images and text.

[0060] Among them, the reference picture is input into the pre-trained CLIP image encoder, and using the powerful image understanding ability of this encoder, the key image features in the reference picture are extracted. These features include visual elements such as colors, shapes, and textures in the picture.

[0061] Step S4: Input the attribute prompt words, depth map, and image features into the IP-Adapter model. Based on text cross-attention, semantic binding of the modifier and noun is performed, and iterative training is carried out in combination with the depth map and image features to output an image with the attribute prompt words bound.

[0062] Among them, the Input-Prompt Adapter is the input prompt adapter model. This model is designed to handle cross-modal tasks and shows strong capabilities especially in binding text attributes to image content. The IP-Adapter model receives the attribute prompt words, depth map, and image features as inputs and uses the built-in text cross-attention mechanism to accurately identify and bind the semantic relationship between the modifier and the noun. In this process, the depth map provides the spatial structure information of the 3D model, while the image features capture the visual details of the reference picture, and both jointly assist the IP-Adapter model in performing more accurate attribute binding.

[0063] It should be noted that in the embodiments of this application, by extracting the depth of each face in the target 3D model and extracting the image features of the reference picture through the pre-trained CLIP image encoder, the attribute prompt words, depth map, and image features are input into the IP-Adapter model. Based on text cross-attention, semantic binding of the modifier and noun is performed, forcing the modifier and noun to be related, promoting the semantic association of related attributes with each other, so as to avoid the problem of attribute leakage when binding attributes to different parts of the same object, and achieving attribute binding between different parts of the same object. At the same time, since there is no need to calculate the loss and there is no need to backpropagate to update the variables in the latent space, the computational amount is significantly reduced.

[0064] In some embodiments, this method further includes:

[0065] Using the spaCY library to extract the modifier and noun in the attribute prompt words.

[0066] Among them, through the natural language processing ability of the spaCY library, the grammatical structure in the attribute prompt words can be efficiently parsed, and the modifier and noun can be accurately distinguished. As an advanced natural language processing tool, the spaCY library provides rich language models and algorithms, which can accurately identify information such as part-of-speech and syntactic relationships in the text.

[0067] In the embodiment of the present application, using the spaCY library to parse the attribute prompt words can ensure the accurate extraction of modifiers and nouns, providing a reliable basis for the subsequent attribute binding step. Through the processing of the spaCY library, the accuracy and efficiency of attribute binding can be further improved, enabling the 3D model to more accurately reflect the attribute characteristics expected by the user.

[0068] In some embodiments, deep extraction is performed on each face of the target 3D model to obtain a depth map corresponding to the target 3D model, including:

[0069] Step S201: Perform deep extraction on each face of the target 3D model to obtain depth maps of each face in the 3D model;

[0070] Step S202: For each face in the 3D model, perform normalization processing on the depth map to obtain a depth map corresponding to the target 3D model.

[0071] Exemplarily, after obtaining the depth maps of each face in the target 3D model through pytorch3D, in order to ensure that the depth information ranges in the depth maps of each face of the 3D model are consistent and reduce the situation of generating image errors due to depth range differences, it is necessary to perform normalization processing on the depth maps to convert them into relative depth maps to unify the input information.

[0072] Specifically, calculate the minimum and maximum values of the current depth value ( ).

[0073]

[0074] In the formula, is the minimum depth value, is the maximum depth value.

[0075] Reverse the depth value so that the farther depth value becomes smaller and the closer depth value becomes larger; and normalize the depth value to the interval [0, 1], that is:

[0076]

[0077] In the formula, is the normalized depth value.

[0078] Map the normalized depth value to the target range [50, 255].

[0079]

[0080] In the formula, is the depth value mapped to the target range.

[0081] In some embodiments, extracting the image features of the reference picture based on a pre-trained CLIP image encoder includes:

[0082] Step S301: Input the reference picture and the attribute prompt words corresponding to the reference picture into the pre-trained CLIP image encoder to obtain a picture embedding and a prompt word embedding.

[0083] Among them, the CLIP image encoder is Contrastive Language–Image Pre-training, that is, a contrastive language-image pre-training model. Taking the reference picture and the attribute prompt words corresponding to the reference picture as inputs, and utilizing the powerful capabilities of the pre-trained CLIP image encoder, the picture is deeply understood and analyzed. Through joint training of a large number of text-image pairs, this encoder has learned rich visual and semantic feature representations. Therefore, it can accurately capture the key information in the reference picture. During the encoding process, the picture is converted into a vector representation in a high-dimensional space, that is, the picture embedding, which can represent the rich content and style of the image, and the attribute prompt words are also converted into corresponding vector representations, that is, the prompt word embeddings.

[0084] Step S302: Divide the cross-attention of the Unet network of the pre-trained CLIP image encoder into a picture cross-attention mechanism and a text cross-attention mechanism.

[0085] Step S303: Based on the picture cross-attention mechanism and the text cross-attention mechanism, perform cross-attention allocation on the picture embedding and the prompt word embedding respectively with the same latent variable to obtain the image features of the reference picture.

[0086] Among them, the process of performing cross-attention allocation on the picture embedding and the prompt word embedding respectively with the same latent variable through the picture cross-attention mechanism and the text cross-attention mechanism is:

[0087]

[0088] In the formula, is the Softmax function, is the query vector, is the key vector of the text feature, T is the matrix transpose, is the value vector of the text feature, d is the dimension of the key vector, is the key vector of the picture feature, The value vector of the image features.

[0089] Among them, by performing cross-attention allocation on the image embedding and the prompt embedding, the effective association and matching of image and text information can be achieved. Under the image cross-attention mechanism, each element in the image embedding will interact with the latent variable to evaluate their correlation. Similarly, under the text cross-attention mechanism, each element in the prompt embedding will also interact with the latent variable. This two-way cross-attention mechanism ensures the comprehensive and in-depth interaction between image and text information, enabling the more accurate extraction of key image features in the reference image. These features not only reflect the visual elements in the image, such as color, shape, and texture, but also embody the semantic association with the attribute prompts.

[0090] In some embodiments, the attribute prompt, the depth map, and the image features are input into the IP-Adapter model. Based on the text cross-attention, the modifier and the noun are semantically bound, and iterative training is performed in combination with the depth map and the image features to output the image with the bound attribute prompt, including:

[0091] Step S401: Input the attribute prompt, the depth map, and the image features into the IP-Adapter model. Based on the text cross-attention in the denoising network Unet in the IP-Adapter model, attention allocation is performed on the modifier and the noun, so that the modifier and the noun are semantically bound.

[0092] In the text cross-attention of the IP-Adapter denoising network Unet, the original attention mechanism calculation formula is:

[0093]

[0094] Among them, it is found that the key (Key) and value (Value) in the cross-attention layer have strong semantic meanings related to the object layout and content. At the same time, we find that the product of the attention score ( ) and the value (V) is the multiplication of each attribute with itself and is not associated with other attributes. We believe that this is the main reason why the attribute cannot be correctly bound to the object.

[0095] Based on this, the application embodiment proposes a new attention calculation formula for text cross-attention as:

[0096]

[0097] In the formula, D is the attention score, is the Softmax function, is the query vector, is the key vector, and T is the matrix transpose. is a value vector, d is the dimension of the key vector, i is the noun index, n is the number of nouns, and K i is the key vector of the i-th noun, is the value vector of the modifier corresponding to the i-th noun.

[0098] Among them, the first term in the numerator is the original formula, and the second term forces the modifier token to be associated with the entity noun attribute. Among them, K i only contains entity nouns, Vi only contains modifiers. By separately splitting entity nouns and modifiers, tokens that are semantically related are forced to be associated with each other, prompting the attention of the modifier to fall within the scope of the entity noun.

[0099] Exemplarily, the value of Q is [16, 9216, 40], and the value of K is [16, 77, 40]. Among them, 77 in K represents 77 tokens, and each prompt word occupies one token. If the number of prompt word tokens is less than 77, it will be padded with [PAD] to a fixed length. Therefore, the values of the attention map obtained by Q and K through Softmax are [16, 9216, 77]. Similarly, V is also a text feature like K, and its value is [16, 77, 40]. It is realized that tokens that are semantically related are forced to be associated with each other, prompting the attention of the modifier to fall within the scope of the entity noun. Specifically, first, on the attention map [16, 9216, 77], I only use entity nouns for calculation, and the extracted dimension is [16, 9216, 1], while V also only uses modifiers for calculation, and the extracted dimension is [16, 1, 40]. Then, through the attention formula calculation, according to the principle of matrix calculation, performing matrix multiplication in this way can force the association between the modifier and the entity noun.

[0100] At the same time, since the improvement is made to the forward propagation process of the model and no training is required, in order to retain the original capabilities of the model, we retain the original formula and balance the resulting changes by taking the average of the cumulative sum of the second term, thereby ensuring that the original performance of the model is not affected.

[0101] Step S402: Under the constraints of the attribute prompt words, depth map, and image features after attention allocation, predict the noise at the current step through the denoising network Unet.

[0102] Step S403: Update the current latent variable based on the noise at the current step, predict the noise at the next step according to the current latent variable, and repeat this step until the convergence condition is reached, and then generate an image under the current latent variable;

[0103] Step S404: Decode the image under the current latent variable to obtain an image bound with attribute prompt words.

[0104] Among them, the UNet network consists of a downsampling module, an intermediate module, and an upsampling module, and each module contains an attention formula. The process of generating an image is a process of denoising. In the first cycle, the latent variable of random noise that follows a Gaussian distribution is input into the UNet network. Each time the UNet network runs, it can predict the noise at the current step. Then, the latent variable input into the UNet network at the current step is adjusted according to the noise predicted by the UNet network to gradually approach the target image. After the UNet network is executed in several cycles, the decoder of the VAE can decode the latent variable into a normal image.

[0105] In an exemplary embodiment, in order to more clearly clarify a three-dimensional model attribute binding method provided by the embodiments of the present application, the following uses a specific embodiment to specifically describe the three-dimensional model attribute binding method.

[0106] First, input the attribute prompt (prompt: a black chair with orange legs, hint: a black chair with orange legs), the reference picture as Figure 3 shown, and the three-dimensional model.

[0107] Then, use the spaCY library to extract the modifier and noun pairs in the attribute prompt, and distinguish the modifier and noun in the attribute prompt. That is, prompt: a black chair with orange legs will be extracted as [black, chair] and [orange, legs]. Avoid using manual annotation of modifiers and nouns.

[0108] After obtaining the depth map of the three-dimensional model as Figure 4 shown through pytorch3D, in order to ensure that the depth information ranges in the depth maps of each face of the three-dimensional model are consistent and reduce the situation of generating image errors due to depth range differences, it is necessary to normalize the depth map and convert it into a relative depth map to unify the input information.

[0109] Extract image features from the reference picture using a pre-trained CLIP image encoder. The CLIP model is a multi-modal model trained through contrastive learning on a large dataset containing image-text pairs. The image embedding obtained using the CLIP model can represent the rich content and style of the image. In the Unet network, the image embedding can make the generated picture carry the style and features in the reference picture through a cross-attention mechanism with the latent variable Z. The specific operation design is as follows: In the cross-attention of the Unet network, the cross-attention mechanism is divided into a picture cross-attention mechanism and a text cross-attention mechanism. Both the image embedding and the prompt embedding perform a cross-attention mechanism with the same latent variable Z. The formula is:

[0110]

[0111] In the formula, is the Softmax function, is the query vector, is the key vector of the text feature, T is the matrix transpose, is the value vector of the text feature, d is the dimension of the key vector, is the key vector of the picture feature, is the value vector of the picture feature.

[0112] Improve the text cross-attention mechanism in the denoising network Unet of IP-Adapter: Input the prompt, reference picture, and depth map into IP-Adapter to generate a picture. The attention calculation formula for the proposed new text cross-attention is:

[0113]

[0114] In the formula, D is the attention score, is the Softmax function, is the query vector, is the key vector, T is the matrix transpose, is the value vector, d is the dimension of the key vector, i is the noun index, n is the number of nouns, Ki is the key vector of the i-th noun, is the value vector of the modifier corresponding to the i-th noun.

[0115] Among them, the first term of the numerator is the original formula, and the second term forces the modifier token to be associated with the entity noun attribute. Among them, Ki only contains entity nouns, and Vi only contains modifiers. By separately splitting entity nouns and modifiers, it forces tokens that are semantically related to be associated with each other, prompting the attention of the modifier to fall within the scope of the entity noun.

[0116] As Figure 5 shown in the figure generated before modification, Figure 6 is the figure generated after modification, Figure 7 and Figure 8 are the attention maps of the tokens "orange" and "legs" at the tenth denoising step before modification, respectively. Figure 9 and Figure 10 are the attention maps of the tokens "orange" and "legs" at the tenth denoising step after modification, respectively. By comparison, it is found that in the attention map, the brighter the color, the higher the similarity degree of the token to this area. Figure 7 and Figure 8 are the attention maps of the tokens "orange" and "legs" at the tenth denoising step before modification, respectively, indicating the degree of attention of each word in the text feature K to different regions in the image feature Q, that is, the attention Figure 7 and Figure 8 respectively represent the degree of attention of "orange" to the image feature Q and the degree of attention of "legs" to the image feature Q.

[0117] Compared with Figure 7 and Figure 8 before modification, Figure 9 and Figure 10 after modification, the color on the legs of the chair is brighter, which can illustrate that the similarity degree of the tokens "orange" and "legs" to the image feature Q in this area is higher, that is, the modifier orange is successfully bound to the legs.

[0118] Based on the same inventive concept, the embodiment of the present application also provides a three-dimensional model attribute binding system for implementing the above-mentioned three-dimensional model attribute binding method.

[0119] The implementation solution provided by this system to solve the problem is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the three-dimensional model attribute binding system provided below can refer to the limitations on the three-dimensional model attribute binding method in the above text, and will not be repeated here.

[0120] As Figure 11 shown, the embodiment of the present application provides a three-dimensional model attribute binding system, including:

[0121] The data acquisition module 100 is used to acquire a target 3D model, as well as the corresponding attribute prompt words and reference pictures for the target 3D model; among them, the attribute prompt words include modifiers and nouns;

[0122] The depth extraction module 200 is used to perform depth extraction on each face of the target 3D model to obtain the depth map corresponding to the target 3D model;

[0123] The feature extraction module 300 is used to extract the image features of the reference pictures based on the pre-trained CLIP image encoder;

[0124] The attribute binding module 400 is used to input the attribute prompt words, depth map and image features into the IP-Adapter model, semantically bind the modifiers and nouns based on text cross-attention, and perform iterative training in combination with the depth map and image features, and output the image with the bound attribute prompt words.

[0125] In some embodiments, the system further includes: a prompt word extraction module, which is used to extract the modifiers and nouns in the attribute prompt words using the spaCY library.

[0126] In some embodiments, the depth extraction module 200 is used to:

[0127] Perform depth extraction on each face of the target 3D model to obtain the depth map of each face in the 3D model;

[0128] For each face in the 3D model, perform normalization processing on the depth map to obtain the depth map corresponding to the target 3D model.

[0129] In some embodiments, the feature extraction module 300 is used to:

[0130] Input the reference picture and the corresponding attribute prompt words of the reference picture into the pre-trained CLIP image encoder to obtain picture embeddings and prompt word embeddings;

[0131] Divide the cross-attention of the Unet network of the pre-trained CLIP image encoder into a picture cross-attention mechanism and a text cross-attention mechanism;

[0132] Based on the picture cross-attention mechanism and the text cross-attention mechanism, perform cross-attention distribution on the picture embeddings and prompt word embeddings respectively with the same latent variable to obtain the image features of the reference picture.

[0133] In some embodiments, the attribute binding module 400 is used to:

[0134] Input the attribute prompt words, depth map, and image features into the IP-Adapter model. Based on the text cross-attention in the denoising network Unet in the IP-Adapter model, allocate attention between the modifier and the noun, so that the modifier and the noun are semantically bound;

[0135] Under the constraints of the attribute prompt words, depth map, and image features after attention allocation, predict the noise at the current step through the denoising network Unet;

[0136] Update the current latent variable based on the noise at the current step, predict the noise at the next step according to the current latent variable, and repeat this step until the convergence condition is reached, then generate the image under the current latent variable;

[0137] Decode the image under the current latent variable to obtain the image with the bound attribute prompt words.

[0138] In some embodiments, the attention calculation formula of the text cross-attention is:

[0139]

[0140] In the formula, D is the attention score, is the Softmax function, is the query vector, is the key vector, T is the matrix transpose, is the value vector, d is the dimension of the key vector, i is the noun index, n is the number of nouns, K i is the key vector of the i-th noun, is the value vector of the modifier corresponding to the i-th noun.

[0141] As Figure 12 shown, an embodiment of the present application provides an electronic device. The electronic device 10 includes a memory 20 and a processor 30. A computer program is stored in the memory 20. When the computer program is executed by the processor 30, the processor 30 is caused to execute the steps of the three-dimensional model attribute binding method in the above embodiment.

[0142] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the three-dimensional model attribute binding method in the above embodiment are implemented.

[0143] An embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. Wherein, when the program instructions are executed by a computer, the computer is caused to execute the steps of the three-dimensional model attribute binding method in the above embodiment.

[0144] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, electronic devices, computer storage media, and computer program products described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0145] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0146] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0147] In several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media, computer program products, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0150] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs that can store program codes.

[0151] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A three-dimensional model attribute binding method, characterized in that: include: Acquire a target three-dimensional model, and attribute prompt words and reference pictures corresponding to the target three-dimensional model; wherein the attribute prompt words include modifiers and nouns; Performing depth extraction on each surface in the target three-dimensional model to obtain a depth map corresponding to the target three-dimensional model; Extracting image features of the reference image based on a pre-trained CLIP image encoder; The attribute prompt word, the depth map and the image features are input into the IP-Adapter model, the modifier is semantically bound to the noun based on text cross-attention, and iterative training is performed in combination with the depth map and the image features to output an image bound with the attribute prompt word.

2. The three-dimensional model attribute binding method according to claim 1, characterized in that: Also includes: The spaCY library is used to extract modifiers and nouns in the attribute prompt words.

3. The three-dimensional model attribute binding method according to claim 1, characterized in that: The step of performing depth extraction on each surface in the target three-dimensional model to obtain a depth map corresponding to the target three-dimensional model includes: Performing depth extraction on each surface in the target three-dimensional model to obtain a depth map of each surface in the three-dimensional model; For each face in the three-dimensional model, the depth map is normalized to obtain a depth map corresponding to the target three-dimensional model.

4. The three-dimensional model attribute binding method according to claim 1, characterized in that: The pre-trained CLIP image encoder extracts image features of the reference image, including: Inputting the reference picture and the attribute prompt word corresponding to the reference picture into the pre-trained CLIP image encoder to obtain picture embedding and prompt word embedding; Dividing the cross attention of the Unet network of the pre-trained CLIP image encoder into a picture cross attention mechanism and a text cross attention mechanism; Based on the picture cross-attention mechanism and the text cross-attention mechanism, cross-attention allocation is performed on the picture embedding and the prompt word embedding respectively with the same latent variable to obtain the image features of the reference picture.

5. The three-dimensional model attribute binding method according to claim 1, characterized in that: The method of inputting the attribute prompt word, the depth map and the image feature into the IP-Adapter model, semantically binding the modifier word with the noun based on text cross attention, iteratively training the depth map and the image feature, and outputting the image bound with the attribute prompt word includes: Inputting the attribute prompt word, the depth map and the image feature into the IP-Adapter model, allocating attention to the modifier and the noun based on the text cross attention in the denoising network Unet in the IP-Adapter model, so that the modifier and the noun are semantically bound; Based on the attribute prompt words after attention allocation, the depth map and the constraints of the image features, the noise at the current step is predicted by the denoising network Unet; Update the current latent variable based on the noise at the current step number, predict the noise at the next step number based on the current latent variable, and repeat this step until the convergence condition is reached, and generate the image under the current latent variable; The image under the current latent variable is decoded to obtain an image bound to the attribute prompt word.

6. The three-dimensional model attribute binding method according to claim 5, characterized in that: The attention calculation formula of the text cross attention is: Where D is the attention score, is the Softmax function, is the query vector, is the key vector, T is the matrix transpose, is the value vector, d is the dimension of the key vector, i is the noun index, n is the number of nouns, K i is the key vector of the i-th noun, is the value vector of the modifier corresponding to the i-th noun.

7. A 3D model attribute binding system, characterized in that: include: A data acquisition module, used to acquire a target three-dimensional model, and attribute prompt words and reference pictures corresponding to the target three-dimensional model; wherein the attribute prompt words include modifiers and nouns; A depth extraction module, used to perform depth extraction on each surface in the target three-dimensional model to obtain a depth map corresponding to the target three-dimensional model; A feature extraction module, used for extracting image features of the reference image based on a pre-trained CLIP image encoder; The attribute binding module is used to input the attribute prompt word, the depth map and the image feature into the IP-Adapter model, semantically bind the modifier word with the noun based on text cross attention, and iteratively train in combination with the depth map and the image feature to output the image bound with the attribute prompt word.

8. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the three-dimensional model attribute binding method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the three-dimensional model attribute binding method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the steps of the three-dimensional model attribute binding method as described in any one of claims 1-6.