Three-dimensional model rendering method, related device, equipment and storage medium

Through the end-to-end texture map generation process, UV control signals, text control signals and noise images are used, combined with the backbone network and control network, the texture map of the three-dimensional model is generated, solving the problem of time-consuming and high learning cost of manually generating texture maps, and achieving efficient and simple texture map generation.

CN119941952APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510020466.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, it takes a long time to generate texture maps manually, especially for complex three-dimensional models, which requires a lot of time to design and adjust the UV layout. At the same time, mastering professional software requires a certain learning cost, which is not friendly for beginners or when project cycles are tight.

Method used

It provides an end-to-end texture map generation process, by obtaining the texture map coordinate UV control signals and text control signals of the basic three-dimensional model, combined with noise images, a map generation model (including backbone network and control network) is used to generate texture maps for the basic three-dimensional model.

Benefits of technology

It realizes simple and direct operation and high efficiency texture maps, reduces the time for manual design and adjustment of UV layout, and reduces the cost of learning professional software, which is suitable for beginners and tight project cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941952A_ABST
    Figure CN119941952A_ABST
Patent Text Reader

Abstract

The invention discloses a rendering method of a three-dimensional model, a related device, equipment and a storage medium. The method comprises the following steps: acquiring a texture mapping coordinate UV control signal corresponding to a basic three-dimensional model; acquiring a text control signal; based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained through a map generation model, the map generation model comprises a backbone network and a control network, the control network is used for providing UV tensor generated based on the UV control signal for the backbone network, and the control network is used for providing the texture map for the basic three-dimensional model; the backbone network is used for outputting a texture mapping graph according to the UV tensor, the text control signal and the noise image; and rendering the basic three-dimensional model by using the texture mapping graph to obtain a target three-dimensional model. According to the scheme provided by the invention, an end-to-end texture map generation process is realized, the operation is simple and direct, and the generation efficiency is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a rendering method, related apparatus, device and storage medium for a three-dimensional model. Background Art

[0002] In today's digital age, computer graphics has helped the development of many industries. Early 3D models were rough and lacked texture details. In order to meet people's high requirements for visual effects, texture mapping came into being. It makes characters more realistic in games, helps with special effects production in movies and TV, gives products an appearance in industrial design, and shows details in the field of architecture.

[0003] Currently, users can create texture maps using professional software. First, import the 3D model. Then, use automatic mapping to initially unfold the UVs, that is, map the surface of the 3D model to a 2D plane space, which is represented by U and V coordinate axes. Finally, users need to manually adjust the UV block layout and export it.

[0004] However, the inventors have found that the current solutions have at least the following problems: manually generating texture maps takes a long time, especially for complex three-dimensional models, which requires a lot of time to design and adjust the UV layout. In addition, mastering professional software requires a certain learning cost, which is not friendly to beginners or those with tight project cycles. In this regard, an effective method is urgently needed to solve such problems. Summary of the invention

[0005] The embodiments of the present application provide a three-dimensional model rendering method, related devices, equipment and storage medium, which realize an end-to-end texture mapping process, which is not only simple and direct to operate but also has high generation efficiency.

[0006] In view of this, the present application provides a method for rendering a three-dimensional model, comprising:

[0007] Obtain the texture mapping coordinate UV control signal corresponding to the basic 3D model;

[0008] Get text control signal;

[0009] Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained through a mapping generation model, wherein the mapping generation model includes a backbone network and a control network, the control network is used to provide the backbone network with a UV tensor generated based on the UV control signal, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image;

[0010] The basic 3D model is rendered using a texture map to obtain a target 3D model.

[0011] Another aspect of the present application provides a three-dimensional model rendering device, comprising:

[0012] An acquisition module is used to obtain a texture mapping coordinate UV control signal corresponding to a basic three-dimensional model;

[0013] The acquisition module is also used to acquire the text control signal;

[0014] The acquisition module is further used to acquire a texture map for the basic three-dimensional model through a mapping generation model based on the UV control signal, the text control signal and the noise image, wherein the mapping generation model includes a backbone network and a control network, the control network is used to provide the backbone network with a UV tensor generated based on the UV control signal, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image;

[0015] The rendering module is used to render the basic three-dimensional model using a texture map to obtain a target three-dimensional model.

[0016] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0017] An acquisition module, specifically used to acquire a basic three-dimensional model;

[0018] Perform UV unfolding on the basic 3D model to obtain a 2D UV geometric image;

[0019] The feature extraction is performed on the two-dimensional UV geometric image to obtain the UV control signal.

[0020] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0021] An acquisition module, specifically used to acquire text prompt information in response to a text input operation;

[0022] Segment the text prompt information to obtain at least one sentence;

[0023] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0024] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0025] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0026] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0027] An acquisition module, specifically used to acquire image prompt information in response to an image input operation;

[0028] Based on the image prompt information, generate text prompt information through the image-to-text model;

[0029] Segment the text prompt information to obtain at least one sentence;

[0030] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0031] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0032] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0033] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0034] An acquisition module, specifically configured to acquire first text prompt information in response to a text input operation;

[0035] In response to the image input operation, acquiring image prompt information;

[0036] Based on the image prompt information, generate second text prompt information through the image-to-text model;

[0037] Segment processing is performed on the first text prompt information and the second text prompt information to obtain at least one sentence;

[0038] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0039] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0040] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0041] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0042] An acquisition module, specifically used to acquire a random control signal when no text input operation or image input operation is detected;

[0043] Use random control signals as text control signals.

[0044] In one possible design, in another implementation of another aspect of the embodiment of the present application, the backbone network includes at least two backbone modules, and the control network includes a first control module;

[0045] An acquisition module, specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0046] Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through a first control module included in the control network;

[0047] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0048] A texture map for the base three-dimensional model is generated according to the second noise tensor.

[0049] In one possible design, in another implementation of another aspect of the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module;

[0050] An acquisition module, specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0051] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0052] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0053] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0054] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0055] In one possible design, in another implementation of another aspect of the embodiment of the present application, the backbone network includes at least four backbone modules, and the control network includes a first control module and a second control module;

[0056] An acquisition module, specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0057] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0058] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0059] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0060] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0061] Based on the text control signal, the third noise tensor and the second UV tensor, a fourth noise tensor is obtained through a fourth backbone module included in the backbone network;

[0062] A texture map for the base three-dimensional model is generated according to the fourth noise tensor.

[0063] In one possible design, in another implementation of another aspect of the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module and a second control module;

[0064] An acquisition module, specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0065] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0066] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0067] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0068] Based on the text control signal, the second noise tensor and the second UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0069] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0070] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0071] An acquisition module, specifically used to generate a first key matrix and a first value matrix according to the text control signal;

[0072] generating a first query matrix according to the noise image;

[0073] Multiply and scale the first query matrix by the transpose of the first key matrix to obtain a first attention score matrix;

[0074] Normalize the first attention score matrix to obtain the first attention weight matrix;

[0075] A first noise tensor is generated according to the first attention weight matrix and the first value matrix.

[0076] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0077] An acquisition module, specifically used to generate a second key matrix and a second value matrix according to a UV control signal, a text control signal and a noise image;

[0078] Generate a second query matrix according to the text control signal and the noise image;

[0079] Multiply and scale the second query matrix by the transpose of the second key matrix to obtain a second attention score matrix;

[0080] Normalize the second attention score matrix to obtain the second attention weight matrix;

[0081] Generate a first UV tensor according to the second attention weight matrix and the second value matrix.

[0082] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0083] An acquisition module, specifically used to generate a third key matrix and a third value matrix according to the text control signal, the first noise tensor and the first UV tensor;

[0084] Generate a third query matrix according to the text control signal and the first noise tensor;

[0085] Multiply and scale the third query matrix by the transpose of the third key matrix to obtain a third attention score matrix;

[0086] Normalize the third attention score matrix to obtain the third attention weight matrix;

[0087] A second noise tensor is generated according to the third attention weight matrix and the third value matrix.

[0088] Another aspect of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned methods when executing the computer program.

[0089] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned methods are implemented.

[0090] Another aspect of the present application provides a computer program product, including a computer program, which implements the above-mentioned methods when executed by a processor.

[0091] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0092] In an embodiment of the present application, a rendering method for a three-dimensional model is provided, which obtains a UV control signal corresponding to a basic three-dimensional model on the one hand, and obtains a text control signal on the other hand. Thus, a mapping map generation model processes a noise image based on the UV control signal and the text control signal to obtain a texture mapping map. The target three-dimensional model can be obtained by rendering the texture mapping map to the basic three-dimensional model. In the above manner, an end-to-end texture mapping map generation process is realized, that is, a rich variety of target three-dimensional models can be generated by inputting different basic three-dimensional models, without spending a lot of time designing and adjusting the UV layout, thereby accelerating the generation efficiency and diversity of assets and reducing the cost of asset production. It can be seen that not only is the operation simple and direct, but the generation efficiency is also high. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 A schematic diagram of a virtual try-on in an embodiment of the present application;

[0094] Figure 2 A schematic diagram for use in teaching material production in an embodiment of the present application;

[0095] Figure 3 A schematic diagram of an embodiment of the present application applied to game design;

[0096] Figure 4 A schematic diagram of an implementation environment of the three-dimensional model rendering method in an embodiment of the present application;

[0097] Figure 5 A schematic diagram of another implementation environment of the three-dimensional model rendering method in an embodiment of the present application;

[0098] Figure 6 A schematic diagram of a flow chart of a three-dimensional model rendering method in an embodiment of the present application;

[0099] Figure 7 A schematic diagram of the structure of the DiT model in the embodiment of the present application;

[0100] Figure 8 A schematic diagram of a structure of a mapping diagram generation model in an embodiment of the present application;

[0101] Fig. 9 Another structural schematic diagram of the mapping diagram generation model in the embodiment of the present application;

[0102] Fig.10 Another structural schematic diagram of the mapping diagram generation model in the embodiment of the present application;

[0103] Fig.11 Another structural schematic diagram of the mapping diagram generation model in the embodiment of the present application;

[0104] Fig.12 A schematic diagram of feature processing based on a backbone module and a control module in an embodiment of the present application;

[0105] Fig.13 A schematic diagram of a process flow of a three-dimensional model rendering device in an embodiment of the present application;

[0106] Fig.14 A schematic diagram of the structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0107] The embodiments of the present application provide a three-dimensional model rendering method, related devices, equipment and storage medium, which realize an end-to-end texture mapping process, which is not only simple and direct to operate but also has high generation efficiency.

[0108] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0109] Texture mapping is a technique that applies an image or pattern to the surface of a 3D model. It adds rich details and realism to the originally monotonous geometric model. Through texture mapping, the material characteristics of objects, such as wood grain, stone texture, fabric texture, etc., can be simulated to make the virtual object closer to the appearance of objects in the real world. It can be seen that texture mapping is extremely important in computer graphics and related fields.

[0110] Currently, users can create texture maps through professional software, or generate texture maps with the help of a model based on a U-net. Among them, the encoder of UNet will reduce the resolution of the noise image (latent) and restore it through the decoder. Although such an operation can retain globality and can greatly reduce the amount of calculation, it makes it difficult to generate high-resolution images due to the reduction in resolution. In addition, the generated texture map is a two-dimensional image and can only reflect one perspective. For texture coloring of three-dimensional models, if two-dimensional images of multiple perspectives are generated and baked, there is a problem of inconsistent results of multi-perspective generation.

[0111] Based on this, in an embodiment of the present application, a mapping generation model for generating a texture map is provided. The mapping generation model adds a control network on the basis of the diffusion transformer (DiT) model, which not only retains the generalization of big data, but also realizes the geometric control of the model on the texture mapping coordinates (UV). As a result, fine textures can be generated directly in the UV space, avoiding occlusion and multi-view inconsistency problems, and a complete texture can be generated at one time without baking, so that it can be directly provided to users as a three-dimensional asset.

[0112] Before introducing the specific method of the present application, the application scenario of the present application is first exemplified. It should be noted that the following application scenario is only for illustration and is not limited thereto.

[0113] 1. Applied to virtual try-on scenarios;

[0114] Users' virtual try-on of clothing allows them to feel the effects of products more accurately before purchasing, reduce shopping risks, and improve user satisfaction and willingness to buy.

[0115] For example, see Figure 1 , Figure 1 This is a schematic diagram of a virtual try-on in an embodiment of the present application. As shown in the figure, a basic three-dimensional model (e.g., a human body model) is displayed on the virtual try-on interface. Based on this, the user can enter text prompt information, that is, a description of the clothing. In addition, the user can also upload pictures related to clothing, and the background generates relevant text prompt information based on the pictures. Thus, the mapping map generation model is called to generate the target three-dimensional model, that is, to show the effect after wearing the clothing. The user can also choose accessories and other objects for matching, so as to present a better try-on effect.

[0116] 2. Applied to teaching material production scenarios;

[0117] In today's digital teaching era, making virtual architectural 3D models for teaching plays an important role. Relying on advanced 3D modeling software, it can accurately construct the three-dimensional form of various buildings. For teachers, the model can intuitively display the spatial structure, internal layout and decorative details of the building, making abstract architectural knowledge visual and significantly improving teaching efficiency.

[0118] For example, see Figure 2 , Figure 2 This is a schematic diagram of the teaching material production in the embodiment of the present application. As shown in the figure, a basic three-dimensional model (for example, a building model) is displayed on the teaching material production interface. Based on this, the user can enter text prompt information, that is, a relevant description of the building. Thus, the mapping map generation model is called to generate a target three-dimensional model corresponding to the building, thereby presenting the appearance of the building to be displayed.

[0119] 3. Applied to game design scenarios;

[0120] The 3D models used in games are very important. They can create realistic game scenes and enhance the immersion and sense of involvement in the game. At the same time, models of different styles can enrich the game content and meet the needs of diversified gameplay. Exquisite models can also improve the visual quality of the game and attract more players, thereby improving the competitiveness and commercial value of the game.

[0121] For example, see Figure 3 , Figure 3 This is a schematic diagram of a game design in an embodiment of the present application. As shown in the figure, a basic three-dimensional model (for example, a vehicle model) is displayed on the game production interface. Based on this, the user can enter text prompt information, that is, a description of the vehicle model. Thus, the mapping map generation model is called to generate a target three-dimensional model corresponding to the vehicle model, thereby presenting the appearance of the vehicle to be displayed.

[0122] It should be noted that the above application scenarios are only examples, and the 3D model rendering method provided in this embodiment can also be applied to other scenarios, which are not limited here.

[0123] The method provided in this application can be applied to Figure 4 The implementation environment shown includes a terminal 401. The terminal 401 involved in this application includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, virtual reality devices, smart home appliances, vehicle terminals, aircraft, etc. Among them, the client is deployed on the terminal 401, and the client can be run on the terminal 401 in the form of a browser, or in the form of an independent application (application, APP) or a small program, etc.

[0124] In combination with the above implementation environment, in step A1, the user inputs the basic three-dimensional model and text prompt information through the terminal 401. In step A2, the terminal 401 generates a UV control signal based on the basic three-dimensional model. In step A3, the terminal 401 generates a text control signal based on the text prompt information. In step A4, the terminal 401 calls the mapping map generation model to process the UV control signal, the text control signal and the noise image to obtain a texture mapping map. In step A5, the basic three-dimensional model is rendered using the texture mapping map to obtain a target three-dimensional model. In step A6, the target three-dimensional model is displayed.

[0125] The method provided in this application can be applied to Figure 5 The implementation environment shown in the figure includes a terminal 501 and a server 502, and the terminal 501 and the server 502 can communicate with each other through a network 503. The network 503 uses standard communication technology and / or protocols, usually the Internet, but can also be any network, including but not limited to Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, dedicated network or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technology can be used to replace or supplement the above data communication technology.

[0126] The terminal 501 involved in this application is Figure 4 The terminal 401 in is similar, so it will not be described here.

[0127] The server 502 involved in the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence (AI) platforms.

[0128] It should be noted that this application takes the configuration of the mapping generation model deployed on the server 502 as an example for explanation. In some embodiments, the configuration of the mapping generation model can also be deployed on the terminal 501. In some embodiments, part of the configuration of the mapping generation model is deployed on the terminal 501, and part of the configuration is deployed on the server 502.

[0129] In combination with the above implementation environment, in step B1, the user inputs the basic three-dimensional model and text prompt information through the terminal 501. In step B2, the terminal 501 sends the basic three-dimensional model and text prompt information to the server 502 through the network 503. In step B3, the server 502 generates a UV control signal based on the basic three-dimensional model. In step B4, the server 502 generates a text control signal based on the text prompt information. In step B5, the server 502 calls the mapping map generation model to process the UV control signal, the text control signal and the noise image to obtain a texture mapping map. In step B6, the server 502 uses the texture mapping map to render the basic three-dimensional model to obtain a target three-dimensional model. In step B7, the server 502 sends the target three-dimensional model to the terminal 501 through the network 503. In step B8, the terminal 501 displays the target three-dimensional model.

[0130] In combination with the above introduction, the rendering method of the three-dimensional model in this application will be introduced below. Figure 6 The rendering method of the three-dimensional model in the embodiment of the present application can be completed independently by the server, or by the terminal, or by the terminal and the server in cooperation. The method provided by the present application includes:

[0131] S601, obtaining a texture mapping coordinate UV control signal corresponding to a basic three-dimensional model;

[0132] In one or more embodiments, a basic three-dimensional model is first obtained, and then a corresponding UV control signal may be generated based on the basic three-dimensional model, wherein the UV control signal may be represented as a three-dimensional tensor.

[0133] It is understood that the basic 3D model can also be called a "white model", "plain model", etc., which refers to a 3D model that only contains the basic geometric shape and structure of the model without adding details such as material, texture, color, etc. The basic 3D model is mainly composed of polygons (for example, triangles, quadrilaterals, etc.) and is used to construct the general shape of the model.

[0134] S602, obtaining a text control signal;

[0135] In one or more embodiments, a text control signal is obtained, and the text control signal can be represented as a three-dimensional tensor, i.e., (B, Kt, C). For example, "B" represents the number of sentences (e.g., B=10), "Kt" represents the maximum number of words (e.g., Kt=154), and "C" represents the word vector dimension (e.g., C=4096).

[0136] S603, based on the UV control signal, the text control signal and the noise image, obtaining a texture map for the basic three-dimensional model through a mapping generation model, wherein the mapping generation model includes a backbone network and a control network, the control network is used to provide a UV tensor generated based on the UV control signal to the backbone network, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image;

[0137] In one or more embodiments, the UV control signal, the text control signal, and the noise image are used as inputs of a mapping generation model, and a texture mapping map is outputted through the mapping generation model. A texture mapping map refers to a two-dimensional image mapped on the surface of a three-dimensional object, which is used to enhance the visual effect of the object and make it look more real and detailed. Texture mapping maps usually include attributes such as color, smoothness, and reflectivity, which play a key role in the three-dimensional rendering process.

[0138] Specifically, the mapping graph generation model includes a backbone network and a control network, wherein the backbone network includes M backbone modules, and the control network includes N control modules, where M is greater than N, and both M and N are positive integers. The backbone network can adopt a DiT model with a transformer structure as the core, such as Figure 7 As shown, Figure 7 This is a schematic diagram of the structure of the DiT model in the embodiment of the present application, in which the DiT model collects billions of image-text pairs as training data and is trained with a large number of graphics cards. During inference, the input text is used as a condition, and a prediction image is obtained after multiple rounds of denoising process. It has strong generalization and text response capabilities and can retain the details of high-resolution images.

[0139] The control network is similar to the encoder of UNet, which is used to receive UV control signals, text control signals and random noise signals corresponding to the noise image. Based on this, the output of each control module is used as the input of the backbone module, so that the backbone module can learn the UV characteristics of the basic 3D model without changing the original structure, thereby achieving fine-tuning training and reducing training costs.

[0140] S604: Rendering the basic three-dimensional model using the texture map to obtain a target three-dimensional model.

[0141] In one or more embodiments, in 3D modeling and rendering, texture mapping is the process of attaching a 2D texture image to the surface of a 3D model. That is, texture mapping coordinates (i.e., UV coordinates) are required to determine how each pixel on the texture map is associated with the vertices of the base 3D model. Among them, UV coordinates are 2D coordinates, which are stored together with the vertices of the base 3D model and are used to interpolate and calculate texture coordinates during the rendering process to obtain the target 3D model.

[0142] In an embodiment of the present application, a method for rendering a three-dimensional model is provided. Through the above method, an end-to-end texture map generation process is realized, that is, the user is given a basic three-dimensional model without texture, and can quickly generate high-definition and delicate texture maps with different styles through different prompts, and usually obtain a target three-dimensional model with texture within 15 seconds. Since the image is generated based on UV coordinates, ambiguity and incompleteness caused by occlusion, missing perspective, etc. can be avoided.

[0143] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining the texture mapping coordinate UV control signal corresponding to the basic three-dimensional model specifically includes:

[0144] Obtain a basic 3D model;

[0145] Perform UV unfolding on the basic 3D model to obtain a 2D UV geometric image;

[0146] The feature extraction is performed on the two-dimensional UV geometric image to obtain the UV control signal.

[0147] In one or more embodiments, a method for generating a UV control signal is introduced. As can be seen from the above embodiments, the basic three-dimensional model includes the basic structure of the model, and a two-dimensional UV geometric image can be obtained after UV unfolding the basic three-dimensional model. Based on this, the two-dimensional UV geometric image is used as the input of a feature extraction model (for example, a convolutional neural network (CNN)), and an image encoding matrix is ​​output through the model. Combined with the number of sentences of the text control signal, the same number of image encoding matrices are generated, thereby obtaining a UV control signal.

[0148] Specifically, UV unfolding of the base 3D model means the process of mapping each point on the surface of the base 3D model to a 2D plane. That is, each 3D vertex on the base 3D model is mapped to a pixel on the 2D UV geometric image, and the position xyz vertex attributes are rasterized to obtain a 2D UV geometric map. In 3D modeling and rendering, UV unfolding is crucial because it determines how the texture is applied to the 3D model. Each vertex has a corresponding UV coordinate in the 3D model, and these coordinates define the position of each pixel on the texture image. It can be seen that the 2D UV geometric image refers to the result of mapping the surface of the 3D model to a 2D plane. The 2D UV geometric map is specifically manifested as a series of meshes, each of which represents a face of the model and corresponds to a pixel on the texture image through UV coordinates.

[0149] It should be noted that the number of parameters in the control network is about 1 / 10 of that in the backbone network to support fine-tuning on small-scale data. In the 3D data of hundreds of thousands, each object is retopologically, UV-unwrapped, and texture-mapped, and hundreds of thousands of pairs of 2D UV geometric images and texture-mapped training data are obtained. Training with an 8-card cluster takes about 3 days. The trained model will not generate images in the original rendering perspective, but in the UV space.

[0150] Secondly, in an embodiment of the present application, a method for generating a UV control signal is provided. Through the above method, the geometric prior of the basic three-dimensional model as a whole (i.e., geometric information such as depth, normal, symmetry, and regional similarity) is fully utilized, thereby guiding the model to directly generate a texture map instead of a two-dimensional image of a single perspective. At the same time, the texture map refers to the geometric distribution information in the two-dimensional UV geometric image, retaining features such as coherence and symmetry, so the target three-dimensional model obtained is consistent with human intuition and the physical characteristics of the input geometry.

[0151] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining a text control signal specifically includes:

[0152] In response to a text input operation, obtaining text prompt information;

[0153] Segment the text prompt information to obtain at least one sentence;

[0154] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0155] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0156] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0157] In one or more embodiments, a method for generating a text control signal is introduced. As can be seen from the above embodiments, one of the inputs of the mapping generation model is a text control signal. The text control signal can be derived from text prompt information input by the user. The following will introduce a process of generating a text control signal based on the text prompt information.

[0158] Specifically, the user can enter a paragraph as text prompt information. First, the text prompt information is processed by sentence segmentation to obtain B sentences (B is an integer greater than or equal to 1). Punctuation marks can usually be used as the basis for sentence segmentation. After the sentence segmentation is completed, each sentence is segmented to obtain at least one word included in each sentence. Next, taking a sentence as an example, all the words in the sentence are encoded separately, for example, the words are encoded using a bidirectional encoder representations from transformers (BERT) model to obtain a word vector. Based on this, for each sentence, a matrix (Kt, C) composed of word vectors is constructed, and the matrix composed of all sentences is the text control signal (B, Kt, C). That is, "B" represents the number of sentences, assuming that each word is represented as a token, "Kt" represents the maximum number of tokens, and "C" represents the word vector dimension.

[0159] It should be noted that the term "in response to" in this application is used to indicate the conditions or states on which the execution of an operation depends, and one or more operations can be executed when certain conditions or states are met. These operations can be real-time or have a certain delay.

[0160] Secondly, in the embodiment of the present application, a method for generating a text control signal is provided. Through the above method, the user can input text prompt information according to actual needs, so that the model can focus on the content of the text prompt information and generate a texture map that meets the user's needs, thereby improving the product's user experience.

[0161] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining a text control signal specifically includes:

[0162] In response to the image input operation, acquiring image prompt information;

[0163] Based on the image prompt information, generate text prompt information through the image-to-text model;

[0164] Segment the text prompt information to obtain at least one sentence;

[0165] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0166] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0167] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0168] In one or more embodiments, another method of generating a text control signal is introduced. As can be seen from the above embodiments, one of the inputs of the mapping generation model is a text control signal. The text control signal can be derived from image prompt information input by the user. The following will introduce the process of generating a text control signal based on the image prompt information.

[0169] Specifically, the user can input a picture as image prompt information. First, the image-to-text model is called to analyze the image prompt information to obtain the text prompt information. Then, the text prompt information is segmented. After the sentence segmentation is completed, each sentence is segmented to obtain at least one word included in each sentence. Next, taking a sentence as an example, all the words in the sentence are encoded to obtain the word vector corresponding to each word. Based on this, for each sentence, a matrix composed of word vectors is constructed, and the matrix composed of all sentences is the text control signal (B, Kt, C).

[0170] Exemplarily, the image-to-text model used in this application can be a vision caption model. The vision caption model uses CNN to analyze image prompt information. For example, for a landscape photo, CNN can identify objects in the image (e.g., mountains, rivers), scene layout (e.g., the positional relationship of objects), and some details (e.g., the color of the sky), etc., and convert these visual information into image feature vectors. Then, using the image feature vector as input, a recurrent neural network (RNN) or transformer architecture is used to output appropriate words to generate text prompt information.

[0171] Exemplarily, the image-to-text model used in this application can also be a generative pretrained transformer 4vision (GPT-4) model. The GPT-4 model performs feature analysis on the input image prompt information, encodes information such as visual elements (e.g., objects, scenes, actions) in the image, and then uses the language model part to generate text prompt information for the image prompt information with reference to the learned knowledge and patterns.

[0172] Secondly, in the embodiment of the present application, another method of generating a text control signal is provided. Through the above method, the user can input image prompt information according to actual needs, and generate corresponding text prompt information based on the image prompt information. In this way, it is possible to capture very detailed content in the image prompt information, thereby providing more accurate guidance, and then enabling the model to generate a texture map that meets the user's needs, thereby improving the product's user experience.

[0173] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining a text control signal specifically includes:

[0174] In response to a text input operation, obtaining first text prompt information;

[0175] In response to the image input operation, acquiring image prompt information;

[0176] Based on the image prompt information, generate second text prompt information through the image-to-text model;

[0177] Segment processing is performed on the first text prompt information and the second text prompt information to obtain at least one sentence;

[0178] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0179] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0180] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0181] In one or more embodiments, another method of generating a text control signal is introduced. As can be seen from the above embodiments, one of the inputs of the mapping generation model is a text control signal. The text control signal can be derived from text prompt information and image prompt information input by the user. The following will introduce a process of generating a text control signal based on the text prompt information and the image prompt information.

[0182] Specifically, on the one hand, the user can input a paragraph as the first text prompt information. On the other hand, the user can input a picture as the image prompt information. Based on this, the image prompt information is analyzed by calling the image-to-text model to obtain the second text prompt information. Then, the first text prompt information and the second text prompt information are spliced ​​into a target text prompt information. Therefore, the target text prompt information is sentence-processed. After the sentence is completed, each sentence is segmented separately to obtain at least one word included in each sentence. Next, taking a sentence as an example, all the words in the sentence are encoded separately to obtain the word vector corresponding to each word. Based on this, for each sentence, a matrix composed of word vectors is constructed separately, and the matrix composed of all sentences is the text control signal (B, Kt, C).

[0183] It should be noted that the image-generated text model used in this application can be a vision caption model or a GPT-4 model, or other image-generated text models, which are not limited here.

[0184] Secondly, in the embodiment of the present application, another method of generating a text control signal is provided. Through the above method, the user can input image prompt information and text prompt information at the same time according to actual needs, and use the image prompt information together with the text prompt information after converting it into text, so as to combine visual details with abstract concepts. In this way, the generated content can be made more three-dimensional, and the model can generate a texture map that meets the user's needs, thereby improving the product's user experience.

[0185] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining a text control signal specifically includes:

[0186] acquiring a random control signal when no text input operation or image input operation is detected;

[0187] Use random control signals as text control signals.

[0188] In one or more embodiments, another method of generating a text control signal is introduced. As can be seen from the above embodiments, one of the inputs of the mapping generation model is a text control signal. If the user does not input text or an image, the text control signal can be generated in the following manner.

[0189] (1) Use the random word generation tool;

[0190] Specifically, words are randomly selected based on a given vocabulary to construct text prompt information. Usually, a certain number of words are randomly selected from a predefined word set, and they are combined according to certain grammatical rules (for example, noun + verb + adjective) to obtain text prompt information. Based on this, the text prompt information is sentence-by-sentence processed. After the sentence is completed, each sentence is word-by-sentence processed to obtain at least one word included in each sentence. Next, taking a sentence as an example, all the words in the sentence are encoded separately to obtain the word vector corresponding to each word. Based on this, for each sentence, a matrix composed of word vectors is constructed separately, and the matrix composed of all sentences is a random control signal, that is, a text control signal.

[0191] (2) Use random functions;

[0192] Specifically, the deep learning framework has a random function for initializing tensors. Based on this, the random function can be used in combination with a suitable shape to generate a random control signal. Taking PyTorch as an example, the torch.randn function can generate a normally distributed random tensor (i.e., a random control signal), so the random tensor can be directly used as a text control signal.

[0193] Secondly, in the embodiment of the present application, another method of generating a text control signal is provided. By using the random control signal as the text control signal in the above manner, the model can be prompted to develop a more flexible processing mechanism. In terms of content generation, the input of the random control signal can make the generated content more diverse.

[0194] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the backbone network includes at least two backbone modules, and the control network includes a first control module;

[0195] Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained by generating a model through a mapping map, specifically including:

[0196] Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network;

[0197] Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through a first control module included in the control network;

[0198] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0199] A texture map for the base three-dimensional model is generated according to the second noise tensor.

[0200] In one or more embodiments, a method for generating a texture map based on a map generation model is introduced. From the above embodiments, it can be seen that the main backbone network includes M backbone modules, and the control network includes N control modules. The following will be described by taking M greater than 2 (for example, M equals 30) and N equals 1 as an example.

[0201] Specifically, for ease of understanding, see Figure 8 , Figure 8 This is a structural schematic diagram of the mapping image generation model in the embodiment of the present application. As shown in the figure, a two-dimensional UV geometric image is obtained after UV expansion of a given basic three-dimensional model, and then a UV control signal is generated. Feature extraction is performed on the randomly initialized noise image to obtain a random noise signal. The text control signal and the random noise signal are input together into the first trunk module (i.e., "trunk module 1"), and the attention calculation is performed by the first trunk module to obtain a first noise tensor. The UV control signal, the text control signal and the random noise signal are input together into the first control module (i.e., "control module 1"), and the first control module performs attention calculation to obtain a first UV tensor.

[0202] The first noise tensor and the first UV tensor are input into the second backbone module (i.e., "backbone module 2"), and the text control signal is input, and the attention calculation is performed through the second backbone module to obtain the second noise tensor. After being processed by multiple backbone modules, the model gradually converts the noise image into a texture map, and finally renders the texture map to the basic 3D model to obtain the target 3D model.

[0203] It should be noted that the DiT model divides the image into blocks to obtain a random noise signal, but does not change the size of the random noise signal. Instead, each backbone module uses a random noise signal of the same size for interactive calculations, thereby retaining the high resolution of the image. Based on this, the backbone network copies several multimodal diffusion transformer modules (MM-Dit-Block) from the DiT model and trains them with the same weights as the initial values. Among them, the backbone module involved in the present application may be a multimodal diffusion transformer module (MM-Dit-Block).

[0204] It can be seen that the data flow of the mapping graph generation model is unidirectional. The results of the control network will be accumulated in the backbone network, but the results of the backbone network will not be input into the control network. This way, the control network will learn only the control function as much as possible, and the generalization of the backbone network will be retained to the greatest extent.

[0205] Secondly, in the embodiment of the present application, a method for generating a texture map based on a map generation model is provided. In the above method, since there is less UV data, the data set cannot be too large, and the overall task and the texture map are in the same generation domain, only one control module can be trained. That is, using a control module to connect with the backbone network can not only provide geometric priors, but also reduce the amount of calculation and the number of parameters of the model, thereby reducing the computational cost.

[0206] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module;

[0207] Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained by generating a model through a mapping map, specifically including:

[0208] Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network;

[0209] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0210] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0211] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0212] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0213] In one or more embodiments, a method for generating a texture map based on a map generation model is introduced. From the above embodiments, it can be seen that the main backbone network includes M backbone modules, and the control network includes N control modules. The following will be described by taking M greater than 2 (for example, M equals 30) and N equals 1 as an example.

[0214] Specifically, for ease of understanding, see Fig. 9 , Fig. 9Another structural schematic diagram of the mapping diagram generation model in the embodiment of the present application, as shown in the figure, a two-dimensional UV geometric image is obtained after UV expansion of a given basic three-dimensional model, and then a UV control signal is generated. Feature extraction is performed on the randomly initialized noise image to obtain a random noise signal. The text control signal and the random noise signal are input together into the first trunk module (i.e., "trunk module 1"), and the attention calculation is performed by the first trunk module to obtain a first noise tensor. The first noise tensor and the text control signal are input together into the second trunk module (i.e., "trunk module 2"), and the attention calculation is performed by the second trunk module to obtain a second noise tensor. The UV control signal, the text control signal and the random noise signal are input together into the first control module (i.e., "control module 1"), and the first control module performs attention calculation to obtain a first UV tensor.

[0215] The first UV tensor output by the first control module can be used as the input of a backbone module. For example, the first UV tensor and the second noise tensor are input into the third backbone module (i.e., "backbone module 3"), and the text control signal is also input into the third backbone module, and the third backbone module performs attention calculation to obtain the third noise tensor. After being processed by multiple backbone modules, the model gradually converts the noise image into a texture map, and finally renders the texture map to the basic three-dimensional model to obtain the target three-dimensional model.

[0216] Secondly, in an embodiment of the present application, another method for generating a texture map based on a mapping generation model is provided. Through the above method, since there is less UV data, the data set cannot be too large, and the overall task and the texture map are in the same generation domain, only one control module can be trained. That is, by connecting a control module with any backbone module, on the one hand, it can not only provide geometric priors, but also reduce the amount of calculation and the number of parameters of the model, thereby reducing the computational cost. On the other hand, it can improve the diversity of the model structure and enhance the generalization ability of the model.

[0217] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the backbone network includes at least four backbone modules, and the control network includes a first control module and a second control module;

[0218] Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained by generating a model through a mapping map, specifically including:

[0219] Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network;

[0220] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0221] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0222] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0223] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0224] Based on the text control signal, the third noise tensor and the second UV tensor, a fourth noise tensor is obtained through a fourth backbone module included in the backbone network;

[0225] A texture map for the base three-dimensional model is generated according to the fourth noise tensor.

[0226] In one or more embodiments, another method for generating a texture map based on a map generation model is introduced. As can be seen from the above embodiments, the main backbone network includes M backbone modules, and the control network includes N control modules. The following will be described by taking M greater than 4 (for example, M equals 30) and N equals 2 as an example.

[0227] Specifically, for ease of understanding, see Fig.10 , Fig.10 Another structural schematic diagram of the mapping diagram generation model in the embodiment of the present application, as shown in the figure, a two-dimensional UV geometric image is obtained after UV expansion of a given basic three-dimensional model, and then a UV control signal is generated. Feature extraction is performed on the randomly initialized noise image to obtain a random noise signal. The text control signal and the random noise signal are input together into the first trunk module (i.e., "trunk module 1"), and the attention calculation is performed by the first trunk module to obtain a first noise tensor. The first noise tensor and the text control signal are input together into the second trunk module (i.e., "trunk module 2"), and the attention calculation is performed by the second trunk module to obtain a second noise tensor. The UV control signal, the text control signal and the random noise signal are input together into the first control module (i.e., "control module 1"), and the first control module performs attention calculation to obtain a first UV tensor.

[0228] The first UV tensor output by the first control module can be used as the input of a certain backbone module. Take the first UV tensor and the second noise tensor as an example, which are input into the third backbone module (i.e., "backbone module 3") together, and the text control signal is also input into the third backbone module, and the attention calculation is performed through the third backbone module to obtain the third noise tensor. In addition, the first UV tensor output by the first control module also needs to be input into the second control module (i.e., "control module 2"). Based on this, the UV control signal, the text control signal and the first UV tensor are input into the second control module together, and the second control module performs attention calculation to obtain the second UV tensor.

[0229] The text control signal, the third noise tensor, and the second UV tensor are input to the fourth backbone module (i.e., "backbone module 4"), and the fourth backbone module performs attention calculation to obtain the fourth noise tensor. After being processed by multiple backbone modules, the model gradually converts the noise image into a texture map, and finally renders the texture map to the basic 3D model to obtain the target 3D model.

[0230] Secondly, in an embodiment of the present application, another method for generating a texture map based on a mapping generation model is provided. Through the above method, since there is less UV data, the data set cannot be too large, and the overall task and the texture map are in the same generation domain, only two control modules can be trained. By connecting the two control modules with the backbone network, on the one hand, it can not only provide geometric priors, but also capture more comprehensive features. On the other hand, it can improve the diversity of the model structure and enhance the generalization ability of the model.

[0231] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module and a second control module;

[0232] Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained by generating a model through a mapping map, specifically including:

[0233] Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network;

[0234] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0235] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0236] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0237] Based on the text control signal, the second noise tensor and the second UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0238] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0239] In one or more embodiments, another method for generating a texture map based on a map generation model is introduced. As can be seen from the above embodiments, the main backbone network includes M backbone modules, and the control network includes N control modules. The following will be described by taking M greater than 3 (for example, M equals 30) and N equals 2 as an example.

[0240] Specifically, for ease of understanding, see Fig.11 , Fig.11 Another structural schematic diagram of the mapping image generation model in the embodiment of the present application, as shown in the figure, a two-dimensional UV geometric image is obtained after UV expansion of a given basic three-dimensional model, and then a UV control signal is generated. Feature extraction is performed on the randomly initialized noise image to obtain a random noise signal. The text control signal and the random noise signal are input together into the first trunk module (i.e., "trunk module 1"), and the attention calculation is performed by the first trunk module to obtain a first noise tensor. The UV control signal, the text control signal and the random noise signal are input together into the first control module (i.e., "control module 1"), and the first control module performs attention calculation to obtain a first UV tensor.

[0241] The first UV tensor, the text control signal, and the first noise tensor are used as inputs of the second backbone module (i.e., "backbone module 2"), and the attention calculation is performed by the second backbone module to obtain the second noise tensor. In addition, the first UV tensor output by the first control module also needs to be input to the second control module (i.e., "control module 2"). Based on this, the UV control signal, the text control signal, and the first UV tensor are input into the second control module together, and the second control module performs attention calculation to obtain the second UV tensor.

[0242] The second noise tensor, the second UV tensor, and the text control signal are input into the third backbone module (i.e., "backbone module 3"), and the third noise tensor is obtained by performing attention calculation through the third backbone module. After being processed by multiple backbone modules, the model gradually converts the noise image into a texture map, and finally renders the texture map to the basic 3D model to obtain the target 3D model.

[0243] Secondly, in the embodiment of the present application, another method for generating a texture map based on a map generation model is provided. In the above method, since there is less UV data, the data set cannot be too large, and the overall task and the texture map are in the same generation domain, only two control modules can be trained. Using two control modules connected to the backbone network can not only provide geometric priors, but also capture more comprehensive features.

[0244] Based on the text control signal and the noise image, a first noise tensor is obtained through a first backbone module included in the backbone network, specifically including:

[0245] Generate a first key matrix and a first value matrix according to the text control signal;

[0246] generating a first query matrix according to the noise image;

[0247] Multiply and scale the first query matrix by the transpose of the first key matrix to obtain a first attention score matrix;

[0248] Normalize the first attention score matrix to obtain the first attention weight matrix;

[0249] A first noise tensor is generated according to the first attention weight matrix and the first value matrix.

[0250] In one or more embodiments, a method for performing attention calculation based on a backbone module is introduced. As can be seen from the above embodiments, each backbone module uses an attention mechanism to calculate the input data. In one case, the backbone network needs to perform attention calculation on two input data. For ease of understanding, please refer to Fig.11 , below, we will take the first backbone module (i.e., “backbone module 1”) as an example to introduce the process of attention calculation through the first backbone module.

[0251] Specifically, assuming that the text control signal is represented by (B, Kt, C), and the random noise signal corresponding to the noise image is represented by (B, Kx, C). For ease of calculation, the dimension "B" can be ignored first, that is, the part (Kt, C) is extracted from the text control signal as the first key matrix and the first value matrix, and (Kx, C) is extracted from the random noise signal as the first query matrix.

[0252] First, the first attention score matrix can be calculated as follows:

[0253]

[0254] Among them, scaled scores represents the scaled attention score (i.e., the first attention score matrix), QK Trepresents the original attention score (i.e., the original attention score matrix). Q represents the first query matrix, K represents the first key matrix, and d k Represents the feature dimension.

[0255] Then, the first attention weight matrix can be calculated as follows:

[0256]

[0257] Where P represents the first attention weight matrix. Softmax() represents normalization.

[0258] Next, the first noise matrix can be calculated as follows:

[0259] Attention(Q,K,V)=P*V;Formula (3)

[0260] Among them, Attention(Q,K,V) represents the first noise matrix, and V represents the first value matrix.

[0261] Finally, the first noise matrix restores the dimension "B", that is, the first noise tensor is obtained.

[0262] Again, in the embodiment of the present application, a method for performing attention calculation based on the backbone module is provided. Through the above method, the attention mechanism can help the model distinguish between noise and useful image features, more effectively retain key features and remove noise during the denoising process, thereby improving the quality of the generated image.

[0263] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, based on the UV control signal, the text control signal and the noise image, obtaining the first UV tensor through the first control module specifically includes:

[0264] Generate a second key matrix and a second value matrix according to the UV control signal, the text control signal and the noise image;

[0265] Generate a second query matrix according to the text control signal and the noise image;

[0266] Multiply and scale the second query matrix by the transpose of the second key matrix to obtain a second attention score matrix;

[0267] Normalize the second attention score matrix to obtain the second attention weight matrix;

[0268] Generate a first UV tensor according to the second attention weight matrix and the second value matrix.

[0269] In one or more embodiments, a method for performing attention calculation based on a control module is introduced. As can be seen from the above embodiments, each control module uses an attention mechanism to calculate the input data. The control module needs to perform attention calculation on the three input data. For ease of understanding, please refer to Fig.11 , below, we will take the first control module (ie, “control module 1”) as an example to introduce the process of performing attention calculation through the first control module.

[0270] Specifically, see Fig.12 , Fig.12 This is a diagram of feature processing based on the backbone module and the control module in an embodiment of the present application. As shown in the figure, it is assumed that the text control signal is represented by (B, Kt, C), the random noise signal corresponding to the noise image is represented by (B, Kx, C), and the UV control signal is represented by (B, Kx, C) after passing through a trainable linear network. For ease of calculation, the dimension "B" can be ignored first. Based on this, (Kx, C) extracted from the UV control signal after the linear operation is spliced ​​with (Kt, C) extracted from the text control signal and (Kx, C) extracted from the random noise signal to obtain (Kt+2Kx, C), which is used as the second key matrix and the second value matrix. The (Kt, C) extracted from the text control signal and the (Kx, C) extracted from the random noise signal are spliced ​​to obtain (Kt+Kx, C), which is used as the second query matrix.

[0271] First, the second attention score matrix can be calculated as follows:

[0272]

[0273] Among them, scaled scores represents the scaled attention scores (e.g., the second attention score matrix). represents a query matrix (eg, a second query matrix), represents a key matrix (eg, a second key matrix), d k represents the feature dimension. k q =Kt+Kx,k k =Kt+2Kx.

[0274] Then, the second attention weight matrix can be calculated as follows:

[0275]

[0276] in, represents an attention weight matrix (e.g., a second attention weight matrix). softmax() represents a normalization process.

[0277] Next, the first UV matrix can be calculated as follows:

[0278]

[0279] in, represents the output matrix (e.g., the first UV matrix), represents a value matrix (eg, a second value matrix).

[0280] Finally, the first UV matrix restores the dimension "B", that is, the first UV tensor is obtained.

[0281] It can be seen that during the attention calculation process, the text control signal interacts globally with the random noise signal.

[0282] Again, in the embodiment of the present application, a method for performing attention calculation based on a control module is provided. Through the above method, the attention mechanism can help the model distinguish between noise and useful image features, more effectively retain key features and remove noise during the denoising process, thereby improving the quality of the generated image.

[0283] Optionally, in the above Figure 6 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, based on the text control signal, the first noise tensor and the first UV tensor, a second noise tensor is obtained through a second backbone module included in the backbone network, specifically including:

[0284] Generate a third key matrix and a third value matrix according to the text control signal, the first noise tensor and the first UV tensor;

[0285] Generate a third query matrix according to the text control signal and the first noise tensor;

[0286] Multiply and scale the third query matrix by the transpose of the third key matrix to obtain a third attention score matrix;

[0287] Normalize the third attention score matrix to obtain the third attention weight matrix;

[0288] A second noise tensor is generated according to the third attention weight matrix and the third value matrix.

[0289] In one or more embodiments, a method of fusing the control network output to the backbone network is introduced. As can be seen from the above embodiments, each backbone module uses an attention mechanism to calculate the input data. In another case, the backbone network needs to perform attention calculations on three input data. For ease of understanding, please refer to Fig.11, the following will introduce the process of integrating the output results of the first control network (ie, "control module 1") and the output results of the first trunk module (ie, "trunk module 1") and inputting them into the second trunk module (ie, "trunk module 2").

[0290] Specifically, see Fig.12 , Fig.12 This is a diagram of feature processing based on the backbone module and the control module in an embodiment of the present application. As shown in the figure, it is assumed that the text control signal is represented by (B, Kt, C), the first noise tensor is represented by (B, Kx, C), and the first UV tensor is represented by (B, Kx, C). For ease of calculation, the dimension "B" can be ignored first. Based on this, (Kx, C) extracted from the first UV tensor is concatenated with (Kt, C) extracted from the text control signal and (Kx, C) extracted from the first noise tensor to obtain (Kt+2Kx, C), which is used as the third key matrix and the third value matrix. The (Kt, C) extracted from the text control signal and the (Kx, C) extracted from the first noise tensor are concatenated to obtain (Kt+Kx, C), which is used as the third query matrix.

[0291] First, the third attention score matrix can be calculated as follows:

[0292]

[0293] Among them, scaled scores represents the scaled attention scores (e.g., the third attention score matrix). represents a query matrix (eg, the third query matrix), represents a bond matrix (e.g., the third bond matrix), d k represents the feature dimension. k q =Kt+Kx,k k =Kt+2Kx.

[0294] Then, the third attention weight matrix can be calculated as follows:

[0295]

[0296] in, represents an attention weight matrix (e.g., the third attention weight matrix). softmax() represents a normalization process.

[0297] Next, the second noise matrix can be calculated as follows:

[0298]

[0299] in, represents the output matrix (e.g., the second noise matrix), represents a value matrix (eg, a third value matrix).

[0300] Finally, the second noise matrix restores the dimension "B", that is, the second noise tensor is obtained.

[0301] It can be seen that during the attention calculation process, the text control signal interacts globally with the first noise tensor. At the same time, without adding a control network, k q =k k =Kt+Kx, when adding control network, k q =Kt+Kx,k k =Kt+2Kx. The dimension of the final output UV matrix is ​​still k q *c. After restoring the dimension "B", it is still (B,Kt+Kx,C). This can minimize the impact on the backbone network to maintain generalization.

[0302] Again, in the embodiment of the present application, a method of fusing the control network output with the value backbone network is provided. In the above method, since the dimension of the query matrix is ​​not changed, only the dimension of the key matrix and the value matrix is ​​changed, the original dimension can be maintained after the attention calculation, thereby minimizing the impact on the backbone network to maintain the generalization of the backbone network.

[0303] The following is a detailed description of the 3D model rendering device in this application. Fig.13 , Fig.13 This is a schematic diagram of an embodiment of a 3D model rendering device in an embodiment of the present application. The 3D model rendering device 130 includes:

[0304] An acquisition module 1301 is used to acquire a texture mapping coordinate UV control signal corresponding to a basic three-dimensional model;

[0305] The acquisition module 1301 is also used to acquire a text control signal;

[0306] The acquisition module 1301 is further used to acquire a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image, wherein the map generation model includes a backbone network and a control network, the control network is used to provide the backbone network with a UV tensor generated based on the UV control signal, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image;

[0307] The rendering module 1302 is used to render the basic three-dimensional model using a texture map to obtain a target three-dimensional model.

[0308] Optionally, in the above Fig.13Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0309] An acquisition module 1301 is specifically used to acquire a basic three-dimensional model;

[0310] Perform UV unfolding on the basic 3D model to obtain a 2D UV geometric image;

[0311] The feature extraction is performed on the two-dimensional UV geometric image to obtain the UV control signal.

[0312] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0313] The acquisition module 1301 is specifically configured to acquire text prompt information in response to a text input operation;

[0314] Segment the text prompt information to obtain at least one sentence;

[0315] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0316] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0317] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0318] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0319] The acquisition module 1301 is specifically configured to acquire image prompt information in response to an image input operation;

[0320] Based on the image prompt information, generate text prompt information through the image-to-text model;

[0321] Segment the text prompt information to obtain at least one sentence;

[0322] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0323] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0324] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0325] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0326] The acquisition module 1301 is specifically configured to acquire first text prompt information in response to a text input operation;

[0327] In response to the image input operation, acquiring image prompt information;

[0328] Based on the image prompt information, generate second text prompt information through the image-to-text model;

[0329] Segment processing is performed on the first text prompt information and the second text prompt information to obtain at least one sentence;

[0330] For each sentence in at least one sentence, perform word segmentation on the sentence to obtain at least one word;

[0331] For each word in each sentence, encode the word and obtain the word vector corresponding to the word;

[0332] Generate text control signals based on the word vectors corresponding to each word in each sentence.

[0333] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0334] The acquisition module 1301 is specifically used to acquire a random control signal when no text input operation or image input operation is detected;

[0335] Use random control signals as text control signals.

[0336] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the three-dimensional model rendering device 130 provided in the embodiment of the present application, the backbone network includes at least two backbone modules, and the control network includes a first control module;

[0337] An acquisition module 1301 is specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0338] Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through a first control module included in the control network;

[0339] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0340] A texture map for the base three-dimensional model is generated according to the second noise tensor.

[0341] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the three-dimensional model rendering device 130 provided in the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module;

[0342] An acquisition module 1301 is specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0343] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0344] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0345] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0346] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0347] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the three-dimensional model rendering device 130 provided in the embodiment of the present application, the backbone network includes at least four backbone modules, and the control network includes a first control module and a second control module;

[0348] An acquisition module 1301 is specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0349] Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0350] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0351] Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0352] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0353] Based on the text control signal, the third noise tensor and the second UV tensor, a fourth noise tensor is obtained through a fourth backbone module included in the backbone network;

[0354] A texture map for the base three-dimensional model is generated according to the fourth noise tensor.

[0355] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the three-dimensional model rendering device 130 provided in the embodiment of the present application, the backbone network includes at least three backbone modules, and the control network includes a first control module and a second control module;

[0356] An acquisition module 1301 is specifically configured to acquire a first noise tensor through a first backbone module included in the backbone network based on the text control signal and the noise image;

[0357] Based on the UV control signal, the text control signal and the noise image, a first UV tensor is acquired through a first control module;

[0358] Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network;

[0359] Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through a second control module;

[0360] Based on the text control signal, the second noise tensor and the second UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network;

[0361] A texture map for the base three-dimensional model is generated according to the third noise tensor.

[0362] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0363] An acquisition module 1301 is specifically configured to generate a first key matrix and a first value matrix according to a text control signal;

[0364] generating a first query matrix according to the noise image;

[0365] Multiply and scale the first query matrix by the transpose of the first key matrix to obtain a first attention score matrix;

[0366] Normalize the first attention score matrix to obtain the first attention weight matrix;

[0367] A first noise tensor is generated according to the first attention weight matrix and the first value matrix.

[0368] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0369] The acquisition module 1301 is specifically used to generate a second key matrix and a second value matrix according to the UV control signal, the text control signal and the noise image;

[0370] Generate a second query matrix according to the text control signal and the noise image;

[0371] Multiply and scale the second query matrix by the transpose of the second key matrix to obtain a second attention score matrix;

[0372] Normalize the second attention score matrix to obtain the second attention weight matrix;

[0373] Generate a first UV tensor according to the second attention weight matrix and the second value matrix.

[0374] Optionally, in the above Fig.13 Based on the corresponding embodiment, in another embodiment of the 3D model rendering device 130 provided in the embodiment of the present application,

[0375] The acquisition module 1301 is specifically used to generate a third key matrix and a third value matrix according to the text control signal, the first noise tensor and the first UV tensor;

[0376] Generate a third query matrix according to the text control signal and the first noise tensor;

[0377] Multiply and scale the third query matrix by the transpose of the third key matrix to obtain a third attention score matrix;

[0378] Normalize the third attention score matrix to obtain the third attention weight matrix;

[0379] A second noise tensor is generated according to the third attention weight matrix and the third value matrix.

[0380] Fig.1414 is a schematic diagram of a computer device structure provided by an embodiment of the present application. The computer device involved in the present application may be a server or a terminal, which is not limited here. The computer device 1400 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1422 (for example, one or more processors) and a memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 may be short-term storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device. Furthermore, the central processing unit 1422 may be configured to communicate with the storage medium 1430 to execute a series of instruction operations in the storage medium 1430 on the computer device 1400.

[0381] The computer device 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server 2003 or Windows Server 2008. TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM etc.

[0382] The steps performed by the computer device in the above embodiment can be based on the Fig.14 The computer device structure shown.

[0383] A computer-readable storage medium is also provided in an embodiment of the present application, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the above embodiments are implemented.

[0384] A computer program product is also provided in an embodiment of the present application, including a computer program, which, when executed by a processor, implements the steps of the methods described in the above embodiments.

[0385] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0386] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0387] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0388] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0389] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0390] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a server or terminal device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store computer programs.

[0391] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for rendering a three-dimensional model, characterized in that: include: Obtain the texture mapping coordinate UV control signal corresponding to the basic 3D model; Get text control signal; Based on the UV control signal, the text control signal and the noise image, a texture map for the basic three-dimensional model is obtained through a map generation model, wherein the map generation model includes a backbone network and a control network, the control network is used to provide the backbone network with a UV tensor generated based on the UV control signal, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image; The basic three-dimensional model is rendered using the texture map to obtain a target three-dimensional model.

2. The rendering method according to claim 1, characterized in that: The step of obtaining a texture mapping coordinate UV control signal corresponding to the basic three-dimensional model includes: Acquire the basic three-dimensional model; Performing UV unfolding on the basic three-dimensional model to obtain a two-dimensional UV geometric image; Feature extraction is performed on the two-dimensional UV geometric image to obtain the UV control signal.

3. The rendering method according to claim 1, characterized in that: The obtaining of the text control signal comprises: In response to a text input operation, obtaining text prompt information; Segment the text prompt information to obtain at least one sentence; For each sentence in the at least one sentence, perform word segmentation processing on the sentence to obtain at least one word; For each word in each sentence, encode the word to obtain a word vector corresponding to the word; The text control signal is generated according to the word vector corresponding to each word in each sentence.

4. The rendering method according to claim 1, characterized in that: The obtaining of the text control signal comprises: In response to the image input operation, acquiring image prompt information; Based on the image prompt information, generate text prompt information through an image-to-text model; Segment the text prompt information to obtain at least one sentence; For each sentence in the at least one sentence, perform word segmentation processing on the sentence to obtain at least one word; For each word in each sentence, encode the word to obtain a word vector corresponding to the word; The text control signal is generated according to the word vector corresponding to each word in each sentence.

5. The rendering method according to claim 1, characterized in that: The obtaining of the text control signal comprises: In response to a text input operation, obtaining first text prompt information; In response to the image input operation, acquiring image prompt information; Based on the image prompt information, generating second text prompt information through an image-to-text model; Segment the first text prompt information and the second text prompt information to obtain at least one sentence; For each sentence in the at least one sentence, perform word segmentation processing on the sentence to obtain at least one word; For each word in each sentence, encode the word to obtain a word vector corresponding to the word; The text control signal is generated according to the word vector corresponding to each word in each sentence.

6. The rendering method according to claim 1, characterized in that: The obtaining of the text control signal comprises: acquiring a random control signal when no text input operation or image input operation is detected; The random control signal is used as the text control signal.

7. The rendering method according to any one of claims 1 to 6, characterized in that: The backbone network includes at least two backbone modules, and the control network includes a first control module; The method of acquiring a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image includes: Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network; Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through the first control module included in the control network; Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network; A texture map for the base three-dimensional model is generated according to the second noise tensor.

8. The rendering method according to any one of claims 1 to 6, characterized in that: The backbone network includes at least three backbone modules, and the control network includes a first control module; The method of acquiring a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image includes: Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network; Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network; Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through the first control module; Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third trunk module included in the trunk network; A texture map for the basic three-dimensional model is generated according to the third noise tensor.

9. The rendering method according to any one of claims 1 to 6, characterized in that: The backbone network includes at least four backbone modules, and the control network includes a first control module and a second control module; The method of acquiring a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image includes: Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network; Based on the text control signal and the first noise tensor, obtaining a second noise tensor through a second backbone module included in the backbone network; Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through the first control module; Based on the text control signal, the second noise tensor and the first UV tensor, a third noise tensor is obtained through a third backbone module included in the backbone network; Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through the second control module; Based on the text control signal, the third noise tensor and the second UV tensor, obtaining a fourth noise tensor through a fourth trunk module included in the trunk network; A texture map for the basic three-dimensional model is generated according to the fourth noise tensor.

10. The rendering method according to any one of claims 1 to 6, characterized in that: The backbone network includes at least three backbone modules, and the control network includes a first control module and a second control module; The method of acquiring a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image includes: Based on the text control signal and the noise image, obtaining a first noise tensor through a first backbone module included in the backbone network; Based on the UV control signal, the text control signal and the noise image, obtaining a first UV tensor through the first control module; Based on the text control signal, the first noise tensor and the first UV tensor, obtaining a second noise tensor through a second backbone module included in the backbone network; Based on the UV control signal, the text control signal and the first UV tensor, obtaining a second UV tensor through the second control module; Based on the text control signal, the second noise tensor and the second UV tensor, obtaining a third noise tensor through a third backbone module included in the backbone network; A texture map for the basic three-dimensional model is generated according to the third noise tensor.

11. The rendering method according to claim 10, characterized in that: The acquiring a first noise tensor based on the text control signal and the noise image through a first trunk module included in the trunk network includes: Generate a first key matrix and a first value matrix according to the text control signal; generating a first query matrix according to the noise image; Multiplying and scaling the first query matrix by the transpose of the first key matrix to obtain a first attention score matrix; Normalizing the first attention score matrix to obtain a first attention weight matrix; Generate the first noise tensor according to the first attention weight matrix and the first value matrix.

12. The rendering method according to claim 10, characterized in that: The acquiring a first UV tensor through the first control module based on the UV control signal, the text control signal and the noise image includes: Generate a second key matrix and a second value matrix according to the UV control signal, the text control signal and the noise image; generating a second query matrix according to the text control signal and the noise image; Multiplying and scaling the second query matrix by the transpose of the second key matrix to obtain a second attention score matrix; Normalizing the second attention score matrix to obtain a second attention weight matrix; Generate the first UV tensor according to the second attention weight matrix and the second value matrix.

13. The rendering method according to any one of claims 10 to 12, characterized in that: The acquiring, based on the text control signal, the first noise tensor and the first UV tensor, a second noise tensor through a second trunk module included in the trunk network comprises: Generate a third key matrix and a third value matrix according to the text control signal, the first noise tensor and the first UV tensor; Generate a third query matrix according to the text control signal and the first noise tensor; Multiplying and scaling the third query matrix and the transpose of the third key matrix to obtain a third attention score matrix; Normalizing the third attention score matrix to obtain a third attention weight matrix; Generate the second noise tensor according to the third attention weight matrix and the third value matrix.

14. A three-dimensional model rendering device, characterized in that: include: An acquisition module is used to obtain a texture mapping coordinate UV control signal corresponding to a basic three-dimensional model; The acquisition module is also used to acquire the text control signal; The acquisition module is further used to acquire a texture map for the basic three-dimensional model through a map generation model based on the UV control signal, the text control signal and the noise image, wherein the map generation model includes a backbone network and a control network, the control network is used to provide the backbone network with a UV tensor generated based on the UV control signal, and the backbone network is used to output a texture map according to the UV tensor, the text control signal and the noise image; A rendering module is used to render the basic three-dimensional model using the texture map to obtain a target three-dimensional model.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of the three-dimensional model rendering method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the three-dimensional model rendering method according to any one of claims 1 to 13 are implemented.

17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the three-dimensional model rendering method described in any one of claims 1 to 13 are implemented.

Citation Information

Cited By

  • Data-driven UV mapping generation method, electronic device and program product

    CN120976397A