A model determination method and related apparatus
By acquiring three sets of mutually perpendicular two-dimensional features of a 3D object and introducing position encoding, the initial encoder and decoder are trained to generate a more detailed 3D object, solving the problem of lost spatial information in existing models and improving generation efficiency.
Patent Information
- Application Number
- CN202510099342.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-01-20
AI Technical Summary
In existing technologies, trained models struggle to generate accurate 3D objects because spatial information is lost during the encoding process.
By acquiring three sets of mutually perpendicular two-dimensional features from a three-dimensional object sample and introducing position encoding, the initial encoder and decoder are trained to obtain the target encoder and decoder. The target encoder is then used to further train the initial model to generate a target model with rich spatial information.
It generates more detailed 3D objects, solving the problem of models failing to generate accurate 3D objects due to the loss of spatial information during the encoding process, and improving the quality and efficiency of 3D object generation.
Smart Images

Figure CN120032048B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a model determination method and related apparatus. Background Technology
[0002] Three-dimensional objects are used to simulate and represent the shape and structure of objects in three-dimensional space, such as three-dimensional buildings and figures, and have wide applications in games, animation, virtual reality, design, architecture and other fields.
[0003] In practical applications, the ability to quickly construct 3D objects has become crucial. Related technologies involve encoding the 3D data of a 3D object into features using an encoder, then reconstructing the 3D object from these features using a decoder, and training both the encoder and decoder based on this process. In the application phase, a diffusion model is used to generate features, which are then fed into the trained decoder to generate the desired 3D object. This eliminates the need for manual design of 3D objects, thus enabling more efficient construction.
[0004] However, models trained using relevant techniques (including the aforementioned diffusion model and decoder) are difficult to generate accurate 3D objects. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a model determination method and related apparatus that can determine more precise 3D objects, thus solving the technical problem in related technologies where the trained model is difficult to generate accurate 3D objects due to the loss of spatial information during the encoding process.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On one hand, embodiments of this application provide a model determination method, the method comprising:
[0008] Obtain three sets of two-dimensional features corresponding to a three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other.
[0009] Based on the three sets of two-dimensional features, the position codes corresponding to the three sets of two-dimensional features are determined by the first network layer in the initial encoder. The position codes are used to indicate the feature relationships between the two-dimensional features.
[0010] Based on the three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features, the predicted encoded features corresponding to the three-dimensional object sample are obtained by encoding through the second network layer in the initial encoder.
[0011] Based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, the initial encoder and the initial decoder are trained to obtain the target encoder and the target decoder.
[0012] Based on the three sets of two-dimensional features, the target encoding features corresponding to the three-dimensional object sample are determined by the target encoder.
[0013] Based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample, the initial model is trained to obtain the target model, and the first control data is used to describe the three-dimensional object sample.
[0014] In response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. The target features are then decoded using the target decoder to determine the 3D object corresponding to the generation request. The second control data is used to describe the 3D object generation requirements.
[0015] In another aspect, embodiments of this application provide a model determination apparatus, the apparatus comprising an acquisition unit, a determination unit, an encoding unit, and a training unit:
[0016] The acquisition unit is used to acquire three sets of two-dimensional features corresponding to the three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other.
[0017] The determining unit is used to determine the position codes corresponding to the three sets of two-dimensional features respectively through the first network layer in the initial encoder, based on the three sets of two-dimensional features, wherein the position codes are used to indicate the feature relationships between the two-dimensional features;
[0018] The encoding unit is used to encode the predicted encoded features corresponding to the three-dimensional object sample by encoding the three sets of two-dimensional features and the position encodings corresponding to the three sets of two-dimensional features through the second network layer in the initial encoder.
[0019] The training unit is used to train the model of the initial encoder and the initial decoder based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, so as to obtain the target encoder and the target decoder.
[0020] The determining unit is further configured to determine the target encoding features corresponding to the three-dimensional object sample by means of the target encoder based on the three sets of two-dimensional features;
[0021] The training unit is further configured to train the initial model based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample to obtain the target model, wherein the first control data is used to describe the three-dimensional object sample.
[0022] The determining unit is further configured to, in response to a generation request, perform denoising processing on the initial noise data based on the second control data in the generation request using the target model to determine the target features corresponding to the second control data, and decode the target features using the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.
[0023] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:
[0024] The memory is used to store computer programs and to transfer the computer programs to the processor;
[0025] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.
[0026] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program, which, when run by a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.
[0027] On the other hand, embodiments of this application provide a computer program product, including a computer program that, when run on a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.
[0028] As can be seen from the above technical solution, in the first part of model training, the positional encodings corresponding to the three sets of two-dimensional features are first determined by the first network layer in the initial encoder. Then, subsequent processing is performed based on the three sets of two-dimensional features and their respective positional encodings. The positional encoding is used to indicate the feature relationships between the two-dimensional features. Since the three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each set of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three sets of two-dimensional features are mutually perpendicular, the three-dimensional data of any point on the three-dimensional object sample is split into three sets of two-dimensional features. Therefore, the feature relationships between these three sets of two-dimensional features can reflect the absolute spatial information of that point in the three-dimensional object sample. Similarly, the feature relationships between the three sets of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically include the relative relationship between two two-dimensional features in the same feature plane within that feature plane, and the relative relationship between two two-dimensional features in different feature planes. Therefore, it can also reflect the relative spatial information of these two points in the three-dimensional object sample. Based on this, by using three sets of two-dimensional features to indicate three-dimensional data and introducing positional encoding, the spatial information carried by the three-dimensional data can be reflected based on rich feature relationships. Furthermore, training uses 3D object samples as supervision. Therefore, the initial encoder learns the rich spatial information from these samples to obtain predictive coding features with rich spatial information. The initial decoder learns how to decode these predictive coding features to determine more detailed predicted 3D objects. The resulting target encoder has better encoding capabilities, able to encode three sets of 2D features into target coding features with rich spatial information. The target decoder has better decoding capabilities to determine more detailed 3D objects. However, in practical applications, the generated 3D objects often lack 3D data, making it difficult to directly utilize the trained target encoder in the application phase. Therefore, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model, resulting in a target model that replaces the target encoder in the application phase.
[0029] In the application phase, in response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. Then, the target decoder decodes the target features to determine the 3D object corresponding to the generation request. Because the target model is trained using a target encoder, the target features determined using the target model still possess rich spatial information. Furthermore, the target decoder can better decode these spatially rich target features, thus enabling the determination of a more detailed 3D object for the generation request. Based on this, the technical problem in related technologies where the loss of spatial information during the encoding process prevents the trained model from generating detailed 3D objects is solved. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a schematic diagram illustrating an application scenario of a model determination method provided in an embodiment of this application.
[0032] Figure 2 This application provides a schematic diagram illustrating the generation of a 3D object during the model application stage.
[0033] Figure 3 A flowchart illustrating a model determination method provided in this application embodiment;
[0034] Figure 4 A schematic diagram of three feature planes provided for an embodiment of this application;
[0035] Figure 5 A schematic diagram illustrating the generation of a 3D object based on descriptive text, provided in an embodiment of this application;
[0036] Figure 6 A schematic diagram illustrating the generation of a 3D object based on a reference image, provided in an embodiment of this application;
[0037] Figure 7 A schematic diagram of the network structure of an initial encoder and an initial decoder provided in an embodiment of this application;
[0038] Figure 8 A comparative schematic diagram of two different attention mechanisms provided in the embodiments of this application;
[0039] Figure 9 A schematic diagram of a model framework provided for an embodiment of this application;
[0040] Figure 10 A schematic diagram illustrating the generation of a three-dimensional object geometry, provided as an embodiment of this application;
[0041] Figure 11 A structural diagram of a model determining device provided in an embodiment of this application;
[0042] Figure 12 A structural diagram of a terminal provided in an embodiment of this application;
[0043] Figure 13 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0044] The embodiments of this application will now be described with reference to the accompanying drawings.
[0045] To construct 3D objects, some related technologies utilize 2D images (such as a single-view image of a 3D object as training samples, or multiple-view images used together as training samples) to train a model for generating 3D objects from images. In the application stage, a 2D image can be input, and the model can generate a 3D object.
[0046] However, this method is highly dependent on image quality. Low image quality (e.g., low resolution) or inconsistencies between multiple viewpoints (e.g., the same part appearing in multiple images from different perspectives) can easily lead to distortion of the generated 3D object. Specifically, an image only provides a limited indication of what a 3D object might be, so directly constructing a 3D object from an image may not result in the desired object. For example, the generated 3D object might resemble the image from the front, but have a distorted side profile. Furthermore, inconsistencies between multiple viewpoints can lead to oddly shaped 3D objects. For instance, if a squirrel's head appears in multiple viewpoints, the resulting 3D squirrel object will have multiple heads superimposed on each other.
[0047] In other related technologies, models are trained using 3D data of 3D objects (such as point cloud data). Specifically, an encoder encodes the 3D data of the 3D object into features, and a decoder reconstructs the 3D object from these features. The encoder and decoder are then trained based on this process. In the application phase, a diffusion model is used to generate features, which are then fed into the trained decoder to generate the desired 3D object.
[0048] While this method utilizes 3D data to train the model, thus addressing some of the problems in the first related technique, the second method suffers from a loss of spatial information about the 3D object due to its encoding method. This results in the trained model failing to generate detailed 3D objects. For example, the generated 3D object may exhibit misalignment between multiple parts or lack certain details.
[0049] To address this, this application provides a model determination method and related apparatus. First, using three sets of two-dimensional features from a 3D object sample and introducing position encoding, an initial encoder and an initial decoder are trained. The position encoding indicates the feature relationships between the two-dimensional features, and the three sets of two-dimensional features indicate the 3D data. Compared to the 3D data itself, because each set of two-dimensional features indicates two dimensions of the 3D data, and the three feature planes corresponding to the three sets of two-dimensional features are mutually perpendicular, the 3D data of any point on the 3D object sample is split into three sets of two-dimensional features. Therefore, the feature relationships between these three sets of two-dimensional features can reflect the absolute spatial information of that point in the 3D object sample. Similarly, the feature relationships between the three sets of two-dimensional features corresponding to any two points on the 3D object sample specifically include the relative relationships between two two-dimensional features within the same feature plane and the relative relationships between two two-dimensional features within different feature planes, thus also reflecting the relative spatial information of these two points in the 3D object sample.
[0050] Based on this, by using three sets of two-dimensional features to indicate three-dimensional data and introducing positional encoding, the spatial information carried by the three-dimensional data can be reflected based on the rich relationships between features. In this way, during model training, the initial encoder can fully learn the spatial information carried by the three-dimensional data to obtain predictive encoded features with rich spatial information. Based on these predictive encoded features, the initial decoder can then learn how to decode them to determine more precise predicted three-dimensional objects. Correspondingly, the resulting target encoder has better encoding capabilities, able to encode the three sets of two-dimensional features into target encoded features with rich spatial information, and the target decoder has better decoding capabilities to determine more precise three-dimensional objects.
[0051] Next, the trained target encoder is used to train the initial model. Specifically, the target encoder encodes three sets of two-dimensional features into target encoded features with rich spatial information. Based on this, the initial model is trained, enabling it to learn how to perform denoising to determine features with equally rich spatial information, thus obtaining a target model that can replace the target encoder in the application stage.
[0052] Correspondingly, in the application stage, the target model can be used to determine the target model features that meet the generation request and have rich spatial information. The target decoder can then better decode the target features with rich spatial information, thereby determining a more detailed 3D object.
[0053] Compared to the first related technique mentioned above, this method relies on 3D data rather than 2D images during the model training phase. Since 3D data can accurately reflect the overall situation of a 3D object (such as front view, side view, back view, etc.), it avoids the technical problem of failing to generate accurate 3D objects by relying solely on 2D images. Furthermore, compared to the second related technique mentioned above, spatial information is no longer lost, thus solving the technical problem of difficulty in generating detailed 3D objects due to the loss of spatial information.
[0054] The model determination method provided in this application can be implemented using a computer device, which can be a terminal or a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include, but are not limited to, smartphones, computers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. Terminals and servers can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this connection.
[0055] The embodiments of this application can be specifically applied to various scenarios that require the generation of three-dimensional objects. For example, in applications such as games and animation, it can generate three-dimensional virtual objects. Another example is in virtual reality applications, it can generate three-dimensional virtual models in virtual space. Yet another example is in fields such as design and architecture, it can generate three-dimensional objects (such as vases, furniture, etc.) and three-dimensional buildings.
[0056] In game applications, generated 3D virtual objects are typically referred to as 3D game assets. These are the basic building blocks of game applications and can include 3D virtual characters, 3D game scenes, and 3D items within those scenes. Looking at the details of a 3D object, it usually includes geometric shapes (which can also be represented by geometric patches) and texture maps (such as materials and colors).
[0057] It should be noted that in the specific embodiments of this application, the process of determining the model and constructing a three-dimensional object using the model may involve user information and other related data. When the above embodiments of this application are applied to specific products or technologies, separate consent or permission from the user is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0058] To better understand, Figure 1 This illustration shows an application scenario of the model determination method provided in the embodiments of this application. Figure 1 In the scenario shown, taking server 100 as an example of the aforementioned computer equipment, it can include two parts: ① the model training phase and ② the model application phase. Specifically:
[0059] Server 100 can acquire three sets of two-dimensional features corresponding to a 3D object sample and use these three sets of two-dimensional features for model training. Specifically, the model training phase (①) can include two parts:
[0060] In the first part of model training, server 100 can input three sets of two-dimensional features into the first network layer of the initial encoder to determine the positional encodings corresponding to the three sets of two-dimensional features. Then, based on the three sets of two-dimensional features and their respective positional encodings, it inputs them into the second network layer of the initial encoder to determine the predicted encoded features corresponding to the three-dimensional object sample. Next, server 100 can input the predicted encoded features into the initial decoder to determine the predicted three-dimensional object.
[0061] Among them, three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features indicates two dimensions of the three-dimensional data, and the three feature planes corresponding to the three sets of two-dimensional features are mutually perpendicular. Position encoding can be used to indicate the feature relationships between the two-dimensional features. It can be understood that three-dimensional data can reflect the spatial information of a three-dimensional object sample. After using three sets of two-dimensional features to indicate the three-dimensional data, the spatial information carried by the three-dimensional data can be reflected by utilizing the feature relationships between different feature planes and the feature relationships within the same feature plane. Based on this, by using three sets of two-dimensional features and introducing position encoding, the spatial information carried by the three-dimensional data can be reflected based on rich relationships between features.
[0062] Thus, using 3D object samples as supervision, server 100 can train the initial encoder and initial decoder based on the differences between the predicted 3D object and the 3D object samples. During training, the initial encoder learns the rich spatial information in the 3D object samples to obtain predictive encoded features with rich spatial information. The initial decoder learns how to decode the predictive encoded features to determine a more detailed and accurate predicted 3D object. Ultimately, this results in a target encoder with better encoding capabilities and a target decoder with better decoding capabilities.
[0063] In the second part of model training, server 100 can utilize the target encoder to determine the target encoded features based on three sets of two-dimensional features. These target encoded features are those with rich spatial information, enabling the reconstruction of detailed 3D object samples. However, in practical applications, the desired 3D objects often lack 3D data, making it difficult to directly utilize the trained target encoder in the application phase. Therefore, in the second part of model training, the initial model is trained using the noisy encoded features corresponding to the target encoded features and the first control data to obtain the target model, which replaces the target encoder in the application phase. The first control data describes the 3D object sample. That is, it enables the initial model to learn how to denoise the noisy encoded features based on the description of the first control data, ultimately determining features with rich spatial information that are close to the target encoded features.
[0064] Accordingly, in the model application stage (②), server 100 can input the second control data and initial noise data into the target model. The second control data is used to describe the requirements for generating the 3D object. The target model then performs denoising processing on the initial noise data based on the second control data to determine target features with rich spatial information. Furthermore, server 100 can input the target features into the target decoder and decode the target features to determine the corresponding 3D object.
[0065] Because positional encoding is introduced during model training, the model can learn the spatial information carried in the 3D data. The resulting combination of "target model + target decoder" can generate detailed 3D objects that meet the generation requirements described by the second control data. This solves the technical problem in related technologies where spatial information is lost during the encoding process, resulting in models that cannot generate detailed 3D objects and therefore cannot generate 3D objects that meet the requirements.
[0066] To better understand, embodiments of this application also provide, as follows: Figure 2 The diagram illustrates the generation of 3D objects during a model application phase. A connection (such as a network connection) exists between the terminal 201 (e.g., the aforementioned computer) and the server 202. Specifically, in the actual application phase:
[0067] When a user wants to generate a 3D object, they can input second control data A through terminal 201 to describe the 3D object generation requirements. For example, in this embodiment, the second control data A is "a girl with long black hair, standing with her arms outstretched, wearing a long polka-dot puffy dress with a bow at the neckline and wavy hem".
[0068] Next, terminal 201 can send a generation request carrying second control data A to server 202. Correspondingly, server 202, in response to the generation request, can begin implementing the aforementioned... Figure 1 The processing steps in the model application stage of the example are as follows: the target model performs denoising processing on the initial noise data according to the second control data A to determine the target feature A, and the target decoder decodes the target feature A to determine the three-dimensional object A.
[0069] Furthermore, server 202 can return the generated 3D object A to terminal 201 so that the user can view it on terminal 201. For example, the final generated 3D object A can be displayed on the computer screen of terminal 201.
[0070] As can be seen, when a user wants to generate a 3D object, they only need to input control data to describe the specific generation requirements. The trained target model and target decoder can then be used to quickly generate a detailed 3D object that meets the generation requirements. This not only improves the quality of the 3D object, but also, because it is an efficient generation process, it helps to improve the generation efficiency of 3D objects.
[0071] Understandably, in practical applications, users can also perform various operations such as dragging, rotating, and scaling on the computer page of terminal 201 to flexibly view different faces, local details, and so on of the 3D object A.
[0072] Figure 3 A flowchart of a model determination method provided in this application embodiment, using a server as an example of the aforementioned computer device, is used for illustration. The method includes S301-S307:
[0073] S301: Obtain the three sets of two-dimensional features corresponding to the three-dimensional object sample.
[0074] Among them, three sets of two-dimensional features can be used to indicate the three-dimensional data of a three-dimensional object sample, each set of two-dimensional features can be used to indicate two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other.
[0075] It should be noted that this application does not impose any limitations on how the three sets of two-dimensional features are determined. In practical applications, three mutually perpendicular feature planes can be constructed, and the three-dimensional data can be mapped and projected onto these three feature planes.
[0076] For example, three mutually perpendicular feature planes can be seen in... Figure 4As shown. Typically, the 3D data of a 3D object sample refers to the values of each point on the 3D object sample in a 3D coordinate system, such as (x, y, z). Correspondingly, the three feature planes can be denoted as the (x, y) feature plane, the (x, z) feature plane, and the (y, z) feature plane, respectively. In practical applications, the three sets of 2D features can also be called triplane features. Furthermore, the 3D data of a 3D object sample can also be point cloud data, etc.
[0077] It is understandable that 3D data can represent the spatial information of a 3D object sample. After using three sets of 2D features to indicate the 3D data, the 3D data of any point on the 3D object sample is split into three sets of 2D features. For example, point A(x1, y1, z1) will be split into three sets of 2D features belonging to three feature planes: (x1, y1), (x1, z1), and (y1, z1). Therefore, the absolute spatial information of point A in the 3D object sample can be represented by the feature relationship between these three sets of 2D features.
[0078] Similarly, any two points on a 3D object sample, such as point A(x1, y1, z1) and point B(x2, y2, z2), will be split into two-dimensional features (x1, y1) and (x2, y2) belonging to the (x, y) feature plane, two-dimensional features (x1, z1) and (x2, z2) belonging to the (x, z) feature plane, and two-dimensional features (y1, z1) and (y2, z2) belonging to the (y, z) feature plane.
[0079] That is, the feature relationship between the three sets of two-dimensional features corresponding to any two points on a 3D object sample. Specifically, this includes the relative relationship between two two-dimensional features within the same feature plane and the relative relationship between two two-dimensional features in different feature planes. In this way, not only can the absolute spatial information of a point be reflected through the feature relationship between its three sets of two-dimensional features, but also the relative spatial information between two points can be reflected through the feature relationship between two two-dimensional features within the same feature plane.
[0080] Based on this, by using three sets of two-dimensional features to represent three-dimensional data, we can utilize richer feature relationships between two-dimensional features (including feature relationships between two-dimensional features in the same feature plane and feature relationships between two-dimensional features in different feature planes) to fully reflect the spatial information carried in the three-dimensional data, so that the model can learn spatial information better during model training.
[0081] It is understandable that different 3D object samples differ in their geometry, the number of features they contain, etc., thus resulting in different 3D data and correspondingly different dimensions of the three feature planes. It should also be noted that this application does not impose any limitations on the dimensions of the feature planes. For ease of understanding, the following examples are provided in this application's embodiments:
[0082] Typically, various 3D object samples are used for model training to enable the model to learn fully and avoid underfitting. To facilitate model learning (specifically, to facilitate the learning of the initial encoder, etc.), one possible implementation is to use three feature planes of uniform size to represent the 3D data of each 3D object sample. Similarly, for each feature, it can also be represented using a vector with the same feature dimension, thus achieving uniformity.
[0083] In practice, the plane size corresponding to each of the three feature planes is N, and each feature in the three sets of two-dimensional features ∈ R. N*d R represents the set of real numbers, and d represents the feature dimension of each feature, where d is a positive integer. N can be represented as u*v, where u represents the length of the feature plane, and v represents the width of the feature plane; both u and v are positive integers.
[0084] Based on this, for any 3D object sample with arbitrary geometric shape and an arbitrary number of features, its corresponding 3D data can be transformed into three feature planes of uniform size, and the number of feature dimensions for each feature is also consistent. This data is then fed into the initial encoder, facilitating learning from a wide variety of 3D object samples.
[0085] S302: Based on the three sets of two-dimensional features, determine the position codes corresponding to the three sets of two-dimensional features through the first network layer in the initial encoder.
[0086] S303: Based on three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features respectively, the prediction coding features corresponding to the three-dimensional object samples are obtained by encoding through the second network layer in the initial encoder.
[0087] Positional encoding can be used to indicate the feature relationships between two-dimensional features. As explained above, the feature relationships between two-dimensional features can reflect spatial information; therefore, positional encoding can also reflect spatial information. Based on this, the initial encoder can learn the spatial information of the three-dimensional object sample by incorporating the indications of positional encoding during the encoding process, and then determine the corresponding predictive encoded features.
[0088] In practical applications, encoding is a technique that processes arbitrary data into a certain feature representation (such as a vector-like feature representation) so that the model can understand it. Encoding can also compress and reduce the dimensionality of 3D data, thereby reducing redundant information. Correspondingly, decoding is a technique that uses the feature representation obtained from encoding to reconstruct the data. Accordingly, the predicted encoded features can refer to the feature representation corresponding to the three sets of 2D features of the 3D object sample determined by the initial encoder. These predicted encoded features are the features that the initial encoder believes can indicate the 3D object sample, and can be used for subsequent reconstruction.
[0089] From the network structure of the initial encoder, it includes a first network layer and a second network layer. The first network layer is responsible for determining positional encoding, while the second network layer is responsible for encoding. Encoders used in related technologies directly encode three-dimensional data into one-dimensional features. In contrast, the encoder used in this application adds a first network layer to introduce positional encoding, thereby solving the technical problem of losing spatial information during the encoding process in related technologies.
[0090] Furthermore, because the positional encoding is determined by the first network layer in the initial encoder, it becomes learnable through model parameter adjustments during training. In other words, the positional encoding determined in this application is a relative positional encoding. Thus, during model training, for different 3D object samples (e.g., different geometric shapes, details), i.e., for different sets of two-dimensional features, the initial encoder can continuously learn which three sets of two-dimensional features to target and which positional encoding to determine, ensuring that the determined predicted encoding features can reconstruct the 3D object sample. This allows it to learn the ability to determine appropriate and accurate positional encodings for different sets of two-dimensional features, improving the initial encoder's spatial information perception capability.
[0091] Furthermore, compared to related technologies that assign a unique code to each location (i.e., absolute positional coding), this method eliminates the need for manual annotation of the positional codes corresponding to the 3D and 2D features of each 3D object sample. This not only saves on manual annotation costs but also avoids introducing human experience into the model training process, which could lead to model learning errors. For example, incorrect positional codes annotated based on human experience can skew the model's learning direction.
[0092] It should be noted that this application does not impose any limitations on the network structure of the initial encoder, apart from including the aforementioned first and second network layers. For example, it does not limit the inclusion of other network layers (such as residual networks), nor the number of the first and second network layers. These can be flexibly set according to actual needs.
[0093] S304: Based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, the initial encoder and initial decoder are trained to obtain the target encoder and target decoder.
[0094] The initial decoder is used to decode the predictive coding features determined by the initial decoder in order to reconstruct the object. Correspondingly, the predicted 3D object refers to the 3D object reconstructed by the initial decoder based on the predictive coding features, that is, the 3D object that the initial decoder believes is indicated by the predictive coding features.
[0095] Correspondingly, the difference between the predicted 3D object and the 3D object sample reflects the bias in the initial encoder's determination of predictive coding features and the bias in the initial decoder's reconstruction of the 3D object based on the predictive coding features. Therefore, training the model based on this difference enables the initial encoder to learn how to determine more accurate predictive coding features, specifically how to determine more accurate positional codes suitable for 3D object samples through the first network layer, how to encode through the second network layer, etc., and enables the initial decoder to learn how to reconstruct the 3D object sample based on the predictive coding features.
[0096] In this way, by using 3D object samples as supervision and training the model, the final target encoder has better encoding capabilities and can encode three sets of 2D features into target encoding features with rich spatial information. The target decoder has better decoding capabilities and can reconstruct and restore the 3D object sample based on the target encoding features, that is, it has the ability to determine the 3D object with precise details.
[0097] It should be noted that the initial decoder has a network structure corresponding to the initial encoder. Therefore, the network structure of the initial decoder can be set with reference to the network structure of the initial encoder. This will not be elaborated here.
[0098] Because the initial decoder has a network structure corresponding to the initial encoder, and the initial encoder of this application introduces positional encoding during the encoding process, spatial information is incorporated into the predicted encoded features. Consequently, when the initial decoder decodes the predicted encoded features, this information can guide the initial decoder to perform better decoding. Specifically, it can guide the initial decoder to determine which features should represent which points on the 3D object, thereby facilitating the identification of 3D objects with accurate details (such as accurate relative relationships between different parts of the 3D object, accurate local details, etc.).
[0099] S305: Based on three sets of two-dimensional features, the target encoding features corresponding to the three-dimensional object sample are determined by the target encoder.
[0100] S306: Based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object samples, train the initial model to obtain the target model.
[0101] Since the 3D objects to be generated in practical applications often lack 3D data, it is difficult to directly utilize a pre-trained target encoder in the application phase. Therefore, after training the target encoder and target decoder, a second part of model training is required. Specifically, this involves training an initial model using the target encoder to obtain a target model that can replace the target encoder in the application phase. Specifically:
[0102] First, the server can determine the corresponding target encoding features based on the three-dimensional and two-dimensional features of the three-dimensional object sample through the target encoder. The target encoding features refer to the encoding features that have rich spatial information and can reconstruct and restore the three-dimensional object sample after being input into the target decoder.
[0103] Next, the server can train the initial model based on the noisy encoded features corresponding to the target encoded features and the first control data to obtain the target model. The noisy encoded features can be obtained by adding noise to the target encoded features; that is, by introducing additional information to the target encoded features to create interference, thereby altering the target encoded features to some extent. The first control data can be used to describe the 3D object sample; in this context, it can also be understood as describing the 3D object that needs to be reconstructed during the model training phase.
[0104] During the training of the initial model based on noisy coded features and first control data, the initial model learns how to denoise the noisy coded features according to the first control data, so that the denoised coded features are close to the target coded features. That is, theoretically, after feeding the denoised coded features into the target decoder, the 3D object sample can be reconstructed. Based on this training, the target model learns how to denoise the control data used to describe the requirements, ensuring that the final determined features are also features with rich spatial information, similar to the features determined by the target encoder, thus replacing the target encoder in the application stage.
[0105] It should be noted that this application does not impose any restrictions on the settings of the initial model. For example, the initial model can be a generative model such as a diffusion model (DM). Because diffusion models are more random in the denoising process, they are more imaginative and can support guidance from multiple types of control data (such as at least one of reference images, descriptive text, etc.), making them more conducive to generating 3D objects that meet the requirements.
[0106] Furthermore, this application does not impose any limitations on the settings of control data. In practical applications, control data is used to express the requirements of 3D objects. During the model training phase, control data expresses the requirements of the 3D object samples that need to be reconstructed and restored. During the model application phase, control data expresses the requirements of the 3D objects that need to be generated. Typically, control data is not the aforementioned 3D data, but rather simpler data that can express the requirements. For example, control data may include at least one of reference images and descriptive text.
[0107] Corresponding to the model training phase, the first control data is used to describe the 3D object sample and may include at least one of a first reference image corresponding to the 3D object sample and a first descriptive text corresponding to the 3D object sample. The first reference image may be a 2D image corresponding to the 3D object sample (such as a front view of the 3D object sample) or a 2D image similar to the 3D object sample (such as similar style, similar geometric shape, etc.). The first descriptive text can be used to describe the specific details of the 3D object sample, such as what the 3D object sample is and what details it possesses.
[0108] In practical applications, if the initial control data includes multiple data types, it can be aligned to the same feature space before being fed into the initial model, such as by using a large diffusion model. Alignment helps to extract realistic descriptions from multiple data types, thus better guiding the initial model's understanding and subsequent denoising processes.
[0109] S307: In response to the generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data, and the target features are decoded using the target decoder to determine the 3D object corresponding to the generation request.
[0110] Based on the solutions provided in S301-S306 above, the final model used to generate 3D objects can be determined through two-part model training, specifically including the target model and the target decoder. The introduction of positional encoding during model training solves the technical problem of lost spatial information in related technologies. Accordingly, in the application phase, when it is necessary to generate 3D objects, the target model and target decoder can be used to quickly generate the required 3D objects. Specifically:
[0111] When a 3D object needs to be generated, the user can initiate a generation request, specifically by inputting second control data describing the requirements for generating the 3D object. Correspondingly, the server can respond to this generation request by using the aforementioned target model to denoise the initial noisy data based on the second control data in the generation request, thereby determining the target features corresponding to the second control data. Then, the server can decode the target features using a target decoder to determine the 3D object corresponding to the generation request.
[0112] The second control data can be used to describe the requirements for generating 3D objects. Specifically, it can describe the desired 3D object (such as what the desired 3D object looks like and what characteristics it has). The target features determined based on the second control data can also be used to indicate the desired 3D object. Because the target model is obtained through model training using a target encoder, the target features determined using the target model still possess rich spatial information. Furthermore, the target decoder can better decode target features with rich spatial information, thus enabling it to determine a more detailed 3D object that meets the generation requirements described by the second control data.
[0113] It is understandable that S307 can be executed only when it is necessary to generate a 3D object, that is, S307 can be omitted when it is not necessary to generate a 3D object.
[0114] It should be noted that this application does not impose any limitations on the settings of the second control data. Similar to the aforementioned description of the first control data, the first control data corresponding to the model application stage may also include at least one of a second reference image used to describe the requirements for generating 3D objects and a second descriptive text used to describe the requirements for generating 3D objects. The second reference image may be used to describe the style, geometry, etc., of the 3D object to be generated, while the second descriptive text may be used to describe what the 3D object to be generated is and what its characteristics are. Similar to the first control data, if the second control data includes multiple types of data, they are aligned to the same feature space before being fed into the target model, such as using the aforementioned diffusion model for alignment. Based on this, it is beneficial to extract the true generation requirements from multiple types of data, thereby better guiding the target model to understand the generation requirements and then perform denoising processing, etc.
[0115] As an example, the second control data could be a second descriptive text. Specifically, it could be, for example, a second descriptive text. Figure 2The descriptive text in the example, "a girl with long black hair, standing with her arms outstretched, wearing a long polka-dot puffy dress with a bow at the neckline and wavy hem," is used to guide the target model in denoising the initial noisy data based on this second control data, in order to determine the target features that can be used to generate a 3D object that conforms to the description of the second control data. Figure 2 For example, the final generated 3D object A conforms to the generation requirements described by the second control data A, and all details (such as a bow at the neckline and wavy lines at the hem) are accurate. Furthermore, it can also be... Figure 5 The descriptive text in the example, "long-haired boy, standing in a T-pose, wearing a striped belt, and high-top sneakers," corresponds to the generated 3D object B, such as... Figure 5 As shown, it conforms to the generation requirements described in the second control data B and has precise details (such as a striped belt).
[0116] As yet another example, the second control data could be a second reference image. Specifically, for example, in... Figure 6 In the example, the second control data C can be a two-dimensional reference image. The corresponding generated three-dimensional object C is as follows: Figure 6 As shown, it meets the requirements and has precise details (such as standing posture, tie, plaid vest, etc.).
[0117] Of course, descriptive text and reference images can also be used together as control data to more precisely describe the requirements for generating 3D objects. This also helps guide the target model in denoising, thereby identifying more accurate target objects and improving the accuracy of the generated 3D objects. Further details on this will not be elaborated upon here.
[0118] It should be noted that, besides not limiting the type of control data, there are also no limitations on the quantity of control data. For example, it can include multiple descriptive text segments or multiple reference images. In practical applications, when generating multiple 3D objects for the same game project, a style reference image X can be set. When generating each 3D object, in addition to the reference image describing that 3D object, this style reference image X can also be used as input. Based on this, image X collectively guides the generation process of these multiple 3D objects, specifically guiding their style, thus facilitating the generation of multiple 3D objects with a unified style.
[0119] It should also be noted that in practical applications, users can also initiate a generation request by inputting secondary control data through the terminal. Furthermore, users can view the generated 3D objects (such as the aforementioned 3D object A, 3D object B, 3D object C, etc.) through the terminal, and perform various operations such as dragging, rotating, and scaling to flexibly view different faces and local details of the 3D objects. In addition, the generated 3D objects can be loaded into other applications (such as design software, texture mapping software, etc.) and combined with other applications to fine-tune the generated 3D objects (such as adjusting local details, applying texture mapping, etc.) to obtain 3D objects with richer details and better meeting the requirements.
[0120] As can be seen from the above technical solution, in the first part of model training, the positional encodings corresponding to the three sets of two-dimensional features are first determined by the first network layer in the initial encoder. Then, subsequent processing is performed based on the three sets of two-dimensional features and their respective positional encodings. The positional encoding is used to indicate the feature relationships between the two-dimensional features. Since the three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each set of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three sets of two-dimensional features are mutually perpendicular, the three-dimensional data of any point on the three-dimensional object sample is split into three sets of two-dimensional features. Therefore, the feature relationships between these three sets of two-dimensional features can reflect the absolute spatial information of that point in the three-dimensional object sample. Similarly, the feature relationships between the three sets of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically include the relative relationship between two two-dimensional features in the same feature plane within that feature plane, and the relative relationship between two two-dimensional features in different feature planes. Therefore, it can also reflect the relative spatial information of these two points in the three-dimensional object sample. Based on this, by using three sets of two-dimensional features to indicate three-dimensional data and introducing positional encoding, the spatial information carried by the three-dimensional data can be reflected based on rich feature relationships. Furthermore, training uses 3D object samples as supervision. Therefore, the initial encoder learns the rich spatial information from these samples to obtain predictive coding features with rich spatial information. The initial decoder learns how to decode these predictive coding features to determine more detailed predicted 3D objects. The resulting target encoder has better encoding capabilities, able to encode three sets of 2D features into target coding features with rich spatial information. The target decoder has better decoding capabilities to determine more detailed 3D objects. However, in practical applications, the generated 3D objects often lack 3D data, making it difficult to directly utilize the trained target encoder in the application phase. Therefore, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model, resulting in a target model that replaces the target encoder in the application phase.
[0121] In the application phase, in response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. Then, the target decoder decodes the target features to determine the 3D object corresponding to the generation request. Because the target model is trained using a target encoder, the target features determined using the target model still possess rich spatial information. Furthermore, the target decoder can better decode these spatially rich target features, thus enabling the determination of a more detailed 3D object for the generation request. Based on this, the technical problem in related technologies where the loss of spatial information during the encoding process prevents the trained model from generating detailed 3D objects is solved.
[0122] The above embodiments illustrate the model determination method provided in this application. It should also be noted that this application does not impose any limitations on the network structure of the model, how to determine positional encoding, how to perform encoding, or how to train the model. For a more comprehensive understanding, this application will be described in detail through the following embodiments.
[0123] (i) Regarding the network structure of the initial encoder, the embodiments of this application provide the following examples:
[0124] Regarding the initial encoder, from a network structure perspective, it includes the aforementioned first and second network layers. Beyond this, this application makes no further limitations. For example, it does not limit the inclusion of other network layers, nor the number of the first and second network layers. These can be flexibly configured according to actual needs.
[0125] To enhance the encoder's coding ability and determine more accurate coding features, one possible implementation is to add multiple first network layers and multiple second network layers to the initial encoder. This allows the encoder to fully learn the information carried in the input data (i.e., the three sets of two-dimensional features) during the model training phase, thereby improving its ability to determine more accurate coding features.
[0126] In specific implementation, the aforementioned initial encoder may include M groups of networks connected in sequence. In the M groups of networks, each of the first M-1 groups of networks may include a first network layer, a second network layer, and a third network layer connected in sequence, and the Mth group of networks may include a first network layer and a second network layer connected in sequence. Here, M is a positive integer greater than 1.
[0127] It is understandable that the encoding methods will differ depending on the initial encoder network structure. In this embodiment, within a set of networks, position encoding is determined first, then encoded, and then fed into the next set of networks to continue determining position encoding and encoding, until the last set of networks completes both position encoding and encoding. That is, processing is performed sequentially through M sets of networks. In practical applications, since the encoded features are mostly one-dimensional features, to facilitate the determination of position encoding through the next set of networks, each of the first M-1 sets of networks includes the aforementioned third network layer. This third layer transforms the encoded features obtained from the second network layer before feeding them into the next set of networks for processing. For ease of understanding, the following detailed explanation is provided:
[0128] That is, in a specific implementation of S302, the server can determine the positional encoding corresponding to each of the three sets of two-dimensional features through the first network layer of the first network in the M-group network. And, in a specific implementation of S303, the server can encode the three sets of two-dimensional features and their corresponding positional encodings through the second network layer of the first network to obtain the predicted encoding features corresponding to the first network. Then, the server can transform the predicted encoding features corresponding to the first network using the third network layer of the first network to obtain the three sets of two-dimensional encoding features corresponding to the first network.
[0129] Based on this, the processing of the first group of networks is completed. Next, the data is sent to the next group of networks for processing. In specific implementation:
[0130] The server can determine the positional codes corresponding to the three sets of two-dimensional coding features of the (i-1)th network in the M network groups by using the first network layer in the i-th network group. Here, i is an integer greater than 1 and less than or equal to M. Furthermore, the server can encode the predicted coding features corresponding to the i-th network group by using the three sets of two-dimensional coding features and their corresponding positional codes through the second network layer in the i-th network group. And, if the i-th network group includes a third network layer, the server transforms the predicted coding features corresponding to the i-th network group using the third network layer to obtain the three sets of two-dimensional coding features corresponding to the i-th network group.
[0131] Based on this, processing the three sets of two-dimensional features of the 3D object sample sequentially through M sets of networks helps the initial encoder better capture the information carried in the three sets of two-dimensional features for deep learning. Furthermore, the positional encoding is redefined in each set of networks, thus enabling better capture of the spatial information carried in the three sets of two-dimensional features during deep learning. Ultimately, this leads to the formation of more accurate encoded features. Accordingly, the predicted encoded features corresponding to the 3D object sample can be determined based on the predicted encoded features corresponding to the Mth set of networks.
[0132] It should be noted that this application does not impose any limitations on how to determine the final predictive coding features of the 3D object sample based on the predictive coding features corresponding to the Mth group of networks. It is understood that this is related to the network structure, such as whether the Mth group of networks includes the aforementioned third network layer, or whether the initial encoder includes other networks (such as an output layer) after the Mth group of networks. For ease of understanding, the embodiments of this application provide the following examples:
[0133] In some embodiments, the Mth network group may not include the aforementioned third network layer; correspondingly, the predictive coding features corresponding to the 3D object sample are the predictive coding features corresponding to the Mth network group. This simplifies the model structure.
[0134] In some embodiments, the structure of the Mth network group is the same as that of the previous M-1 networks, i.e., it also includes a third network layer. Accordingly, the predictive coding features corresponding to the Mth network group can be transformed based on the third network layer in the Mth network group to obtain three sets of two-dimensional predictive coding features corresponding to the Mth network group. At this time, the predictive coding features corresponding to the three-dimensional object sample are the three sets of two-dimensional predictive coding features corresponding to the Mth network group. Based on this, the final predicted coding features are three sets of two-dimensional features that can express three-dimensional spatial information. Therefore, they can explicitly inform the initial decoder of the spatial information in the predictive coding features, thereby guiding the initial decoder to perform better decoding. In particular, compared with the encoding method in related technologies that directly encodes three-dimensional data into one-dimensional features, this application can better preserve the three-dimensional spatial information, thereby enabling better model training and improving model capabilities. That is to say, in this application, the expression form of three sets of two-dimensional features is always maintained during the encoding and decoding processes.
[0135] Furthermore, other network layers can be added as needed. For example, setting a ResNet before the first network group helps alleviate problems such as gradient vanishing during model training, thereby ensuring the stability of model training and accelerating training. In practical applications, the ResNet can be composed of multiple stacked Res Blocks.
[0136] It should also be noted that the above embodiments are illustrated using an encoder as an example. Since decoding is the reverse process of encoding, the decoder has a network structure corresponding to the encoder. Therefore, after determining the network structure of the initial encoder, the network structure of the initial decoder can be adapted, and the steps regarding how to process data through the initial decoder can be referred to the descriptions above, which will not be repeated here.
[0137] In practical applications, a generative model such as a variational autoencoder (VAE) can be used to construct the aforementioned initial encoder and decoder. Specifically, it consists of an encoder and a decoder, and the model structure is U-shaped.
[0138] For ease of understanding, embodiments of this application also provide, as follows: Figure 7 The diagram shown illustrates a network structure for an initial encoder and an initial decoder. Figure 7 In the example, taking M=3 as an example, it can be seen that the initial encoder encodes the three sets of two-dimensional features corresponding to the 3D object sample to obtain the corresponding predicted encoded features. Then, the initial decoder processes the predicted encoded features to obtain the corresponding three sets of decoded two-dimensional features. These three sets of decoded two-dimensional features are used for reconstruction to obtain the aforementioned predicted 3D object.
[0139] It should also be noted that this application does not impose any limitations on the specific settings of each network layer. Typically, each network layer can be set based on the transformer, such as setting the transformer as the aforementioned second network layer. Correspondingly, when the initial encoder includes multiple sets of networks, it can also be considered that the initial encoder consists of multiple stacked transformer blocks.
[0140] (II) Regarding how to perform encoding, the embodiments of this application provide the following examples:
[0141] As can be seen from the above embodiments, how to perform encoding is related to the network structure of the encoder.
[0142] In addition, to better understand the encoding process, the embodiments of this application also provide the following examples:
[0143] In practical applications, in order to facilitate the initial encoder (specifically the second network layer in the initial encoder) to better capture the feature relationships between features based on positional encoding, one possible implementation is to use attention for encoding. In the attention mechanism, any feature can be calculated separately from other features, which is beneficial for better learning the relationships between features.
[0144] Accordingly, in the specific implementation of S303, for each of the three sets of two-dimensional features, the server can concatenate the two-dimensional features and their corresponding positional codes to obtain the concatenated features corresponding to the three sets of two-dimensional features respectively. Then, the server can encode the concatenated features corresponding to the three sets of two-dimensional features through the second network layer in the initial encoder using an attention mechanism to obtain the predicted encoded features corresponding to the three-dimensional object sample. The attention mechanism-based encoding process is used to indicate that the concatenated features belonging to the first feature plane and the concatenated features belonging to the second feature plane are dot-products, and that the first and second feature planes are any of the three feature planes mentioned above; that is, they can be the same feature plane or different feature planes.
[0145] Based on this, by concatenating the positional encoding with the original features, the two are fed into the second network layer as a whole and encoded based on the attention mechanism, so that the dot product can be calculated for any two features (including the dot product between two features in the same feature plane and the dot product between two features in different feature planes), thereby better capturing the relationship between features.
[0146] Furthermore, this application utilizes three sets of two-dimensional features to represent three-dimensional data; in short, it uses three feature planes to represent three-dimensional data. Therefore, this approach can simultaneously consider the relative positional relationships within the same feature plane, as well as the relative positional relationships between different feature planes, thereby helping the initial encoder to better learn the spatial information carried by the three-dimensional data and obtain more accurate predictive coding features.
[0147] In this application, the input to the initial encoder consists of three sets of two-dimensional features, which can be simply described as three feature planes. The relative relationships between features within the same feature plane are stronger than those between features in different feature planes, and the three feature planes are mutually perpendicular. Therefore, to facilitate the initial encoder's learning process by reducing the relative relationships between different feature planes when capturing spatial information based on position encoding, one possible implementation is that the position encodings corresponding to the three sets of two-dimensional features can satisfy an orthogonal relationship.
[0148] Correspondingly, in the process of encoding based on the attention mechanism, that is, in the process of performing dot product, the orthogonality relation can be used to indicate that if the first feature plane and the second feature plane are different feature planes, then the dot product result between the positional encoding in the concatenated features belonging to the first feature plane and the positional encoding in the concatenated features belonging to the second feature plane is zero.
[0149] Based on this, by setting the positional encodings of the three sets of two-dimensional features to satisfy an orthogonal relationship, the computation between features in different feature planes can be weakened during the attention-based encoding process. This helps the initial encoder to capture more accurate spatial information better and faster, thereby accelerating the model training of the initial encoder.
[0150] Typically, to further accelerate model training, while weakening the correlation between different feature planes, the correlation within the same feature plane can be strengthened. That is, during the dot product process, if the first and second feature planes are the same feature plane, the dot product between the positional codes in the concatenated features belonging to the first feature plane and the positional codes in the concatenated features belonging to the second feature plane is greater than zero. Thus, in model training, the goals of strengthening the correlation within the same feature plane and weakening the correlation between different feature planes are simultaneously achieved, thereby facilitating further acceleration of model training.
[0151] To better understand the encoding process, the embodiments of this application also provide the following examples:
[0152] On the one hand, in practical implementation, the aforementioned splicing characteristics can be represented by the following formula:
[0153] L ′ 1 = L1 + P1, L ′ 2 = L² + P², L ′ 3 = L3 + P3
[0154] In the above formula, L1 represents the first set of two-dimensional features, P1 represents the positional encoding corresponding to the first set of two-dimensional features, and L ′ 1 is used to represent the corresponding concatenated feature; similarly, L2 is used to represent the first set of two-dimensional features, P2 is used to represent the positional encoding corresponding to the first set of two-dimensional features, and L... ′ 2 is used to represent the corresponding concatenated features; and L3 is used to represent the first set of two-dimensional features, P3 is used to represent the positional encoding corresponding to the first set of two-dimensional features, L ′ 3 is used to represent the corresponding splicing feature.
[0155] Accordingly, during the dot product process, the dot product result between two concatenated features can include four items: the dot product result between the two original features, the dot product result between the positional encodings of one original feature and another, and the dot product result between the two positional encodings.
[0156] On the other hand, the dot product between two positional codes can be expressed by the following formula:
[0157] P m [i,:]P n[j,:]=0,m≠n
[0158] P m [i,:]P n [j,:]>0,m=n
[0159] In the above formula, m and n can both be any numbers from 1, 2, and 3, used to represent one of the three characteristic planes, and P m [i,:] is used to represent the position code corresponding to the feature [i,:] belonging to the m-th feature plane, P n [j,:] is used to represent the position code corresponding to the feature [j,:] belonging to the nth feature plane, P m [i,:]P n [j,:] is used to represent the dot product result between these two positional codes.
[0160] As can be seen from the above formula, the dot product between the positional codes of any two features belonging to the same feature plane is zero, while the dot product between the positional codes of any two features belonging to different feature planes is greater than zero.
[0161] It should also be noted that this application does not impose any limitations on the setting of the attention mechanism. In practical applications, the attention mechanism can include self-attention, linear attention, etc. For ease of understanding, the embodiments of this application provide the following examples:
[0162] Taking the aforementioned Transformer structure as an example with the second network layer, based on the self-attention mechanism, the input data is calculated into a query (Q), key (K), and value (V) matrix, and the calculation results are represented as follows:
[0163]
[0164] In the above formula, Attention(Q,K,V) is used to represent the result calculated based on the self-attention mechanism, which can be used to represent the aforementioned predictive encoding features. softmax refers to the normalized exponential function. Q, K, and V represent three matrices, which are matrices obtained by mapping the concatenated features corresponding to the three sets of two-dimensional features mentioned above to the input data. T is used to represent the matrix transpose.
[0165] Furthermore, the planar dimension of each of the three aforementioned feature planes can be denoted as N, and each feature ∈ R in the three sets of two-dimensional features. N*d For example, corresponding to the above formula, Where, d k and dv Used to represent the size of the matrix obtained after mapping based on the attention mechanism.
[0166] In this attention mechanism, QK T The complexity of matrix multiplication increases quadratically with the increase of N. Therefore, when N is large, that is, when the resolution of the three sets of two-dimensional features in the input is high, the computational difficulty and cost will increase, and more computational resources will be required.
[0167] In another implementation, the aforementioned attention mechanism can also be a linear attention mechanism, especially when N is relatively large, in order to reduce computational complexity, save computational resources, and improve coding efficiency. The embodiments of this application provide the following examples to illustrate this:
[0168] In practical applications, the computation results of the linear attention mechanism can be represented as follows:
[0169]
[0170] In the above formula, LinearAttention(Q,K,V) represents the result calculated based on the linear attention mechanism, that is, it can be used to represent the aforementioned predictive encoding features. The meanings of Q, K, V, and T are the same as above. And, It refers to the matrix after mapping Q. It refers to the matrix after mapping K. It refers to the matrix after mapping V.
[0171] In the linear attention mechanism, first calculate... This avoids calculating the complete N×N matrix. Therefore, while ensuring computational efficiency, the computational complexity of the attention mechanism is reduced to linear, making it more efficient at handling long sequences of data, and thus more suitable for cases where N is relatively large.
[0172] In practical applications, time complexity can be used to describe the relationship between the time required for an algorithm to execute and the size of the input data; that is, time complexity can be used to intuitively represent computational complexity. Typically, time complexity is represented using Big O notation, i.e., O(f(n)), where n is the size of the input data, and f(n) is the function relating the algorithm's running time to the size of the input data. Correspondingly, in the example above, from self-attention mechanism to linear attention mechanism, the time complexity changes from O(N^2) to O(N^2). 2 The value was reduced to O(N).
[0173] In this regard, embodiments of this application also provide, as follows: Figure 8The diagram illustrates a comparison of two different attention mechanisms. Accordingly, in practical applications, a suitable attention mechanism can be flexibly selected based on the specific circumstances. For example, when the resolution of the three sets of input 2D features is relatively high, resulting in a large N, a linear attention mechanism can be used. This allows for the generation of more detailed 3D objects from the input high-resolution 2D features while reducing computational complexity and resource consumption, thus solving the problem of efficient encoding and decoding of high-resolution 3D object data.
[0174] The encoding process has been described in detail through the above embodiments. It should also be noted that decoding is the reverse process of encoding; therefore, for decoding, please refer to the descriptions of the various embodiments for encoding described above, which will not be repeated here.
[0175] (III) Regarding how to determine the location code, the embodiments of this application provide the following examples:
[0176] For ease of understanding, this embodiment uses the example of the position codes corresponding to the aforementioned three sets of two-dimensional features satisfying an orthogonal relationship, and provides the following illustration:
[0177] On the one hand, since orthogonality is used to indicate that the dot product between the position codes of two features is zero when the features belong to different feature planes, in order to quickly determine the position codes that satisfy the orthogonality relation, one possible implementation can be to use the property that the multiplication of matrix elements that satisfy the orthogonality relation is zero to determine the position codes required in this application.
[0178] On the other hand, since positional encoding is learnable during model training, and model training essentially involves adjusting model parameters, a matrix can be constructed based on the model parameters of the first network layer, and then the positional encoding can be determined. In this way, the values of the model parameters can be directly reflected in the positional encoding, and correspondingly, during model training, the positional encoding is updated directly by adjusting the model parameters.
[0179] To better understand the method of determining positional encoding by leveraging matrix properties and combining the model parameters of the first network layer, the embodiments of this application provide the following specific explanations:
[0180] If the model parameters of the first network layer can include the first parameter, then in the specific implementation of S302, the server can construct an initial matrix through the first network layer based on the three sets of two-dimensional features and the first parameter. In the initial matrix, the elements of the first column and the elements of the second column are orthogonal, and the elements of the first column and the elements of the second column are two different columns of elements in the initial matrix. Therefore, the server can determine the position codes corresponding to the three sets of two-dimensional features based on the initial matrix, so as to ensure that the position codes corresponding to the three sets of two-dimensional features satisfy the orthogonality relationship.
[0181] Accordingly, during model training, adjusting the first parameter changes the initial matrix, thereby determining a new positional encoding. Ultimately, at the end of model training, the first parameter of the first network layer in the target encoder can determine the precise positional encoding.
[0182] It should be noted that this application does not impose any limitations on how the position encoding is determined based on the initial matrix. For ease of understanding, the following examples are provided in the embodiments of this application:
[0183] Based on the above method of determining positional encoding, it can be ensured that the dot product between the positional encodings of two features belonging to different feature planes is zero, thus facilitating the initial encoder to capture the relative relationship between any two features belonging to different feature planes. In practical applications, a feature plane may include multiple two-dimensional features. Therefore, in order to facilitate the initial encoder to better capture the relative relationship between any two features belonging to the same feature plane, a corresponding positional encoding can also be determined for each feature.
[0184] In one possible implementation, the server can determine the target matrices corresponding to the three feature planes based on the initial matrix and the plane dimensions corresponding to each feature plane. For each feature plane, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane. In the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of a two-dimensional feature within the feature plane. Furthermore, the target matrices corresponding to the three feature planes can be determined as the position codes corresponding to three sets of two-dimensional features.
[0185] Based on this, each two-dimensional feature in each set of two-dimensional features has a corresponding positional encoding, which enables the initial encoder to better capture the relative relationship between any two features belonging to the same feature plane during the encoding process based on the attention mechanism.
[0186] To better understand, in this embodiment, the plane size corresponding to the three aforementioned feature planes is N, and each feature in the three sets of two-dimensional features ∈ R. N*dFor example, the following example illustrates the specific method for determining the location encoding:
[0187] Since the three feature planes have the same planar dimensions, the planar dimensions can be represented using the model parameters in the first network layer to facilitate determining the positional encoding. That is, the model parameters of the first network layer can also include a second parameter. Wherein, the first parameter ∈ R... d*1 And not zero, thus the constructed initial matrix ∈ R d*d and the second parameter ∈R 1*N Based on this, in order to construct ∈R N*d Location encoding.
[0188] In practice, since the elements of any two columns in the initial matrix satisfy an orthogonal relationship, the three columns of the initial matrix can first be defined as three first matrices, each of which ∈ R. d*1 Next, the three first matrices can be multiplied by the second matrix indicated by the second parameter, respectively, to obtain three matrices to be determined, where the second matrix ∈ R. 1*N Each of the three undetermined matrices ∈ R d*N Finally, the three undetermined matrices can be transposed to obtain the target matrices corresponding to the three feature planes, each target matrix ∈ R. N*d The target matrices corresponding to the three feature planes are the position codes corresponding to the three sets of two-dimensional features.
[0189] Based on this, the required positional encoding can be determined through simple matrix operations. Furthermore, when the three feature planes are of equal size, the second parameter in the first network layer directly represents the plane size and participates in the process of determining the target matrix (i.e., determining the positional encoding). This allows the updated positional encoding to be directly obtained by adjusting the second parameter during model training. Consequently, after model training, the first and second parameters of the second network layer in the target encoder can determine the accurate positional encoding.
[0190] In the above embodiments, it should also be noted that this application does not impose any limitations on the method of constructing the initial matrix. For ease of understanding, the embodiments of this application provide the following examples:
[0191] Since the initial matrix needs to satisfy the aforementioned orthogonality between the elements of any two columns, one possible implementation is to construct a householder matrix as the initial matrix. Specifically, the householder matrix can be denoted as H and expressed by the following formula:
[0192]
[0193] In the above formula, H represents the householder matrix, which is the initial matrix mentioned above, and h represents the first parameter mentioned above, where h∈R. d*1 And not zero, T is used to denote the matrix transpose, and I is used to denote the identity matrix.
[0194] As can be seen from the properties of the householder matrix, if we take the first three columns of this matrix as the three first matrices mentioned above, we can denote them as h1, h2, h3 ∈ R. d*1 This ensures that h1, h2, and h3 are mutually orthogonal.
[0195] Next, h1, h2, h3 and another learnable network parameter (i.e., the aforementioned second parameter) t∈R are... 1*N Perform matrix multiplication, and then transpose the resulting matrix to obtain three target matrices P1, P2, P3 ∈ R. N*d , used to represent the positional codes corresponding to the three sets of two-dimensional features respectively. Since h1, h2, and h3 are mutually orthogonal, it can be guaranteed that P1, P2, and P3 are also mutually orthogonal.
[0196] Based on this, by directly using the model parameters of the first network layer (such as the first and second parameters mentioned above) to determine the target matrix (i.e., the position encoding), the determined position encoding is continuously adjusted during the model training process. In other words, the aforementioned position encoding is learnable, and the position encoding determined in this application is a relative position encoding.
[0197] Compared to the method of assigning a unique code to each position in related technologies (i.e., absolute position coding), this method eliminates the need for manual assignment of position codes, reducing costs and avoiding subjective bias. Furthermore, it can adaptively determine precise position codes for each 3D object sample, enabling the encoder to better learn how to determine position codes from a variety of 3D object samples. This allows the encoder to better perceive spatial information, ensuring that the target encoded features maintain the relative spatial relationship between the three sets of 2D features of the 3D object sample.
[0198] (iv) Regarding how to train the model, the embodiments of this application provide the following examples:
[0199] 1. For the first part of model training, namely training the initial encoder and initial decoder, the following example is provided:
[0200] In one possible implementation, S304 described above can first determine the training loss based on the difference between the predicted 3D object and the 3D object sample. This training loss can be used to indicate the bias of the initial encoder in determining the predicted encoded features and the bias of the initial decoder in reconstructing the 3D object based on the predicted encoded features. Accordingly, the initial encoder and initial decoder can be trained based on the training loss. For example, during model training, the model training can be terminated when the training loss meets the training termination condition (i.e., the indicated bias is within an acceptable range), thus obtaining the corresponding target encoder and target decoder.
[0201] To improve model training efficiency, in another possible implementation, if the positional encodings corresponding to the three sets of two-dimensional features satisfy an orthogonal relationship, this orthogonality can be used as the optimization objective for model training. Specifically, in the implementation of S304, firstly, the server can determine the training loss based on the difference between the predicted 3D object determined by the initial decoder according to the predicted encoded features and the 3D object sample. Then, the server can adjust the model parameters of the initial encoder and initial decoder according to the optimization objective and the training loss to obtain the target encoder and target decoder.
[0202] Among them, the training loss can be used to indicate the bias of the initial encoder in determining the predicted coding features and the bias of the initial decoder in reconstructing and restoring the 3D object based on the predicted coding features. The optimization objective can be used to indicate that the position codes corresponding to the three sets of 2D features determined by the first network layer after the model parameters are adjusted still satisfy the orthogonality relationship.
[0203] Therefore, the requirement to satisfy orthogonality is used as a constraint for adjusting model parameters, ensuring that the positional encoding that satisfies the orthogonality relationship can still be determined based on the adjusted model parameters. This avoids ineffective parameter tuning when the positional encoding determined by the adjusted model parameters does not satisfy the orthogonality relationship, thus improving model training efficiency.
[0204] Furthermore, as can be seen from the foregoing description, the positional encoding is learnable during the model training process of this application. In the method of training the model based on the training loss and the optimization objective, the initial encoder, when learning how to determine the positional encoding, can both adjust the model parameters of the first network layer according to the magnitude of the training loss and avoid unnecessary parameter adjustments based on the optimization objective as a constraint, thus ensuring that the positional encoding is learnable and always maintains an orthogonal relationship.
[0205] 2. Regarding the second part, model training, which involves training the initial model, the following example is provided:
[0206] Typically, the process of adding noise (i.e., increasing noise) can be defined as forward propagation, while the process of denoising can be considered as backward propagation. In this application, the initial model is used to denoise the noisy encoded features corresponding to the target encoded features, and the target model trained by the model is used to denoise the initial noisy data. Therefore, the process of training the initial model can be considered as the initial model learning the inverse process of forward propagation, i.e., learning how to denoise.
[0207] In practical applications, the backpropagation process (i.e., the denoising process) can be used to predict how much noise has been added during forward propagation, or to predict features before forward propagation (i.e., before noise was added). It is understandable that different prediction methods indicate that the initial model has learned different knowledge and capabilities, which is related to the model training method. To better understand this, the embodiments of this application provide the following examples corresponding to these two different prediction methods:
[0208] (1) In order to obtain a target model that can predict how much noise has been added, the embodiments of this application provide the following model training method examples:
[0209] First, in the specific implementation of S306 mentioned above, the server can use the initial model to denoise the noisy coding features based on the first control data to obtain the predicted noise corresponding to the noisy coding features. That is, the predicted noise can refer to the predicted value of the noise added to the noisy coding features as determined by the initial model. Next, the server can train the initial model based on the difference between the predicted noise corresponding to the noisy coding features and the sample noise to obtain the target model.
[0210] Here, sample noise refers to the noise added to the noisy encoded features compared to the target encoded features; that is, the true value of the noise added during the noisy processing. Therefore, the difference between the two can indicate the bias of the initial model in predicting noise. Based on this, the model is trained so that the initial model can learn how to accurately predict the noise added during the noisy processing.
[0211] As can be seen, in this training method, the target encoding features determined by the trained target encoder are used as indirect supervision to train the initial model, enabling it to learn how to perform denoising processing to accurately predict noise. Since the first control data describes the 3D object sample, and the target encoding features also correspond to the features of the 3D object sample, when the initial model can accurately predict noise, it is considered that the initial model has learned how to perform denoising processing under the guidance of the first control data. If the difference between the obtained predicted noise and the noisy encoding features is calculated, the target encoding features can be restored. Therefore, it can be considered that the initial model has learned how to perform denoising processing based on the control data, and by combining the difference method, it can determine the features that conform to the control data.
[0212] Accordingly, in the application phase, specifically in the implementation of S307 mentioned above, the server can respond to the generation request by denoising the initial noise data using the target model based on the second control data, thereby obtaining the predicted noise corresponding to the initial noise data. Furthermore, the server can determine the difference between the initial noise data and the predicted noise corresponding to the initial noise data as the target feature. The target feature has a consistent correspondence with the input second control data and conforms to the generation requirements described by the second control data. Therefore, a 3D object conforming to the generation requirements can be subsequently determined based on the target feature. Based on this, the initial model and the target model only need to predict the noise, which simplifies the model structure.
[0213] It should be noted that this application does not impose any limitations on the method of adding noise to obtain noisy coded features. In practical applications, noise (such as Gaussian noise) can be continuously added to the target coded features at a certain noise intensity to obtain the noisy coded features. It is understandable that different noise intensities result in different differences between the noisy coded features and the target coded features; generally, the greater the noise intensity, the greater the difference between the two.
[0214] To improve model training efficiency, during the model training phase, the noise intensity can be input into the initial model. This explicitly informs the initial model of the noise intensity at which the currently input noisy encoded features were obtained. Based on this, the initial model is guided to perform more reasonable denoising, avoiding insufficient or excessive denoising, thereby accelerating model training and improving its efficiency.
[0215] Furthermore, corresponding to the first model training method, in practical implementation, the initial model training loss can be determined by the following formula:
[0216] L diff =||ε θ -ε0||2
[0217] In the above formula, L diff ε is used to represent the training loss of the initial model. θ The ε0 is used to represent the predicted noise corresponding to the noisy coding feature, and the ||||2 is used to represent the sample noise.
[0218] (2) In order to obtain a target model that can predict the features before noise is added, the embodiments of this application provide the following model training method examples:
[0219] First, in the specific implementation of S306 mentioned above, the server can use the initial model to denoise the noisy encoded features based on the first control data to obtain the denoised features corresponding to the noisy encoded features. Furthermore, the server can train the initial model based on the difference between the denoised features corresponding to the noisy encoded features and the target encoded features to obtain the target model.
[0220] Here, the denoised features corresponding to the noisy encoded features can refer to the features determined by the initial model before the noisy processing. The difference between these features and the target encoded features indicates the deviation of the initial model in predicting the features before the noisy processing. Based on this, the model is trained so that the initial model can learn how to accurately predict the features before the noisy processing.
[0221] Correspondingly, in the application phase, that is, in the specific implementation of the aforementioned S307, the server can respond to the generation request, and perform denoising processing on the initial noise data according to the second control data through the target model to obtain the denoising features corresponding to the initial noise data, and the server can determine the denoising features corresponding to the initial noise data as the target features.
[0222] As can be seen, in this training method, the target encoded features determined by the trained target encoder are used as direct supervision to train the initial model, enabling it to learn how to perform denoising and directly determine accurate target features. Based on this, the required target features can be directly determined using the target model, simplifying the processing steps.
[0223] Furthermore, corresponding to the second model training method, in practical implementation, the initial model training loss can be determined by the following formula:
[0224] L diff =||z θ -z0||2
[0225] In the above formula, L diff z is used to represent the training loss of the initial model. θ z0 is used to represent the denoised feature corresponding to the noisy coded feature, z0 is used to represent the target coded feature, and ||||2 is used to represent the mean square error.
[0226] The model determination method provided in this application has been described in detail through the above embodiments. Overall, it includes an initial encoder, an initial decoder, and an initial model. After model training is completed, it includes a target encoder, a target decoder, and a target model. In practical applications, other model structures can be flexibly set according to actual needs. For ease of understanding, the following examples are provided in the embodiments of this application:
[0227] Typically, 3D objects possess a certain geometric shape. Therefore, to improve the accuracy of the constructed 3D object's geometry, one possible implementation is to add a geometric representation module after the decoder, such as a network module based on the Signed Distance Field (SDF) algorithm. SDF describes the object's geometry by calculating the distance from any point in space to the nearest object, and is a method for representing geometric shapes. Consequently, its introduction facilitates the generation of more accurate 3D objects.
[0228] In related technologies, density-based methods are used to represent spatial geometry. However, the sparsity of the density distribution can affect the final determined geometry, easily leading to uneven geometry in the generated 3D objects. Therefore, this method is not suitable for representing the geometry of 3D objects. In contrast, this application uses an SDF-based method to represent geometry, which is beneficial for generating smoother geometry. For example, by inputting three sets of high-resolution 2D features and combining them with SDF, a geometrically smooth, complete, and detailed 3D object can be generated, resulting in a more refined 3D object.
[0229] To further improve the detail accuracy of the generated 3D objects, another possible implementation involves adding a Multilayer Perceptron (MLP) after the decoder. An MLP is a feedforward artificial neural network model where neurons in each layer are fully connected to those in the next layer, meaning each neuron receives input signals from all neurons in the previous layer. Therefore, MLPs can better capture deep information. Correspondingly, after its introduction, further processing of the decoded features improves the understanding of those features, leading to the construction of more accurate 3D objects.
[0230] Of course, in practical applications, the aforementioned multilayer perceptron and SDF-based network module can be introduced simultaneously. For ease of understanding, embodiments of this application also provide, as shown below... Figure 9 The diagram shows a model framework, where solid arrows represent the model training process and dashed arrows represent the model application process.
[0231] Understandably, before model training, Figure 9 The encoder, decoder, and diffusion model mentioned above are the same as the initial encoder and decoder mentioned above. The multilayer perceptron and geometric representation module are also untrained. After model training is completed, that is, during the model application process indicated by the dashed line, Figure 9 The diffusion model (i.e., the aforementioned target model), the decoder (i.e., the aforementioned target decoder), the multilayer perceptron, and the geometric representation module are all trained using the 3D object sample data from this application. Furthermore, Figure 9 The target 3D object in the context refers to the 3D object corresponding to the aforementioned generation request.
[0232] Accordingly, in the first part of the model training process, that is, when training the model based on the difference between the predicted 3D object and the 3D object sample, the initial encoder, the initial decoder, the multilayer perceptron, and the geometric representation module are trained together so that each part of the network module can be used for the task of generating 3D objects in this application.
[0233] Furthermore, to improve the model's generalization ability, another possible implementation may employ batch normalization and scaling operations. The embodiments of this application provide the following examples to illustrate this:
[0234] In practical applications, such as the VAE structure described above, the encoder outputs a feature distribution to represent the input data; that is, it outputs three sets of two-dimensional features as a single feature distribution. The mean of this feature distribution is μ, and its variance is σ. 2 These two parameters define the distribution p(z|x) of the latent variable z, which is typically assumed to be Gaussian. During model training, the feature distributions corresponding to the three sets of two-dimensional features x form the data distribution plane used to reconstruct and predict the 3D object. Correspondingly, the decoder reconstructs the data x from the latent variable z and outputs the reconstructed data p(x|z) to reconstruct the 3D object.
[0235] However, the amount of high-quality training samples is usually relatively small, which results in an insufficient quantity of high-quality three sets of two-dimensional features x. This leads to generalization problems and degradation in the trained encoder and decoder. Specifically, this manifests in the variance σ of the generated feature distribution. 2 If the latent variables are too small, their distribution is concentrated in a small area used to reconstruct the training data itself, making it difficult to generate new 3D objects through sampling. This results in a lack of imagination in the model application stage, potentially leading to the output of repetitive 3D objects.
[0236] To address this, this application proposes performing batch normalization and rescaling operations before the encoder's feature output layer (i.e., the last layer of the encoder), and accordingly, the decoder's network structure is adaptively modified. In specific implementation, the mean μ and variance σ of the output feature distribution are... 2 The specific steps are as follows:
[0237] μ ′ =scaler(norm(μ))
[0238] σ ′ =scaler(norm(σ))
[0239] In the above formula, μ represents the mean, σ represents the standard deviation, norm indicates the batch normalization operation, and scaler is used to further scale the batch-normalized mean and standard deviation. The scaler is defined as follows:
[0240]
[0241] In the above formula, τ and θ are both scaling factors used to indicate the degree of scaling operation.
[0242] Experimental data show that, after the above processing, the variance σ of the generated feature distribution is... 2 This significantly increases the capacity, meaning that the three sets of two-dimensional features of a 3D object sample can be encoded into a high-quality, continuous latent space representation. This also facilitates better training of the diffusion model. Based on this, the model can not only reconstruct the 3D object samples seen in the training data, but also learn a more general feature distribution from them, thereby improving the model degradation problem and increasing the diversity of generated 3D objects.
[0243] In practical applications, the encoder in the VAE structure outputs a feature distribution from the input data (i.e., the three sets of two-dimensional features in this application), and the decoder reconstructs the input data based on this feature distribution to reconstruct the three-dimensional object. Therefore, when training the initial encoder and decoder based on the difference between the predicted three-dimensional object and the sample three-dimensional object, the loss function of the VAE part can be constructed using the following formula:
[0244] L_vae=|(|Φ(x│z)-S|)|1+β(D_KL(q_φ(z|π)||p(z)))
[0245] In the above formula, L_vae refers to the loss function of the VAE part, which is used to indicate the difference between the predicted 3D object and the 3D object sample. |(|Φ(x│z)-S|)|1 is used to represent the reconstruction loss, that is, the deviation of the initial decoder when reconstructing the 3D object sample. Φ is used to represent the model parameters of the initial decoder. |()|1 is used to represent the L1 norm. S is used to represent the 3D object sample.
[0246] D_KL(q_φ(z|π)||p(z)) is the KL divergence (Kullback-Leibler Divergence). φ is used to represent the model parameters of the initial encoder. This part is used to constrain the feature distribution q_φ(z|π) of the initial encoder output to align to the target normal distribution p(z) in order to determine the spatial continuity of the feature distribution.
[0247] π is used to represent the input data, namely the three sets of two-dimensional features mentioned above. z can represent the feature distribution of the initial encoder output. x can be used to represent the data reconstructed by the decoder based on z, that is, it can represent the decoded features.
[0248] Furthermore, β is a hyperparameter used to balance the relative importance between the reconstruction loss term and the KL divergence term. By adjusting β, a balance can be achieved between the importance attached to the quality of the reconstructed 3D object and the importance attached to the learning feature distribution.
[0249] The model determination method provided in this application has been described in detail through the above embodiments. To better understand this method, the embodiments of this application summarize and explain the model application stage using the model determined in this application as an example:
[0250] When a user needs to generate 3D game assets (such as during game development), they can use reference images, descriptive text, etc., to describe the desired 3D game assets and then initiate a generation request. Next, the corresponding target features are automatically determined through the target model, and the corresponding 3D game assets (3D objects with geometric shapes and texture maps) are automatically generated through the target decoder.
[0251] In practical applications, some scenarios only require generating the geometry of a 3D object, while others require generating a 3D object with both geometric shape and material texture. This can be controlled as needed. For example, generation requirements can be specified by controlling data. Furthermore, supervision can be controlled as needed during the model training phase, i.e., whether to use the geometry of the 3D object sample directly as supervision, or to use the geometry of the 3D object sample combined with the material texture as supervision. Exemplarily, embodiments of this application also provide... Figure 10The diagram illustrates a method for generating the geometry of three-dimensional objects. By specifying the required geometry in the control data, both the generated three-dimensional objects D and E possess precise geometric shapes. This can be controlled according to actual needs, and this application does not impose any limitations.
[0252] In practical applications, automatically generated 3D game assets can be directly imported into game development engines (such as Unreal, Maya, Unity, etc.) for game development and production. Compared to related technologies that rely on manual labor, the creation of 3D game assets involves numerous tedious steps, from concept art design to high-precision model sculpting, model simplification, texture baking, skinning, and skeletal rigging. This complex and time-consuming process significantly slows down game development. In contrast, adopting this application can greatly reduce the costs of game development and production, shortening the development cycle. This is especially beneficial in developing large-scale AAA games, making game production faster and more convenient.
[0253] Furthermore, the models provided in this application can be packaged into a toolkit, allowing gamers to use this toolkit to independently generate 3D game assets and participate in new game modes such as game content creation. Based on this, a richer game generation pipeline can be provided to players, resulting in a more diverse gaming experience.
[0254] As can be seen, this application provides an efficient game asset generation technology, specifically generating adapted 3D game assets based on user-inputted reference images, descriptive text, etc., greatly reducing the barrier and cost of game production and shortening the game production time. For example, see the aforementioned... Figure 2 , Figure 5 as well as Figure 6 The generation effects shown in various embodiments.
[0255] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0256] based on Figure 3 Corresponding to the model determination method provided in the embodiments, this application also provides a model determination device 1100, which includes an acquisition unit 1101, a determination unit 1102, an encoding unit 1103, and a training unit 1104.
[0257] The acquisition unit 1101 is used to acquire three sets of two-dimensional features corresponding to the three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other.
[0258] The determining unit 1102 is used to determine the position codes corresponding to the three sets of two-dimensional features respectively through the first network layer in the initial encoder, based on the three sets of two-dimensional features, wherein the position codes are used to indicate the feature relationships between the two-dimensional features;
[0259] The encoding unit 1103 is used to encode the three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features respectively through the second network layer in the initial encoder to obtain the predicted encoded features corresponding to the three-dimensional object sample.
[0260] The training unit 1104 is used to train the model of the initial encoder and the initial decoder based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, so as to obtain the target encoder and the target decoder.
[0261] The determining unit 1102 is also used to determine the target encoding features corresponding to the three-dimensional object sample by the target encoder based on the three sets of two-dimensional features;
[0262] The training unit 1104 is further configured to train the initial model based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample to obtain the target model, wherein the first control data is used to describe the three-dimensional object sample.
[0263] The determining unit 1102 is further configured to, in response to a generation request, perform denoising processing on the initial noise data based on the second control data in the generation request using the target model to determine the target features corresponding to the second control data, and decode the target features using the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.
[0264] In one possible implementation, if the position codes corresponding to the three sets of two-dimensional features satisfy an orthogonal relationship, the encoding unit is further used to:
[0265] For each of the three sets of two-dimensional features, the two-dimensional feature and its corresponding position code are concatenated to obtain the concatenated features corresponding to the three sets of two-dimensional features respectively.
[0266] Based on the spliced features corresponding to the three sets of two-dimensional features, the second network layer in the initial encoder is used to encode the predicted encoded features corresponding to the three-dimensional object sample using an attention mechanism. The encoding process based on the attention mechanism is used to indicate that the spliced features belonging to the first feature plane and the spliced features belonging to the second feature plane are dot products. During the dot product process, the orthogonality relation is used to indicate that if the first feature plane and the second feature plane are different feature planes, the dot product result between the positional encoding in the spliced features belonging to the first feature plane and the positional encoding in the spliced features belonging to the second feature plane is zero.
[0267] In one possible implementation, the training unit is further used for:
[0268] The training loss is determined based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample.
[0269] Based on the optimization objective and the training loss, the model parameters of the initial encoder and the initial decoder are adjusted to obtain the target encoder and the target decoder. The optimization objective is used to indicate that the positional encodings corresponding to the three sets of two-dimensional features determined by the first network layer after the model parameter adjustment still satisfy the orthogonal relationship.
[0270] In one possible implementation, if the model parameters of the first network layer include the first parameter, the determining unit is further configured to:
[0271] Based on the three sets of two-dimensional features and the first parameter, an initial matrix is constructed through the first network layer. In the initial matrix, the elements of the first column and the elements of the second column satisfy the orthogonal relationship. The elements of the first column and the elements of the second column are two different columns of matrix elements in the initial matrix.
[0272] Based on the initial matrix, determine the position codes corresponding to the three sets of two-dimensional features.
[0273] In one possible implementation, the determining unit is further configured to:
[0274] Based on the initial matrix and the plane dimensions corresponding to the three feature planes, the target matrices corresponding to the three feature planes are determined. For each of the three feature planes, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane. In the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of a two-dimensional feature within the feature plane.
[0275] The target matrices corresponding to the three feature planes are determined as the position codes corresponding to the three sets of two-dimensional features.
[0276] In one possible implementation, the three feature planes each correspond to a plane size of N, and each feature in the three sets of two-dimensional features ∈ R. N*d R represents the set of real numbers, and d represents the feature dimension of each feature.
[0277] In one possible implementation, if the position encodings corresponding to the three sets of two-dimensional features are respectively the target matrices corresponding to the three feature planes, then the first parameter ∈ R d*1 And not zero, the initial matrix ∈ R d*d The model parameters of the first network layer also include a second parameter, the second parameter ∈ R. 1*N The determining unit is further configured to:
[0278] The three columns of matrix elements in the initial matrix are respectively determined as three first matrices, each of the three first matrices ∈ R. d*1 ;
[0279] Multiply the three first matrices by the second matrix indicated by the second parameter, respectively, to obtain three undetermined matrices, where the second matrix ∈ R. 1*N Each of the three undetermined matrices ∈ R d*N ;
[0280] Perform matrix transposes on the three undetermined matrices respectively to obtain the target matrices corresponding to the three feature planes, where each target matrix ∈ R. N*d .
[0281] In one possible implementation, the attention mechanism is a linear attention mechanism.
[0282] In one possible implementation, if the initial encoder comprises M groups of networks connected in sequence, wherein each of the first M-1 groups of networks comprises a first network layer, a second network layer, and a third network layer connected in sequence, and the Mth group of networks comprises a first network layer and a second network layer connected in sequence, the determining unit is further configured to:
[0283] Based on the three sets of two-dimensional features, the position codes corresponding to the three sets of two-dimensional features are determined by the first network layer of the first network in the M-group network.
[0284] The encoding unit is also used to encode the predicted encoding features corresponding to the first group of networks by encoding the three sets of two-dimensional features and the position encodings corresponding to the three sets of two-dimensional features through the second network layer in the first group of networks;
[0285] The determining unit is further configured to:
[0286] The predicted coding features corresponding to the first group of networks are transformed according to the third network layer in the first group of networks to obtain three sets of two-dimensional coding features corresponding to the first group of networks;
[0287] Based on the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks in the M groups of networks, the position codes corresponding to the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks are determined by the first network layer in the i-th group of networks, where i is an integer greater than 1 and less than or equal to M.
[0288] Based on the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks, and the position codes corresponding to the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks, the prediction coding features corresponding to the i-th group of networks are obtained by encoding through the second network layer in the i-th group of networks.
[0289] If the i-th network group includes a third network layer, then the predictive coding features corresponding to the i-th network group are transformed according to the third network layer in the i-th network group to obtain three sets of two-dimensional coding features corresponding to the i-th network group.
[0290] The predictive coding features corresponding to the three-dimensional object sample are determined based on the predictive coding features corresponding to the Mth group of networks.
[0291] In one possible implementation, the first control data includes at least one of a first reference image corresponding to the 3D object sample and a first descriptive text corresponding to the 3D object sample, and the second control data includes at least one of a second reference image for describing the 3D object generation requirements and a second descriptive text for describing the 3D object generation requirements.
[0292] In one possible implementation, the training unit is further used for:
[0293] The initial model performs denoising processing on the noisy coding features based on the first control data to obtain the predicted noise corresponding to the noisy coding features;
[0294] Based on the difference between the predicted noise and the sample noise corresponding to the noisy coding feature, the initial model is trained to obtain the target model, wherein the sample noise is the noise added by the noisy coding feature compared to the target coding feature;
[0295] In response to the generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data, including:
[0296] In response to the generation request, the initial noise data is denoised using the target model based on the second control data to obtain the predicted noise corresponding to the initial noise data;
[0297] The difference between the initial noise data and the predicted noise corresponding to the initial noise data is determined as the target feature.
[0298] As can be seen from the above technical solution, in the first part of model training, the positional encodings corresponding to the three sets of two-dimensional features are first determined by the first network layer in the initial encoder. Then, subsequent processing is performed based on the three sets of two-dimensional features and their respective positional encodings. The positional encoding is used to indicate the feature relationships between the two-dimensional features. Since the three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each set of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three sets of two-dimensional features are mutually perpendicular, the three-dimensional data of any point on the three-dimensional object sample is split into three sets of two-dimensional features. Therefore, the feature relationships between these three sets of two-dimensional features can reflect the absolute spatial information of that point in the three-dimensional object sample. Similarly, the feature relationships between the three sets of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically include the relative relationship between two two-dimensional features in the same feature plane within that feature plane, and the relative relationship between two two-dimensional features in different feature planes. Therefore, it can also reflect the relative spatial information of these two points in the three-dimensional object sample. Based on this, by using three sets of two-dimensional features to indicate three-dimensional data and introducing positional encoding, the spatial information carried by the three-dimensional data can be reflected based on rich feature relationships. Furthermore, training uses 3D object samples as supervision. Therefore, the initial encoder learns the rich spatial information from these samples to obtain predictive coding features with rich spatial information. The initial decoder learns how to decode these predictive coding features to determine more detailed predicted 3D objects. The resulting target encoder has better encoding capabilities, able to encode three sets of 2D features into target coding features with rich spatial information. The target decoder has better decoding capabilities to determine more detailed 3D objects. However, in practical applications, the generated 3D objects often lack 3D data, making it difficult to directly utilize the trained target encoder in the application phase. Therefore, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model, resulting in a target model that replaces the target encoder in the application phase.
[0299] In the application phase, in response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. Then, the target decoder decodes the target features to determine the 3D object corresponding to the generation request. Because the target model is trained using a target encoder, the target features determined using the target model still possess rich spatial information. Furthermore, the target decoder can better decode these spatially rich target features, thus enabling the determination of a more detailed 3D object for the generation request. Based on this, the technical problem in related technologies where the loss of spatial information during the encoding process prevents the trained model from generating detailed 3D objects is solved.
[0300] This application also provides a computer device, which can be a terminal, taking a smartphone as an example:
[0301] Figure 12 The diagram shown is a block diagram of a portion of the structure of a smartphone provided in an embodiment of this application. (Reference) Figure 12 The smartphone includes components such as a radio frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a Wi-Fi module 1170, a processor 1180, and a power supply 1190. The input unit 1130 may include a touch panel 1131 and other input devices 1132, the display unit 1140 may include a display panel 1141, and the audio circuit 1160 may include a speaker 1161 and a microphone 1162. Those skilled in the art will understand that... Figure 12 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0302] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0303] The processor 1180 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1120 and calling data stored in the memory 1120. Optionally, the processor 1180 may include one or more processing units; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1180.
[0304] In this embodiment, the steps performed by the processor 1180 in the smartphone can be based on Figure 12 The structure shown is implemented.
[0305] The computer device provided in this application embodiment can also be a server. Please refer to [link / reference]. Figure 13 As shown, Figure 13 This is a structural diagram of the server 1200 provided in this application embodiment. The server 1200 can vary significantly due to different configurations or performance. It may include one or more processors, such as a central processing unit (CPU) 1222, and a memory 1232, and one or more storage media 1230 (e.g., one or more mass storage devices) for storing application programs 1242 or data 1244. The memory 1232 and storage media 1230 can be temporary or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1222 may be configured to communicate with the storage media 1230 and execute the series of instruction operations in the storage media 1230 on the server 1200.
[0306] Server 1200 may also include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems 1241, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.
[0307] In this embodiment, the central processing unit 1222 in server 1200 can perform the following steps:
[0308] Obtain three sets of two-dimensional features corresponding to a three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other.
[0309] Based on the three sets of two-dimensional features, the position codes corresponding to the three sets of two-dimensional features are determined by the first network layer in the initial encoder. The position codes are used to indicate the feature relationships between the two-dimensional features.
[0310] Based on the three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features, the predicted encoded features corresponding to the three-dimensional object sample are obtained by encoding through the second network layer in the initial encoder.
[0311] Based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, the initial encoder and the initial decoder are trained to obtain the target encoder and the target decoder.
[0312] Based on the three sets of two-dimensional features, the target encoding features corresponding to the three-dimensional object sample are determined by the target encoder.
[0313] Based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample, the initial model is trained to obtain the target model, and the first control data is used to describe the three-dimensional object sample.
[0314] In response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. The target features are then decoded using the target decoder to determine the 3D object corresponding to the generation request. The second control data is used to describe the 3D object generation requirements.
[0315] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program that, when executed by a computer device, causes the computer device to perform the model determination method described in the foregoing embodiments.
[0316] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0317] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0318] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0319] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0320] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0321] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0322] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0323] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0324] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A model determination method, characterized in that, The method includes: Obtain three sets of two-dimensional features corresponding to a three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other. Based on the three sets of two-dimensional features, the position codes corresponding to the three sets of two-dimensional features are determined by the first network layer in the initial encoder. The position codes are used to indicate the feature relationships between the two-dimensional features. Based on the three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features, the predicted encoded features corresponding to the three-dimensional object sample are obtained by encoding through the second network layer in the initial encoder. Based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, the initial encoder and the initial decoder are trained to obtain the target encoder and the target decoder. Based on the three sets of two-dimensional features, the target encoding features corresponding to the three-dimensional object sample are determined by the target encoder. Based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample, the initial model is trained to obtain the target model, and the first control data is used to describe the three-dimensional object sample. In response to a generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data. The target features are then decoded using the target decoder to determine the 3D object corresponding to the generation request. The second control data is used to describe the 3D object generation requirements.
2. The method according to claim 1, characterized in that, If the position codes corresponding to the three sets of two-dimensional features satisfy an orthogonal relationship, the step of encoding the predicted coding features corresponding to the three-dimensional object sample by passing the second network layer in the initial encoder based on the three sets of two-dimensional features and their corresponding position codes includes: For each of the three sets of two-dimensional features, the two-dimensional feature and its corresponding position code are concatenated to obtain the concatenated features corresponding to the three sets of two-dimensional features respectively. Based on the spliced features corresponding to the three sets of two-dimensional features, the second network layer in the initial encoder is used to encode the predicted encoded features corresponding to the three-dimensional object sample using an attention mechanism. The encoding process based on the attention mechanism is used to indicate that the spliced features belonging to the first feature plane and the spliced features belonging to the second feature plane are dot products. During the dot product process, the orthogonality relation is used to indicate that if the first feature plane and the second feature plane are different feature planes, the dot product result between the positional encoding in the spliced features belonging to the first feature plane and the positional encoding in the spliced features belonging to the second feature plane is zero.
3. The method according to claim 2, characterized in that, The step of training a model on the initial encoder and the initial decoder based on the difference between the predicted 3D object determined by the initial decoder according to the predicted encoding features and the 3D object sample to obtain the target encoder and the target decoder includes: The training loss is determined based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample. Based on the optimization objective and the training loss, the model parameters of the initial encoder and the initial decoder are adjusted to obtain the target encoder and the target decoder. The optimization objective is used to indicate that the positional encodings corresponding to the three sets of two-dimensional features determined by the first network layer after the model parameter adjustment still satisfy the orthogonal relationship.
4. The method according to claim 2, characterized in that, If the model parameters of the first network layer include a first parameter, the step of determining the position codes corresponding to the three sets of two-dimensional features respectively through the first network layer in the initial encoder, based on the three sets of two-dimensional features, includes: Based on the three sets of two-dimensional features and the first parameter, an initial matrix is constructed through the first network layer. In the initial matrix, the elements of the first column and the elements of the second column satisfy the orthogonal relationship. The elements of the first column and the elements of the second column are two different columns of matrix elements in the initial matrix. Based on the initial matrix, determine the position codes corresponding to the three sets of two-dimensional features.
5. The method according to claim 4, characterized in that, The step of determining the position codes corresponding to the three sets of two-dimensional features based on the initial matrix includes: Based on the initial matrix and the plane dimensions corresponding to the three feature planes, the target matrices corresponding to the three feature planes are determined. For each of the three feature planes, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane. In the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of a two-dimensional feature within the feature plane. The target matrices corresponding to the three feature planes are determined as the position codes corresponding to the three sets of two-dimensional features.
6. The method according to any one of claims 1-5, characterized in that, The three feature planes each correspond to a plane size of N, and each feature in the three sets of two-dimensional features ∈ R. N*d R represents the set of real numbers, and d represents the feature dimension of each feature.
7. The method according to claim 6, characterized in that, If the position encodings corresponding to the three sets of two-dimensional features are respectively the target matrices corresponding to the three feature planes, then the first parameter ∈R d*1 And not zero, the initial matrix ∈ R d*d The model parameters of the first network layer also include a second parameter, the second parameter ∈ R. 1*N The target matrices corresponding to the three feature planes are determined in the following way: The three columns of matrix elements in the initial matrix are respectively determined as three first matrices, each of the three first matrices ∈ R. d*1 ; Multiply the three first matrices by the second matrix indicated by the second parameter, respectively, to obtain three undetermined matrices, where the second matrix ∈ R. 1*N Each of the three undetermined matrices ∈ R d*N ; Perform matrix transposes on the three undetermined matrices respectively to obtain the target matrices corresponding to the three feature planes, where each target matrix ∈ R. N*d .
8. The method according to claim 2, characterized in that, The attention mechanism is a linear attention mechanism.
9. The method according to claim 1, characterized in that, If the initial encoder comprises M groups of networks connected in sequence, wherein each of the first M-1 groups comprises a first network layer, a second network layer, and a third network layer connected in sequence, and the Mth group comprises a first network layer and a second network layer connected in sequence, the step of determining the position codes corresponding to the three groups of two-dimensional features respectively through the first network layer in the initial encoder, based on the three groups of two-dimensional features, includes: Based on the three sets of two-dimensional features, the position codes corresponding to the three sets of two-dimensional features are determined by the first network layer of the first network in the M-group network. The prediction coding features corresponding to the three-dimensional object sample are obtained by encoding the three sets of two-dimensional features and their corresponding position codes through the second network layer in the initial encoder, including: Based on the three sets of two-dimensional features and the position codes corresponding to the three sets of two-dimensional features, the prediction coding features corresponding to the first set of networks are obtained by encoding through the second network layer in the first set of networks; The method further includes: The predicted coding features corresponding to the first group of networks are transformed according to the third network layer in the first group of networks to obtain three sets of two-dimensional coding features corresponding to the first group of networks; Based on the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks in the M groups of networks, the position codes corresponding to the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks are determined by the first network layer in the i-th group of networks, where i is an integer greater than 1 and less than or equal to M. Based on the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks, and the position codes corresponding to the three sets of two-dimensional coding features corresponding to the (i-1)th group of networks, the prediction coding features corresponding to the i-th group of networks are obtained by encoding through the second network layer in the i-th group of networks. If the i-th network group includes a third network layer, then the predictive coding features corresponding to the i-th network group are transformed according to the third network layer in the i-th network group to obtain three sets of two-dimensional coding features corresponding to the i-th network group. The predictive coding features corresponding to the three-dimensional object sample are determined based on the predictive coding features corresponding to the Mth group of networks.
10. The method according to claim 1, characterized in that, The first control data includes at least one of a first reference image corresponding to the three-dimensional object sample and a first descriptive text corresponding to the three-dimensional object sample; the second control data includes at least one of a second reference image used to describe the requirements for generating the three-dimensional object and a second descriptive text used to describe the requirements for generating the three-dimensional object.
11. The method according to claim 1, characterized in that, The step of training the initial model to obtain the target model based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample includes: The initial model performs denoising processing on the noisy coding features based on the first control data to obtain the predicted noise corresponding to the noisy coding features; Based on the difference between the predicted noise and the sample noise corresponding to the noisy coding feature, the initial model is trained to obtain the target model, wherein the sample noise is the noise added by the noisy coding feature compared to the target coding feature; In response to the generation request, the initial noise data is denoised using the target model based on the second control data in the generation request to determine the target features corresponding to the second control data, including: In response to the generation request, the initial noise data is denoised using the target model based on the second control data to obtain the predicted noise corresponding to the initial noise data; The difference between the initial noise data and the predicted noise corresponding to the initial noise data is determined as the target feature.
12. A model determining device, characterized in that, The device includes an acquisition unit, a determination unit, an encoding unit, and a training unit: The acquisition unit is used to acquire three sets of two-dimensional features corresponding to the three-dimensional object sample. The three sets of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample. Each set of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data. The three feature planes corresponding to the three sets of two-dimensional features are perpendicular to each other. The determining unit is used to determine the position codes corresponding to the three sets of two-dimensional features respectively through the first network layer in the initial encoder, based on the three sets of two-dimensional features, wherein the position codes are used to indicate the feature relationships between the two-dimensional features; The encoding unit is used to encode the predicted encoded features corresponding to the three-dimensional object sample by encoding the three sets of two-dimensional features and the position encodings corresponding to the three sets of two-dimensional features through the second network layer in the initial encoder. The training unit is used to train the model of the initial encoder and the initial decoder based on the difference between the predicted 3D object determined by the initial decoder according to the predicted coding features and the 3D object sample, so as to obtain the target encoder and the target decoder. The determining unit is further configured to determine the target encoding features corresponding to the three-dimensional object sample by means of the target encoder based on the three sets of two-dimensional features; The training unit is further configured to train the initial model based on the noisy coding features corresponding to the target coding features and the first control data corresponding to the three-dimensional object sample to obtain the target model, wherein the first control data is used to describe the three-dimensional object sample. The determining unit is further configured to, in response to a generation request, perform denoising processing on the initial noise data based on the second control data in the generation request using the target model to determine the target features corresponding to the second control data, and decode the target features using the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.
13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-11 according to instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a computer device, causes the computer device to perform the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When it is run on a computer device, it causes the computer device to perform the method according to any one of claims 1-11.
Citation Information
Patent Citations
Three-dimensional model generation method and device, equipment, storage medium and program product
CN117252984A
Encoder training method and related device
CN117456102A