Model determination method and related device

By obtaining three sets of two-dimensional features of three-dimensional object samples and introducing position encoding, model training is solved, and the problem of low three-dimensional object generation accuracy caused by spatial information loss in the existing technology is solved, and more accurate three-dimensional object generation is achieved.

CN120032048AActive Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510099342.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

In the prior art, the trained model is difficult to generate accurate three-dimensional objects, mainly due to the loss of spatial information during the encoding process.

Method used

By obtaining three sets of two-dimensional features of the three-dimensional object sample and introducing position encoding, model training is performed to obtain the target encoder and the target decoder. This method can better retain and utilize the spatial information of three-dimensional objects.

Benefits of technology

It realizes the generation of three-dimensional objects with more accurate details, solving the problem of low model generation accuracy caused by spatial information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032048A_ABST
    Figure CN120032048A_ABST
Patent Text Reader

Abstract

The invention discloses a model determination method and a related device, in a first part of model training, three groups of two-dimensional features of a three-dimensional object sample are utilized to represent three-dimensional data, and position codes are introduced to indicate a feature relationship among the two-dimensional features, so that spatial information of a lace tied by the three-dimensional data can be reflected through rich relationships among the features. And the three-dimensional object sample is used as supervision, so that an encoder can fully learn spatial information to determine more accurate encoding features, and the decoder can reconstruct the three-dimensional object sample. Thirdly, encoding the three groups of two-dimensional features into precise target encoding features by using the trained encoder, and performing model training on the initial model by using the precise target encoding features as supervision to obtain a target model so as to replace the trained encoder in an application stage; according to the method, position coding is introduced, so that the related problem of detail missing of the generated three-dimensional object caused by spatial information loss in coding in the related technology is solved, and the three-dimensional object with accurate details can be determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a model determination method and related devices. Background Art

[0002] Three-dimensional objects are used to simulate and express the shape and structure of objects in three-dimensional space, such as three-dimensional buildings and characters. They are widely used in games, animation, virtual reality, design, architecture and other fields.

[0003] In practical applications, how to quickly construct three-dimensional objects has gradually become the key to constructing three-dimensional objects. In related technologies, the three-dimensional data of a three-dimensional object is encoded into features through an encoder, and then the features are reconstructed into a three-dimensional object through a decoder, and the encoder and decoder are trained based on this process. In the application stage, the diffusion model is used to generate features, and then the features are sent to the trained decoder to generate the required three-dimensional objects. Based on this, it is no longer necessary to rely on manual design of three-dimensional objects, so three-dimensional objects can be constructed more efficiently.

[0004] However, models trained using related techniques (including the aforementioned diffusion model and decoder) are difficult to generate accurate three-dimensional objects. Summary of the invention

[0005] In order to solve the above technical problems, the present application provides a model determination method and related devices, which can determine three-dimensional objects with more precise details, and solve the technical problem in the related technology that the trained model is difficult to generate accurate three-dimensional objects due to the loss of spatial information during the encoding process.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] On the one hand, an embodiment of the present application provides a model determination method, the method comprising:

[0008] Acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other;

[0009] According to the three groups of two-dimensional features, determining, by a first network layer in an initial encoder, position codes corresponding to the three groups of two-dimensional features respectively, wherein the position codes are used to indicate a feature relationship between the two-dimensional features;

[0010] Based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, encoding is performed through the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample;

[0011] Based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample, the initial encoder and the initial decoder are model trained to obtain a target encoder and a target decoder;

[0012] Determining, by the target encoder, target coding features corresponding to the three-dimensional object samples according to the three groups of two-dimensional features;

[0013] Performing model training on an initial model according to a noise-added coding feature corresponding to the target coding feature and first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample;

[0014] In response to a generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

[0015] On the other hand, an embodiment of the present application provides a model determination device, the device comprising an acquisition unit, a determination unit, an encoding unit and a training unit:

[0016] The acquisition unit is used to acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other;

[0017] The determining unit is used to determine, according to the three groups of two-dimensional features, position codes respectively corresponding to the three groups of two-dimensional features through the first network layer in the initial encoder, wherein the position codes are used to indicate a feature relationship between the two-dimensional features;

[0018] The encoding unit is used to perform encoding through the second network layer in the initial encoder based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, so as to obtain the predicted coding features corresponding to the three-dimensional object sample;

[0019] The training unit is used to perform model training on the initial encoder and the initial decoder based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample to obtain a target encoder and a target decoder;

[0020] The determining unit is further configured to determine, through the target encoder, a target coding feature corresponding to the three-dimensional object sample according to the three groups of two-dimensional features;

[0021] The training unit is further used to perform model training on the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample;

[0022] The determination unit is also used to respond to the generation request, denoise the initial noise data through the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and decode the target features through the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

[0023] On the other hand, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory:

[0024] The memory is used to store a computer program and transmit the computer program to the processor;

[0025] The processor is configured to execute the method described in any one of the preceding aspects according to instructions in the computer program.

[0026] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program. When the computer program is executed by a computer device, the computer device executes the method described in any of the above aspects.

[0027] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which, when executed on a computer device, enables the computer device to execute the method described in any one of the aforementioned aspects.

[0028] It can be seen from the above technical solution that in the first part of the model training, the position codes corresponding to the three groups of two-dimensional features are first determined by the first network layer in the initial encoder, and then subsequent processing is performed according to the three groups of two-dimensional features and their respective position codes. Among them, the position code is used to indicate the feature relationship between the two-dimensional features. Since the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other, so that the three-dimensional data of any point on the three-dimensional object sample is split into three groups of two-dimensional features, so the feature relationship between the three groups of two-dimensional features can reflect the absolute spatial information of the point in the three-dimensional object sample. Similarly, the feature relationship between the three groups of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically includes the relative relationship between the two two-dimensional features in the same feature plane in the feature plane, and the relative relationship between the two two-dimensional features in different feature planes, so it can also reflect the relative spatial information of the two points in the three-dimensional object sample. Based on this, after using three groups of two-dimensional features to indicate the three-dimensional data and introducing the position code, the spatial information carried by the three-dimensional data can be reflected based on the rich relationship between the features. And the training is supervised by three-dimensional object samples, so the initial encoder can learn the rich spatial information in the three-dimensional object samples to obtain prediction coding features with rich spatial information, and the initial decoder can learn how to decode the prediction coding features to determine the predicted three-dimensional objects with more precise details. The target encoder finally obtained has better encoding ability and can encode three sets of two-dimensional features into target coding features with rich spatial information. The target decoder has better decoding ability to determine the three-dimensional objects with more precise details. In practical applications, the three-dimensional objects that need to be generated often do not have three-dimensional data, so it is difficult to directly use the trained target encoder in the application stage. For this reason, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model to obtain the target model to replace the target encoder in the application stage.

[0029] Correspondingly, in the application stage, in response to the generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and then the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request. Because the target model is obtained by model training using the target encoder, the target features determined using the target model are still features with rich spatial information, and the target decoder can better decode the target features with rich spatial information, so it can determine a three-dimensional object with more precise details for the generation request. Based on this, the technical problem in the related art that the trained model cannot generate a three-dimensional object with precise details due to the loss of spatial information during the encoding process is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technical members in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0031] Figure 1 A schematic diagram of an application scenario of a model determination method provided in an embodiment of the present application;

[0032] Figure 2 A schematic diagram of generating a three-dimensional object in a model application phase provided in an embodiment of the present application;

[0033] Figure 3 A flow chart of a model determination method provided in an embodiment of the present application;

[0034] Figure 4 A schematic diagram of three characteristic planes provided in an embodiment of the present application;

[0035] Figure 5 A schematic diagram of generating a three-dimensional object based on a description text provided in an embodiment of the present application;

[0036] Figure 6 A schematic diagram of generating a three-dimensional object based on a reference image provided in an embodiment of the present application;

[0037] Figure 7 A schematic diagram of a network structure of an initial encoder and an initial decoder provided in an embodiment of the present application;

[0038] Figure 8 A schematic diagram comparing two different attention mechanisms provided in an embodiment of the present application;

[0039] Fig. 9 A schematic diagram of a model framework provided for an embodiment of the present application;

[0040] Fig.10 A schematic diagram of generating a three-dimensional object geometric shape provided in an embodiment of the present application;

[0041] Fig.11 A structural diagram of a model determination device provided in an embodiment of the present application;

[0042] Fig.12 A structural diagram of a terminal provided in an embodiment of the present application;

[0043] Fig.13 A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] The embodiments of the present application are described below in conjunction with the accompanying drawings.

[0045] In order to construct a three-dimensional object, in some related technologies, a model for generating three-dimensional objects based on images is trained using two-dimensional images (such as a single-view image of a three-dimensional object as a training sample, or multiple-view images as a training sample, etc.). In the application stage, a two-dimensional image can be input and a three-dimensional object can be generated through the model.

[0046] However, this method is very dependent on the quality of the image. When the image quality is low (such as low resolution), or when there are inconsistencies between multi-view images (such as the same part appears in multiple images from different perspectives), it is very easy to cause the generated three-dimensional object to be distorted. Specifically, an image only indicates in a very one-sided way what kind of three-dimensional object it might be, so directly constructing a three-dimensional object from an image may not be the desired three-dimensional object. For example, the front of the generated three-dimensional object is similar to the image, but the side is distorted. In addition, inconsistencies between multi-view images will result in the generation of strange-shaped three-dimensional objects. For example, in a multi-view image, parts of a squirrel's head appear in all multi-view images, and the final generated three-dimensional object of a squirrel will have multiple heads superimposed on each other.

[0047] In other related technologies, the model is trained using 3D data of 3D objects (such as point cloud data). Specifically, the 3D data of the 3D object is encoded into features through an encoder, and then the features are reconstructed into 3D objects through a decoder. The encoder and decoder are trained based on this process. In the application stage, the diffusion model is used to generate features, and then the features are sent to the trained decoder to generate the required 3D object.

[0048] Although this method uses three-dimensional data to train the model, it can solve the problems in the first related technology to a certain extent. However, in the second method, the encoding method used will lose the spatial information of the three-dimensional object, which makes it difficult for the trained model to generate a three-dimensional object with precise details. For example, multiple parts of the generated three-dimensional object are misaligned, and some details are missing.

[0049] To this end, the embodiment of the present application provides a model determination method and related devices. First, the initial encoder and the initial decoder are trained by using three groups of two-dimensional features of the three-dimensional object sample and introducing position coding. Among them, the position coding is used to indicate the feature relationship between the two-dimensional features, and the three groups of two-dimensional features are used to indicate the three-dimensional data. Compared with the three-dimensional data, because each group of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other, the three-dimensional data of any point on the three-dimensional object sample is split into three groups of two-dimensional features, so the feature relationship between the three groups of two-dimensional features can reflect the absolute spatial information of the point in the three-dimensional object sample. Similarly, the feature relationship between the three groups of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically includes the relative relationship between two two-dimensional features in the same feature plane and the relative relationship between two two-dimensional features in different feature planes, so it can also reflect the relative spatial information of the two points in the three-dimensional object sample.

[0050] Based on this, by using three sets of two-dimensional features to indicate three-dimensional data and introducing position coding, the spatial information carried by the three-dimensional data can be reflected based on the rich relationship between the features. In this way, when training the model, the initial encoder can fully learn the spatial information carried by the three-dimensional data to obtain prediction coding features with rich spatial information, and then based on this prediction coding feature, the initial decoder can learn how to decode the prediction coding features to determine the predicted three-dimensional object with more precise details. Accordingly, the obtained target encoder has better encoding capabilities and can encode the three sets of two-dimensional features into target coding features with rich spatial information. The target decoder has better decoding capabilities to determine the three-dimensional object with more precise details.

[0051] Next, the trained target encoder is used to train the initial model. Specifically, the target encoder is used to encode the three sets of two-dimensional features into target encoding features with rich spatial information, and then the initial model is trained based on this, so that the initial model can learn how to perform denoising to determine features with rich spatial information, thereby obtaining a target model that can replace the target encoder in the application stage.

[0052] Correspondingly, in the application stage, the target model can be used to determine the target model features that meet the generation request and have rich spatial information, and the target decoder can be used to better decode the target features with rich spatial information, thereby determining a three-dimensional object with more precise details.

[0053] Compared with the first related technology mentioned above, the model training stage relies on three-dimensional data rather than two-dimensional images. Since three-dimensional data can accurately reflect the overall situation of a three-dimensional object (such as the front situation, side situation, back situation, etc.) compared with two-dimensional images, it can avoid the technical problem of being unable to generate accurate three-dimensional objects by relying solely on two-dimensional images. In addition, compared with the second related technology mentioned above, spatial information is no longer lost, thereby solving the technical problem of difficulty in generating three-dimensional objects with precise details due to the loss of spatial information.

[0054] The model determination method provided in the embodiment of the present application can be implemented by a computer device, which can be a terminal or a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include but are not limited to smart phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.

[0055] The embodiments of the present application can be specifically applied to various scenarios where three-dimensional objects need to be generated, for example, in games, animations and other applications, three-dimensional virtual objects are generated. In another example, in virtual reality applications, three-dimensional virtual models in virtual space are generated. In another example, in the fields of design and architecture, three-dimensional objects (such as vases, furniture, etc.), three-dimensional buildings, etc. are generated.

[0056] Generally, in game applications, the generated 3D virtual objects can also be called 3D game assets, which are the basic components of game applications. For example, they can include 3D virtual characters, 3D game scenes, and 3D objects in game scenes, etc. From the details of 3D objects, they usually include geometric shapes (which can also be represented by geometric patches), texture maps (such as materials and colors), etc.

[0057] It should be noted that in the specific implementation of the present application, the process of determining the model and the process of using the model to construct a three-dimensional object may involve relevant data such as user information. When the above embodiments of the present application are applied to specific products or technologies, the user's separate consent or separate permission is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0058] For better understanding, Figure 1 The following is a schematic diagram showing an application scenario of the model determination method provided in the embodiment of the present application. Figure 1 In the illustrated scenario, the server 100 is used as an example of the aforementioned computer device for illustration, which may include two parts: ① a model training phase and ② a model application phase. Specifically:

[0059] The server 100 can obtain three sets of two-dimensional features corresponding to the three-dimensional object samples, and use the three sets of two-dimensional features corresponding to the three-dimensional object samples to perform model training. Regarding the ① model training stage, the model training can specifically include two parts:

[0060] In the first part of model training, the server 100 may input the three sets of two-dimensional features into the first network layer of the initial encoder to determine the position codes corresponding to the three sets of two-dimensional features. And according to the three sets of two-dimensional features and their corresponding position codes, they are input into the second network layer of the initial encoder to determine the predicted coding features corresponding to the three-dimensional object sample. Then, the server 100 may input the predicted coding features into the initial decoder to determine the predicted three-dimensional object.

[0061] Among them, three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other, and the position coding can be used to indicate the feature relationship between the two-dimensional features. It can be understood that the three-dimensional data can reflect the spatial information of the three-dimensional object sample. After using the three groups of two-dimensional features to indicate the three-dimensional data, the feature relationship between different feature planes and the feature relationship within the same feature plane can be used to reflect the spatial information carried by the three-dimensional data. Based on this, after using three groups of two-dimensional features and introducing position coding, the spatial information carried by the three-dimensional data can be reflected based on the rich relationship between features.

[0062] In this way, with the three-dimensional object samples as supervision, the server 100 can perform model training on the initial encoder and the initial decoder based on the difference between the predicted three-dimensional object and the three-dimensional object sample, so that during the training process, the initial encoder can learn the rich spatial information in the three-dimensional object sample to obtain the prediction coding features with rich spatial information, and the initial decoder can learn how to decode the prediction coding features to determine the predicted three-dimensional object with more precise details. Finally, a target encoder with better encoding capability and a target decoder with better decoding capability are obtained.

[0063] In the second part of the model training, the server 100 can use the target encoder to determine the target coding feature based on the three groups of two-dimensional features. The target coding feature refers to a coding feature with rich spatial information, based on which a three-dimensional object sample with precise details can be reconstructed. However, in actual applications, the three-dimensional object to be generated often does not have three-dimensional data, so it is difficult to directly use the trained target encoder in the application stage. Therefore, in the second part of the model training, the initial model is trained using the noise coding feature corresponding to the target coding feature and the first control data to obtain a target model to replace the target encoder in the application stage. Among them, the first control data is used to describe the three-dimensional object sample. That is, the initial model can learn how to denoise the noise coding feature based on the description of the first control data, so as to finally determine the feature with rich spatial information and close to the target coding feature.

[0064] Correspondingly, in the ② model application stage, the server 100 can input the second control data and the initial noise data into the target model, the second control data is used to describe the three-dimensional object generation requirements, and the initial noise data is denoised according to the second control data through the target model to determine the target features with rich spatial information. In addition, the server 100 can input the target features into the target decoder, and decode the target features through the target decoder to determine the corresponding three-dimensional object.

[0065] Because position encoding is introduced in model training, the model can learn the spatial information carried in the three-dimensional data, and the corresponding combination of "target model + target decoder" can generate a three-dimensional object with precise details that meets the generation requirements described by the second control data. Based on this, the technical problem in the related art that the model cannot generate a three-dimensional object with precise details and cannot generate a three-dimensional object that meets the requirements due to the loss of spatial information in the encoding process is solved.

[0066] For better understanding, the present application also provides the following examples: Figure 2 The schematic diagram of generating a three-dimensional object in a model application stage is shown, and there is a connection relationship (such as a network connection, etc.) between the terminal 201 (such as the aforementioned computer) and the server 202. In the actual application stage, specifically:

[0067] When the user wants to generate a three-dimensional object, the user can input the second control data A through the terminal 201 to describe the three-dimensional object generation requirement. For example, the second control data A in this embodiment is "a girl with long black hair, standing with arms open, wearing a long polka dot puffy dress with a bow at the neckline and wavy patterns on the hem".

[0068] Then, the terminal 201 may send a generation request carrying the second control data A to the server 202. Accordingly, the server 202 may start to implement the aforementioned Figure 1 The processing steps of the ② model application stage in the example are, that is, denoising the initial noise data according to the second control data A through the target model to determine the target feature A, and decoding the target feature A through the target decoder to determine the three-dimensional object A.

[0069] And, the server 202 may return the generated three-dimensional object A to the terminal 201 so that the user can view it on the terminal 201. For example, the finally generated three-dimensional object A may be displayed on the computer page of the terminal 201.

[0070] It can be seen that when users want to generate a 3D object, they only need to input control data to describe the specific generation requirements, and then they can use the trained target model and target decoder to quickly generate a 3D object with precise details that meets the generation requirements. This not only improves the quality of the 3D object, but also helps to improve the generation efficiency of the 3D object because it is effectively generated.

[0071] It is understandable that in actual applications, the user can also perform various operations such as dragging, rotating, and zooming in the computer page of the terminal 201 to flexibly view different surfaces, local details, etc. of the three-dimensional object A.

[0072] Figure 3 A flowchart of a model determination method provided in an embodiment of the present application is described by taking a server as an example of the aforementioned computer device. The method includes S301-S307:

[0073] S301: Obtain three groups of two-dimensional features corresponding to a three-dimensional object sample.

[0074] Among them, three groups of two-dimensional features can be used to indicate three-dimensional data of a three-dimensional object sample, each group of two-dimensional features can be used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other.

[0075] It should be noted that this application does not impose any limitation on how to determine the three sets of two-dimensional features. In practical applications, three mutually perpendicular feature planes can be constructed, and the three-dimensional data can be mapped and projected onto the three feature planes.

[0076] For example, three mutually perpendicular characteristic planes can be seen in Figure 4As shown. Usually, the three-dimensional data of a three-dimensional object sample may refer to the value of each point on the three-dimensional object sample in a three-dimensional coordinate system, such as (x, y, z). Correspondingly, the three feature planes may be respectively recorded as (x, y) feature plane, (x, z) feature plane, and (y, z) feature plane, and in practical applications, the three groups of two-dimensional features may also be referred to as three-plane features, and may be recorded as triplane. In addition, the three-dimensional data of a three-dimensional object sample may also be point cloud data, etc.

[0077] It can be understood that three-dimensional data can reflect the spatial information of a three-dimensional object sample. After the three-dimensional data is indicated by three groups of two-dimensional features, the three-dimensional data of any point on the three-dimensional object sample is split into three groups of two-dimensional features. For example, point A (x1, y1, z1) will be split into three groups of two-dimensional features (x1, y1), (x1, z1) and (y1, z1) belonging to three feature planes. Therefore, the absolute spatial information of point A in the three-dimensional object sample can be reflected through the feature relationship between the three groups of two-dimensional features.

[0078] Similarly, any two points on the corresponding three-dimensional object sample, such as point A (x1, y1, z1) and point B (x2, y2, z2), will be split into two-dimensional features (x1, y1) and (x2, y2) belonging to the (x, y) feature plane, two-dimensional features (x1, z1) and (x2, z2) belonging to the (x, z) feature plane, and two-dimensional features (y1, z1) and (y2, z2) belonging to the (y, z) feature plane.

[0079] That is, the characteristic relationship between the three groups of two-dimensional features corresponding to any two points on the three-dimensional object sample, specifically including the relative relationship between two two-dimensional features in the same characteristic plane, and the relative relationship between two two-dimensional features in different characteristic planes. In this way, not only the absolute spatial information of a single point can be reflected through the characteristic relationship between the three groups of two-dimensional features of the single point itself, but also the relative spatial information between two two-dimensional features in the same characteristic plane can be reflected.

[0080] Based on this, after using three sets of two-dimensional features to express three-dimensional data, we can utilize richer feature relationships between two-dimensional features (including feature relationships between two-dimensional features in the same feature plane and feature relationships between two-dimensional features in different feature planes) to fully reflect the spatial information carried in the three-dimensional data, so that the model can better learn spatial information during model training.

[0081] It is understandable that different three-dimensional object samples have different geometric shapes, the number of features included, etc., so the three-dimensional data is different, and the sizes of the corresponding three feature planes will also be different. In this regard, it should be noted that this application does not impose any restrictions on the sizes of the feature planes. For ease of understanding, the following examples are provided in the embodiments of this application:

[0082] Usually, various three-dimensional object samples are used for model training so that the model can fully learn and avoid underfitting. In order to facilitate model learning (specifically, to facilitate learning of the initial encoder, etc.), in a possible implementation, three feature planes of uniform size can be used to express the three-dimensional data of each three-dimensional object sample. Similarly, for each feature, a vector with the same number of feature dimensions can also be used to represent it, thereby achieving unification.

[0083] In the specific implementation, the plane size corresponding to the three feature planes is N, and each feature ∈R N*d , R is used to represent the real number set, d is used to represent the feature dimension of each feature, and d is a positive integer. Among them, N can be expressed as u*v, u is used to represent the length of the feature plane, and v is used to represent the width of the feature plane. Both u and v are positive integers.

[0084] Based on this, for any geometric shape and any number of features of a 3D object sample, its corresponding 3D data can be converted into three feature planes of uniform size, and the number of feature dimensions of each feature is also consistent. Then it is sent to the initial encoder, which makes it easier for the initial encoder to learn using a variety of 3D object samples.

[0085] S302: According to the three groups of two-dimensional features, determine the position codes corresponding to the three groups of two-dimensional features respectively through the first network layer in the initial encoder.

[0086] S303: Based on the three groups of two-dimensional features and the position codes corresponding to the three groups of two-dimensional features, encoding is performed through the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample.

[0087] Among them, the position coding can be used to indicate the feature relationship between two-dimensional features. Combined with the above description, the feature relationship between two-dimensional features can reflect spatial information, so the position coding can reflect spatial information. Based on this, the initial encoder can learn the spatial information of the three-dimensional object sample in combination with the indication of the position coding during the encoding process, and then determine the corresponding prediction coding features.

[0088] In practical applications, encoding is a technology that processes arbitrary data into a certain feature expression (such as a feature expression in vector form) so that the model can understand it, and through encoding, three-dimensional data can be compressed and reduced in dimension, thereby reducing redundant information. Corresponding to this is decoding, which is a technology that uses the feature expression obtained by encoding to reconstruct and restore. Accordingly, the predicted coding feature may refer to the feature expression corresponding to the data of the three groups of two-dimensional features of the three-dimensional object sample determined by the initial encoder, which is the feature that the initial encoder believes can indicate the three-dimensional object sample, and the predicted coding feature can be used for subsequent reconstruction and restoration.

[0089] From the network structure of the initial encoder, it includes the first network layer and the second network layer, and the function of the first network layer is to determine the position coding, and the function of the second network layer is to encode. The network structure of the encoder used in the related art will directly encode the three-dimensional data into one-dimensional features. In contrast, the encoder used in this application adds a first network layer to introduce position coding, thereby solving the technical problem of the related art losing spatial information during the encoding process.

[0090] Moreover, because the position coding is determined by the first network layer in the initial encoder, the position coding is learnable by adjusting the model parameters of the initial encoder during the model training process, that is, the position coding determined in this application is a relative position coding. In this way, during the model training process, for different three-dimensional object samples (such as different geometric shapes, details, etc.), that is, for different three groups of two-dimensional features, the initial encoder can continuously learn what kind of three groups of two-dimensional features and what kind of position coding to determine, so as to ensure that the determined prediction coding features can reconstruct this three-dimensional object sample. In this way, the ability to determine adaptive and accurate position coding for different three groups of two-dimensional features is learned, and the spatial information perception ability of the initial encoder is improved.

[0091] Compared with the method of assigning a unique code to each position in the related art (i.e., absolute position coding), there is no need to manually label the position codes corresponding to the three-dimensional and two-dimensional features of each three-dimensional object sample, which not only saves the cost of manual labeling, but also avoids manual labeling from bringing human experience into the model training process, thereby causing model learning errors. For example, if the position code labeled by human experience is wrong, it will lead to a deviation in the direction of model learning.

[0092] It should be noted that, for the network structure of the initial encoder, except for the aforementioned first network layer and second network layer, this application does not make any restrictions. For example, it is not limited to include other network layers (such as residual networks, etc.), and the number of first network layers and second network layers. In this regard, it can be flexibly set according to actual needs.

[0093] S304: Based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding features and the three-dimensional object sample, the initial encoder and the initial decoder are model trained to obtain the target encoder and the target decoder.

[0094] The initial decoder is used to decode the predicted coding features determined by the initial decoder for reconstruction. Correspondingly, the predicted three-dimensional object refers to the three-dimensional object reconstructed and restored by the initial decoder based on the predicted coding features, that is, the three-dimensional object indicated by the predicted coding features as considered by the initial decoder.

[0095] Accordingly, the difference between the predicted three-dimensional object and the three-dimensional object sample can reflect the deviation of the initial encoder when determining the predicted coding features, and the deviation of the initial decoder when reconstructing and restoring the three-dimensional object based on the predicted coding features. Therefore, model training is performed based on this difference, so that the initial encoder can learn to determine more accurate prediction coding features, specifically how to determine more accurate position coding suitable for the three-dimensional object sample through the first network layer, how to encode through the second network layer, etc., and the initial decoder can learn how to reconstruct and restore the three-dimensional object sample based on the predicted coding features.

[0096] In this way, with three-dimensional object samples as supervision and through model training, the final target encoder has better encoding ability, and can encode three groups of two-dimensional features into target coding features with rich spatial information. The target decoder has better decoding ability and can reconstruct and restore the three-dimensional object samples based on the target coding features, that is, it has the ability to determine the three-dimensional objects with precise details.

[0097] It should be noted that the initial decoder has a network structure corresponding to the initial encoder, so the network structure of the initial decoder can be set with reference to the network structure of the initial encoder, which will not be repeated here.

[0098] It is also because the initial decoder has a network structure corresponding to the initial encoder, and the initial encoder of the present application introduces position coding in the encoding process, so that the predicted coding features have spatial information. Accordingly, when the initial decoder decodes the predicted coding features, this information can guide the initial decoder to decode better. Specifically, it can guide the initial decoder to determine which features should represent which point on the three-dimensional object, which is conducive to determining the three-dimensional object with accurate details (such as the relative relationship between the parts of the three-dimensional object is accurate, the local details are accurate, etc.).

[0099] S305: Determine target coding features corresponding to the three-dimensional object samples through a target encoder according to the three groups of two-dimensional features.

[0100] S306: Perform model training on the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data corresponding to the three-dimensional object sample to obtain the target model.

[0101] In practical applications, the three-dimensional objects to be generated often do not have three-dimensional data, so it is difficult to directly use the trained target encoder in the application stage. Therefore, after training the target encoder and target decoder, the second part of model training is required, which is to use the target encoder to train the initial model to obtain the target model that can replace the target encoder in the application stage. Specifically:

[0102] First, the server can determine the corresponding target coding features through the target encoder according to the three-dimensional and two-dimensional features of the three-dimensional object sample. The target coding features refer to coding features with rich spatial information that can reconstruct and restore the three-dimensional object sample after inputting into the target decoder.

[0103] Next, the server can perform model training on the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data to obtain the target model. The noise-added coding feature can be obtained by performing noise processing on the target coding feature, that is, on the basis of the target coding feature, additional information is introduced to cause interference, thereby changing the target coding feature to a certain extent. In addition, the first control data can be used to describe a three-dimensional object sample, which can also be understood here as the first control data being used to describe the situation of a three-dimensional object that needs to be reconstructed and restored during the model training stage.

[0104] In the process of training the initial model based on the noise-added coding features and the first control data, the initial model can learn how to denoise the noise-added coding features according to the first control data, so that the denoised coding features corresponding to the noise-added coding features can be close to the target coding features. That is, in theory, after the denoised coding features are sent to the target decoder, the three-dimensional object samples can be reconstructed and restored. Based on this training, the target model can learn how to denoise according to the control data used to describe the requirements, so as to ensure that the features finally determined are also features with rich spatial information, that is, they are similar to the features determined by the target encoder, so that they can replace the target encoder in the application stage.

[0105] It should be noted that the present application does not impose any limitation on the setting of the initial model. For example, the initial model may be a generative model such as a diffusion model (DM). Since the diffusion model is more random in the denoising process, it is more imaginative and can support the guidance of multiple types of control data (such as at least one of a reference image, a description text, etc.), which is more conducive to generating a three-dimensional object that meets the requirements.

[0106] Furthermore, this application does not impose any limitation on the setting of control data. In practical applications, control data is used to express the requirements of three-dimensional objects. In the model training phase, control data expresses the requirements of three-dimensional object samples that need to be reconstructed and restored. In the model application phase, control data expresses the requirements of three-dimensional objects that need to be generated. Usually, control data is not the aforementioned three-dimensional data, but simpler data that can express the requirements. For example, control data can include at least one of a reference image and a description text.

[0107] Corresponding to the model training stage, the first control data is used to describe the three-dimensional object sample, and may include at least one of a first reference image corresponding to the three-dimensional object sample and a first description text corresponding to the three-dimensional object sample. The first reference image may be, for example, a two-dimensional image corresponding to the three-dimensional object sample (such as a front view of the three-dimensional object sample), or a two-dimensional image similar to the three-dimensional object sample (such as a similar style, similar geometric shape, etc.). In addition, the first description text may be used to describe the specific situation of the three-dimensional object sample, such as describing what the three-dimensional object sample is and what details it has.

[0108] In practical applications, if the first control data includes multiple types of data, they can be aligned to the same feature space before being fed into the initial model, such as by using a diffusion model to achieve alignment. Alignment is helpful in mining real descriptions from multiple types of data, which is more conducive to guiding the understanding of the initial model and then performing denoising.

[0109] S307: In response to the generation request, the initial noise data is denoised through the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and the target features are decoded through the target decoder to determine the three-dimensional object corresponding to the generation request.

[0110] Based on the solutions provided by S301-S306 above, the model used to generate the three-dimensional object can be determined through two-part model training, including the target model and the target decoder. It is precisely because position encoding is introduced in the model training process that the technical problem of losing spatial information in related technologies can be solved. Correspondingly, in the application stage, when a three-dimensional object needs to be generated, the target model and the target decoder can be used to quickly generate the required three-dimensional object. Specifically:

[0111] When a three-dimensional object needs to be generated, the user can initiate a generation request, which can be specifically inputting second control data for describing the three-dimensional object generation requirement. Accordingly, the server can respond to the generation request, perform denoising on the initial noise data according to the second control data in the generation request through the aforementioned target model to determine the target feature corresponding to the second control data, and decode the target feature through the target decoder to determine the three-dimensional object corresponding to the generation request.

[0112] Among them, the second control data can be used to describe the requirements for generating a three-dimensional object, specifically, the situation of the three-dimensional object to be generated (such as what the three-dimensional object to be generated is like, what characteristics it has, etc.), and the target features determined based on the second control data can be used to indicate the three-dimensional object to be generated. Because the target model is obtained by model training using the target encoder, the target features determined using the target model are still features with rich spatial information. And the target decoder can better decode the target features with rich spatial information, so it can determine the three-dimensional object that meets the generation requirements described by the second control room data and has more precise details for the generation request.

[0113] It is understandable that S307 may be executed only when a three-dimensional object needs to be generated, that is, when a three-dimensional object does not need to be generated, S307 may not be implemented.

[0114] It should be noted that the present application does not impose any limitation on the setting of the second control data. Similar to the aforementioned description of the first control data, the first control data corresponding to the model application stage may also include at least one of a second reference image for describing the generation requirements of the three-dimensional object and a second description text for describing the generation requirements of the three-dimensional object. Among them, the second reference image can be used to describe the style, geometry, etc. of the three-dimensional object to be generated, and the second description text can be used to describe what the three-dimensional object to be generated is, what characteristics it has, etc. Similar to the previous first control data, if the second control data includes multiple types of data, it is aligned to the same feature space before being sent to the target model, such as using the aforementioned diffusion large model to achieve alignment. Based on this, it is beneficial to mine the real generation requirements from multiple types of data, which is more conducive to guiding the target model to understand the generation requirements, and then perform denoising processing, etc.

[0115] As an example, the second control data may be a second description text. Specifically, for example, it may be Figure 2The description text "a girl with long black hair, standing with arms outstretched, wearing a long polka dot puffy dress with a bow tie at the neckline and wavy hems" in the example is used to guide the target model to perform denoising on the initial noise data based on the second control data to determine the target features that can be used to generate a three-dimensional object that meets the description of the second control data. Figure 2 For example, the finally generated three-dimensional object A meets the generation requirements described by the second control data A, and the details (such as a bow tie at the neckline, wavy patterns at the hem, etc.) are also accurate. Figure 5 In the example, the description text "long-haired boy, standing in T-pose, wearing a twill belt and high-top sneakers" corresponds to the generated three-dimensional object B as follows: Figure 5 As shown, it meets the generation requirements described by the second control data B and has precise details (such as a striped belt).

[0116] As another example, the second control data may be a second reference image. Figure 6 In the example, the second control data C may be a two-dimensional reference image. The correspondingly generated three-dimensional object C is as follows: Figure 6 As shown, it meets the requirements and has accurate details (such as standing posture, tie, plaid vest, etc.).

[0117] Of course, the description text and the reference image can also be used as control data together to more precisely describe the three-dimensional object generation requirements, which is also conducive to guiding the target model to perform denoising processing to determine a more accurate target object, thereby improving the accuracy of the generated three-dimensional object. This will not be elaborated here.

[0118] It should be noted that, in addition to no restrictions on the type of control data, there is no restriction on the amount of control data, such as multiple description texts and multiple reference images. For example, in the actual application stage, when multiple three-dimensional objects need to be generated for the same game project, an image X can be set as a style reference. When generating each three-dimensional object, in addition to the reference image used to describe the three-dimensional object, the image X used as a style reference can also be used as input. Based on this, the image X jointly guides the generation process of the multiple three-dimensional objects, specifically guides the style of the multiple three-dimensional objects, so it is conducive to generating multiple three-dimensional objects with a unified style.

[0119] It should also be noted that in actual applications, the user can also input the second control data through the terminal to initiate a generation request. Also, the generated three-dimensional object (such as the aforementioned three-dimensional object A, three-dimensional object B, three-dimensional object C, etc.) can be viewed through the terminal, and various operations such as dragging, rotating, and scaling can be performed to flexibly view different faces, local details, etc. of the three-dimensional object. In addition, the generated three-dimensional object can also be loaded into other applications (such as design software, texture mapping software, etc.), and combined with other applications, the generated three-dimensional object can be fine-tuned (such as adjusting local details, performing texture mapping, etc.) to obtain a three-dimensional object with richer details and more in line with the requirements.

[0120] It can be seen from the above technical solution that in the first part of the model training, the position codes corresponding to the three groups of two-dimensional features are first determined by the first network layer in the initial encoder, and then subsequent processing is performed according to the three groups of two-dimensional features and their respective position codes. Among them, the position code is used to indicate the feature relationship between the two-dimensional features. Since the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other, so that the three-dimensional data of any point on the three-dimensional object sample is split into three groups of two-dimensional features, so the feature relationship between the three groups of two-dimensional features can reflect the absolute spatial information of the point in the three-dimensional object sample. Similarly, the feature relationship between the three groups of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically includes the relative relationship between the two two-dimensional features in the same feature plane in the feature plane, and the relative relationship between the two two-dimensional features in different feature planes, so it can also reflect the relative spatial information of the two points in the three-dimensional object sample. Based on this, after using three groups of two-dimensional features to indicate the three-dimensional data and introducing the position code, the spatial information carried by the three-dimensional data can be reflected based on the rich relationship between the features. And the training is supervised by three-dimensional object samples, so the initial encoder can learn the rich spatial information in the three-dimensional object samples to obtain prediction coding features with rich spatial information, and the initial decoder can learn how to decode the prediction coding features to determine the predicted three-dimensional objects with more precise details. The target encoder finally obtained has better encoding ability and can encode three sets of two-dimensional features into target coding features with rich spatial information. The target decoder has better decoding ability to determine the three-dimensional objects with more precise details. In practical applications, the three-dimensional objects that need to be generated often do not have three-dimensional data, so it is difficult to directly use the trained target encoder in the application stage. For this reason, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model to obtain the target model to replace the target encoder in the application stage.

[0121] Correspondingly, in the application stage, in response to the generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and then the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request. Because the target model is obtained by model training using the target encoder, the target features determined using the target model are still features with rich spatial information, and the target decoder can better decode the target features with rich spatial information, so it can determine a three-dimensional object with more precise details for the generation request. Based on this, the technical problem in the related art that the trained model cannot generate a three-dimensional object with precise details due to the loss of spatial information during the encoding process is solved.

[0122] Through the above embodiments, the model determination method provided by the present application is illustrated. It should also be noted that the present application does not make any limitation on the network structure of the model, how to determine the position encoding, how to encode, and how to train the model. For a more comprehensive understanding, the present application will be introduced one by one through the following embodiments.

[0123] (I) Regarding the network structure of the initial encoder, the present application embodiment provides the following example description:

[0124] For the initial encoder, from the perspective of network structure, it includes the aforementioned first network layer and second network layer. In addition, this application does not make any restrictions. For example, it does not limit the number of other network layers, the first network layer, and the second network layer. In this regard, it can be flexibly set according to actual needs.

[0125] In order to enhance the encoding ability of the encoder and to determine more accurate encoding features, in a possible implementation method, multiple first network layers and multiple second network layers can be added to the initial encoder to fully learn the information carried in the input data (i.e., three sets of two-dimensional features) during the model training stage, thereby improving the encoding ability to determine more accurate encoding features.

[0126] In a specific implementation, the aforementioned initial encoder may include M groups of networks connected in sequence, in which each group of networks in the first M-1 groups of networks may include a first network layer, a second network layer, and a third network layer connected in sequence, and the Mth group of networks may include a first network layer and a second network layer connected in sequence. Wherein, M is a positive integer greater than 1.

[0127] It is understandable that when the network structure of the initial encoder is different, how to encode and other methods will also be different. In this embodiment, in one group of networks, the position coding is first determined, then encoded, and then sent to the next group of networks, and the position coding continues to be determined, and then encoded, until the last group of networks completes the position coding and encoding. That is, it is processed in sequence through M groups of networks. In practical applications, since the features obtained by encoding are mostly one-dimensional features, in order to facilitate the determination of position coding, etc. by the next group of networks, the aforementioned third network layer is added to each group of networks in the first M-1 groups of networks, and its function is to convert the coding features obtained by the second network layer, and then send it to the next group of networks for processing. For ease of understanding, the following specific explanation is given:

[0128] That is, when the aforementioned S302 is specifically implemented, the server can determine the position codes corresponding to the three groups of two-dimensional features through the first network layer in the first group of networks in the M groups of networks according to the three groups of two-dimensional features. And, when the aforementioned S303 is specifically implemented, the server can encode the three groups of two-dimensional features and the position codes corresponding to the three groups of two-dimensional features through the second network layer in the first group of networks to obtain the predicted coding features corresponding to the first group of networks. Then, the server can convert the predicted coding features corresponding to the first group of networks according to the third network layer in the first group of networks to obtain three groups of two-dimensional coding features corresponding to the first group of networks.

[0129] Based on this, the processing of the first group of networks is completed. Then, it is sent to the next group of networks for processing. In this regard, in the specific implementation:

[0130] The server can determine the position codes corresponding to the three groups of two-dimensional coding features corresponding to the i-1th group of networks through the first network layer in the i-th group of networks based on the three groups of two-dimensional coding features corresponding to the i-1th group of networks in the M groups of networks. Wherein, i is an integer greater than 1 and less than or equal to M. And, the server can encode the three groups of two-dimensional coding features corresponding to the i-1th group of networks and the three groups of two-dimensional coding features corresponding to the i-1th group of networks through the second network layer in the i-th group of networks to obtain the predicted coding features corresponding to the i-th group of networks. And, if the i-th group of networks includes a third network layer, the predicted coding features corresponding to the i-th group of networks are converted according to the third network layer in the i-th group of networks to obtain the three groups of two-dimensional coding features corresponding to the i-th group of networks.

[0131] Based on this, the three groups of two-dimensional features of the three-dimensional object sample are processed in turn through M groups of networks, which is conducive to the initial encoder to better capture the information carried in the three groups of two-dimensional features and perform deep learning. In addition, the position encoding is re-determined in each group of networks, so that in the deep learning process, the spatial information carried in the three groups of two-dimensional features can be better captured. Ultimately, it is conducive to forming more accurate coding features. Correspondingly, the predicted coding features corresponding to the three-dimensional object sample can be determined based on the predicted coding features corresponding to the Mth group of networks.

[0132] Among them, it should be noted that the present application does not make any limitation on how to determine the final prediction coding features of the three-dimensional object sample according to the prediction coding features corresponding to the Mth group of networks. It is understandable that this is related to the network structure, such as whether the Mth group of networks includes the aforementioned third network layer, and whether other networks (such as output layers, etc.) are included in the initial encoder after the Mth group of networks. For ease of understanding, the present application embodiment provides the following example description:

[0133] In some embodiments, the Mth group of networks may not include the aforementioned third network layer, and accordingly, the predicted coding features corresponding to the three-dimensional object samples are the predicted coding features corresponding to the Mth group of networks. Based on this, it is helpful to simplify the model structure.

[0134] In some other embodiments, the structure of the Mth group of networks is the same as that of the first M-1 groups of networks, that is, it also includes a third network layer. Accordingly, the prediction coding features corresponding to the Mth group of networks can be transformed according to the third network layer in the Mth group of networks to obtain three groups of two-dimensional prediction coding features corresponding to the Mth group of networks. At this time, the prediction coding features corresponding to the three-dimensional object samples are the three groups of two-dimensional prediction coding features corresponding to the Mth group of networks. Based on this, the prediction coding features finally obtained are three groups of two-dimensional features, which can express three-dimensional spatial information, so they can explicitly inform the initial decoder of the spatial information in the prediction coding features, thereby guiding the initial decoder to perform better decoding. In particular, compared with the encoding method of directly encoding three-dimensional data into one-dimensional features in the related art, the present application can better retain the three-dimensional spatial information, thereby performing better model training and improving model capabilities. In other words, in the present application, the expression form of three groups of two-dimensional features is always maintained during the encoding and decoding process.

[0135] In addition, other network layers can be set as needed. For example, a residual network (ResNet) is set before the first group of networks, which is helpful to alleviate problems such as gradient disappearance during model training, thereby ensuring the stability of model training and accelerating training. In practical applications, the residual network can be composed of multiple stacked residual blocks (Res Block).

[0136] It should also be noted that the above embodiment is described by taking the encoder as an example. Since decoding is the inverse process of encoding, the decoder has a network structure corresponding to the encoder. Therefore, after determining the network structure of the initial encoder, the network structure of the initial decoder can be adaptively set, and steps such as how to process by the initial decoder can all refer to the above descriptions, which will not be repeated here.

[0137] In practical applications, a generative model such as a variational auto encoder (VAE) can be used to construct the aforementioned initial encoder and initial decoder. Specifically, it consists of an encoder and a decoder, and the model structure is U-shaped.

[0138] For ease of understanding, the present application also provides Figure 7 A schematic diagram of the network structure of an initial encoder and an initial decoder is shown in Figure 7 In the example, taking the aforementioned M=3 as an example, it can be seen that the initial encoder encodes the three sets of two-dimensional features corresponding to the three-dimensional object sample to obtain the corresponding predicted coding features, and then the initial decoder processes the predicted coding features to obtain the corresponding three sets of two-dimensional features after decoding, and the three sets of two-dimensional features after decoding are used to reconstruct and restore to obtain the aforementioned predicted three-dimensional object.

[0139] It should also be noted that this application does not impose any restrictions on the specific settings of each network layer. Generally, each network layer can be set based on a transformer, such as setting the transformer as the aforementioned second network layer. Correspondingly, when multiple groups of networks are included in the initial encoder, the initial encoder can also be considered to be composed of multiple stacked transformer blocks.

[0140] (II) Regarding how to perform encoding, the present application embodiment provides the following example description:

[0141] It can be seen from the above embodiments that how to perform encoding is related to the network structure of the encoder.

[0142] In addition, in order to better understand the encoding process, the present application embodiment also provides the following example description:

[0143] In practical applications, in order to facilitate the initial encoder (specifically the second network layer in the initial encoder) to better capture the feature relationship between features based on position encoding, in a possible implementation method, attention can be used for encoding. In the attention mechanism, any feature can be calculated separately from other features, which is conducive to better learning the relationship between features.

[0144] Accordingly, in the specific implementation of the aforementioned S303, for each group of two-dimensional features in the three groups of two-dimensional features, the server can respectively splice the two-dimensional features and the position codes corresponding to the two-dimensional features to obtain the spliced ​​features corresponding to the three groups of two-dimensional features. Then, the server can encode based on the attention mechanism according to the spliced ​​features corresponding to the three groups of two-dimensional features through the second network layer in the initial encoder to obtain the predicted coding features corresponding to the three-dimensional object samples. Among them, the process of encoding based on the attention mechanism is used to indicate that the spliced ​​features belonging to the first feature plane are dot-producted with the spliced ​​features belonging to the second feature plane, and the first feature plane and the second feature plane are any feature planes among the aforementioned three feature planes, that is, the two can be the same feature plane or different feature planes.

[0145] Based on this, the position code is concatenated with the original feature and sent to the second network layer as a whole. The code is encoded based on the attention mechanism, so that the dot product can be calculated for any two features (including the dot product between two features in the same feature plane and the dot product between two features in different feature planes), thereby better capturing the relationship between features.

[0146] Moreover, the present application uses three sets of two-dimensional features to express three-dimensional data, in short, it uses three feature planes to express three-dimensional data. Therefore, this method can simultaneously take into account the relative position relationship within the same feature plane and the relative position relationship between different feature planes, thereby helping the initial encoder to better learn the spatial information carried by the three-dimensional data and obtain more accurate prediction coding features.

[0147] Since in this application, the input of the initial encoder is three groups of two-dimensional features, in short, three feature planes, and the relative relationship between features in the same feature plane is stronger than the relative relationship between features in different feature planes, and the three feature planes are perpendicular to each other. Therefore, in order to facilitate the initial encoder to weaken the relative relationship between different feature planes when capturing spatial information based on position coding, thereby accelerating the learning process of the initial encoder, in a possible implementation method, the position coding corresponding to the three groups of two-dimensional features can satisfy an orthogonal relationship.

[0148] Correspondingly, in the process of encoding based on the attention mechanism, that is, in the process of performing dot product, the orthogonal relationship can be used to indicate that if the first feature plane and the second feature plane are different feature planes, the dot product result between the position encoding in the spliced ​​feature belonging to the first feature plane and the position encoding in the spliced ​​feature belonging to the second feature plane is zero.

[0149] Based on this, by setting the position codes corresponding to the three groups of two-dimensional features to satisfy an orthogonal relationship, the calculation between features in different feature planes can be weakened during the encoding process based on the attention mechanism. This helps the initial encoder to capture more accurate spatial information better and faster, thereby accelerating the model training of the initial encoder.

[0150] Usually, in order to further accelerate model training, when weakening the correlation between different feature planes, the correlation within the same feature plane can also be enhanced. That is, in the process of performing dot product, if the first feature plane and the second feature plane are the same feature plane, the dot product result between the position code in the spliced ​​feature belonging to the first feature plane and the position code in the spliced ​​feature belonging to the second feature plane is greater than zero. Thus, in model training, the goal of enhancing the correlation within the same feature plane and weakening the correlation between different feature planes is achieved at the same time, which is conducive to further accelerating model training.

[0151] In order to better understand the encoding process, the present application embodiment also provides the following example description:

[0152] On the one hand, in a specific implementation, the aforementioned splicing feature can be expressed by the following formula:

[0153] L ′ 1 =L 1 +P 1 ,L ′ 2 =L 2 +P 2 ,L ′ 3 =L 3 +P 3

[0154] In the above formula, L 1 Used to represent the first set of two-dimensional features, P 1 It is used to represent the position code corresponding to the first set of two-dimensional features, L ′ 1 Used to represent the corresponding splicing features; similarly, L 2 Used to represent the first set of two-dimensional features, P 2 It is used to represent the position code corresponding to the first set of two-dimensional features, L ′ 2 Used to represent the corresponding splicing features; and, L 3 Used to represent the first set of two-dimensional features, P 3 It is used to represent the position code corresponding to the first set of two-dimensional features, L ′ 3 Used to indicate the corresponding splicing features.

[0155] Correspondingly, in the process of performing dot product, the dot product result between two concatenated features may include four items: the dot product result between two original features, the dot product result between one original feature and the position encoding of another original feature, and the dot product result between two position encodings.

[0156] On the other hand, the dot product result between two position codes can be expressed by the following formula:

[0157] P m [i,:]P n [j,:]=0,m≠n

[0158] P m [i,:]P n [j,:]>0,m=n

[0159] In the above formula, m and n can be any number between 1, 2, and 3, which is used to represent one of the three characteristic planes, and P m [i,:] is used to represent the position code corresponding to the feature [i,:] in the mth feature plane, P n [j,:] is used to represent the position code corresponding to the feature [j,:] in the nth feature plane, P m [i,:]P n [j,:] is used to represent the dot product result between these two position encodings.

[0160] It can be seen from the above formula that the dot product result between the position codes of any two features belonging to the same feature plane is zero, and the dot product result between the position codes of any two features belonging to different feature planes is greater than zero.

[0161] It should also be noted that this application does not make any limitation on the setting of the attention mechanism. In practical applications, the attention mechanism (Attention) may include a self-attention mechanism (Self-Attention), a linear attention mechanism (Linear Attention), etc. For ease of understanding, the embodiments of this application provide the following example description:

[0162] Taking the Transformer structure mentioned above as the second network layer as an example, based on the self-attention mechanism, the input data will be calculated into query (Q), key (K) and value (V) matrices, and the calculation results will be expressed as follows:

[0163]

[0164] In the above formula, Attention(Q,K,V) is used to represent the result of calculation based on the self-attention mechanism, that is, it can be used to represent the aforementioned predictive coding features, softmax refers to the normalized exponential function, Q, K, and V represent three matrices, which are the matrices obtained by mapping the concatenated features corresponding to the input data, namely the three groups of two-dimensional features mentioned above, and T is used to represent the matrix transpose.

[0165] And, the plane size of each of the three feature planes mentioned above can be recorded as N, and each feature ∈R in the three sets of two-dimensional features N*d For example, corresponding to the above formula, Among them, d k and d v Used to represent the size of the matrix obtained after mapping based on the attention mechanism.

[0166] Corresponding to this attention mechanism, QK T The complexity of matrix multiplication increases quadratically with the increase of N. Therefore, when N is relatively large, that is, when the resolution of the three sets of two-dimensional features input is relatively high, the computational difficulty and cost will increase, and more computing resources will be consumed.

[0167] In this regard, in another embodiment, the aforementioned attention mechanism can also be a linear attention mechanism, especially when N is relatively large, in order to reduce the computational complexity, save computing resources, improve coding efficiency, etc. In this regard, the present application embodiment provides the following example description:

[0168] In practical applications, the linear attention mechanism can express the calculation results in the following way:

[0169]

[0170] In the above formula, LinearAttention(Q,K,V) is used to represent the result of calculation based on the linear attention mechanism, that is, it can be used to represent the aforementioned predictive coding features, and the meanings of Q, K, V and T are the same as above. And, refers to the matrix after Q is mapped, It refers to the matrix after K is mapped. It refers to the matrix after V is mapped.

[0171] In the linear attention mechanism, we first calculate This avoids calculating the complete N×N matrix. Therefore, while ensuring the computational effect, the computational complexity of the attention mechanism is reduced to linear, which can process long sequence data more efficiently, that is, it is more suitable for situations where N is relatively large.

[0172] In practical applications, time complexity can be used to describe the relationship between the time required for the algorithm to execute and the size of the input data. In other words, time complexity can be used to intuitively reflect computational complexity. Usually, time complexity is expressed using Big O notation, that is, O(f(n)), where n is the size of the input data, and f(n) is the relationship function between the algorithm running time and the size of the input data. In the above example, from the self-attention mechanism to the linear attention mechanism, the time complexity increases from O(N 2 ) is reduced to O(N).

[0173] In this regard, the present application also provides the following embodiments: Figure 8 The comparison diagram of two different attention mechanisms is shown in FIG. Accordingly, in practical applications, the appropriate attention mechanism can be flexibly selected according to the actual situation. For example, when the resolution of the three sets of input two-dimensional features is relatively high and N is relatively large, a linear attention mechanism can be used to achieve the purpose of generating more refined three-dimensional objects by inputting three sets of high-resolution two-dimensional features on the basis of reducing computational complexity and computing resource consumption, thereby solving the problem of efficient encoding and decoding of high-resolution three-dimensional object data.

[0174] The encoding process is described in detail through the above embodiments. It should also be noted that decoding is the inverse process of encoding, so for decoding, reference can be made to the description of the above embodiments for encoding, which will not be repeated here.

[0175] (III) Regarding how to determine the position code, the present application embodiment provides the following example description:

[0176] For ease of understanding, in this embodiment, taking the above-mentioned three groups of two-dimensional features respectively corresponding to the position codes satisfying an orthogonal relationship as an example, the following example description is provided:

[0177] On the one hand, since the orthogonal relationship is used to indicate that when two features belong to different feature planes, the dot product result between the position codes of the two features is zero. In order to quickly determine the position code that satisfies the orthogonal relationship, in a possible implementation, the position code required by the present application can be determined by using the property that the matrix elements that satisfy the orthogonal relationship can be multiplied to zero.

[0178] On the other hand, since the position encoding is learnable during the model training process, and the essence of model training is to adjust the model parameters, the matrix can be constructed according to the model parameters of the first network layer, and then the position encoding can be determined. In this way, the value of the model parameters can be directly reflected in the position encoding, and accordingly, during the model training process, the position encoding can be directly updated by adjusting the model parameters.

[0179] For a better understanding, the embodiment of the present application provides the following specific description for determining the position encoding method by means of matrix properties and in combination with the model parameters of the first network layer:

[0180] If the model parameters of the first network layer may include the first parameter, then in the specific implementation of the aforementioned S302, the server may construct an initial matrix through the first network layer according to the three groups of two-dimensional features and the first parameter, in which the first column matrix elements and the second column matrix elements satisfy an orthogonal relationship, and the first column matrix elements and the second column matrix elements are two different columns of matrix elements in the initial matrix. Therefore, the server may determine the position codes corresponding to the three groups of two-dimensional features according to the initial matrix to ensure that the position codes corresponding to the three groups of two-dimensional features satisfy an orthogonal relationship.

[0181] Correspondingly, during the model training process, by adjusting the first parameter, the initial matrix can be changed to determine a new position encoding. Finally, at the end of the model training, the first parameter of the first network layer in the target encoder can determine the precise position encoding.

[0182] It should be noted that the present application does not impose any limitation on how to determine the position encoding method according to the initial matrix. For ease of understanding, the present application provides the following example description:

[0183] Based on the above method of determining the position code, it is possible to ensure that the dot product result between the position codes of two features belonging to different feature planes is zero, so that the initial encoder can capture the relative relationship between any two features belonging to different feature planes. In practical applications, a feature plane may include multiple two-dimensional features, so in order to facilitate the initial encoder to better capture the relative relationship between any two features belonging to the same feature plane, the corresponding position code can also be determined for each feature.

[0184] In this regard, in a possible implementation, the server can determine the target matrices corresponding to the three feature planes according to the initial matrix and the plane sizes corresponding to the three feature planes. For each of the three feature planes, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane. In the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of a two-dimensional feature in the feature plane. Furthermore, the target matrices corresponding to the three feature planes can be determined as the position codes corresponding to the three groups of two-dimensional features.

[0185] Based on this, each two-dimensional feature in each group of two-dimensional features has a corresponding position encoding, so that in the process of encoding based on the attention mechanism, the initial encoder can better capture the relative relationship between any two features belonging to the same feature plane.

[0186] For better understanding, in the embodiment of the present application, the plane size corresponding to the three characteristic planes mentioned above is N, and each feature ∈R in the three sets of two-dimensional features is N*d For example, the following example is provided to illustrate the specific method of determining the position code:

[0187] Since the plane sizes of the three feature planes are the same, in order to facilitate the determination of the position coding, the model parameters in the first network layer can be used to represent the plane size. That is, the model parameters of the first network layer can also include a second parameter. Among them, the first parameter ∈ R d*1 , and is not zero, so the constructed initial matrix ∈R d*d , and the second parameter ∈R 1*N , based on this, in order to construct ∈R N*d The position code.

[0188] In the specific implementation, since any two columns of matrix elements in the initial matrix satisfy an orthogonal relationship, the three columns of matrix elements in the initial matrix can be firstly determined as three first matrices, each of which is ∈R d*1 Then, the three first matrices can be multiplied by the second matrix indicated by the second parameter to obtain three undetermined matrices, the second matrix ∈ R 1*N , each of the three undetermined matrices ∈ R d*N Finally, the three undetermined matrices can be transposed to obtain the target matrices corresponding to the three feature planes. Each target matrix ∈ R N*d Among them, the target matrices corresponding to the three feature planes are the position codes corresponding to the three groups of two-dimensional features.

[0189] Based on this, through simple matrix operations, the position coding that meets the requirements can be determined. And when the three feature planes are the same size, the second parameter in the first network layer is directly used to express the plane size and participate in the process of determining the target matrix (i.e. determining the position coding), so that during the model training process, the updated position coding can be directly obtained by adjusting the second parameter. Correspondingly, after the model training is completed, the first parameter and the second parameter of the second network layer in the target encoder can determine the precise position coding.

[0190] In the above embodiments, it should be noted that the present application does not impose any limitation on how to construct the initial matrix. For ease of understanding, the present application provides the following examples:

[0191] Since the initial matrix needs to satisfy the aforementioned orthogonal relationship between any two columns of matrix elements, in a possible implementation, a householder matrix can be constructed as the aforementioned initial matrix. Specifically, the householder matrix can be denoted as H and expressed by the following formula:

[0192]

[0193] In the above formula, H is used to represent the householder matrix, which is the initial matrix mentioned above, and h is used to represent the first parameter mentioned above, h∈R d*1 , and is not zero, T is used to represent matrix transpose, and I is used to represent the identity matrix.

[0194] From the properties of the householder matrix, we can know that if we take the first three columns of the matrix as the three first matrices mentioned above, we can record them as h 1 ,h 2 ,h 3 ∈R d*1 , to ensure 1 ,h 2 ,h 3 are mutually orthogonal.

[0195] Next, h 1 ,h 2 ,h 3 and another learnable network parameter (i.e. the second parameter mentioned above) t∈R 1*N Perform matrix multiplication and then transpose the matrix obtained by multiplication to obtain three target matrices P 1 , P 2 , P 3 ∈R N*d , used to represent the position codes corresponding to the three sets of two-dimensional features. 1 ,h 2 ,h 3 are mutually orthogonal, so it can ensure that P 1 , P 2 , P 3 They are also mutually orthogonal.

[0196] Based on this, the model parameters of the first network layer (such as the first and second parameters mentioned above) are directly used to determine the target matrix (i.e., the position encoding), so that during the model training process, the determined position encoding is continuously adjusted, that is, the aforementioned position encoding is learnable, that is, the position encoding determined in this application is a relative position encoding.

[0197] Compared with the method of assigning a unique code to each position in the related art (i.e., absolute position coding), on the one hand, there is no need to manually assign position codes to each position, which reduces costs and avoids the introduction of artificial subjective bias; on the other hand, it can adaptively determine the precise position code for each three-dimensional object sample, so that the encoder can better learn how to determine the position code from a variety of three-dimensional object samples, so as to determine the ability of accurate coding features, that is, the initial encoder can better perceive spatial information, so that the target coding features obtained by encoding can maintain the relative spatial relationship between the three groups of two-dimensional features of the three-dimensional object sample.

[0198] (IV) Regarding how to perform model training, the present application embodiment provides the following example description:

[0199] 1. For the first part of model training, that is, model training for the initial encoder and initial decoder, the following example is provided:

[0200] In a possible implementation, the aforementioned S304 can first determine the training loss based on the difference between the predicted three-dimensional object and the three-dimensional object sample during implementation. The training loss can be used to indicate the deviation of the initial encoder when determining the predicted coding features, and the deviation of the initial decoder when reconstructing and restoring the three-dimensional object based on the predicted coding features. Accordingly, the initial encoder and the initial decoder can be model trained according to the training loss. Exemplarily, during the model training process, when the training loss meets the training end condition (i.e., the indicated deviation is within an acceptable range), the model training can be terminated, and the corresponding target encoder and target decoder can be obtained.

[0201] In order to improve the efficiency of model training, in another possible implementation, if the position codes corresponding to the three sets of two-dimensional features respectively satisfy an orthogonal relationship, the need to satisfy the orthogonal relationship can also be used as an optimization target for model training. That is, in the specific implementation of the aforementioned S304, first, the server can determine the training loss based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding features and the three-dimensional object sample. Then, the server can adjust the model parameters of the initial encoder and the initial decoder according to the optimization target and the training loss to obtain the target encoder and the target decoder.

[0202] Among them, the training loss can be used to indicate the deviation of the initial encoder when determining the predicted coding features, and the deviation of the initial decoder when reconstructing and restoring the three-dimensional object based on the predicted coding features, and the optimization target can be used to indicate that the position codes corresponding to the three groups of two-dimensional features determined by the first network layer after the model parameter adjustment are still orthogonal.

[0203] Based on this, the need to satisfy the orthogonal relationship is used as a constraint condition for adjusting the model parameters, so that based on the adjusted model parameters, the position encoding that satisfies the orthogonal relationship can still be determined. Based on this, invalid parameter adjustments that need to be readjusted when the position encoding determined based on the adjusted model parameters does not satisfy the orthogonal relationship are avoided, which is conducive to improving the model training efficiency.

[0204] And, combined with the above description, it can be seen that in the model training process of this application, the position encoding is learnable. In the method of model training based on training loss and optimization target, when the initial encoder learns how to determine the position encoding, it can adjust the model parameters of the first network layer according to the size of the training loss, and can avoid unnecessary parameter adjustments based on the optimization target as a constraint condition, ensuring that the position encoding is learnable and always maintains an orthogonal relationship.

[0205] 2. For the second part of model training, that is, model training of the initial model, the following example description is provided:

[0206] Generally, the process of adding noise (i.e., adding noise) can be defined as forward propagation, and the process of denoising can be considered as back propagation. In the present application, the initial model is used to denoise the noisy coding features corresponding to the target coding features, and the target model obtained by model training is used to denoise the initial noise data. Therefore, the process of model training the initial model can be considered as the inverse process of the initial model learning forward propagation, that is, the process of learning how to denoise.

[0207] In practical applications, the process of back propagation (i.e., the process of denoising) can be used to predict how much noise is added during the forward propagation process, and can also be used to predict the features before the forward propagation (i.e., without adding noise). It is understandable that different prediction methods indicate that the knowledge and abilities learned by the initial model are different, which is related to the way the model is trained. In order to better understand, corresponding to these two different prediction methods, the embodiments of the present application provide the following examples:

[0208] (1) In order to obtain a target model that can predict how much noise is added, the present application embodiment provides the following model training method example:

[0209] First, in the specific implementation of the aforementioned S306, the server may perform denoising on the noisy coding feature according to the first control data through the initial model to obtain the predicted noise corresponding to the noisy coding feature, that is, the predicted noise may refer to the predicted value of the noise added to the noisy coding feature determined by the initial model. Then, the server may perform model training on the initial model according to the difference between the predicted noise corresponding to the noisy coding feature and the sample noise to obtain the target model.

[0210] The sample noise is the noise added by the noise-added coding feature compared to the target coding feature, that is, the true value of the noise added during the noise-adding process. Therefore, the difference between the two can indicate the deviation of the initial model in predicting noise. Based on this, the model training is carried out so that the initial model can learn how to accurately predict the noise added during the noise-adding process.

[0211] It can be seen that in this training method, the target coding features determined by the trained target encoder are used as indirect supervision to model the initial model so that the initial model learns how to perform denoising in order to accurately predict the noise. Since the first control data is used to describe the three-dimensional object sample and the target coding features are also features corresponding to the three-dimensional object samples, when the initial model can accurately predict the noise, it is considered that the initial model has learned how to perform denoising under the guidance of the first control data. If the predicted noise is subtracted from the noisy coding features, the target coding features can be restored. Therefore, it can be considered that the initial model has learned how to perform denoising based on the control data, and then combined with the difference method, it can determine the features that meet the control data.

[0212] Correspondingly, in the application stage, that is, when the aforementioned S307 is specifically implemented, the server can respond to the generation request and perform denoising on the initial noise data according to the second control data through the target model to obtain the predicted noise corresponding to the initial noise data. And, the server can determine the difference between the initial noise data and the predicted noise corresponding to the initial noise data as the target feature. Among them, the target feature has a consistent corresponding relationship with the input second control data and meets the generation requirements described by the second control data, so that the three-dimensional object that meets the generation requirements can be determined based on the target feature later. Based on this, the initial model and the target model only need to predict the noise, which is conducive to simplifying the model structure.

[0213] It should be noted that the present application does not impose any limitation on how to perform the noise addition process to obtain the noisy coding feature. In practical applications, the target coding feature may be continuously added with noise (such as adding Gaussian noise) at a certain noise addition intensity to obtain the noisy coding feature after the noise addition. It is understandable that the difference between the noisy coding feature and the target coding feature is different with different noise addition intensities. Generally, the greater the noise addition intensity, the greater the difference between the two.

[0214] In order to improve the model training efficiency, during the model training phase, the noise intensity can also be input into the initial model to explicitly inform the initial model of the noise intensity at which the currently input noise coding feature is obtained after noise processing. Based on this, the initial model is guided to perform more reasonable denoising to avoid too little or too much denoising, which is conducive to accelerating model training and improving model training efficiency.

[0215] And, corresponding to the first model training method, in the specific implementation, the training loss of the initial model can be determined by the following formula:

[0216] L diff =||ε θ -ε 0 || 2

[0217] In the above formula, L diff Used to represent the training loss of the initial model, ε θ Used to represent the prediction noise corresponding to the noise-added coding feature, ε 0 Used to represent sample noise, |||| 2 It is used to represent mean square error.

[0218] (2) In order to obtain a target model that can predict the features before noise addition, the present application embodiment provides the following model training method examples:

[0219] First, in the specific implementation of the aforementioned S306, the server may perform denoising on the noisy coding feature according to the first control data through the initial model to obtain the denoised feature corresponding to the noisy coding feature. Also, the server may perform model training on the initial model according to the difference between the denoised feature corresponding to the noisy coding feature and the target coding feature to obtain the target model.

[0220] The denoising feature corresponding to the noisy coding feature may refer to the feature determined by the initial model before the denoising process, and the difference between the feature and the target coding feature may indicate the deviation of the initial model in predicting the feature before the denoising process. Based on this, the model training is performed so that the initial model can learn how to accurately predict the feature before the denoising process.

[0221] Correspondingly, in the application stage, that is, when the aforementioned S307 is specifically implemented, the server can respond to the generation request and denoise the initial noise data according to the second control data through the target model to obtain the denoising features corresponding to the initial noise data, and the server can determine the denoising features corresponding to the initial noise data as the target features.

[0222] It can be seen that in this training method, the target encoding features determined by the trained target encoder are used as direct supervision to model the initial model, so that the initial model learns how to perform denoising and directly determine the accurate target features. Based on this, the target model can be used to directly determine the required target features, and the processing steps are simpler.

[0223] And, corresponding to the second model training method, in the specific implementation, the training loss of the initial model can be determined by the following formula:

[0224] L diff =||z θ -z 0 || 2

[0225] In the above formula, L diff Used to represent the training loss of the initial model, z θ It is used to represent the denoising feature corresponding to the denoising coded feature, z 0 Used to represent the target encoding features, |||| 2 It is used to represent mean square error.

[0226] Through the above embodiments, the model determination method provided by the present application is described in detail. Overall, it includes an initial encoder, an initial decoder and an initial model, and after the model training is completed, it includes a target encoder, a target decoder and a target model. In practical applications, other model structures can also be flexibly set according to actual needs. For ease of understanding, the embodiments of the present application provide the following examples:

[0227] Generally, a three-dimensional object has a certain geometric shape. Therefore, in order to improve the accuracy of the geometric shape of the constructed three-dimensional object, in a possible implementation method, a geometric expression module can be added after the decoder, such as a network module based on a signed distance field (SDF) algorithm. Among them, SDF describes the geometric shape of an object by calculating the distance from any point in space to the nearest object. It is a method for representing geometric shapes. Accordingly, after introduction, it is conducive to generating three-dimensional objects with more accurate geometric shapes.

[0228] In the related art, the geometric shapes in space are expressed based on density. Since the sparseness of density distribution will affect the final geometric shape, it is very easy to have the problem of uneven geometry of the generated three-dimensional object. Therefore, this method is not suitable for expressing the geometric shape of the three-dimensional object. In contrast, the present application expresses geometric shapes based on SDF, which is conducive to generating smoother geometry. For example, by inputting three sets of high-resolution two-dimensional features and combining SDF, a three-dimensional object with smooth, complete and detailed geometry is generated, that is, a more refined three-dimensional object is generated.

[0229] In order to further improve the accuracy of the details of the generated three-dimensional objects, in another possible implementation, a multilayer perceptron (MLP) can be added after the decoder. MLP is a feedforward artificial neural network model. Since the neurons in each layer are fully connected to the neurons in the next layer, each neuron will receive input signals from all neurons in the previous layer. Therefore, MLP can better capture deep information. Accordingly, after the introduction, further processing of the decoded features is conducive to improving the understanding of the features, thereby constructing a more accurate three-dimensional object.

[0230] Of course, in practical applications, the above-mentioned multilayer perceptron and the SDF-based network module can also be introduced at the same time. Fig. 9 A schematic diagram of a model framework is shown, in which solid arrows are used to represent the model training process and dotted arrows are used to represent the model application process.

[0231] It is understandable that before model training, Fig. 9 The encoder in is the aforementioned initial encoder, the decoder is the aforementioned initial decoder, the diffusion model is also the aforementioned initial model, and the multilayer perceptron and geometric expression module are not model trained. After the model training is completed, that is, in the model application process represented by the dotted line, Fig. 9 The diffusion model, i.e. the aforementioned target model, the decoder, i.e. the aforementioned target decoder, the multi-layer perceptron and the geometric expression module are all trained using the data of the three-dimensional object sample of the present application. And, Fig. 9 The target three-dimensional object refers to the three-dimensional object corresponding to the aforementioned generation request.

[0232] Correspondingly, in the first part of the model training process, that is, when the model training is performed based on the difference between the predicted three-dimensional object and the three-dimensional object sample, the initial encoder, initial decoder, multi-layer perceptron, and geometric expression module are trained together so that each part of the network module can be suitable for the task of generating three-dimensional objects in this application.

[0233] In addition, in order to improve the generalization ability of the model, in another possible implementation, batch normalization and scaling operations may also be used. In this regard, the embodiment of the present application provides the following example description:

[0234] In practical applications, such as in the aforementioned VAE structure, the encoder will output a feature distribution to express the input data, that is, output three sets of two-dimensional features as a feature distribution. The mean of the feature distribution is μ and the variance is σ 2 , these two parameters define the distribution p(z|x) of the latent variable z, which is usually assumed to be a Gaussian distribution. During the model training phase, the feature distributions corresponding to the three sets of two-dimensional features x form the data distribution plane used to reconstruct and predict the three-dimensional object. Correspondingly, the decoder reconstructs the data x from the latent variable z and outputs the reconstructed data p(x|z) for reconstructing the three-dimensional object.

[0235] However, the amount of high-quality training samples is usually small, which results in insufficient high-quality three-dimensional features x, leading to generalization problems in the trained encoder and decoder, and degradation. This is reflected in the variance σ of the generated feature distribution. 2 If the size is too small, the distribution of latent variables is only concentrated in a small area for reconstructing the training data itself, and it is difficult to generate new 3D objects through sampling. This makes the model lack imagination in the model application stage and may output repeated 3D objects.

[0236] In this regard, the present application proposes to perform batch normalization and rescaling operations before the feature output layer of the encoder (i.e., the last layer of the encoder), and accordingly, the network structure of the decoder is adaptively modified. In the specific implementation, the mean μ and variance σ of the output feature distribution are 2 , the specific operations are as follows:

[0237] μ ′ =scaler(norm(μ))

[0238] σ ′ =scaler(norm(σ))

[0239] In the above formula, μ is used to represent the mean, σ is used to represent the standard deviation, norm is used to indicate the batch normalization operation, scaler is used to further scale the mean and standard deviation after batch normalization, and scaler is defined as follows:

[0240]

[0241] In the above formula, τ and θ are both scaling factors, which are used to indicate the degree of scaling operation.

[0242] The experimental data show that after the above processing, the variance σ of the generated feature distribution 2 Significantly increased, that is to say, the three sets of two-dimensional features of the three-dimensional object sample can be encoded into a high-quality continuous latent space expression, which is also conducive to better model training of the diffusion model. Based on this, the model can not only reconstruct the three-dimensional object samples seen in the training data, but also learn a more general feature distribution from it, thereby improving the problem of model degradation and increasing the diversity of the generated three-dimensional objects.

[0243] In practical applications, the encoder in the VAE structure outputs the input data (i.e., the three sets of two-dimensional features in this application) as a feature distribution, and the decoder reconstructs the input data based on the feature distribution to reconstruct the three-dimensional object. On this basis, when the initial encoder and the initial decoder are trained based on the difference between the predicted three-dimensional object and the three-dimensional object sample, the loss function of the VAE part can be constructed by the following formula:

[0244] L_vae=|(|Φ(x│z)-S|)| 1 +β(D_KL(q_φ(z|π)||p(z)))

[0245] In the above formula, L_vae refers to the loss function of the VAE part, which is used to indicate the difference between the predicted 3D object and the 3D object sample, |(|Φ(x│z)-S|)| 1 is used to represent the reconstruction loss, i.e., the deviation of the initial decoder in reconstructing 3D object samples, Φ is used to represent the model parameters of the initial decoder, |()| 1 Used to represent L 1 Norm, S is used to represent three-dimensional object samples.

[0246] D_KL(q_φ(z|π)||p(z)) is the KL divergence (Kullback-Leibler Divergence), where φ is used to represent the model parameters of the initial encoder. This part is used to constrain the feature distribution q_φ(z|π) output by the initial encoder to align with the target normal distribution p(z) to determine the spatial continuity of the feature distribution.

[0247] π is used to represent the input data, that is, the three sets of two-dimensional features mentioned above, z can represent the feature distribution output by the initial encoder, and x can be used to represent the data reconstructed by the decoder based on z, that is, it can represent the decoded features.

[0248] And, β is a hyperparameter used to balance the relative importance between the reconstruction loss term and the KL divergence term. By adjusting β, a balance can be achieved between the emphasis on the quality of the reconstructed three-dimensional object and the emphasis on the learning feature distribution.

[0249] Through the above embodiments, the model determination method provided by the present application is described in detail. For better understanding, the present application embodiment takes the use of the model determined by the present application in a game application as an example to summarize the model application stage:

[0250] When users need to generate 3D game assets (such as in the game development stage), they can use reference images, description texts, etc. to describe the desired 3D game assets and initiate a generation request. Then, the corresponding target features are automatically determined through the target model, and the corresponding 3D game assets (3D objects with geometric shapes and material maps) are automatically generated through the target decoder.

[0251] In practical applications, some scenes only need to generate the geometric shape of a three-dimensional object, while some scenes need to generate a three-dimensional object that has both a geometric shape and a material map. For this, it can be controlled on demand. For example, by specifying the generation requirements through control data. For another example, supervision can also be controlled on demand during the model training phase, that is, whether the geometric shape of the three-dimensional object sample is directly used as supervision, or the geometric shape of the three-dimensional object sample combined with the material map is used as supervision. Exemplarily, the embodiments of the present application also provide such Fig.10 The schematic diagram of generating the geometric shape of a three-dimensional object is shown. By specifying that the generation requirement is a geometric shape in the control data, the corresponding generated three-dimensional object D and three-dimensional object E both have precise geometric shapes. This can be controlled according to actual needs, and this application does not make any restrictions.

[0252] In practical applications, the automatically generated three-dimensional game assets can be directly imported into game development engines (such as Unreal, Maya, Unity, etc.) for game development and production. Compared with the method of relying on manual labor in related technologies, from original painting design to high-precision model carving, model simplification, map baking, skinning and skeleton binding and other cumbersome steps, the creation of three-dimensional game assets requires complex and time-consuming processes, which consumes great costs and time, which also makes this labor-intensive work mode greatly slow down the development progress of the game. Compared with this, after adopting this application, the cost of game development and production can be greatly reduced, which is conducive to shortening the development cycle. Especially in the scene of developing large-scale 3A games, game production can be made faster and more convenient.

[0253] In addition, the model provided by this application can be packaged into a toolkit, and game players can use the toolkit to independently generate three-dimensional game assets and participate in new game modes such as game content creation. Based on this, a richer game generation pipeline can be provided to players to bring them a richer gaming experience.

[0254] It can be seen that this application provides an efficient game asset generation technology, specifically generating adaptive three-dimensional game assets based on reference images, description texts, etc. input by users, which greatly reduces the threshold and cost of game production and shortens the game production time. Figure 2 , Figure 5 as well as Figure 6 The generation effects shown in the above embodiments.

[0255] It should be noted that, based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.

[0256] based on Figure 3 Corresponding to the model determination method provided in the embodiment, the embodiment of the present application further provides a model determination device 1100, wherein the model determination device 1100 includes an acquisition unit 1101, a determination unit 1102, an encoding unit 1103 and a training unit 1104:

[0257] The acquisition unit 1101 is used to acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other;

[0258] The determining unit 1102 is used to determine, according to the three groups of two-dimensional features, position codes corresponding to the three groups of two-dimensional features respectively through the first network layer in the initial encoder, wherein the position codes are used to indicate a feature relationship between the two-dimensional features;

[0259] The encoding unit 1103 is used to perform encoding through the second network layer in the initial encoder based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, so as to obtain the predicted coding features corresponding to the three-dimensional object sample;

[0260] The training unit 1104 is used to perform model training on the initial encoder and the initial decoder based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample to obtain a target encoder and a target decoder;

[0261] The determining unit 1102 is further used to determine the target coding features corresponding to the three-dimensional object sample through the target encoder according to the three groups of two-dimensional features;

[0262] The training unit 1104 is further configured to perform model training on the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample;

[0263] The determination unit 1102 is also used to respond to the generation request, denoise the initial noise data through the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and decode the target features through the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

[0264] In a possible implementation manner, if the position codes corresponding to the three groups of two-dimensional features respectively satisfy an orthogonal relationship, the encoding unit is further used to:

[0265] For each group of two-dimensional features in the three groups of two-dimensional features, respectively splicing the two-dimensional features and the position codes corresponding to the two-dimensional features to obtain splicing features corresponding to the three groups of two-dimensional features respectively;

[0266] According to the stitching features corresponding to the three groups of two-dimensional features, encoding is performed based on the attention mechanism by the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample, wherein the encoding process based on the attention mechanism is used to indicate the dot product of the stitching features belonging to the first feature plane and the stitching features belonging to the second feature plane, and in the process of performing the dot product, the orthogonal relationship is used to indicate that if the first feature plane and the second feature plane are different feature planes, then the dot product result between the position code in the stitching features belonging to the first feature plane and the position code in the stitching features belonging to the second feature plane is zero.

[0267] In a possible implementation manner, the training unit is further used for:

[0268] determining a training loss based on a difference between a predicted three-dimensional object determined by the initial decoder based on the predictive coding features and the three-dimensional object sample;

[0269] According to the optimization target and the training loss, the model parameters of the initial encoder and the initial decoder are adjusted to obtain the target encoder and the target decoder, and the optimization target is used to indicate that the position codes corresponding to the three groups of two-dimensional features determined by the first network layer after the model parameter adjustment are still orthogonal.

[0270] In a possible implementation manner, if the model parameter of the first network layer includes a first parameter, the determining unit is further configured to:

[0271] According to the three groups of two-dimensional features and the first parameter, an initial matrix is ​​constructed by the first network layer, in which the matrix elements in a first column and the matrix elements in a second column satisfy the orthogonal relationship, and the matrix elements in the first column and the matrix elements in the second column are two different matrix elements in the initial matrix;

[0272] According to the initial matrix, position codes corresponding to the three groups of two-dimensional features are determined.

[0273] In a possible implementation manner, the determining unit is further configured to:

[0274] According to the initial matrix and the plane sizes corresponding to the three feature planes, target matrices corresponding to the three feature planes are determined, wherein for each feature plane in the three feature planes, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane, and in the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of one two-dimensional feature in the feature plane;

[0275] The target matrices corresponding to the three feature planes are determined as the position codes corresponding to the three groups of two-dimensional features.

[0276] In a possible implementation, the plane sizes corresponding to the three feature planes are all N, and each feature ∈R N*d , R is used to represent a real number set, and d is used to represent the feature dimension of each feature.

[0277] In a possible implementation, if the positions corresponding to the three sets of two-dimensional features are encoded as target matrices corresponding to the three feature planes, the first parameter ∈ R d*1 , and is not zero, the initial matrix ∈R d*d , and the model parameters of the first network layer also include a second parameter, the second parameter ∈ R 1*N , the determining unit is further used for:

[0278] The three columns of matrix elements in the initial matrix are respectively determined as three first matrices, each of which ∈R d*1 ;

[0279] The three first matrices are respectively multiplied by the second matrix indicated by the second parameter to obtain three undetermined matrices, where the second matrix ∈ R 1*N, each of the three undetermined matrices ∈ R d*N ;

[0280] Perform matrix transposition on the three undetermined matrices respectively to obtain the target matrices corresponding to the three feature planes. Each target matrix ∈ R N*d .

[0281] In one possible implementation, the attention mechanism is a linear attention mechanism.

[0282] In a possible implementation manner, if the initial encoder includes M groups of networks connected in sequence, in the M groups of networks, each group of networks in the first M-1 groups of networks includes one first network layer, one second network layer, and one third network layer connected in sequence, and the Mth group of networks includes one first network layer and one second network layer connected in sequence, the determining unit is further used to:

[0283] According to the three groups of two-dimensional features, determining position codes respectively corresponding to the three groups of two-dimensional features through a first network layer pair in a first group of networks in the M groups of networks;

[0284] The encoding unit is further used to encode through the second network layer in the first group of networks based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, so as to obtain the predicted coding features corresponding to the first group of networks;

[0285] The determining unit is further configured to:

[0286] Converting the prediction coding features corresponding to the first group of networks according to the third network layer in the first group of networks to obtain three groups of two-dimensional coding features corresponding to the first group of networks;

[0287] According to the three groups of two-dimensional coding features corresponding to the i-1th group of networks in the M groups of networks, the position codes corresponding to the three groups of two-dimensional coding features corresponding to the i-1th group of networks are determined by the first network layer in the i-th group of networks, where i is an integer greater than 1 and less than or equal to M;

[0288] Based on the three groups of two-dimensional coding features corresponding to the i-1th group of networks and the position codes respectively corresponding to the three groups of two-dimensional coding features corresponding to the i-1th group of networks, encoding is performed through the second network layer in the i-th group of networks to obtain the predicted coding features corresponding to the i-th group of networks;

[0289] If the i-th group of networks includes a third network layer, the prediction coding features corresponding to the i-th group of networks are converted according to the third network layer in the i-th group of networks to obtain three groups of two-dimensional coding features corresponding to the i-th group of networks;

[0290] Among them, the predicted coding features corresponding to the three-dimensional object samples are determined based on the predicted coding features corresponding to the Mth group of networks.

[0291] In one possible implementation, the first control data includes at least one of a first reference image corresponding to the three-dimensional object sample and a first description text corresponding to the three-dimensional object sample, and the second control data includes at least one of a second reference image for describing the three-dimensional object generation requirements and a second description text for describing the three-dimensional object generation requirements.

[0292] In a possible implementation manner, the training unit is further used for:

[0293] Performing denoising processing on the noisy coding feature according to the first control data by using the initial model to obtain predicted noise corresponding to the noisy coding feature;

[0294] According to the difference between the predicted noise corresponding to the noise-added coding feature and the sample noise, the initial model is trained to obtain the target model, wherein the sample noise is the noise added by the noise-added coding feature compared to the target coding feature;

[0295] In response to the generation request, performing denoising processing on the initial noise data according to the second control data in the generation request by using the target model to determine the target feature corresponding to the second control data, including:

[0296] In response to the generation request, performing denoising processing on the initial noise data according to the second control data through the target model to obtain predicted noise corresponding to the initial noise data;

[0297] The difference between the initial noise data and the predicted noise corresponding to the initial noise data is determined as the target feature.

[0298] It can be seen from the above technical solution that in the first part of the model training, the position codes corresponding to the three groups of two-dimensional features are first determined by the first network layer in the initial encoder, and then subsequent processing is performed according to the three groups of two-dimensional features and their respective position codes. Among them, the position code is used to indicate the feature relationship between the two-dimensional features. Since the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features indicates two-dimensional data in the three-dimensional data, and the three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other, so that the three-dimensional data of any point on the three-dimensional object sample is split into three groups of two-dimensional features, so the feature relationship between the three groups of two-dimensional features can reflect the absolute spatial information of the point in the three-dimensional object sample. Similarly, the feature relationship between the three groups of two-dimensional features corresponding to any two points on the three-dimensional object sample specifically includes the relative relationship between the two two-dimensional features in the same feature plane in the feature plane, and the relative relationship between the two two-dimensional features in different feature planes, so it can also reflect the relative spatial information of the two points in the three-dimensional object sample. Based on this, after using three groups of two-dimensional features to indicate the three-dimensional data and introducing the position code, the spatial information carried by the three-dimensional data can be reflected based on the rich relationship between the features. And the training is supervised by three-dimensional object samples, so the initial encoder can learn the rich spatial information in the three-dimensional object samples to obtain prediction coding features with rich spatial information, and the initial decoder can learn how to decode the prediction coding features to determine the predicted three-dimensional objects with more precise details. The target encoder finally obtained has better encoding ability and can encode three sets of two-dimensional features into target coding features with rich spatial information. The target decoder has better decoding ability to determine the three-dimensional objects with more precise details. In practical applications, the three-dimensional objects that need to be generated often do not have three-dimensional data, so it is difficult to directly use the trained target encoder in the application stage. For this reason, in the second part of model training, the target coding features determined by the target encoder are used to train the initial model to obtain the target model to replace the target encoder in the application stage.

[0299] Correspondingly, in the application stage, in response to the generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and then the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request. Because the target model is obtained by model training using the target encoder, the target features determined using the target model are still features with rich spatial information, and the target decoder can better decode the target features with rich spatial information, so it can determine a three-dimensional object with more precise details for the generation request. Based on this, the technical problem in the related art that the trained model cannot generate a three-dimensional object with precise details due to the loss of spatial information during the encoding process is solved.

[0300] The embodiment of the present application further provides a computer device, which may be a terminal. For example, the terminal is a smart phone:

[0301] Fig.12 The block diagram shows a partial structure of a smart phone provided in an embodiment of the present application. Fig.12 The smartphone includes: a radio frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a wireless fidelity (WiFi) module 1170, a processor 1180, and a power supply 1190. The input unit 1130 may include a touch panel 1131 and other input devices 1132, the display unit 1140 may include a display panel 1141, and the audio circuit 1160 may include a speaker 1161 and a microphone 1162. Those skilled in the art will appreciate that Fig.12 The structure of the smartphone shown in the figure does not constitute a limitation of the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0302] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, a phone book, etc.), etc. In addition, the memory 1120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0303] The processor 1180 is the control center of the smartphone, which uses various interfaces and lines to connect various parts of the entire smartphone, and executes various functions of the smartphone and processes data by running or executing software programs and / or modules stored in the memory 1120, and calling data stored in the memory 1120. Optionally, the processor 1180 may include one or more processing units; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1180.

[0304] In this embodiment, the steps performed by the processor 1180 in the smartphone may be based on Fig.12 The structure shown is implemented.

[0305] The computer device provided in the embodiment of the present application may also be a server. Fig.13 As shown, Fig.13 The structural diagram of the server 1200 provided in the embodiment of the present application, the server 1200 may have relatively large differences due to different configurations or performances, and may include one or more processors, such as a central processing unit (CPU) 1222, and a memory 1232, one or more storage media 1230 (such as one or more mass storage devices) storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage medium 1230 can be short-term storage or permanent storage. The program stored in the storage medium 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1222 can be configured to communicate with the storage medium 1230 and execute a series of instruction operations in the storage medium 1230 on the server 1200.

[0306] The server 1200 may also include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input and output interfaces 1258, and / or one or more operating systems 1241, such as Windows Server 2000. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0307] In this embodiment, the central processor 1222 in the server 1200 may perform the following steps:

[0308] Acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other;

[0309] According to the three groups of two-dimensional features, determining, by a first network layer in an initial encoder, position codes corresponding to the three groups of two-dimensional features respectively, wherein the position codes are used to indicate a feature relationship between the two-dimensional features;

[0310] Based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, encoding is performed through the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample;

[0311] Based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample, the initial encoder and the initial decoder are model trained to obtain a target encoder and a target decoder;

[0312] Determining, by the target encoder, target coding features corresponding to the three-dimensional object samples according to the three groups of two-dimensional features;

[0313] Performing model training on an initial model according to a noise-added coding feature corresponding to the target coding feature and first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample;

[0314] In response to a generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

[0315] According to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store a computer program. When the computer program is executed by a computer device, the computer device executes the model determination method described in the aforementioned embodiments.

[0316] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in various optional implementations of the above-mentioned embodiments.

[0317] The descriptions of the processes or structures corresponding to the above-mentioned figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.

[0318] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0319] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0320] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0321] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0322] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.

[0323] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0324] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technical members in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model determination method, characterized in that: The method comprises: Acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other; According to the three groups of two-dimensional features, determining, by a first network layer in an initial encoder, position codes corresponding to the three groups of two-dimensional features respectively, wherein the position codes are used to indicate a feature relationship between the two-dimensional features; Based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, encoding is performed through the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample; Based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample, the initial encoder and the initial decoder are model trained to obtain a target encoder and a target decoder; Determining, by the target encoder, target coding features corresponding to the three-dimensional object samples according to the three groups of two-dimensional features; Performing model training on an initial model according to a noise-added coding feature corresponding to the target coding feature and first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample; In response to a generation request, the initial noise data is denoised by the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and the target features are decoded by the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

2. The method according to claim 1, characterized in that If the position codes respectively corresponding to the three groups of two-dimensional features satisfy an orthogonal relationship, the encoding based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features is performed through the second network layer in the initial encoder to obtain the predicted coding features corresponding to the three-dimensional object sample, including: For each group of two-dimensional features in the three groups of two-dimensional features, respectively splicing the two-dimensional features and the position codes corresponding to the two-dimensional features to obtain splicing features corresponding to the three groups of two-dimensional features respectively; According to the stitching features corresponding to the three groups of two-dimensional features, encoding is performed based on the attention mechanism by the second network layer in the initial encoder to obtain predicted coding features corresponding to the three-dimensional object sample, wherein the encoding process based on the attention mechanism is used to indicate the dot product of the stitching features belonging to the first feature plane and the stitching features belonging to the second feature plane, and in the process of performing the dot product, the orthogonal relationship is used to indicate that if the first feature plane and the second feature plane are different feature planes, then the dot product result between the position code in the stitching features belonging to the first feature plane and the position code in the stitching features belonging to the second feature plane is zero.

3. The method according to claim 2, characterized in that The method of performing model training on the initial encoder and the initial decoder based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample to obtain a target encoder and a target decoder includes: determining a training loss based on a difference between a predicted three-dimensional object determined by the initial decoder based on the predictive coding features and the three-dimensional object sample; According to the optimization target and the training loss, the model parameters of the initial encoder and the initial decoder are adjusted to obtain the target encoder and the target decoder, and the optimization target is used to indicate that the position codes corresponding to the three groups of two-dimensional features determined by the first network layer after the model parameter adjustment are still orthogonal.

4. The method according to claim 2, characterized in that: If the model parameter of the first network layer includes the first parameter, determining the position codes respectively corresponding to the three groups of two-dimensional features through the first network layer in the initial encoder according to the three groups of two-dimensional features, comprises: According to the three groups of two-dimensional features and the first parameter, an initial matrix is ​​constructed by the first network layer, in which the matrix elements in a first column and the matrix elements in a second column satisfy the orthogonal relationship, and the matrix elements in the first column and the matrix elements in the second column are two different matrix elements in the initial matrix; According to the initial matrix, position codes corresponding to the three groups of two-dimensional features are determined.

5. The method according to claim 4, characterized in that The step of determining, according to the initial matrix, the position codes respectively corresponding to the three groups of two-dimensional features comprises: According to the initial matrix and the plane sizes corresponding to the three feature planes, target matrices corresponding to the three feature planes are determined, wherein for each feature plane in the three feature planes, the matrix size of the target matrix corresponding to the feature plane is the same as the plane size of the feature plane, and in the target matrix corresponding to the feature plane, one matrix element is used to represent the position code of one two-dimensional feature in the feature plane; The target matrices corresponding to the three feature planes are determined as the position codes corresponding to the three groups of two-dimensional features.

6. The method according to any one of claims 1 to 5, characterized in that The plane sizes corresponding to the three feature planes are all N, and each feature ∈R N*d , R is used to represent a real number set, and d is used to represent the feature dimension of each feature.

7. The method according to claim 6, characterized in that If the positions corresponding to the three sets of two-dimensional features are encoded as target matrices corresponding to the three feature planes, the first parameter ∈ R d*1 , and is not zero, the initial matrix ∈R d*d , and the model parameters of the first network layer also include a second parameter, the second parameter ∈ R 1*N , the target matrices corresponding to the three feature planes are determined by the following method: The three columns of matrix elements in the initial matrix are respectively determined as three first matrices, each of which ∈R d*1 ; The three first matrices are respectively multiplied by the second matrix indicated by the second parameter to obtain three undetermined matrices, where the second matrix ∈ R 1*N , each of the three undetermined matrices ∈ R d*N ; Perform matrix transposition on the three undetermined matrices respectively to obtain the target matrices corresponding to the three feature planes. Each target matrix ∈ R N*d .

8. The method according to claim 2, characterized in that: The attention mechanism is a linear attention mechanism.

9. The method according to claim 1, characterized in that: If the initial encoder includes M groups of networks connected in sequence, in the M groups of networks, each group of networks in the first M-1 groups of networks includes one of the first network layers, one of the second network layers, and one of the third network layers connected in sequence, and the Mth group of networks includes one of the first network layers and one of the second network layers connected in sequence, the determining, according to the three groups of two-dimensional features, position codes corresponding to the three groups of two-dimensional features respectively through the first network layer in the initial encoder includes: According to the three groups of two-dimensional features, determining position codes corresponding to the three groups of two-dimensional features respectively through a first network layer pair in a first group of networks in the M groups of networks; The step of encoding based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, and obtaining the predicted coding features corresponding to the three-dimensional object samples through the second network layer in the initial encoder, comprises: Based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, encoding is performed through the second network layer in the first group of networks to obtain prediction coding features corresponding to the first group of networks; The method further comprises: Converting the prediction coding features corresponding to the first group of networks according to the third network layer in the first group of networks to obtain three groups of two-dimensional coding features corresponding to the first group of networks; According to the three groups of two-dimensional coding features corresponding to the i-1th group of networks in the M groups of networks, the position codes corresponding to the three groups of two-dimensional coding features corresponding to the i-1th group of networks are determined by the first network layer in the i-th group of networks, where i is an integer greater than 1 and less than or equal to M; Based on the three groups of two-dimensional coding features corresponding to the i-1th group of networks and the position codes respectively corresponding to the three groups of two-dimensional coding features corresponding to the i-1th group of networks, encoding is performed through the second network layer in the i-th group of networks to obtain the predicted coding features corresponding to the i-th group of networks; If the i-th group of networks includes a third network layer, the prediction coding features corresponding to the i-th group of networks are converted according to the third network layer in the i-th group of networks to obtain three groups of two-dimensional coding features corresponding to the i-th group of networks; Among them, the predicted coding features corresponding to the three-dimensional object samples are determined based on the predicted coding features corresponding to the Mth group of networks.

10. The method according to claim 1, characterized in that The first control data includes at least one of a first reference image corresponding to the three-dimensional object sample and a first description text corresponding to the three-dimensional object sample, and the second control data includes at least one of a second reference image for describing the three-dimensional object generation requirements and a second description text for describing the three-dimensional object generation requirements.

11. The method according to claim 1, characterized in that: The step of training the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data corresponding to the three-dimensional object sample to obtain the target model includes: Performing denoising processing on the noisy coding feature according to the first control data by using the initial model to obtain predicted noise corresponding to the noisy coding feature; According to the difference between the predicted noise corresponding to the noise-added coding feature and the sample noise, the initial model is trained to obtain the target model, wherein the sample noise is the noise added by the noise-added coding feature compared to the target coding feature; In response to the generation request, performing denoising processing on the initial noise data according to the second control data in the generation request by using the target model to determine the target feature corresponding to the second control data, including: In response to the generation request, performing denoising processing on the initial noise data according to the second control data through the target model to obtain predicted noise corresponding to the initial noise data; The difference between the initial noise data and the predicted noise corresponding to the initial noise data is determined as the target feature.

12. A model determination device, characterized in that: The device comprises an acquisition unit, a determination unit, an encoding unit and a training unit: The acquisition unit is used to acquire three groups of two-dimensional features corresponding to the three-dimensional object sample, the three groups of two-dimensional features are used to indicate the three-dimensional data of the three-dimensional object sample, each group of two-dimensional features is used to indicate two-dimensional data in the three-dimensional data, and three feature planes corresponding to the three groups of two-dimensional features are perpendicular to each other; The determining unit is used to determine, according to the three groups of two-dimensional features, position codes respectively corresponding to the three groups of two-dimensional features through the first network layer in the initial encoder, wherein the position codes are used to indicate a feature relationship between the two-dimensional features; The encoding unit is used to perform encoding through the second network layer in the initial encoder based on the three groups of two-dimensional features and the position codes respectively corresponding to the three groups of two-dimensional features, so as to obtain the predicted coding features corresponding to the three-dimensional object sample; The training unit is used to perform model training on the initial encoder and the initial decoder based on the difference between the predicted three-dimensional object determined by the initial decoder according to the predicted coding feature and the three-dimensional object sample to obtain a target encoder and a target decoder; The determining unit is further configured to determine, through the target encoder, a target coding feature corresponding to the three-dimensional object sample according to the three groups of two-dimensional features; The training unit is further used to perform model training on the initial model according to the noise-added coding feature corresponding to the target coding feature and the first control data corresponding to the three-dimensional object sample to obtain a target model, wherein the first control data is used to describe the three-dimensional object sample; The determination unit is also used to respond to the generation request, denoise the initial noise data through the target model according to the second control data in the generation request to determine the target features corresponding to the second control data, and decode the target features through the target decoder to determine the three-dimensional object corresponding to the generation request, wherein the second control data is used to describe the three-dimensional object generation requirements.

13. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 11 according to the instructions in the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a computer device, the computer device executes the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program, characterized in that When the method is executed on a computer device, the computer device is enabled to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Three-dimensional model generation method and device, equipment, storage medium and program product

    CN117252984A

  • Scene model generation method and related device

    CN117437361A

  • Encoder training method and related device

    CN117456102A