Method and system for adjusting three-dimensional model and related equipment
Through the model adjustment system, the feature adjustment operation is received and the target three-dimensional model is generated, which solves the problem of low editing of three-dimensional model and realizes efficient and accurate three-dimensional model adjustment.
Patent Information
- Application Number
- CN202410381692.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-03-30
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the production and modification of three-dimensional models need to be carried out by professionals, resulting in higher time and cost and lower efficiency.
Provide a three-dimensional model adjustment method, which obtains the appearance characteristics of the three-dimensional model through the model adjustment system, receives feature adjustment operations, and generates a target three-dimensional model based on the adjustment parameters, so as to realize the adjustment of the three-dimensional model without directly editing the model.
It reduces the difficulty of editing three-dimensional models, improves editing efficiency, enables ordinary users to adjust the three-dimensional model according to their needs, and realizes linkage between feature adjustments, improving the accuracy of adjustments.
Smart Images

Figure CN120219679A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, system, and related devices for three-dimensional model adjustment. Background Art
[0002] A three-dimensional model is a three-dimensional model generated using digital technology and a computer. Three-dimensional models have a wide range of applications in fields such as live broadcasts, exhibitions, animations, and advertisements, and there will be a large demand in the future. However, the production of three-dimensional models requires professionals to model using modeling software, which takes a long time and has a high cost. Moreover, when the three-dimensional model needs to be modified, professionals are also required to make adjustments, resulting in low efficiency in the modeling and modification of three-dimensional models. Summary of the Invention
[0003] This application provides a method, system, and related devices for three-dimensional model adjustment, which can achieve the adjustment of a three-dimensional model without directly editing the three-dimensional model, reduce the difficulty of editing the three-dimensional model, and improve the efficiency of editing the three-dimensional model.
[0004] In a first aspect, this application provides a method for three-dimensional model adjustment, including: a model adjustment system obtains a first three-dimensional model with multiple appearance features, where the multiple appearance features include a first appearance feature and a second appearance feature having an association relationship; the model adjustment system receives a feature adjustment operation, and then determines the input parameters of the model generation model for the three-dimensional model according to the adjustment parameters indicated by the feature adjustment operation, and generates a target three-dimensional model according to the input parameters and the model generation model for the three-dimensional model; wherein, the feature adjustment operation indicates to use the adjustment parameters to adjust the first appearance feature among the multiple appearance features, and the target three-dimensional model is a three-dimensional model obtained by adjusting the first appearance feature and the second appearance feature of the first three-dimensional model; the adjustment parameters include the adjustment intensity and / or adjustment direction of the first appearance feature.
[0005] The fact that the above appearance features have an association relationship means that after one of the features is adjusted, the adjustment operation will also affect other features having an association relationship with this feature. For example, for a human face, the age feature is associated with the number of wrinkles, the depth of facial skin color, the degree of facial skin relaxation, etc. When the age is small, there are fewer wrinkles, the facial skin color is light, and the facial skin is more firm and plump; when the age is large, there are more wrinkles, the facial skin color is deep, and the facial skin is more relaxed, enhancing the realism of the three-dimensional model.
[0006] When it is necessary to adjust a 3D model, only the adjustment parameters for a feature of the 3D model to be adjusted need to be input. There is no need for the user to directly edit the 3D model. The model adjustment system can obtain the input parameters for generating the target 3D model after adjustment based on the adjustment parameters. By using these input parameters and the 3D model generation model, the target 3D model can be generated, thus realizing the adjustment of the 3D model. This method enables ordinary users to adjust the 3D model according to their needs and achieve the expected effect, reduces the difficulty of editing the 3D model, and improves the efficiency of model adjustment. Since the first appearance feature and the second appearance feature are related, adjusting the first appearance feature will also affect the second appearance feature, realizing the linkage between feature adjustments and making the adjustment of the 3D model more accurate and more in line with the actual adjusted situation.
[0007] For example, if the first 3D model is a digital human and the age feature of the digital human is adjusted. If the above feature adjustment operation is to adjust the digital human in the direction of a greater age, the adjusted digital human will appear older than the initial digital human, with more wrinkles, a darker skin color, and looser facial skin; if the above feature adjustment operation is to adjust the digital human in the direction of a younger age, the adjusted digital human will appear younger than the initial digital human, with fewer wrinkles, a lighter skin color, and tighter facial skin.
[0008] In a possible implementation, the above 3D model generation model is used to generate a 3D model using the latent vector corresponding to the 3D model. After the model adjustment system obtains the first 3D model, it further includes determining the first latent vector corresponding to the first 3D model, then adjusting the first latent vector according to the adjustment parameters to obtain a second latent vector, and inputting the second latent vector into the above 3D model generation model to generate the above target 3D model; where the above input parameters include the second latent vector.
[0009] By defining the features that the 3D model can be adjusted (such as the above first appearance feature and second appearance feature) and representing the 3D model with a latent vector, when it is necessary to adjust a 3D model, only the features that the 3D model needs to be adjusted are selected, and the adjustment parameters for the latent vector are input. The model adjustment system can adjust the latent vector of the 3D model according to the adjustment parameters to obtain the adjusted latent vector, and then convert the adjusted latent vector into the adjusted 3D model through the 3D model generation model, thus realizing the adjustment of the 3D model. Through the above 3D model adjustment method, the user does not need to directly edit the 3D model on the 3D model. Only by selecting the features that need to be adjusted and the adjustment parameters for the features, the adjustment of the 3D model can be realized, enabling ordinary users to adjust the 3D model according to their needs and achieve the expected effect, reducing the difficulty of editing the 3D model, and improving the efficiency of model adjustment.
[0010] Further, when representing a 3D model with a latent vector, the latent vector can represent multiple appearance features included in the 3D model and can represent the correlation relationships between the appearance features. Therefore, when the first latent vector is adjusted according to the adjustment parameter to obtain the second latent vector, although the adjustment parameter is for the first appearance feature, the second appearance feature associated with the first appearance feature is also adjusted while adjusting the first appearance feature, realizing the linkage between feature adjustments, making the adjustment of the 3D model more accurate and more in line with the change relationships between the features of the physical object corresponding to the 3D model in reality.
[0011] In a possible implementation manner, the model adjustment system determines the first latent vector corresponding to the first 3D model by converting the first 3D model into the first latent vector through a first encoder. Converting the high-dimensional 3D model into a low-dimensional latent vector and representing the 3D model with the low-dimensional latent vector facilitates the processing of the 3D model by the computing device and improves the efficiency of editing the 3D model.
[0012] By using an encoder based on artificial intelligence to convert the 3D model into a latent vector, during the process of training the encoder to convert the 3D model into a latent vector, through a 3D model, the characteristics of different features of the 3D model can be learned, that is, a latent vector can represent multiple features of the 3D model and the correlation relationships between the multiple features. For example, if the training set used to train the encoder is digital humans, during the training process of the encoder, for digital humans of different ages in the dataset, the characteristics of digital humans of different ages can be learned. For example, digital humans with a younger age have fewer wrinkles, lighter skin color, and tighter facial skin, while digital humans with an older age have more wrinkles, darker skin color, and looser facial skin; another example is for the nose bridge, the nose bridge of a digital human with a lower nose bridge is wider than that of a digital human with a higher nose bridge.
[0013] Therefore, if there are other features among the multiple appearance features of the first 3D model that are associated with the first appearance feature, for example, the second appearance feature is associated with the first appearance feature, then when adjusting the first appearance feature of the first 3D model, the second appearance feature will also be adjusted, realizing the linkage between feature adjustments, making the adjustment of the 3D model more accurate and more in line with the actual adjusted situation. For example, when the nose bridge is lowered, the nose bridge will become wider synchronously, and when the age is adjusted to be younger, the wrinkles will decrease, the skin color will become lighter, and the facial skin will become more plump.
[0014] In a possible implementation, the above-mentioned process of inputting the second latent vector into the 3D model generation model to generate the target 3D model includes: first, inputting the first latent vector into the 3D model generation model to obtain a second 3D model; inputting the second latent vector into the 3D model generation model to obtain a third 3D model; then determining the adjustment difference based on the third 3D model and the second 3D model; and finally determining the target 3D model based on the adjustment difference and the first 3D model; wherein, the adjustment difference is the difference between the third 3D model and the second 3D model.
[0015] Since the first latent vector is obtained based on the first 3D model M1, there is a conversion error δ = G(w1) - M1 during the conversion of the second 3D model G(w1) generated according to the first latent vector w1. Therefore, there is also the above-mentioned conversion error δ between the third 3D model G(w2) generated according to the second latent vector w2 and the 3D model that the actual user wants to obtain. To reduce the above error, the above-mentioned conversion error is eliminated based on the third 3D model G(w2), so that the target 3D model M0 = G(w2) - δ = G(w2) - [G(w1) - M1] = M1 + G(w2) - G(w1). Through the above method, the error generated during the conversion of the latent vector and the 3D model can be reduced, and the gap between the obtained target 3D model and the 3D model that the user actually wants to obtain can be reduced.
[0016] In a possible implementation, the above-mentioned process of converting the first 3D model into the first latent vector through the first encoder includes: inputting the first 3D model into the first encoder to obtain a third latent vector; then inputting the third latent vector into the 3D model generation model to obtain a fourth 3D model; determining the conversion error based on the first 3D model and the fourth 3D model; the conversion error includes the error generated during the process of converting the first 3D model into the third latent vector; and then optimizing the third latent vector according to the conversion error to determine the first latent vector.
[0017] The three-dimensional model is converted into a latent vector that can represent the three-dimensional model and each feature of the three-dimensional model through an encoder. During the conversion process, the first three-dimensional model M1 is first converted into a third latent vector w3, and then the third latent vector w3 is converted into a fourth three-dimensional model G(w3) through a three-dimensional model generation model. There must be a difference between the fourth three-dimensional model G(w3) and the first three-dimensional model M1, and this difference includes the difference generated during the process of converting the first three-dimensional model M1 into the third latent vector w3. In this application, the value of the loss function is calculated based on the fourth three-dimensional model G(w3) and the first three-dimensional model M1, and the third latent vector w3 is optimized based on the value of the loss function. Through the above method, the third latent vector is optimized one or more times until the first latent vector that can finally represent the first three-dimensional model is obtained, which can reduce the error generated during the process of converting the first three-dimensional model into the first latent vector, and make the three-dimensional model generated by the three-dimensional model generation model based on the first latent vector closer to the first three-dimensional model, that is, the first latent vector obtained by the above method can better represent the first three-dimensional model.
[0018] In a possible implementation manner, the above-mentioned adjusting the first latent vector according to the adjustment parameter to obtain a second latent vector includes: determining the feature direction of the first appearance feature in the latent space, and determining the second latent vector according to the first latent vector, the feature direction corresponding to the first appearance feature, and the above-mentioned adjustment parameter; wherein, the feature direction indicates the change direction of the latent vector corresponding to the three-dimensional model when the three-dimensional model is adjusted based on the first appearance feature.
[0019] By determining the feature direction corresponding to the first appearance feature, it is possible to determine the change direction of the first latent vector when the first appearance feature of the first three-dimensional model is adjusted. When adjusting the first appearance feature of the first three-dimensional model, adjusting the first latent vector in this feature direction can achieve the adjustment of the first appearance feature of the first three-dimensional model.
[0020] In a possible implementation manner, before the above-mentioned determining the second latent vector according to the first latent vector, the feature direction corresponding to the first appearance feature, and the adjustment parameter, it further includes: randomly sampling in the latent space through a three-dimensional model generation model to generate a plurality of sample three-dimensional models; respectively converting the plurality of sample three-dimensional models into their corresponding latent vectors through the first encoder; adding labels to the plurality of sample three-dimensional models respectively according to the first appearance feature to obtain a feature training set; and finally determining the feature direction of the first appearance feature based on a machine learning algorithm and the feature training set. Wherein, the label is used to indicate the intensity and / or direction of the first appearance feature of the corresponding sample three-dimensional model, and the feature training set includes the latent vector corresponding to each sample three-dimensional model and the label corresponding to each sample three-dimensional model.
[0021] In this application, when a user needs to edit and adjust a 3D model, the user will first determine the feature to be adjusted. Therefore, by annotating a feature in the training set, such as the above-mentioned first appearance feature, and annotating the data in the training set according to this feature, the data in the training set can be divided into multiple categories according to this feature. Then, by training a machine learning model with the data in the training set, the feature direction corresponding to the first appearance feature can be obtained. In this way, when the user selects to adjust the first appearance feature, the first latent vector is adjusted in the feature direction of the first appearance feature, so as to realize the adjustment of the first appearance feature of the 3D model.
[0022] In a possible implementation manner, before converting the first 3D model into the first latent vector through the first encoder, it further includes: inputting the fifth 3D model into the second encoder to obtain the latent vector corresponding to the fifth 3D model; where the fifth 3D model is one in the second training set, and the second training set includes multiple 3D models; generating the sixth 3D model according to the 3D model generation model and the latent vector corresponding to the fifth 3D model; determining the value of the loss function according to the fifth 3D model and the sixth 3D model, and updating the second encoder based on the value of the loss function to obtain the first encoder. By training the first encoder with 3D models of the same type as the first 3D model, the trained first encoder can be obtained, so that the latent vector converted by the first encoder can more accurately represent the 3D model.
[0023] In a possible implementation manner, the above-mentioned first 3D model is the head 3D model of a digital human; the multiple features of the first 3D model include some or all of gender, age, skin color, face shape, eye size, or nose bridge height.
[0024] In a second aspect, this application provides a 3D model adjustment device, including:
[0025] A processing module, configured to determine a first 3D model, where the first 3D model has multiple appearance features, and the multiple appearance features include a first appearance feature and a second appearance feature, and the first appearance feature and the second appearance feature have an association relationship;
[0026] A receiving module, configured to receive a feature adjustment operation, where the feature adjustment operation indicates to adjust the first appearance feature among the multiple appearance features by using adjustment parameters, and the adjustment parameters include an adjustment intensity and / or an adjustment direction;
[0027] The processing module is further configured to determine the input parameters of the 3D model generation model according to the adjustment parameters; generate a target 3D model according to the input parameters and the 3D model generation model, where the target 3D model is the 3D model after adjusting the first appearance feature and the second appearance feature of the first 3D model.
[0028] In a possible implementation, the above three-dimensional model generation model is used to generate a three-dimensional model using the latent vector corresponding to the three-dimensional model. The above processing module is further configured to: determine a first latent vector corresponding to the first three-dimensional model; the processing module determines the input parameters of the three-dimensional model generation model according to the adjustment parameters, and generates a target three-dimensional model according to the input parameters and the three-dimensional model generation model, specifically including: adjusting the first latent vector according to the adjustment parameters to obtain a second latent vector, where the input parameters include the second latent vector; inputting the second latent vector into the three-dimensional model generation model to generate the target three-dimensional model.
[0029] In a possible implementation, the above processing module is specifically configured to convert the first three-dimensional model into a first latent vector through a first encoder.
[0030] In a possible implementation, the above processing module is specifically configured to first input the first latent vector into the three-dimensional model generation model to obtain a second three-dimensional model; input the second latent vector into the three-dimensional model generation model to obtain a third three-dimensional model; then determine an adjustment difference according to the third three-dimensional model and the second three-dimensional model, and finally determine the target three-dimensional model according to the adjustment difference and the first three-dimensional model; where the adjustment difference is the difference between the third three-dimensional model and the second three-dimensional model.
[0031] In a possible implementation, the above processing module is specifically configured to input the first three-dimensional model into a first encoder to obtain a third latent vector; then input the third latent vector into the three-dimensional model generation model to obtain a fourth three-dimensional model; determine a conversion error according to the first three-dimensional model and the fourth three-dimensional model; the conversion error includes the error generated in the process of converting the first three-dimensional model into the third latent vector; and then optimize the third latent vector according to the conversion error to determine the first latent vector.
[0032] In a possible implementation, the processing module is specifically configured to determine the feature direction of the first appearance feature in the latent space, and determine the second latent vector according to the first latent vector, the feature direction corresponding to the first appearance feature, and the above adjustment parameters; where the feature direction indicates the change direction of the latent vector corresponding to the three-dimensional model when adjusting the three-dimensional model based on the first appearance feature.
[0033] In a possible implementation, the above three-dimensional model adjustment device further includes a training module, which is used to randomly sample the model in the latent space through the three-dimensional model to generate multiple sample three-dimensional models; convert the multiple sample three-dimensional models into their respective corresponding latent vectors through the first encoder; add labels to the multiple sample three-dimensional models respectively according to the first appearance feature to obtain a feature training set; and finally determine the feature direction of the first appearance feature based on the machine learning algorithm and the feature training set. Wherein, the label is used to indicate the intensity and / or direction of the first appearance feature of the corresponding sample three-dimensional model, and the feature training set includes the latent vector corresponding to each sample three-dimensional model and the label corresponding to each sample three-dimensional model.
[0034] In a possible implementation, the above processing module is further used to input the fifth three-dimensional model into the second encoder to obtain the latent vector corresponding to the fifth three-dimensional model; wherein, the fifth three-dimensional model is one of the second training set, and the second training set includes multiple three-dimensional models; obtain the sixth three-dimensional model according to the three-dimensional model generation model and the latent vector corresponding to the fifth three-dimensional model; determine the value of the loss function according to the fifth three-dimensional model and the sixth three-dimensional model, and update the second encoder based on the value of the loss function to obtain the first encoder. Training the first encoder with three-dimensional models of the same type as the first three-dimensional model, the trained first encoder can enable the latent vector converted by the first encoder to more accurately represent the three-dimensional model.
[0035] In a possible implementation, the first three-dimensional model is the head three-dimensional model of a digital human; at least one feature of the first three-dimensional model includes some or all of gender, age, skin color, face shape, eye size, or nose bridge height.
[0036] In a third aspect, the present application provides a computing device, including a processor and a memory, and the processor executes instructions stored in the memory to implement the method described in the above first aspect or any possible implementation manner of the first aspect.
[0037] In a fourth aspect, the present application provides a computing device cluster, and the computing device cluster includes at least one computing device. Each computing device includes a processor and a memory. Wherein, the processor of each computing device is used to execute instructions stored in the memory to enable the computing device cluster to implement the method described in the above first aspect or any possible implementation manner of the first aspect.
[0038] In a fifth aspect, the present application provides a computer-readable storage medium, which is characterized in that it includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device is enabled to execute the method described in the above first aspect or any possible implementation manner of the first aspect.
[0039] In a sixth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is run on a computing device, it implements the method described in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings
[0040] Figure 1 FIG. 6 is a schematic structural diagram of an adjustment system provided by the present application;
[0041] Figure 2 FIG. 10 is a schematic flowchart of a three-dimensional model adjustment method provided by the present application;
[0042] Figure 3 FIG. 14 is a schematic diagram of converting a three-dimensional model into a latent vector provided by the present application;
[0043] Figure 4 FIG. 18 is a schematic diagram of a model adjustment interface provided by the present application;
[0044] Figure 5 FIG. 22 is a schematic diagram of determining a third three-dimensional model provided by the present application;
[0045] Figure 6 FIG. 26 is a schematic diagram of a model adjustment system provided by the present application;
[0046] Figure 7 FIG. 30 is a schematic diagram of a computing device provided by the present application;
[0047] Figure 8 FIG. 34 is a schematic diagram of a computing device cluster provided by the present application;
[0048] Figure 9 FIG. 38 is a schematic diagram of network connection between computing devices provided by the present application. Detailed Embodiments
[0049] The technical solutions provided by the present application will be described below with reference to the accompanying drawings.
[0050] A three-dimensional model is a three-dimensional model generated using digital technology and a computer. Three-dimensional models have a wide range of applications in fields such as live broadcasts, exhibitions, animations, and advertisements, and there will be a large demand in the future. However, the production of three-dimensional models requires professionals to model using modeling software, which takes a relatively long time and has a high cost. And when the model needs to be modified, it also requires professionals to make adjustments and modifications, resulting in low efficiency in the modeling and modification of three-dimensional models.
[0051] Exemplarily, for a three-dimensional model of a human head (hereinafter simply referred to as a digital human), a digital human is a human image generated by using digital technology and a computer. Digital humans have a wide range of applications in fields such as virtual reality (VR), augmented reality (AR), animation, live streaming, exhibitions, and advertising. However, the production of digital humans requires professional artists or personnel with a certain background in graphics to model through modeling software. When the digital human needs to be modified, professional personnel are also required to make adjustments. Otherwise, it is difficult to achieve the required effect, resulting in low efficiency in the modeling and modification of digital humans. For example, if the face of a digital human needs to be modified, changing the face of the digital human from a younger state to an older state, if not a professional artist, ordinary users cannot accurately depict the facial details of the person after aging on the model, which will lead to problems such as a longer modification time or a large gap between the modified digital human and the actual required effect.
[0052] To solve the above problems, the present application provides a method for adjusting a three-dimensional model. When a three-dimensional model needs to be adjusted, the model adjustment system only needs to obtain the adjustment parameter of a feature of the three-dimensional model, and then can adjust the feature according to the adjustment parameter. Users do not need to directly adjust the three-dimensional model on the three-dimensional model, so that ordinary users can adjust the three-dimensional model according to their needs and achieve the expected effect, improving the efficiency of model adjustment. And when adjusting a feature of the three-dimensional model, other features associated with this feature can also be adjusted, making the adjustment of the three-dimensional model more accurate and more in line with the change relationship between the features of the physical object corresponding to the three-dimensional model in reality.
[0053] For example, the three-dimensional model is a digital human, and the adjustable features corresponding to the digital human include age, gender, skin color, face shape, eye size, eye socket depth, nose bridge height, nose bridge width, number of wrinkles, skin color depth, facial skin relaxation degree, hair density, etc. If the user needs to adjust the face of the digital human to make the face of the digital human show an older state, the user only needs to select the "age" feature, and then input the text "make the person older" or adjust the three-dimensional model in the direction of increasing age through the control for adjusting age, and the facial features of the digital human can be adjusted in the direction of increasing age to obtain the adjusted digital human. And the adjusted digital human has more wrinkles, darker facial skin color, looser facial skin, and sparser hair.
[0054] See Figure 1 , Figure 1It is a schematic architecture diagram of an adjustment system provided by this application. This adjustment system is used to implement the three-dimensional model adjustment method provided by this application. The adjustment system includes a client 110 and a model adjustment system 120. Among them, the client 110 and the model adjustment system 120 are communicatively connected. This communication connection can be a wireless connection or a wired connection. The client 110 that establishes a communication connection with the model adjustment system 120 can include one or more, and this application does not make specific limitations.
[0055] The model adjustment system 120 can be deployed on a single computing device or on a computing device cluster including multiple computing devices. The computing device can be a bare metal server, a virtual machine, a container, or an edge computing device. A virtual machine refers to a complete computer system with complete hardware system functions simulated by software and running in a completely isolated environment. What can be done on a physical computer can be achieved in a virtual machine. When creating a virtual machine in a computing device, part of the hard disk and memory capacity of the physical machine needs to be used as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent basic input / output system (BIOS), hard disk, and operating system, and can operate the virtual machine just like using a physical machine; A container is a portable software unit that can combine an application and all its dependencies into a software package. This software package is not restricted by the underlying host operating system, so there is no need to build a complex environment, simplifying the process from application development to deployment; An edge computing device refers to a device that is closer to the data source and end user and has the characteristics of low latency and high bandwidth, such as an intelligent router, an edge server, etc. The computing device cluster can include multiple of the above-mentioned computing devices, and this application does not make specific limitations.
[0056] The above-mentioned computing device can be a computing device in a cloud data center, an edge server, or a local server in an enterprise local data center. If the model adjustment system 120 is deployed on a computing device or a computing device cluster in an enterprise local data center, the client 110 can be deployed on an office device within the enterprise, such as an office computer.
[0057] If the model adjustment system 120 can be deployed in a cloud data center, the client 110 can be deployed on a terminal device used by the user. In a specific implementation, the user can obtain the usage permission of the model adjustment system 120 by purchasing a cloud service. The cloud data center provides the user with a usage account for the cloud service, enabling the user to log in to the account using the client 110 and establish a communication connection between the client 110 and the model adjustment system 120.
[0058] The client 110 is used to implement human-computer interaction. The client 110 is used to display the model adjustment interface provided by the model adjustment system 120 to the user. The user can perform model adjustment operations on the model adjustment interface to input adjustment parameters for adjusting the model. After obtaining the adjustment parameters, the client 110 sends the adjustment parameters to the model adjustment system 120 through the terminal device. The model adjustment system 120 adjusts the three-dimensional model according to the received adjustment parameters to obtain an adjusted three-dimensional model, and displays the adjusted three-dimensional model on the model adjustment interface.
[0059] The above-mentioned client 110 can be software or an application program running on a terminal device, such as a client on a personal computer (PC), or a browser-based client, or an application (APP) running on a mobile terminal, or a console of a cloud platform. The present application does not make specific limitations. The terminal device can be an electronic device with computing capabilities such as a PC, a smart phone, a wearable device, a palm processing device, a tablet computer, a notebook computer, an augmented reality (AR) device, a virtual reality (VR) device, or a vehicle-mounted device, etc., and no specific limitations are made here.
[0060] It should be understood that the above examples are for illustration purposes, and the specific deployment of the client 110 and the model adjustment system 120 can be determined according to the actual application scenario, and the present application does not make specific limitations. For example, the client 110 and the model adjustment system 120 can also be deployed on the same computing device to implement the method for adjusting the three-dimensional model provided by the present application through one computing device.
[0061] The method for adjusting the three-dimensional model provided by the present application is introduced below with reference to the accompanying drawings. It should be noted that in the present application, when adjusting a three-dimensional model, the model adjustment system may need to convert the three-dimensional model to be adjusted into a corresponding latent vector through an encoder. Among them, the latent vector can represent the three-dimensional model, so the latent vector can represent each feature included in the three-dimensional model. Then, based on the latent vector, the model adjustment system generates an adjusted latent vector according to the adjustment of a feature by the user. After obtaining the adjusted latent vector, the adjusted latent vector is input into the three-dimensional model generation model, and the three-dimensional model generation model generates an adjusted three-dimensional model according to the adjusted latent vector. Therefore, before adjusting the three-dimensional model, it is necessary to first train the three-dimensional model generation model for converting the latent vector into a three-dimensional model and the encoder for converting the three-dimensional model into a latent vector.
[0062] The training method of the generation model and the training method of the encoder provided by the present application are introduced below respectively.
[0063] I. Training a 3D model generation model
[0064] When training a 3D model generation model for converting a latent vector into a 3D model, first obtain a first training set, which includes multiple 3D models. For example, if the 3D model to be adjusted is a digital human, the first training set includes multiple digital humans with the same geometric topology. Then, train the 3D model generation model with the first training set until the model converges, obtaining a trained 3D model generation model and the latent space corresponding to the 3D model generation model. Among them, the 3D model generation model can be any one of models such as a generative adversarial network (GAN), a variational auto-encoder (VAE), a diffusion model, or a stable diffusion model. In this application, the 3D models in the above first training set are expressed by means of material maps and position maps. Among them, the material maps include diffuse maps, normal maps, etc.
[0065] In a possible implementation, after obtaining the first training set, it is also possible to enhance the training set to obtain more training data, such as performing data augmentation through methods such as Alpha blending, texture transformation, and lighting transformation to obtain more training data.
[0066] II. Training an encoder
[0067] When training an encoder for converting a 3D model into a latent vector, first obtain a second training set, which includes multiple 3D models. For example, if the 3D model to be adjusted is a digital human, the second training set includes multiple digital humans with the same geometric topology. Then, train the second encoder with the second training set until convergence, obtaining a first encoder. Among them, the encoder can be a residual network (ResNet) or a PointNet encoder; the second training set for training the encoder and the first training set for training the generation model can be the same or different, and this application does not make specific limitations.
[0068] Taking the second encoder as ResNet as an example, during the training process of the second encoder, if the training data input to the second encoder is the 3D model a, and the second encoder ResNet is denoted as R, then the hidden vector output after a passes through the second encoder is R(a); when calculating the loss function during the training process, the 3D model generation model (denoted as G) converts the hidden vector R(a) output by ResNet into the 3D model G(R(a)), and the loss function can be denoted as: Loss1 = L(G(R(a)), a). The first encoder is obtained by training the second encoder according to the above loss function and the second training set.
[0069] By using an encoder based on artificial intelligence to convert a 3D model into a hidden vector, during the process of training the encoder to convert a 3D model into a hidden vector, through a 3D model, the characteristics of different features of the 3D model can be learned, that is, a hidden vector can represent multiple features of the 3D model and the correlation between multiple features. For example, if the training set used to train the encoder is digital humans, during the training process of the encoder, for digital humans of different ages in the dataset, the characteristics of digital humans of different ages can be learned. For example, digital humans with a younger age have fewer wrinkles, lighter skin color, and tighter facial skin, while digital humans with an older age have more wrinkles, darker skin color, and looser facial skin; another example is for the nose bridge, the nose bridge of a digital human with a lower nose bridge is wider than that of a digital human with a higher nose bridge.
[0070] After obtaining the above-mentioned trained 3D model generation model and the first encoder, according to the Figure 2 method shown below, the 3D model is adjusted. Figure 2 It is a schematic flowchart of a 3D model adjustment method provided by this application, and this method includes S210 - S250.
[0071] S210. The model adjustment system determines the first 3D model to be adjusted.
[0072] In this application, the first 3D model can be a 3D model in the database. For example, the first 3D model is a 3D model selected by the user from the database that needs to be adjusted. The first 3D model can also be a 3D model uploaded by the user through the client. For example, the model adjustment system provides a model adjustment interface to the client, and the user uploads the first 3D model to be adjusted through the model adjustment interface. The first 3D model can also be a 3D model generated by the model adjustment system through random sampling in the above-mentioned hidden space, and this application does not make specific limitations.
[0073] S220. The model adjustment system determines the first hidden vector corresponding to the first 3D model.
[0074] After the model adjustment system obtains the first 3D model to be adjusted, it converts the first 3D model into a first latent vector. The first latent vector can represent the first 3D model and can also represent each appearance feature included in the first 3D model. It should be noted that if the first 3D model is a 3D model generated by random sampling in the above latent space, there is no need to convert the first 3D model into a vector again; for the sake of description, the latent vector randomly sampled from the latent space is also referred to as the first latent vector.
[0075] In a possible implementation, the model adjustment system converts the first 3D model into a first latent vector through a first encoder, specifically including: inputting the first 3D model into the first encoder, and taking the output of the first encoder as the first latent vector corresponding to the first 3D model.
[0076] In another possible implementation, as Figure 3 shown, Figure 3 is a schematic diagram of converting a 3D model into a latent vector provided by this application. The method for the model adjustment system to convert the first 3D model into a first latent vector through a first encoder includes the following steps:
[0077] (1) Input the first 3D model M1 into the first encoder, and the first encoder outputs a third latent vector w3;
[0078] (2) Input the third latent vector w3 into the 3D model generation model, and the 3D model generation model generates a fourth 3D model according to the third latent vector w3; where, if the 3D model generation model is denoted as G, the fourth 3D model is denoted as G(w3);
[0079] (3) Calculate the loss function Loss2 according to the fourth 3D model G(w3) and the first 3D model M, that is, the conversion error generated in the process of converting the first 3D model into the third latent vector, where the loss function is Loss2 = L(G(w3), M);
[0080] (4) Determine whether the value of the loss function Loss2 is greater than the first threshold. If the loss function Loss2 is less than or equal to the first threshold, execute (5); if the loss function Loss2 is greater than the first threshold, execute (6);
[0081] (5) If the value of the loss function Loss2 is less than or equal to the first threshold, take the above third latent vector w3 as the first latent vector corresponding to the first 3D model;
[0082] (6) If the value of the loss function Loss2 is greater than the first threshold, optimize the third hidden vector w3 according to the value of the loss function Loss2 to obtain the updated hidden vector w, and use the updated hidden vector w as the third hidden vector w3, and then execute the above steps (2) to (4) again until the value of the loss function Loss2 is less than or equal to the second threshold.
[0083] In a possible implementation, since the three-dimensional model is expressed by means of a diffuse texture map, a normal map, etc., the above loss function Loss2 can be the following formula (1):
[0084] Loss = k1 * Ld + k2 * Lm + k3 * Ln (Formula 1)
[0085] Where Ld is the diffuse texture map encoding loss, Lm is the Mesh encoding loss, Ln is the normal map encoding loss, and k1, k2, and k3 are weights. Ld and Ln can be perceptual losses, and Lm can be a distance loss, which is not specifically limited in this application.
[0086] Optionally, when optimizing the third hidden vector according to the value of the loss function Loss2, an adaptive moment estimation (Adam) optimizer, an adaptive gradient algorithm (Adagrad), an Adamax optimizer, or a Nadam optimizer can be used.
[0087] The encoder converts the three-dimensional model into a hidden vector that can represent the three-dimensional model and each feature of the three-dimensional model. During the conversion process, the first three-dimensional model M1 is first converted into the third hidden vector w3, and then the third hidden vector w3 is converted into the fourth three-dimensional model G(w3) through the three-dimensional model generation model. There must be a difference between the fourth three-dimensional model G(w3) and the first three-dimensional model M1, and this difference includes the difference generated during the process of converting the first three-dimensional model M1 into the third hidden vector w3. This application calculates the value of the loss function based on the fourth three-dimensional model G(w3) and the first three-dimensional model M1, optimizes the third hidden vector w3 based on the value of the loss function, and optimizes the third hidden vector one or more times through the above method until the first hidden vector that can finally represent the first three-dimensional model is obtained, which can reduce the error generated during the process of converting the first three-dimensional model into the first hidden vector, making the three-dimensional model generated by the three-dimensional model generation model based on the first hidden vector closer to the first three-dimensional model, that is, the first hidden vector obtained by the above method can better represent the first three-dimensional model.
[0088] S230. The model adjustment system receives a feature adjustment operation and obtains adjustment parameters for adjusting the first appearance feature.
[0089] In this application, the first 3D model has multiple appearance features. For example, if the first 3D model is a digital human, the corresponding appearance features of the first 3D model include but are not limited to gender, age, skin color, face shape, eye size, eye socket depth, nose bridge height, nose bridge width, lip thickness, nose bridge width, number of wrinkles, skin color depth, degree of facial skin relaxation, hair density, etc.
[0090] When it is necessary to adjust the first 3D model, the user can select an appearance feature of the first 3D model, such as the first appearance feature, and adjust the first 3D model by adjusting the first appearance feature. For example, if the first 3D model is a digital human, when the user needs to adjust the digital human to look more mature, the user can select the feature of "age" and adjust this feature. After the model adjustment system receives the adjustment operation of the user, it adjusts the age feature of the digital human to achieve the purpose of making the digital human look more mature.
[0091] After the user adjusts the first appearance feature, the model adjustment system can receive the feature adjustment operation, and this feature adjustment operation instructs to adjust the first appearance feature using the adjustment parameter. Among them, the adjustment parameter includes the adjustment intensity or the adjustment direction, and the adjustment parameter can also include both the adjustment intensity and the adjustment direction. The adjustment intensity indicates the degree of adjustment of the first appearance feature, and the adjustment direction is used to indicate in which direction of the first appearance feature to adjust. For example, if the first 3D model is a digital human and the first appearance feature is the eye size, the adjustment intensity can be the adjusted eye size, or it can be the size or proportion of the increase or decrease of the eyes based on the current eye size of the digital human. The adjustment direction refers to whether to adjust in the direction of larger eyes or smaller eyes based on the current eye size of the digital human.
[0092] In this application, the model adjustment system can provide a model adjustment interface to the user, as Figure 4 shown, Figure 4 is a schematic diagram of a model adjustment interface provided by this application. This model adjustment interface includes a model display area 410, a feature display area 420, and a feature adjustment area 430. Among them, the model display area 410 is used to display the 3D model to be adjusted selected by the user, the feature display area 420 is used to display the adjustable features included in the 3D model to be adjusted, and the feature adjustment area 430 is used to display the controls for adjusting the features.
[0093] After importing the first 3D model in the model adjustment interface, the model adjustment interface can display the first 3D model in the model display area 410 and display multiple appearance features of the first 3D model in the feature display area 420. The user can select a feature of the first 3D model in the feature display area 420, such as the first appearance feature. After the user selects a first appearance feature, the model adjustment system can provide the user with a control for adjusting the first appearance feature and display the control in the feature adjustment area 430. Among them, the control includes at least one adjustment direction; for example, the control includes a line segment and a slider that can move on the line segment, and the user can adjust the position of the slider to adjust the first appearance feature.
[0094] The user adjusts the first appearance feature by adjusting the position of the slider included in the control. After the user determines that the adjustment is completed, the model adjustment system obtains the position of the adjusted slider and determines the adjustment parameter according to the position of the slider.
[0095] Exemplarily, taking the above first 3D model as a digital human as an example, after the user imports the digital human in the model adjustment interface, the model adjustment system displays the digital human in the above model display area 410 and displays the appearance features that the digital human can adjust in the feature display area 420, such as Figure 4 As shown, the appearance features that the digital human can adjust include gender, age, skin color, face shape, eye size, eye socket depth, nose bridge height, nose bridge width, number of wrinkles, skin color depth, facial skin relaxation degree, etc. If the user selects the feature of "age" for adjustment, the model adjustment interface displays a control for adjusting the age in the feature adjustment area 430.
[0096] As Figure 4 As shown, when the user selects the feature of "age" to adjust the digital human, the control for adjusting the age includes two adjustment directions. If the user adjusts the control in the direction indicating "small", it can make the finally generated adjusted digital human look younger than the initial digital human (i.e., the first 3D model). If the user adjusts the control in the direction indicating "big", it can make the finally generated adjusted digital human look older than the initial digital human.
[0097] If the adjustment range of the control in the figure is from -1 to 1, after the user completes the adjustment, the user triggers the "confirm" button in the feature adjustment area 430. After the model adjustment system detects the user's confirmation operation, it obtains the value corresponding to the current position of the slider included in the control, which is the above adjustment parameter α.
[0098] In a possible implementation, the adjustment range of the control can also be other ranges. For example, the adjustment range of the above control can be from -100 to 100. After the model adjustment system obtains the value corresponding to the current position of the slider, it converts the value corresponding to the current position of the slider into a range from -1 to 1 to obtain the above adjustment parameter α. Where α = t / 100, and t is the value corresponding to the current position of the slider. For example, in the above Figure 4 if the value corresponding to the current position of the slider is 20, then the value of the adjustment parameter α is 0.2.
[0099] In a possible implementation, the feature adjustment operation may not be implemented by selecting a feature in the above interface and adjusting the control. When the user needs to adjust a certain feature, the feature adjustment operation can be implemented by inputting a command. For example, the user can input "Adjust the age of the digital human to look about 40 years old", or "Make the eyes of the digital human a little smaller". After receiving the above command, the model adjustment system extracts semantic information from the command to determine the adjustment parameter.
[0100] For example, for each appearance feature of the three-dimensional model, a step size for each adjustment is set. The above command "Make the eyes of the digital human a little smaller" only includes the adjustment direction. The model adjustment system can obtain the current eye size of the first three-dimensional model, and the adjustment intensity is to subtract the step size from the current eye size.
[0101] S240. The model adjustment system determines the input parameters of the three-dimensional model generation model according to the adjustment parameter.
[0102] The three-dimensional model generation model can generate a three-dimensional model through a latent vector. After the model adjustment system determines the input parameters of the three-dimensional model generation model according to the adjustment parameter, it can generate a target three-dimensional model with the first appearance feature adjusted according to the input parameters.
[0103] In this application, since the latent vector of the three-dimensional model can represent each feature included in the three-dimensional model after converting the three-dimensional model into a latent vector through an encoder, the first latent vector also includes various adjustable features of the first three-dimensional model. When adjusting the first appearance feature of the first three-dimensional model, it is actually adjusting the first latent vector.
[0104] After the model adjustment system obtains the adjustment parameter for adjusting the first appearance feature, it adjusts the first latent vector to obtain a second latent vector, and the above three-dimensional model generation model includes this second latent vector.
[0105] When the model adjustment system determines the second hidden vector according to the adjustment parameter and the first hidden vector, the model adjustment system first determines the feature direction of the first appearance feature in the hidden space. This feature direction indicates the change direction of the hidden vector corresponding to the three-dimensional model when the three-dimensional model is adjusted based on the first appearance feature. Then, the model adjustment system determines the second hidden vector according to the first appearance feature, the first hidden vector, and the adjustment parameter.
[0106] For example, if the first hidden vector corresponding to the first three-dimensional model is θ1 and the feature direction corresponding to the first appearance feature is θ0, then based on the first hidden vector, the second hidden vector θ2 after adjusting the first appearance feature is θ2 = θ1 + α * θ0. Here, α is the adjustment parameter, which is used to indicate the adjustment direction and adjustment intensity of the mapping component representing the first appearance feature in the first hidden vector. The adjustment direction includes the positive direction and the negative direction. When α is a positive value, it means adjusting in the positive direction, and when α is a negative value, it means adjusting in the negative direction. The magnitude of the α value determines the degree of semantic adjustment. The larger the absolute value of α, the greater the adjustment degree of the first appearance feature. For example, if the first appearance feature is age, the positive direction is to increase age, and the negative direction is to decrease age. The larger the absolute value of α, the older (or younger) the adjusted digital human is.
[0107] In this application, the feature direction corresponding to each feature of the above three-dimensional model can be determined by machine learning methods. For any feature of the three-dimensional model, such as the first appearance feature, when determining the feature direction corresponding to this feature, the model adjustment system first obtains the third training set, which includes multiple sample three-dimensional models; then converts the multiple sample three-dimensional models into hidden vectors respectively through the above first encoder, and labels each sample three-dimensional model to obtain a feature training set, which includes the hidden vectors corresponding to the multiple sample three-dimensional models after labeling; then determines the feature direction corresponding to the first appearance feature based on the machine learning algorithm and the above feature training set.
[0108] Exemplarily, if the above first three-dimensional model is a digital human and the first appearance feature is gender, after obtaining the feature training set by gender-labeling multiple sample three-dimensional models, using the hidden vector corresponding to each digital human in the feature training set as the input and the gender corresponding to each digital human as the label, train a support vector machine (SVM) to obtain the vector direction of the hyperplane corresponding to the SVM, that is, the feature direction corresponding to the gender feature of the digital human. If the user adjusts the gender feature of a digital human, when adjusting the feature direction of gender to the first direction, the digital human will become more masculine, and if adjusting the feature direction of gender to the second direction, the digital human will become more feminine.
[0109] Optionally, the third training set may be the above-mentioned first training set, or samples generated by the three-dimensional model generation model by sampling from the latent space corresponding to the three-dimensional model generation model.
[0110] Optionally, the above machine learning algorithm may also be algorithms such as support vector regression (SVR), principal component analysis (PCA), or kernel principal component analysis (kernel PCA). It should be understood that when using unsupervised learning algorithms such as PCA, it is not necessary to label the sample three-dimensional models.
[0111] S250. The model adjustment system determines a target three-dimensional model obtained by adjusting the first three-dimensional model according to the input parameters.
[0112] After obtaining the input parameters including the second latent vector, the model adjustment system can determine the three-dimensional model obtained by adjusting the first three-dimensional model according to the second latent vector.
[0113] The model adjustment system can obtain the three-dimensional model obtained by adjusting the first three-dimensional model through the following two methods.
[0114] The first method: The model adjustment system inputs the second latent vector into the three-dimensional model generation model, and decodes the second latent vector through the three-dimensional model generation model to obtain the adjusted three-dimensional model.
[0115] The second method: The model adjustment system inputs the first latent vector into the three-dimensional model generation model to obtain a second three-dimensional model, inputs the second latent vector into the three-dimensional model generation model to obtain a third three-dimensional model, determines the adjustment difference according to the third three-dimensional model and the second three-dimensional model, and then determines the three-dimensional model obtained by adjusting the first three-dimensional model according to the adjustment difference and the first three-dimensional model. Wherein, if the first three-dimensional model is M1, the first latent vector corresponding to the first three-dimensional model is w1, the second latent vector is w2, and the three-dimensional model generation model is denoted as G, the second three-dimensional model can be expressed as G(w1), and the third three-dimensional model can be expressed as G(w2). Then the three-dimensional model M0 obtained by adjusting the first three-dimensional model = M1 + G(w2) - G(w1), where the adjustment difference is G(w2) - G(w1).
[0116] As Figure 5 shown, Figure 5It is a schematic diagram of a method provided by this application for determining a third 3D model. The first 3D model M1 is processed by a first encoder to obtain a first latent vector w1. The first latent vector is adjusted to obtain a second latent vector w2. The second latent vector w2 is processed by a 3D model generation model to obtain G(w2). Among them, the method of converting the first 3D model M1 into the first latent vector w1 can refer to the relevant introduction in S210 above and will not be elaborated here.
[0117] Since there is a conversion error δ = G(w1) - M1 in the process of converting the second 3D model G(w1) generated from the first latent vector w1. Therefore, there is also the above conversion error δ between the third 3D model G(w2) and the model that the actual user wants to obtain. To reduce the above error, the above conversion error is eliminated based on the third 3D model G(w2), so that the target 3D model M0 = G(w2) - δ = G(w2) - [G(w1) - M1] = M1 + G(w2) - G(w1). Through the above method, the error generated in the conversion process between the latent vector and the 3D model can be further reduced, and the gap between the obtained target 3D model and the 3D model that the user actually wants to obtain can be reduced.
[0118] In this application, when a 3D model needs to be adjusted, the model adjustment system only needs to obtain the adjustment parameters for one feature, and does not require the user to directly edit the 3D model. The model adjustment system can obtain the input parameters for generating the adjusted target 3D model according to the adjustment parameters. Through this input parameter and the 3D model generation model, the target 3D model can be generated, thereby realizing the adjustment of the 3D model. This method enables ordinary users to adjust the 3D model according to their needs and achieve the expected effect, reduces the difficulty of editing the 3D model, and improves the efficiency of model adjustment. For example, by providing a model adjustment interface as Figure 4 shown, the user only needs to select the feature to adjust the 3D model, and then input the direction and size of the adjustment for this feature, and the editing of the 3D model can be realized. There is no need to directly edit the 3D model on the 3D model, which reduces the difficulty of editing the 3D model. It enables ordinary users without art foundation to adjust the 3D model according to their needs and achieve the expected effect, and improves the efficiency of model adjustment.
[0119] In addition, the 3D model has multiple appearance features, and the latent vector can represent the correlation relationship between these multiple appearance features. The appearance features having a correlation relationship means that after one of the features is adjusted, the adjustment operation will also adjust other features that have a correlation relationship with this feature.
[0120] Exemplarily, the appearance features included in the digital human include age, gender, skin color, face shape, eye size, eye socket depth, nose bridge height, nose bridge width, number of wrinkles, skin color depth, facial skin relaxation degree, hair density, etc. There are correlation relationships between some of these features. For example, the age feature has a correlation relationship with features such as the number of wrinkles, facial skin color depth, and facial skin relaxation degree. When the age is small, there are fewer wrinkles, the facial skin color is light, the facial skin is firm, and the hair is denser; when the age is large, there are more wrinkles, the facial skin color is dark, the facial skin is loose, and the hair is sparser. Another example is that there is a correlation relationship between the nose bridge height and the nose bridge width. When the nose bridge is high, the nose bridge is relatively narrow; when the nose bridge is low, the nose bridge is relatively wide.
[0121] In this application, since the three-dimensional model is converted into a latent vector by using an artificial intelligence-based encoder, during the process of training the encoder to convert the three-dimensional model into a latent vector, through a three-dimensional model, the characteristics of different features of the three-dimensional model can be learned, that is, a latent vector can represent multiple features of the three-dimensional model and the correlation relationships between multiple features. For example, if the training set used to train the encoder is a digital human, during the training process of the encoder, for digital humans of different ages in the dataset, the characteristics of digital humans of different ages can be learned. For example, digital humans with a small age have fewer wrinkles, light skin color, firm facial skin, etc., and digital humans with a large age have more wrinkles, dark skin color, loose facial skin; another example is for the nose bridge, the nose bridge of a digital human with a low nose bridge is wider than that of a digital human with a high nose bridge.
[0122] Therefore, in the above embodiments, if there are other features among the multiple appearance features of the first three-dimensional model that have a correlation relationship with the first appearance feature, for example, the second appearance feature has a correlation relationship with the first appearance feature, then when adjusting the first appearance feature of the first three-dimensional model, the second appearance feature will also be adjusted. The above target three-dimensional model is the three-dimensional model after adjusting the first appearance feature and the second feature. The linkage between feature adjustments is realized, making the adjustment of the three-dimensional model more accurate and more in line with the change relationship between the various features of the physical object corresponding to the three-dimensional model in reality. For example, when the nose bridge is lowered, the nose bridge will become wider synchronously; when the age is adjusted to be smaller, the wrinkles on the face will decrease, the skin color will become lighter, and the facial skin will become more plump.
[0123] It should be noted that the above Figures 2 - 5 introduces the three-dimensional model adjustment method provided by this application by adjusting the first appearance feature of the first three-dimensional model. That is, it introduces the operations performed by the adjustment system when the user makes a single adjustment to the three-dimensional model. A single adjustment can include when the user is Figure 4After selecting a feature on the model adjustment interface shown and selecting the adjustment direction and size for the feature, click the confirmation button. The user can adjust any feature of a 3D model one or more times, or adjust multiple features one or more times separately. The method for each adjustment can refer to the relevant description in the corresponding embodiment Figures 2 - 5 described above, and will not be elaborated here.
[0124] In other embodiments provided by the present application, when adjusting the first feature, the adjustment degree of the second feature related to the first feature can be inferred and generated by a correlation neural network, and the correlation neural network can be trained by the training sets of the first feature and the second feature; the adjustment degree of the second feature related to the first feature can also be obtained by numerical calculation, and the present application does not limit this.
[0125] For the above method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence; secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the present invention. Other reasonable step combinations that those skilled in the art can think of based on the above description also fall within the protection scope of the present invention.
[0126] See Figure 6 , Figure 6 is a schematic diagram of a model adjustment system provided by the present application. The above model adjustment system 120 includes a 3D model adjustment device 610, and the 3D model adjustment device 610 includes a receiving module 612 and a processing module 613.
[0127] The processing module 613 is used to determine the first 3D model that needs to be adjusted. The first 3D model can be a 3D model in the database. For example, the first 3D model is a 3D model selected by the user from the database that needs to be adjusted. The first 3D model can also be a 3D model uploaded by the user through the client. For example, the model adjustment system provides a model adjustment interface to the client, and the user uploads the first 3D model that needs to be adjusted through the model adjustment interface. The first 3D model can also be a 3D model generated by the model adjustment system after random sampling in the above latent space, and the present application does not make specific limitations.
[0128] The receiving module 612 is configured to receive a feature adjustment operation, which indicates to adjust a first appearance feature among a plurality of appearance features by using adjustment parameters, where the adjustment parameters include an adjustment intensity and / or an adjustment direction. After a user adjusts the first appearance feature of the first 3D model, the model adjustment system can receive the feature adjustment operation, which indicates to adjust the first appearance feature by using the adjustment parameters. Among them, the adjustment parameters include an adjustment intensity or an adjustment direction, and the adjustment parameters may also include both an adjustment intensity and an adjustment direction. The adjustment intensity indicates the degree of adjustment of the first appearance feature, and the adjustment direction is used to indicate the direction in which the first appearance feature is adjusted. For example, if the first 3D model is a digital human and the first appearance feature is the eye size, the adjustment intensity may be the adjusted eye size, or it may be the size or ratio of the increase or decrease of the eyes based on the current eye size of the digital human. The adjustment direction refers to whether to adjust in the direction of larger eyes or smaller eyes based on the current eye size of the digital human. For the relevant descriptions of the receiving module 612 receiving the model adjustment operation and the processing module 613 obtaining the adjustment parameters, reference can be made to the relevant descriptions in S230 above, and details are not repeated here.
[0129] The processing module 613 is configured to determine input parameters of the model for generating a 3D model according to the adjustment parameters; generate a target 3D model according to the input parameters and the 3D model. Specifically, the method for the processing module 613 to determine the target 3D model according to the adjustment parameters can refer to the relevant descriptions in S240 and S250 above, and details are not repeated here.
[0130] Specifically, the method for the 3D model adjustment device 610 to adjust the first 3D model can refer to the description in the corresponding embodiment above Figures 2 - 5 and details are not repeated here.
[0131] Both the above receiving module 612 and processing module 613 can be implemented by software or by hardware. Exemplarily, next, taking the processing module 613 as an example, the implementation manner of the processing module 613 is introduced, and the implementation manners of other modules can refer to the implementation manner of the processing module 613.
[0132] As an example of a software functional unit, the processing module 613 includes code running on a computing instance. The computing instance includes at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance can be one or more. For example, the processing module 613 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code can be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code can be distributed in the same availability zone (AZ), or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. One region can include multiple AZs.
[0133] Similarly, the multiple hosts / virtual machines / containers for running this code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0134] The processing module 613 can be deployed on at least one computing device, such as a server, etc. Alternatively, the processing module 613 can be a device implemented using a central processing unit (CPU), or can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. Among them, the above PLD can be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.
[0135] The multiple computing devices included in the processing module 613 can be distributed in the same region or in different regions. The multiple computing devices included in the processing module 613 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the weight determination unit 133 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and acceleration cards.
[0136] It should be noted that in other embodiments, the above receiving module 612 and processing module 613 can both be used to execute any step of the three-dimensional model adjustment method provided in this application. The steps that these modules are responsible for implementing respectively can be specified as needed, and all functions of the three-dimensional model adjustment are realized by implementing different steps in model adjustment through them respectively.
[0137] It should also be noted that the above three-dimensional model adjustment device 610 is used to execute any embodiment of the three-dimensional model adjustment method provided in this application. For specific details, please refer to Figures 2 - 5 the relevant descriptions in the corresponding embodiments, which will not be elaborated here. Figure 6 The three-dimensional model adjustment device 610 is only illustrated by the above division of each functional module. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the three-dimensional model adjustment device 610 is divided into other different functional modules to complete all or part of the functions described above.
[0138] As Figure 6 shown, the above model adjustment system 120 further includes a training device 620. The training device 620 includes a training module for randomly sampling in the latent space through the three-dimensional model generation model to generate multiple sample three-dimensional models; respectively converting the multiple sample three-dimensional models into their corresponding latent vectors through the first encoder; adding labels to the multiple sample three-dimensional models respectively according to the first appearance feature to obtain a feature training set; and finally determining the feature direction of the first appearance feature based on a machine learning algorithm and the feature training set. Among them, the label is used to indicate the intensity and / or direction of the first appearance feature of the corresponding sample three-dimensional model, and the feature training set includes the latent vector corresponding to each sample three-dimensional model and the label corresponding to each sample three-dimensional model.
[0139] The training device 620 is also used to train the generation model and the encoder. Among them, the generation model is used to convert the latent vector into a three-dimensional model; the encoder is used to convert the three-dimensional model into a latent vector. The process of the training device 620 training the three-dimensional model and the encoder can refer to the relevant introduction in the above method embodiments, which will not be elaborated here.
[0140] The present application also provides a computing device, such as Figure 7 shown Figure 7 is a schematic diagram of a computing device provided by the present application. The computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other through the bus 702. The computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.
[0141] The bus 702 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only one line is shown in , but it does not mean that there is only one bus or one type of bus. The bus 702 can include a path for transmitting information between various components of the computing device 700 (for example, the memory 706, the processor 704, the communication interface 708).
[0142] The above-mentioned processor 704 can be a Central Processing Unit (CPU), or can include a CPU and other hardware chips. There can be various types of the above-mentioned hardware chips. For example, the co-processing unit can include any one of chips such as a graphics processing unit (GPU), a tensor processing unit (TPU), a programmable logic device (PLD), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), or a digital signal processor (DSP). The computing device 700 can include one or more of any one of the above-mentioned types of hardware chips, or can include multiple types of the above-mentioned hardware chips. The present application does not make specific limitations.
[0143] The memory 706 may include volatile memory, such as random access memory (RAM). The memory 706 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0144] Executable program code is stored in the memory 706, and the processor 704 executes the executable program code to implement Figures 2 - 5 the method described in the corresponding method embodiment. That is to say, the memory 706 stores program code for implementing the functions of the model adjustment system 120, thereby implementing Figures 2 - 5 the three-dimensional model adjustment method shown. The program code includes one or more software modules, and the above one or more software modules include Figure 6 the receiving module 612 and the processing module 613 in the three-dimensional model adjustment device 610 shown, etc. The processor 704 executes the executable program code to implement Figures 2 - 5 the method described in the corresponding method embodiment, which will not be elaborated here.
[0145] The communication interface 108 can be a wired interface or a wireless interface for communicating with other modules or devices. For example, receiving the above-mentioned traceability request, receiving configuration information input by the user through the configuration interface, etc. The wired interface can be an Ethernet interface, a local interconnect network (LIN), etc., and the wireless interface can be a cellular network interface or use a wireless local area network interface, etc.
[0146] This application also provides a cluster of computing devices. The cluster of computing devices includes a plurality of computing devices 700. The computing device can be a server, such as a central server, an edge server, a local server in a local data center, or a server in a data center of a cloud environment, etc. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0147] As Figure 8 shown, Figure 8 is a schematic diagram of a cluster of computing devices provided by this application. Instructions for implementing the three-dimensional model adjustment method in the embodiment shown can be stored in the memories 706 of the plurality of computing devices 700 in the cluster of computing devices. Figures 2 - 5
[0148] In some possible implementations, the memories 706 of multiple computing devices 700 in the computing device cluster may also store some instructions for executing the above methods respectively, that is, the memories 706 in different computing devices 700 in the computing device cluster may store different instructions, respectively for executing Figures 2 - 5 partial functions of the three-dimensional model adjustment method shown, and the combination of multiple computing devices 700 can jointly implement Figures 2 - 5 the three-dimensional model adjustment method shown.
[0149] In some possible implementations, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 9 shows a possible implementation. As Figure 9 shown, Figure 9 is a schematic diagram of the connection between computing devices provided by the present application through a network. Two computing devices 700A and 700B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the function of the receiving module 612. At the same time, the memory 706 in the computing device 700B stores instructions for executing the function of the processing module 613.
[0150] It should be understood that Figure 9 the functions of the computing device 700A shown in
[0151] can also be jointly completed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be jointly completed by multiple computing devices 700. Figure 9 The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to
[0152] the connection method of the described computing device cluster. The difference is that the memories 706 in one or more computing devices 700 in this computing device cluster may store the same instructions for executing the three-dimensional model adjustment method.
[0153] The present application also provides a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it enables at least one computing device to implement Figures 2 - 5 the program code of the three-dimensional model adjustment method in the illustrated embodiment.
[0154] The present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to implement Figures 2 - 5 the three-dimensional model adjustment method in the illustrated embodiment.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional model adjustment method, characterized in that: The method comprises: Determine a first three-dimensional model, wherein the first three-dimensional model has a plurality of appearance features, wherein the plurality of appearance features include a first appearance feature and a second appearance feature, and the first appearance feature and the second appearance feature are associated with each other; receiving a feature adjustment operation, the feature adjustment operation indicating to adjust a first appearance feature among the plurality of appearance features using an adjustment parameter, wherein the adjustment parameter comprises an adjustment strength and / or an adjustment direction; Determining input parameters of a three-dimensional model generation model according to the adjustment parameters; A target three-dimensional model is generated according to the input parameters and a three-dimensional model generation model, wherein the target three-dimensional model is a three-dimensional model obtained by adjusting the first appearance feature and the second appearance feature of the first three-dimensional model.
2. The method according to claim 1, characterized in that The three-dimensional model generation model is used to generate a three-dimensional model using latent vectors corresponding to the three-dimensional model, and the method further includes: Determining a first latent vector corresponding to the first three-dimensional model; The step of determining the input parameters of the three-dimensional model generation model according to the adjustment parameters includes: Adjust the first latent vector according to the adjustment parameter to obtain a second latent vector, wherein the input parameter includes the second latent vector; The step of generating a target three-dimensional model according to the input parameters and the three-dimensional model generation model comprises: The second latent vector is input into the three-dimensional model generation model to generate the target three-dimensional model.
3. The method according to claim 2, characterized in that The determining a first latent vector corresponding to the first three-dimensional model includes: The first three-dimensional model is converted into a first latent vector by a first encoder.
4. The method according to claim 2 or 3, characterized in that: The step of inputting the second latent vector into the three-dimensional model generation model to generate the target three-dimensional model includes: Inputting the first latent vector into the three-dimensional model generation model to obtain a second three-dimensional model; Inputting the second latent vector into the three-dimensional model generation model to obtain a third three-dimensional model; determining an adjustment difference based on the third three-dimensional model and the second three-dimensional model; the adjustment difference being a difference between the third three-dimensional model and the second three-dimensional model; The target three-dimensional model is determined according to the adjustment difference and the first three-dimensional model.
5. The method according to claim 3, characterized in that: The converting the first three-dimensional model into a first latent vector by a first encoder includes: Inputting the first three-dimensional model into the first encoder to obtain a third latent vector; Inputting the third latent vector into the three-dimensional model generation model to obtain a fourth three-dimensional model; Determine a conversion error according to the first three-dimensional model and the fourth three-dimensional model, wherein the conversion error includes an error generated in a process of converting the first three-dimensional model into the third latent vector; The third latent vector is optimized according to the conversion error to determine the first latent vector.
6. The method according to any one of claims 2 to 4, characterized in that: The step of adjusting the first latent vector according to the adjustment parameter to obtain a second latent vector includes: determining a feature direction of the first appearance feature in a latent space, the feature direction indicating a direction in which a latent vector corresponding to the three-dimensional model changes when the three-dimensional model is adjusted based on the first appearance feature; The second latent vector is determined according to the first latent vector, a feature direction corresponding to the first appearance feature, and the adjustment parameter.
7. The method according to claim 5, characterized in that Before determining the second latent vector according to the first latent vector, the feature direction corresponding to the first appearance feature, and the adjustment parameter, the method further includes: Randomly sampling in the latent space through the three-dimensional model generation model to generate a plurality of sample three-dimensional models; The plurality of sample three-dimensional models are respectively converted into latent vectors corresponding to the plurality of sample three-dimensional models by the first encoder; adding labels to the plurality of sample three-dimensional models respectively according to the first appearance feature to obtain a feature training set, wherein the label is used to indicate the intensity and / or direction of the first appearance feature of the corresponding sample three-dimensional model, and the feature training set includes a latent vector corresponding to each sample three-dimensional model and a label corresponding to each sample three-dimensional model; Based on a machine learning algorithm and the feature training set, a feature direction of the first appearance feature is determined.
8. The method according to claim 3 or 5, characterized in that: Before converting the first three-dimensional model into a first latent vector by the first encoder, the method further includes: Inputting a fifth three-dimensional model into a second encoder to obtain a latent vector corresponding to the fifth three-dimensional model; the fifth three-dimensional model is one of the second training set, and the second training set includes multiple three-dimensional models; Obtaining a sixth three-dimensional model according to the three-dimensional model generation model and latent vectors corresponding to the fifth three-dimensional model; A value of a loss function is determined according to the fifth three-dimensional model and the sixth three-dimensional model, and the second encoder is updated based on the value of the loss function to obtain the first encoder.
9. The method according to any one of claims 1 to 8, characterized in that: The first three-dimensional model is a three-dimensional head model of a digital human; at least one feature of the first three-dimensional model includes part or all of gender, age, skin color, face shape, eye size, or nose bridge height.
10. A three-dimensional model adjustment device, characterized in that: include: A processing module, configured to determine a first three-dimensional model, wherein the first three-dimensional model has a plurality of appearance features, wherein the plurality of appearance features include a first appearance feature and a second appearance feature, and the first appearance feature and the second appearance feature are associated with each other; A receiving module, configured to receive a feature adjustment operation, wherein the feature adjustment operation indicates to use an adjustment parameter to adjust a first appearance feature among the plurality of appearance features, wherein the adjustment parameter includes an adjustment strength and / or an adjustment direction; The processing module is further used to determine the input parameters of the three-dimensional model generation model according to the adjustment parameters; The processing module is further used to generate a target three-dimensional model according to the input parameters and a three-dimensional model generation model, wherein the target three-dimensional model is a three-dimensional model obtained by adjusting the first appearance feature and the second appearance feature of the first three-dimensional model.
11. The device according to claim 10, characterized in that The three-dimensional model generation model is used to generate a three-dimensional model using latent vectors corresponding to the three-dimensional model, and the processing module is also used to: Determining a first latent vector corresponding to the first three-dimensional model; The processing module is specifically used for: Adjust the first latent vector according to the adjustment parameter to obtain a second latent vector, wherein the input parameter includes the second latent vector; The second latent vector is input into the three-dimensional model generation model to generate the target three-dimensional model.
12. The device according to claim 11, characterized in that The processing module is specifically used for: The first three-dimensional model is converted into a first latent vector by a first encoder.
13. The device according to claim 11 or 12, characterized in that The processing module is specifically used for: Inputting the first latent vector into the three-dimensional model generation model to obtain a second three-dimensional model; Inputting the second latent vector into the three-dimensional model generation model to obtain a third three-dimensional model; determining an adjustment difference based on the third three-dimensional model and the second three-dimensional model; the adjustment difference being a difference between the third three-dimensional model and the second three-dimensional model; The target three-dimensional model is determined according to the adjustment difference and the first three-dimensional model.
14. The device according to claim 12, characterized in that The processing module is specifically used for: Inputting the first three-dimensional model into the first encoder to obtain a third latent vector; Inputting the third latent vector into the three-dimensional model generation model to obtain a fourth three-dimensional model; Determine a conversion error according to the first three-dimensional model and the fourth three-dimensional model, wherein the conversion error includes an error generated in a process of converting the first three-dimensional model into the third latent vector; The third latent vector is optimized according to the conversion error to determine the first latent vector.
15. The device according to any one of claims 11 to 13, characterized in that: The processing module is specifically used for: determining a feature direction of the first appearance feature in the latent space, the feature direction indicating a direction in which a latent vector corresponding to the three-dimensional model changes when the three-dimensional model is adjusted based on the first appearance feature; The second latent vector is determined according to the first latent vector, a feature direction corresponding to the first appearance feature, and the adjustment parameter.
16. The device according to claim 14, characterized in that The device also includes a training module for: Randomly sampling in the latent space through the three-dimensional model generation model to generate a plurality of sample three-dimensional models; The plurality of sample three-dimensional models are respectively converted into latent vectors corresponding to the plurality of sample three-dimensional models by the first encoder; adding labels to the plurality of sample three-dimensional models respectively according to the first appearance feature to obtain a feature training set, wherein the label is used to indicate the intensity and / or direction of the first appearance feature of the corresponding sample three-dimensional model, and the feature training set includes a latent vector corresponding to each sample three-dimensional model and a label corresponding to each sample three-dimensional model; Based on a machine learning algorithm and the feature training set, a feature direction of the first appearance feature is determined.
17. The device according to claim 12 or 14, characterized in that The processing module is also used for: Inputting a fifth three-dimensional model into a second encoder to obtain a latent vector corresponding to the fifth three-dimensional model; the fifth three-dimensional model is one of the second training set, and the second training set includes multiple three-dimensional models; Obtaining a sixth three-dimensional model according to the three-dimensional model generation model and latent vectors corresponding to the fifth three-dimensional model; A value of a loss function is determined according to the fifth three-dimensional model and the sixth three-dimensional model, and the second encoder is updated based on the value of the loss function to obtain the first encoder.
18. The device according to any one of claims 10 to 17, characterized in that: The first three-dimensional model is a three-dimensional head model of a digital human; at least one feature of the first three-dimensional model includes part or all of gender, age, skin color, face shape, eye size, or nose bridge height.
19. A computing device cluster, characterized in that: The system comprises at least one computing device, each computing device comprises a processor and a memory, and the processor of each computing device is used to execute instructions stored in the memory, so that the computing device cluster executes the method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method according to any one of claims 1 to 9.
21. A computer program product, characterized in that When the computer program product is executed by a computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 9.