Training methods for 3D digital human models and methods for generating 3D digital humans
By training and calibrating 3D digital human models, and using sample videos and audio-visual generation models with personalized parameters, the problem of unstable 3D digital human rendering was solved, achieving more stable and accurate 3D digital human generation.
Patent Information
- Application Number
- CN202411643621.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-15
AI Technical Summary
The existing 3D digital human rendering methods are unstable, leading to problems such as head shaking, deformation, and torso separation, which affect the stability of 3D digital humans.
By acquiring sample videos with personalized parameters, extracting sample data and training an audio-visual generation model, calibrating the initial digital human using the head parameters of sample users, generating a more stable personalized and general digital human generation model, and adjusting it in conjunction with the animated initial digital human to generate a more stable 3D digital human.
It improves the stability and accuracy of 3D digital humans, ensuring that head movements are closer to those of users in sample videos, thus enhancing the performance capabilities of digital humans.
Smart Images

Figure CN119516061B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, specifically a method for training a three-dimensional digital human model and a method for generating a three-dimensional digital human. Background Technology
[0002] A 3D digital human is a virtual avatar that can be displayed on a client interface. By inputting voice, the 3D digital human can perform corresponding actions, allowing the audience to visually perceive that the voice is spoken by the 3D digital human.
[0003] In related technologies, the rendering methods used are unstable. For example, the generated 3D digital human head is prone to shaking, deformation, and separation of the head and torso, resulting in poor stability of the 3D digital human obtained by this method. Summary of the Invention
[0004] This application provides a method for training a 3D digital human model and a method for generating a 3D digital human, which can improve the stability of the 3D digital human. The method for training a 3D digital human model and a method for generating a 3D digital human provided in this application are implemented as follows:
[0005] In one aspect of this application, a method for training a three-dimensional digital human model is provided, the method comprising:
[0006] Obtain sample videos corresponding to different personalized parameters, which include at least one of identity information and emotion information;
[0007] The corresponding sample data are obtained from the sample videos corresponding to each personalized parameter. The sample data includes sample speech, sample user emotion, and sample user head parameters. The sample user head parameters are obtained after calibration based on the position of the target head key points of the sample user in the sample video.
[0008] Based on the sample data corresponding to each personalized parameter, train the audio-visual generation model to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0009] Different personalized digital human generation models are normalized to obtain a general digital human generation model.
[0010] In this scheme, the sample data in the sample video includes the head parameters of the sample user. Since these head parameters are obtained by calibrating the points to be calibrated on the initial digital human after establishing the head key points of the calibration digital human based on the sample video, the final personalized digital human generation model and the general digital human generation model can generate digital human heads that are more stable and closer to the head movements of the user in the sample video, thereby enhancing the stability of the digital human.
[0011] In one embodiment, obtaining corresponding sample data from sample videos corresponding to each personalized parameter includes: establishing a corresponding training digital human based on the sample video, and determining the head parameters of the sample user based on the head parameters of the training digital human; identifying the emotions of the sample user through the facial parameters of the training digital human; determining the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0012] In this scheme, by establishing a training digital human, the head parameters and emotions of the sample users can be obtained more quickly and accurately, thus enabling the acquisition of sample data more quickly and accurately.
[0013] In one embodiment, the sample video includes multiple video frames. Establishing a corresponding training digital human based on the sample video includes: establishing a calibration digital human using multiple video frames in the sample video, wherein the calibration digital human includes multiple frames; calibrating the initial digital human based on the target head key points of the calibration digital human to obtain the training digital human.
[0014] In this scheme, a more accurate calibrated digital human can be obtained by using multiple video frames from the sample video. Then, the initial digital human can be calibrated based on the target head key points in the calibrated digital human, resulting in a more accurate trained digital human, which can improve the accuracy and stability of the sample data.
[0015] In one embodiment, the target head key points include a first head key point and a second head key point, wherein the second head key point is determined based on optical flow information in a sample video. The initial digital person is calibrated based on the target head key points of the calibration digital person to obtain a training digital person. This includes: performing a first calibration on a first point to be calibrated on the initial digital person based on a first loss function and the position of the first head key point in the calibration digital person, to obtain a first-calibrated digital person; wherein the first head key point corresponds to the first point to be calibrated on the initial digital person; and performing a second calibration on a second point to be calibrated on the first-calibrated digital person based on a second loss function and the position of the second head key point in the calibration digital person, to obtain a training digital person; wherein the second head key point corresponds to the second point to be calibrated on the first-calibrated digital person.
[0016] In one embodiment, the first calibrated digital human is the digital human obtained when the first loss function is minimized during the first calibration process of the first calibration point on the initial digital human; the training digital human is the digital human obtained when the second loss function is minimized during the second calibration process of the second calibration point on the first calibrated digital human.
[0017] In one embodiment, determining a second head key point based on optical flow information in a sample video includes: acquiring light density and color in the sample video; determining the head region of the sample user based on the light density and color; determining a second head key point from multiple pixels in the head region based on the brightness and / or contrast with adjacent pixels in each pixel in the head region, wherein the brightness and / or contrast with adjacent pixels of the second head key point is greater than a preset threshold.
[0018] In the above scheme, by determining the first loss function and the second loss function in sequence, and then calibrating the initial digital human by calibrating the digital human, a more accurate training digital human can be obtained, and further, sample data can be obtained more accurately.
[0019] In one embodiment, identifying the emotions of sample users by training the facial parameters of a digital human includes: calculating the similarity between the facial parameters of the trained digital human and the facial parameters of various pre-set emotion avatar models; determining the emotion label of the trained digital human based on the similarity, wherein the emotion label includes the emotion and weight of the target emotion avatar model, the target emotion avatar model being the emotion avatar model corresponding to a similarity greater than or equal to a threshold, and the emotion of the sample user including the emotion label of the trained digital human.
[0020] In this scheme, by determining emotion labels, the emotions of each training digital human can be identified in more detail. In addition to the emotions themselves, the weights corresponding to the emotions can also be determined, thus obtaining more accurate and detailed emotion labels.
[0021] In one embodiment, the method further includes: determining the weight of sample data based on the degree of matching between the head parameters of the sample user in the sample video and the sample speech corresponding to the sample video, wherein the weight of the sample data is positively correlated with the degree of matching; and training an audio-visual generation model based on the sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter, including: training an audio-visual generation model based on sample data whose weights corresponding to each personalized parameter are greater than a preset weight threshold to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0022] In this scheme, by setting the weights of the sample data, higher quality and more effective sample data can be obtained, thereby improving the accuracy of the model trained.
[0023] Another aspect of this application embodiment also provides a method for generating a three-dimensional digital human, the method comprising:
[0024] Obtain personalized parameters for the target; personalized parameters for the target include at least one of the target's identity information and the target's emotional information.
[0025] Based on the target personalized parameters and mapping relationship, the target adjustment parameters are determined; the mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters, and the adjustment parameters are used to indicate the deviation of head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model.
[0026] Based on the target adjustment parameters and the initial animation digital human, the target animation digital human is generated. The initial animation digital human is generated based on the first voice input to the general digital human generation model.
[0027] In this scheme, by using a digital human generation model, a more stable 3D digital human can be obtained, thereby improving the stability of the 3D digital human.
[0028] In one possible implementation, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, each frame adjustment parameter corresponding to one speech frame; based on the target adjustment parameters and the initial animation digital man, the animation target digital man is generated, specifically including: based on a frame adjustment parameter corresponding to each speech frame, adjusting each frame of the initial animation digital man corresponding to that speech frame to generate the animation target digital man.
[0029] In this scheme, the adjustment of the three-dimensional digital human can be achieved more comprehensively by adjusting each frame of the initial digital human animation. Furthermore, since each frame adjustment parameter corresponds to a voice frame, adjusting each frame of the initial digital human animation based on the frame adjustment parameters corresponding to each voice frame can also result in a more realistic three-dimensional digital human that better matches the content of the first voice.
[0030] In one possible implementation, each frame of the digital human includes multiple grid cells in three-dimensional space. The adjustment parameters for each frame include: the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted. Based on the adjustment parameters for each frame, each frame of the initial animated digital human corresponding to the speech frame is adjusted to generate the target animated digital human. Specifically, this includes: adjusting the three-dimensional spatial coordinates of the grid cells to be adjusted in the initial animated digital human corresponding to the speech frame based on the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted in the adjustment parameters for each frame to generate the target animated digital human corresponding to the speech frame.
[0031] In this scheme, by using the identifier and offset value of the grid cell to be adjusted, the grid cell that needs to be adjusted and the value that needs to be adjusted for each grid cell can be determined more accurately and quickly. This allows for faster and more accurate adjustment of the grid cells, resulting in a more accurate animated target digital human.
[0032] In one possible implementation, the method further includes sending the animated target digital man to the client so that the client renders the animated target digital man.
[0033] In this solution, by transmitting the target animated digital human, the client can more accurately and quickly render the animated target digital human, and then the client can display the rendered 3D digital human.
[0034] In one possible implementation, sending the animated target digital human to the client specifically includes: sending the target digital human and target deviation parameters corresponding to the first voice frame in the first voice to the client; wherein, the target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent voice frames in the animated target digital human.
[0035] In this scheme, by determining the target digital human and the target deviation parameters, the client can more accurately render the 3D digital human.
[0036] In one possible implementation, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital person corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; sending the target digital person and target deviation parameters corresponding to the first speech frame in the first speech to the client includes: sending multiple target deviation parameters to the client sequentially according to the time sequence of each speech frame in the first speech.
[0037] In this scheme, the transmission of multiple target deviation parameters is carried out in the order of the voice frames in the first voice, which can more accurately realize the rendering of the 3D digital human by the client, thereby improving the real-time performance and accuracy of the 3D digital human rendering by the client.
[0038] In one possible implementation, sending the target deviation parameter to the client includes: determining the target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; and sending the target deviation parameter to the client.
[0039] In this scheme, the number of mesh cells with differences in the deviation parameters can be reduced by determining the target deviation parameter from the deviation parameter, thereby reducing the amount of information transmitted between the client and the server, reducing the amount of computation on the client during the rendering process, and improving the rendering efficiency.
[0040] The computing device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the method of this application.
[0041] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method provided in this application embodiment.
[0042] The training method for the 3D digital human model and the 3D digital human generation method provided in this application embodiment can acquire sample videos corresponding to different personalized parameters; obtain corresponding sample data from the sample videos corresponding to each personalized parameter; train the audio-visual generation model based on the sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter; and normalize the different personalized digital human generation models to obtain a general digital human generation model. Since the head parameters of the sample users are obtained by calibrating the points to be calibrated on the initial digital human based on the target head key points established in the calibration digital human from the sample videos, the audio-visual generation model can be trained based on more accurate and stable sample data, thereby obtaining a more stable and accurate general digital human generation model. This general digital human generation model can then be used to obtain a more stable 3D digital human, thus improving the stability of the 3D digital human. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram showing the interface of the three-dimensional digital human generation method provided in the embodiments of this application;
[0045] Figure 2 This is a schematic diagram illustrating an application scenario of the three-dimensional digital human generation method provided in the embodiments of this application;
[0046] Figure 3 This is a flowchart illustrating the training method for the three-dimensional digital human model provided in the embodiments of this application;
[0047] Figure 4 This is a schematic diagram of the process for determining sample data provided in the embodiments of this application;
[0048] Figure 5 This is a schematic diagram of a process for obtaining a trained digital human, provided in an embodiment of this application.
[0049] Figure 6 This is a schematic diagram of the head point cloud provided in the embodiments of this application;
[0050] Figure 7 This is a schematic diagram of a process for calibrating an initial digital human, as provided in an embodiment of this application.
[0051] Figure 8 This is a schematic diagram of the process for identifying the emotions of sample users provided in the embodiments of this application;
[0052] Figure 9 This is a schematic diagram of the emotion avatar model provided in the embodiments of this application;
[0053] Figure 10 This is a flowchart illustrating the three-dimensional digital human generation method provided in the embodiments of this application;
[0054] Figure 11 This is a schematic diagram of the head mesh of the 3D digital human provided in the embodiments of this application. Please refer to... Figure 11 ;
[0055] Figure 12 This is a schematic diagram illustrating the interface changes for rendering a 3D digital human based on deviation parameters, as provided in the embodiments of this application.
[0056] Figure 13 This is a schematic diagram of the structure of the training device for the three-dimensional digital human model provided in the embodiments of this application;
[0057] Figure 14 This is a schematic diagram of the structure of the three-dimensional digital human generation device provided in the embodiments of this application. Please refer to... Figure 14 ;
[0058] Figure 15 This is a schematic diagram of the structure of the computing device provided in the embodiments of this application. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0061] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0062] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0063] The following explains the three-dimensional digital human provided in the embodiments of this application and its practical application scenarios.
[0064] First, a 3D digital human can be a virtual avatar displayed on a client, such as a virtual anchor or virtual teacher.
[0065] Optionally, a 3D digital human is a virtual character developed based on computer technology and artificial intelligence technology. It features high-quality character image and exquisite visual expression. During use, users can input voice and make the 3D digital human perform corresponding actions and expressions based on the voice, thus making the audience mistakenly believe that the voice is made by the 3D digital human.
[0066] This technology can be used in various fields to meet corresponding needs. For example, 3D digital humans can be used to simulate real live streamers for live e-commerce, live gaming, or live chatting. Alternatively, 3D digital humans can be used to simulate real teachers for online teaching.
[0067] The 3D digital humans used in the above scenarios are mainly generated in real time. In actual use, it may be necessary to generate them in advance and then render them in non-real time. For example, 3D digital humans are used in movies and TV series to replace actors, or videos are made based on 3D digital humans. There are no specific restrictions here. You can choose to render the 3D digital human in real time or render it in advance according to actual needs.
[0068] In real-time rendering scenarios, taking the aforementioned virtual anchor as an example, the 3D digital human can be displayed on the live broadcast screen, and the 3D digital human can display different actions and expressions based on the input voice, such as happy actions and expressions, angry actions and expressions, etc., without specific limitations.
[0069] In the application of 3D digital humans, in order to reduce the pressure on the client, the client can perform the rendering, while the server generates the data needed for the rendering process and sends it to the client.
[0070] For the client, the corresponding 3D digital human can be displayed in the corresponding display interface. The following explains the scenario of the 3D digital human being displayed in the client interface provided in the embodiments of this application.
[0071] Figure 1 This is a schematic diagram illustrating the interface of the 3D digital human generation method provided in this application embodiment. Please refer to... Figure 1 , Figure 1 The interface shown is the live streaming interface, which can display a 3D digital human 110. The appearance of the 3D digital human can be customized by the user, such as simulating a real person or an anime character, etc. There are no specific restrictions here.
[0072] In this live streaming scenario, users can input relevant parameters into the 3D digital human through voice or text input, thereby enabling the 3D digital human to perform corresponding actions.
[0073] For example, if a user inputs a voice expressing surprise, the corresponding 3D digital human can express a surprised expression based on that voice, and the lips can also make a surprised mouth shape.
[0074] In the application of 3D digital humans, in addition to the client needing to render and display the corresponding 3D digital human, the process of calculating the relevant parameters of each action of the 3D digital human can be completed by the server. The following explains the actual application scenario of the 3D digital human generation method provided in the embodiments of this application.
[0075] In this scenario, the server can be a single server, multiple servers, or a server cluster. The client can be various terminal devices, such as mobile phones, personal computers, tablets, etc., without specific limitations. The client and server are connected via a communication link.
[0076] Figure 2 This is a schematic diagram illustrating the application scenario of the 3D digital human generation method provided in the embodiments of this application. Please refer to... Figure 2 This scenario may include a server 210 and multiple clients 220 corresponding to the server 210.
[0077] Among them, the server 210 can be one end running on a server, such as a cloud server, AI (Artificial Intelligence) server, rack server, etc., which can realize the relevant parameter calculation of three-dimensional digital human. The AI server can be a server used to perform artificial intelligence calculations, and the rack server can be a server installed in a standardized network equipment rack, for example, without specific restrictions.
[0078] Client 220 can be a terminal device used by the user, such as including but not limited to mobile phones, wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), tablets, laptops, in-vehicle terminals, PCs (Personal Computers), etc., without specific restrictions.
[0079] It should be noted that in one application scenario, the server 210 can simultaneously perform data interaction for multiple clients 220. For example, the client sends relevant parameters of the 3D digital human to the server, the server calculates the relevant data required for rendering based on the relevant parameters, and then sends this data to the corresponding client so that the client can complete the rendering work.
[0080] In this scenario, this application provides a method for training a three-dimensional digital human model and a method for generating a three-dimensional digital human.
[0081] Figure 3 This is a flowchart illustrating the training method for the 3D digital human model provided in this embodiment. Please refer to... Figure 3 The method includes:
[0082] S310: Obtain sample videos corresponding to different personalized parameters.
[0083] Optionally, the sample videos can be pre-acquired 2D videos (two-dimensional videos), and these sample videos can be pre-classified to obtain sample videos corresponding to each personalized parameter.
[0084] The personalized parameters may include at least one of identity information and emotional information. The identity information may be the identity represented by the three-dimensional digital human, such as age, gender, etc.; the emotional information may be the emotions that the three-dimensional digital human can express, such as happiness, anger, sadness, etc.
[0085] It should be noted that corresponding sample videos can be configured according to the personalized parameters required, or the personalized parameters in these sample videos can be determined by collecting sample videos. No specific restrictions are imposed here.
[0086] For example, if there are 100 different sets of personalized parameters, a corresponding sample video can be set for each set of personalized parameters. Alternatively, if 20 sample videos are obtained in advance, the personalized parameters present in these 20 sample videos can be extracted.
[0087] It should be noted that the two-dimensional video can be a video of a real person speaking, such as a video of a teacher giving a lecture. In this sample video, there needs to be a real person's head image, preferably a facial image.
[0088] In any given two-dimensional video, during the actual implementation process, the real person may exhibit different emotions, such as changing from a happy emotion to a sad emotion, or the real person in the video may change, for example, the first half of the video may be filmed as a man, and the second half as a woman, etc.
[0089] Before obtaining sample videos, these two-dimensional videos can be segmented. For example, clustering can be used to identify segments in the video that have different emotions or identities, resulting in multiple segments. Each segment can then be used as a sample video.
[0090] Optionally, by segmenting the two-dimensional video, multiple sample videos can be obtained, and each sample video can correspond to a set of personalized parameters. For example, the personalized parameters corresponding to one sample video are: male, 50 years old, happy; and the personalized parameters corresponding to another sample video are: sad.
[0091] It should be noted that the correspondence between sample videos and personalized parameters can be presented in Table 1.
[0092] Table 1
[0093] Sample Video Personalized parameters Sample Video 1 male Sample Video 2 female Sample Video 3 Male, 20 years old … … Sample video N-1 Male, 50 years old, happy Sample video N Male, happy
[0094] Please refer to Table 1. Each sample video can correspond to a different set of personalized parameters. There can be one or more personalized parameters, and no specific restrictions are imposed here.
[0095] S320: Obtain the corresponding sample data from the sample videos corresponding to each personalized parameter.
[0096] The sample data includes sample audio, sample user emotions, and sample user head parameters.
[0097] Optionally, after obtaining the sample video, the corresponding sample data can be determined from the sample video. For the same sample video, multiple sets of sample data can be obtained, where each set of sample data can include sample speech, sample user's emotion, and sample user's head parameters.
[0098] The sample speech can be a part of the speech corresponding to the sample video, the sample user's emotion can be the emotion of a real person in the sample video, and the sample user's head parameters can be the head parameters of a three-dimensional model (i.e., the sample user) built in three-dimensional space based on the head image of a real person in the sample video.
[0099] For each personalized parameter, the sample video can yield the aforementioned multiple sets of sample data.
[0100] It should be noted that the head parameters of the sample user can specifically include the head contour parameters, facial parameters, lip parameters, etc. These parameters can be the positions of the corresponding mesh cells. For example, the head contour parameters can be the positions of the mesh cells that make up the head contour of the entire 3D digital human; the facial parameters can be the positions of the mesh cells that make up the face of the 3D digital human; and the lip parameters can be the positions of the mesh cells that make up the lips of the 3D digital human.
[0101] Optionally, the head parameters and emotions of the sample users can correspond to the sample speech, and each speech frame can correspond to the head parameters and emotions of a sample user.
[0102] For example: If a sample speech in a sample data contains 100 speech frames, then there can be 100 emotions and 100 head parameters of the sample user. The emotion and head parameters of each frame of the sample user correspond to one speech frame of the sample speech.
[0103] In one embodiment, the head parameters of the sample user are obtained by calibrating the points to be calibrated on the initial digital human based on the target head key points in the calibration digital human established from the sample video.
[0104] It should be noted that the target head key points of the sample users refer to a subset of key points selected from the heads of the sample users. The target head key points can be multiple key points in the calibration of the digital human head.
[0105] The initial digital human can include multiple calibration points, each of which can correspond to a target head key point. The position of the calibration point corresponding to the target head key point can be calibrated based on the position of the target head key point.
[0106] After calibrating the corresponding calibration points based on the key points of each target head, a calibrated digital human can be obtained. The head parameters of the calibrated digital human can then be used as the head parameters of the sample user.
[0107] It should be noted that a calibration digitized human can be a digitized human model without a specific physical model. For example, it can be a model framework where only the coordinates of key points on the head exist, but there is no corresponding specific appearance or style in three-dimensional space. A calibration digitized human can be obtained based on multiple frames of two-dimensional images from a sample video.
[0108] The initial digital human can be a digital human with a specific appearance and style, and each point of this initial digital human to be calibrated has an actual location in three-dimensional space. Alternatively, the initial digital human can be a digital human with a specific appearance and style pre-generated in three-dimensional space.
[0109] In one embodiment, the positions of the key points of each target head in the calibrated digit human can be used to calibrate the various points to be calibrated of the initial digit human, that is, to adjust the positions of the various points to be calibrated of the initial digit human, thereby obtaining a calibrated digit human with a specific appearance and style in three-dimensional space, which can be used as a training digit human.
[0110] S330: Based on the sample data corresponding to each personalized parameter, train the audio-visual generation model to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0111] Optionally, sample data corresponding to each personalized parameter can be obtained through the above method, and then the audio-visual generation model can be trained based on these sample data.
[0112] The audio-visual generation model can be an audio-visual encoder or a neural network structure. Its input can be a sample sound and the emotion of a sample user, and its output can be a three-dimensional digital human. The model can compare the head parameters of the output three-dimensional digital human with the head parameters of the sample user. After extensive training, if the audio-visual generation model converges, for example, if the matching degree between the head parameters of the output three-dimensional digital human and the head parameters of the sample user is greater than a preset similarity threshold, then the audio-visual generation model can be determined to have converged, thus obtaining the personalized digital human generation model corresponding to the personalized parameters.
[0113] It should be noted that the above training process can be performed on a set of personalized parameters corresponding to each sample video to obtain a personalized digital human generation model corresponding to each set of personalized parameters. In other words, multiple personalized digital human generation models are obtained by training the same audio-visual generation model with different sample data.
[0114] S340: Normalize different personalized digital human generation models to obtain a general digital human generation model.
[0115] It should be noted that after obtaining multiple personalized digital human generation models, these personalized digital human generation models can be normalized.
[0116] It should be noted that normalization methods can include various approaches. For example, input feature normalization can be used, meaning the input data is normalized before training these personalized digital human generation models to ensure similar scale across different sample data. Alternatively, inter-layer normalization can be employed. Since each personalized digital human generation model is trained based on the same audio-visual generation model, they possess a certain degree of similarity. For instance, if the neural network has the same number of layers, normalization can be applied after each layer is generated, especially before the activation function. Another approach is weight normalization, where the weights of each layer in the neural network are reparameterized, decoupling the length and direction of the weight vector, thus normalizing these personalized digital human generation models.
[0117] In practice, any of the above-mentioned normalization methods can be used to normalize multiple personalized digital human generation models, thereby obtaining a general digital human generation model.
[0118] A general digital human generation model can be obtained by normalizing all personalized digital human generation models.
[0119] In practical applications, the input parameter of this general digital human generation model can be a first speech, and the output can be a three-dimensional digital human.
[0120] The training method for the 3D digital human model provided in this application embodiment can acquire sample videos corresponding to different personalized parameters; obtain corresponding sample data from the sample videos corresponding to each personalized parameter; train the audio-visual generation model based on the sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter; and normalize the different personalized digital human generation models to obtain a general digital human generation model. Since the head parameters of the sample users are obtained by calibrating the points to be calibrated on the initial digital human based on the target head key points established in the calibration digital human from the sample videos, the audio-visual generation model can be trained based on more accurate and stable sample data, thereby obtaining a more stable and accurate general digital human generation model. This general digital human generation model can then be used to obtain a more stable 3D digital human, thus improving the stability of the 3D digital human.
[0121] The following explains the implementation process of determining sample data based on sample videos provided in the embodiments of this application.
[0122] Figure 4 This is a flowchart illustrating the process of determining sample data provided in the embodiments of this application. Please refer to... Figure 4 For sample videos, the corresponding sample speech in each frame of the sample video can be determined, and a three-dimensional model can be built based on the video image of the real person. Thus, the emotion and head parameters of the three-dimensional model built in three-dimensional space can be determined. The head parameters of the model are used as the head parameters of the sample user, and the facial emotion of the model is used as the emotion of the sample user.
[0123] In this process, a three-dimensional model can be built based on multiple frames of two-dimensional images in a two-dimensional video, which is called training a digital human. Then, the head parameters of the sample user can be determined based on the head parameters of the training digital human, and the emotions of the sample user can be identified.
[0124] In one embodiment, obtaining corresponding sample data from sample videos corresponding to each personalized parameter includes: establishing a corresponding training digital human based on the sample video, and determining the head parameters of the sample user based on the head parameters of the training digital human; identifying the emotions of the sample user through the facial parameters of the training digital human; determining the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0125] It should be noted that the audio can be extracted from the sample video and used as sample audio in the sample data.
[0126] Optionally, the head region of the real person in the video can be determined from the sample video, and then the head parameters of the sample user in the two-dimensional image can be determined from the head region. A training digital human can be established by converting multiple frames of two-dimensional images to three-dimensional modeling, and the head parameters of the training digital human can be obtained. The head parameters of the training digital human can be considered to be accurate parameters, that is, the head parameters of the training digital human can be used as the head parameters of the sample user in the sample data.
[0127] After obtaining the training digitized human, the emotions of the sample users can be identified based on the facial parameters of the training digitized human. For example, the emotions of the sample users can be determined by comparing the differences between the facial parameters of the 3D digitized human and the facial parameters of a preset 3D digitized human with emotional tendencies.
[0128] For each audio frame, there can be a corresponding training digit. The head parameters of the sample user corresponding to the training digit are the same as the head parameters of the sample user corresponding to the audio frame. Correspondingly, the emotion of the sample user corresponding to the training digit is the same as the emotion of the sample user corresponding to the audio frame.
[0129] The above steps allow us to obtain the head parameters and emotions of the sample user for each audio frame in the sample data.
[0130] It should be noted that the training digital human can be a digital human obtained by transforming two-dimensional graphics into three-dimensional ones. For example, projection modeling can be used to obtain the depth information of two-dimensional images by using different angles of a real person's head in multiple frames of two-dimensional images. Then, the training digital human can be built based on the depth information of the two-dimensional images and the real person's face image displayed in the two-dimensional images.
[0131] In the training method for a 3D digital human model provided in this application embodiment, a corresponding training digital human can be established based on sample videos, and the head parameters of the sample user can be determined based on the head parameters of the training digital human; the emotions of the sample user can be identified through the facial parameters of the training digital human; the sample speech corresponding to each personalized parameter can be determined, as well as the emotion and head parameters of the sample user corresponding to each speech frame of the sample speech. The method of establishing a training digital human allows for faster and more accurate acquisition of the sample user's head parameters and emotions, thus enabling faster and more accurate acquisition of sample data.
[0132] The following explains the specific implementation process for determining and training a three-dimensional digital human as provided in the embodiments of this application.
[0133] Figure 5 This is a schematic diagram of a process for obtaining a trained digital human, provided in an embodiment of this application. Please refer to... Figure 5 , Figure 5 The process shown involves first creating a calibration dummy based on sample videos, then calibrating the initial dummy based on the calibration dummy to obtain the training dummy. The specific steps are as follows:
[0134] In one embodiment, the sample video includes multiple video frames. Establishing a corresponding training digital human based on the sample video includes: establishing a calibration digital human using multiple video frames in the sample video, wherein the calibration digital human includes multiple frames; calibrating the initial digital human based on the target head key points of the calibration digital human to obtain the training digital human.
[0135] It should be noted that the calibration digital human can be built based on sample videos. For example, the depth information of the sample user's head can be calculated from multiple video frames of the sample video, and then the calibration digital human can be built based on the two-dimensional image of the sample user in the sample video and the depth information of the sample user's head. This calibration digital human can be a digital human without a specific appearance or style; for example, it can be a three-dimensional digital human framework.
[0136] In the process of calculating depth information, depth information can be calculated based on the differences between the two-dimensional images of the sample user between multiple video frames. For example, if one frame of the two-dimensional image is the frontal face of the sample user and another frame of the two-dimensional image is the side face of the sample user, the depth information of a portion of the points in the head region of the sample user can be determined based on the positional differences between the frontal and side faces. The above steps can be performed based on any two sample video frames to calculate the depth information of the sample user's head.
[0137] Alternatively, in actual implementation, a three-dimensional model can be set up in space, and multiple frames of two-dimensional images can be mapped onto the three-dimensional model to obtain the depth information of each point in the head region of the sample user, thereby realizing the process of converting two-dimensional images into three-dimensional models. The resulting three-dimensional model framework can then be used as a calibration digital human.
[0138] It should be noted that the depth information of the sample user can refer to the depth of each point in the head region of the sample user's two-dimensional image. After calculating the depth information in the above way, a calibrated digital human can be established based on the sample user's two-dimensional image and the depth information of the sample user's head.
[0139] Optionally, after obtaining the calibrated digital human, an initial digital human can be created. This initial digital human can be a digital human without any personalized parameter bias, such as a digital human without gender bias, emotional bias, or age bias. Alternatively, it can be any digital human with certain emotional information or certain identity information. There are no specific restrictions here, and an initial digital human can be set according to actual needs.
[0140] The initial digital human can be a digital human with a specific model, such as a digital human that can be randomly generated in three-dimensional space, or a digital human that has been pre-built without personalized parameter preferences. When it is necessary to build an initial digital human, the pre-built digital human can be called.
[0141] In one embodiment, the initial digital human can be calibrated based on the position of the target head key point of the calibration digital human, thereby using the calibrated initial digital human as a training digital human.
[0142] In the training method for a 3D digital human model provided in this application embodiment, a calibration digital human can be established using multiple video frames from a sample video. The calibration digital human includes multiple frames. An initial digital human is then calibrated based on the target head key points of the calibration digital human to obtain a training digital human. Specifically, using multiple video frames from a sample video allows for a more accurate calibration digital human. Furthermore, calibrating the initial digital human based on the target head key points of the calibration digital human leads to a more accurate training digital human, thereby improving the accuracy and stability of the sample data.
[0143] In one embodiment, during the process of building a calibrated digital human based on sample videos, head features can be extracted using head point clouds to determine depth information.
[0144] It should be noted that head point clouds can be constructed by using the head region of the sample user in the two-dimensional image of the sample video in multiple frames of the sample video. These head point clouds are a set of points with depth information in the head region of the sample user. For example, the depth information of multiple points can be determined from multiple frames of video, and the set of these points is the head point cloud.
[0145] Head point cloud is a way to represent the head features of a sample user in a two-dimensional image by using multiple points. For example, multiple points can be used to form the head contour and facial details of the sample user.
[0146] After obtaining the head point cloud, these head point clouds can be converted into corresponding mesh cells in three-dimensional space based on the depth information of the head point cloud, thereby realizing the conversion from two-dimensional image to three-dimensional image. The three-dimensional digital human obtained through this conversion method is the aforementioned calibrated digital human.
[0147] To explain head point clouds more clearly, we will now explain an example distribution of head point clouds.
[0148] Figure 6 This is a schematic diagram of the head point cloud provided in the embodiments of this application. Please refer to... Figure 6 , Figure 6 The head contour and facial expressions of the sample user are formed by multiple points in the head point cloud. Based on these points, a two-dimensional space to three-dimensional space conversion can be performed. Since each point has depth information, a corresponding three-dimensional digital human can be built in three-dimensional space based on the distribution of these points in the two-dimensional plane and their corresponding depth information. This three-dimensional digital human is the aforementioned calibration digital human.
[0149] The following explains the specific implementation process of calibrating the initial digital human based on the calibration digital human provided in the embodiments of this application.
[0150] Figure 7 This is a schematic diagram illustrating a process for calibrating an initial digital human, as provided in an embodiment of this application. Please refer to... Figure 7 The target head key points include a first head key point and a second head key point. The second head key point is determined based on optical flow information in the sample video. The initial digital human is calibrated based on the target head key points of the calibration digital human to obtain the training digital human, which includes:
[0151] S710: Based on the first loss function and the position of the target head key points in the calibration digit, perform the first calibration on the first point to be calibrated on the initial digit to obtain the digit after the first calibration.
[0152] Among them, the first head key point corresponds to the first calibration point on the initial digital human.
[0153] Based on the position of the first head key point in the calibration digital man, the corresponding calibration point on the initial digital man is determined. Based on the first loss function, the first calibration point on the initial digital man is calibrated to obtain the digital man after the first calibration.
[0154] It should be noted that the first head key point can be any target head key point. It can be selected randomly from the head grid cells of the calibrated digital human as the first head key point.
[0155] The initial digital human and the calibration digital human can be two 3D digital humans composed of the same number of grid cells. The position of the first head key point corresponding to the first calibration point in the initial digital human can be determined based on the grid cell number. Then, the position of the first calibration point in the initial digital human can be calibrated based on the position of the first head key point in the calibration digital human. The first calibration process can be the process of calculating the result of the first loss function of the initial digital human and the calibration digital human.
[0156] In one embodiment, the first calibrated digital human is the digital human obtained when the first loss function is minimized during the first calibration process of the first calibration point on the initial digital human.
[0157] The specific formula for calculating the first loss function is as follows:
[0158] L init =∑ i ||P i -K i ||2;
[0159] In this formula, L init This is the first loss function, where i is a set of randomly initialized coordinates in 3D space, representing the number of first head keypoints. The position of each first head keypoint in the calibrated digital human is P. i The initial position of the calibration point on the digital human is K. i .
[0160] For each frame of the digital human, calibration can be performed using the aforementioned first loss function, and in each frame, this first loss function L... init In the minimum case, each K i The digital person formed by the values of is the first calibrated digital person, which means that the first calibrated digital person can be determined when the first loss function is minimized.
[0161] S720: Based on the second loss function and the position of the second head key point in the calibration digital man, perform a second calibration on the second point to be calibrated on the first calibrated digital man to obtain the training digital man.
[0162] Among them, the second head key point has a corresponding relationship with the second calibration point on the first calibrated digital human.
[0163] It should be noted that the second head key point can be the target head key point obtained after filtering based on optical flow information. Optical flow information refers to the motion trajectory of key points with obvious positional changes in any two adjacent frames in the sample video. The key point that generates this motion trajectory can be used as the second head key point.
[0164] In this context, "significant positional change" refers to the pixel with the highest brightness in the head region of the sample user in the sample video, or it could be the pixel with the highest contrast.
[0165] The following methods can be used to determine the key points of the second head:
[0166] Optionally, determining the second head key point based on optical flow information in the sample video includes: acquiring light density and color in the sample video, determining the head region of the sample user based on light density and color; determining the second head key point from multiple pixels in the head region based on the brightness and / or contrast with adjacent pixels of each pixel in the head region, wherein the brightness and / or contrast with adjacent pixels of the second head key point is greater than a preset threshold.
[0167] Light density refers to the brightness of different areas in each frame of the sample video, and it is related to the image's contrast and brightness. Color refers to the color of each pixel in each frame of the sample video, which can be represented by the weights of the three colors RGB (red, green, and blue).
[0168] The head region of the sample user can be determined based on the two types of information mentioned above. For example, the light density in the head region is usually within a certain range, and the color of the pixels in this region can also be determined within a certain range based on the real person's skin color and scene lighting. Based on these two ranges, every pixel in the video frame image can be traversed to determine the corresponding head region.
[0169] It should be noted that the second head key point can be a pixel in the sample user's head area in the sample video. For example, it can be the brightest pixel in the sample user's head area in the sample video, or it can be the pixel with the greatest contrast.
[0170] Among them, maximum contrast can refer to the largest difference in color or brightness relative to other pixels around that pixel.
[0171] Optionally, in addition to selecting the method with the highest brightness and the highest contrast, a brightness threshold and a contrast threshold can also be set, and pixels with brightness greater than the brightness threshold and contrast with adjacent pixels greater than the preset threshold can be used as the second head key points.
[0172] Optionally, after determining the second head keypoint, the position of the second head keypoint in the calibrated digital human can be determined.
[0173] Furthermore, based on the position of the second head keypoint in the calibrated digital human and the second loss function, a second calibration can be performed on the position of the second calibration point on the first calibrated digital human. When the second loss function is minimized, each of the second calibration points on the first calibrated digital human is K. i The value can be used as the result of the second calibration, and the trained digital human can be obtained based on the result of the second calibration.
[0174] In one embodiment, the digital person is obtained by minimizing the second loss function during the process of training the digital person to perform a second calibration on the second calibration point on the first calibrated digital person.
[0175] The specific formula for calculating the second loss function is as follows:
[0176] L sec =∑ i ||P i (R,T)-K i ||2;
[0177] In this formula, L sec This is the second loss function, where i is a set of randomly initialized coordinates in 3D space, representing the number of second head keypoints. The position of each second head keypoint in the calibrated digital human is P. i The position of the second calibration point on the first calibrated digital human is K. i .
[0178] It should be noted that R can be a rotation vector and T can be a displacement vector. In the second loss function, the position of point Pi in three-dimensional space can be represented by the coordinate representation of the rotation vector and the position vector.
[0179] For each frame of the digital human, the aforementioned second loss function can be used for calibration, and in each frame, this second loss function L... sec In the minimum case, each K i The initial digital human formed by the values of is the second-calibrated digital human, which is the training digital human. In other words, the training digital human can be determined when the second loss function is minimized.
[0180] Optionally, the number of second head keypoints can be less than the number of first head keypoints.
[0181] In practice, coarse calibration can be achieved through the first calibration method, which involves randomly selecting a large number of head key points for calibration. After the coarse calibration is completed, fine calibration can be achieved through the second calibration method, which involves selectively selecting a small number of head key points for calibration. After performing the first and second calibrations in sequence, the initial digital human can be calibrated.
[0182] It should be noted that after the first and second calibrations are performed in sequence, the digital human after the second calibration, which minimizes the second loss function, can be used as the training digital human.
[0183] In the training of the 3D digital human model provided in this embodiment, a first calibration is performed on the first point to be calibrated on the initial digital human based on a first loss function and the position of the first head keypoint in the calibration digital human, resulting in a first-calibrated digital human; then, a second calibration is performed on the second point to be calibrated on the first-calibrated digital human based on a second loss function and the position of the second head keypoint in the calibration digital human, resulting in a trained digital human. By sequentially determining the first and second loss functions, and then calibrating the initial digital human using the calibration digital human, a more accurate trained digital human can be obtained, further leading to more accurate sample data.
[0184] It should be noted that after obtaining the training digital human, the emotions of the training digital human can be determined, and the following methods can be used for determination.
[0185] Figure 8 This is a flowchart illustrating the process of identifying user emotions in sample cases as provided in this application embodiment. Please refer to... Figure 8 By training the facial parameters of digital humans, the system can identify the emotions of sample users, including:
[0186] S810: Calculate the similarity between the facial parameters of the trained digital human and the facial parameters of various pre-set emotional avatar models.
[0187] Optionally, multiple emotion avatar models can be pre-stored on the server side. Each emotion avatar model can be a 3D digital human with an expression tendency. Each 3D digital human with an expression tendency can represent an emotion, such as happiness, sadness, anger, etc. In addition, the facial expression of each emotion model is matched with the corresponding emotion, such as the expression corresponding to the emotion of happiness, the expression corresponding to the emotion of sadness, the expression corresponding to the emotion of anger, etc., without specific restrictions.
[0188] It should be noted that the facial parameters of each emotion model can be the parameters corresponding to the emotion represented by that model.
[0189] For example, different emotion avatar models can be set for different emotions. For instance, human emotions can be divided into 52 different emotions, and an emotion avatar model can be set for each emotion.
[0190] It should be noted that the number and distribution of grid cells in all emotion avatar models can be the same as the number and distribution of grid cells used in training digital humans.
[0191] After obtaining the training digit, the facial parameters of the training digit can be compared with the facial parameters of the emotion avatar model. For example, similarity can be calculated, where the similarity can be determined by comparing the relative positions of the grid cells corresponding to the facial parameters in the training digit and the emotion avatar model.
[0192] The relative position can be the relative position between multiple grid cells in the training digital human and the emotion avatar model, such as the deviation between the position of any grid cell in the training digital human and its position in the emotion avatar model.
[0193] It should be noted that the number and distribution of grid cells in the emotion avatar model can be the same as those in the training digital human. In practice, the similarity between the facial parameters of the training digital human and each pre-set emotion avatar model can be determined based on the degree of matching of the relative positions.
[0194] S820: Determine the emotion labels for training digital humans based on similarity.
[0195] The emotion tags include the emotion and weight of the target emotion avatar model, which is the emotion avatar model with a similarity greater than or equal to a threshold, and the emotion of the sample user includes the emotion tags of the training digital human.
[0196] In one embodiment, the facial parameters of the trained digital human can be compared with those of each pre-set emotion avatar model, and the target emotion avatar model with the highest similarity in the comparison results can be determined. The emotion and weight of the target emotion avatar model can be used as the emotion in the emotion label of the trained digital human and the weight obtained from the comparison.
[0197] It should be noted that similarity can be represented by weights. For example, if we are comparing the similarity between the training digit and the emotion avatar model corresponding to the emotion of happiness, we can obtain the weight of the emotion of happiness in the training digit. For example, if the similarity is 60%, the weight of the emotion of happiness is 0.6.
[0198] Optionally, the similarity between the trained digital human and each emotion avatar model can be determined in the above manner, thereby obtaining the emotion label of the three-dimensional digital human.
[0199] Among them, emotion labels can be recorded in the form of "emotion + weight". For example, "happy 0.6, sad 0.1, angry 0.1" means that among the expressions corresponding to the emotion label, the weight of the emotion of happiness is 0.6, the weight of the emotion of sadness is 0.1, and the weight of the emotion of anger is 0.1.
[0200] It should be noted that the weight represents the degree of similarity between the trained digital human and the corresponding emotion avatar model. The emotion labels recorded above may indicate that the expression has a similarity of 0.6 to the emotion of happiness, a similarity of 0.1 to the emotion of sadness, and a similarity of 0.1 to the emotion of anger.
[0201] Alternatively, multiple weights can be recorded in the emotion tags according to a certain comparison order with multiple emotion avatar models, such as "0.6, 0.1, 0.1, 0.5, ..., 0.2". If 52 emotion avatar models are compared, then this order can have 52 similarity weights.
[0202] It should be noted that in the actual process of determining emotion labels, cases with low similarity can be discarded, and only results with high similarity can be retained. For example, for the label "happy 0.6, sad 0.1, angry 0.1" in the previous example, if the similarity threshold is 0.5, the emotion label can be changed to "happy 0.6".
[0203] In the actual process of representing emotion labels, if there is only one label and the weight is 100%, the emotion corresponding to that label can be used as the label, such as "happy" or "sad". If there is only one label but the weight is not 100%, the emotion and weight corresponding to that label can also be used as the label, such as "happy 0.6" or "sad 0.5". If there are multiple labels and each label has a corresponding weight, they can be represented by a combination, such as "happy 0.6, sad 0.2" or "angry 0.6, sad 0.5, happy 0.1".
[0204] In other words, after determining the similarity, the target emotional avatar model can be identified from the emotional avatar model, that is, the emotional avatar model with a similarity greater than or equal to the threshold, thereby obtaining the corresponding emotional label.
[0205] Optionally, the emotions of the sample users may include emotion labels used to train the digital human. These emotion labels can be used to represent the emotions of the sample users. Alternatively, the emotion labels can be combined with other relevant data to represent the emotions of the sample users. No specific restrictions are imposed here.
[0206] It should be noted that the emotions of sample users can be represented by emotion labels. That is to say, the emotions of sample users in the same frame are not unique, and may include both "happy" and "angry" emotions at the same time. In this case, the emotions of sample users in a certain frame can be represented by emotion labels with weights. For the emotions of sample users in multiple frames, multiple emotion labels can be used to represent them.
[0207] In the training method for the 3D digital human model provided in this application embodiment, the similarity between the facial parameters of the training digital human and each pre-set emotion avatar model can be calculated; based on the similarity, the emotion label of the training digital human is determined. By determining the emotion label, the emotion of each training digital human can be determined in more detail. In addition to the emotion itself, the weight corresponding to the emotion can also be determined, thus obtaining a more accurate and detailed emotion label.
[0208] In one embodiment, the sample video includes multiple video frames, the training digital human includes multiple frames, each frame of the training digital human corresponds one-to-one with a video frame, and the sample user's emotion includes the emotion label of the training digital human for each frame.
[0209] Optionally, the sample speech in the sample video may include multiple speech frames. Each speech frame can correspond to a video frame in the sample video, and correspondingly, it can also correspond to a training digital human frame. The emotions of the training digital human corresponding to each video frame may differ to some extent. For example, the training digital human may have a happy expression in the first video frame and a sad expression in the tenth video frame. In this case, the emotion label of each video frame can be recorded. These multiple emotion labels together constitute the emotion of the sample user, that is, the emotion of the sample user mentioned above.
[0210] It should be noted that during the training of the model, the emotions of the sample users can be trained using one or more specific emotions, or the weights of emotions can be added to the specific emotions for training. For example, the emotions of the sample users in each frame can be input into the audio-visual generation model as an emotion label for training, without any specific restrictions.
[0211] Through the above steps, sample speech, sample user emotion, and sample user head parameters can be obtained, thus obtaining more accurate and detailed sample data. After training the audio-visual generation model based on this sample data, a personalized digital human generation model can be obtained.
[0212] It should be noted that a general digital human generation model can be obtained by normalizing the personalized digital human generation model.
[0213] Correspondingly, the emotional avatar model explained above can also be a three-dimensional digital human with preset expressions, which can be displayed in the following manner.
[0214] Figure 9 This is a schematic diagram of the emotion avatar model provided in the embodiments of this application. Please refer to... Figure 9 , Figure 9 The 3D digital human on the left could be an avatar model corresponding to the emotion of surprise. Figure 9 The 3D digital human on the right could be a model of an emotional avatar corresponding to the emotion of calmness. Figure 9 The example only uses the two emotion avatar models corresponding to these two emotions. In actual implementation, different emotion avatar models can be set for different emotions.
[0215] In one embodiment, the method further includes: determining the weight of the sample data based on the degree of matching between the head parameters of the sample user in the sample video and the sample audio corresponding to the sample video, wherein the weight of the sample data is positively correlated with the degree of matching.
[0216] The weight of the sample data can be a measure to indicate the accuracy of the sample data. If the weight of the sample data is greater than a preset weight threshold, the accuracy of the sample data can be determined to be high. If the weight of the sample data is less than or equal to the preset weight threshold, the accuracy of the sample data can be determined to be low.
[0217] It should be noted that the degree of matching between the sample user's head parameters and the sample audio corresponding to the sample video can be manually labeled, or it can be obtained through the corresponding neural network model.
[0218] If a human representation method is used, the corresponding sample speech and the lip movements of the sample user in each sample video can be pre-determined to see if they match. If they match, the matching degree can be determined as 1; if they do not match, the matching degree can be determined as 0. An average value can be calculated based on the overall matching degree of a video. For example, in a 1-minute video, if the matching degree is 1 for the first 30 seconds and 0 for the last 30 seconds, the overall matching degree can be determined as 0.5.
[0219] If model recognition is used, the video can be input using a pre-trained recognition model, and the model can determine whether the sound of each frame matches the lip movements of the sample user. Then, a similar annotation method can be used to determine the degree of matching of a video.
[0220] It should be noted that the weight of the sample data can be set according to the degree of matching. The higher the degree of matching, the greater the weight of the sample data; the lower the degree of matching, the smaller the weight of the sample data.
[0221] After obtaining the weights of the sample data, the sample data can be filtered based on the weights during the training process.
[0222] Optionally, an audio-visual generation model is trained based on sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter. This includes: training an audio-visual generation model based on sample data whose weights are greater than a preset weight threshold for each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0223] For example, if the preset weight threshold is 0.8, then sample data with a weight less than 0.8 can be discarded and these samples will not be used in the training process of the personalized digital human generation model.
[0224] In the training method for the 3D digital human model provided in this application embodiment, the weights of sample data can be determined based on the degree of matching between the head parameters of the sample user in the sample video and the sample speech corresponding to the sample video. Based on the sample data whose weights corresponding to each personalized parameter are greater than a preset weight threshold, an audio-visual generation model is trained to obtain a personalized digital human generation model corresponding to each personalized parameter. By setting the weights of the sample data, higher quality and more effective sample data can be obtained, improving the accuracy of the trained model.
[0225] The above process is the training process for a general three-dimensional digital human model. In actual implementation, after obtaining the general three-dimensional digital human model, it can be applied accordingly, such as generating a three-dimensional digital human. The implementation process of the three-dimensional digital human generation method provided in the embodiments of this application will be explained below.
[0226] Figure 10 This is a flowchart illustrating the three-dimensional digital human generation method provided in the embodiments of this application. Please refer to... Figure 10 The method includes:
[0227] S1010: Obtain target personalized parameters.
[0228] It should be noted that the personalized parameters for the target include at least one of the target's identity information and the target's emotional information.
[0229] Optionally, personalized parameters can be parameters corresponding to the selected preset after the user selects from multiple presets in the client, or parameters determined by the user through input. There are no specific restrictions here, and they can be set according to actual needs.
[0230] The target personalized parameters can be personalized parameters sent from the client to the server, or default personalized parameters in the server. There are no specific restrictions here, and they can be set according to the actual scenario.
[0231] For example, if you need to use pre-set expressions from the server to generate a corresponding 3D digital human, you can use the default personalized parameters in the server; if you need to generate 3D digital humans with different expressions according to the client's needs, you can use the personalized parameters sent from the receiving client.
[0232] It should be noted that if personalized parameters sent by the client are used, multiple personalized parameters can be set for each client, which can be selected in advance, generated in advance, or generated in real time. The personalized parameters corresponding to the 3D digital human that are actively selected by the user or displayed by default by the client can be used as the aforementioned target personalized parameters.
[0233] Among the personalized parameters for the target, the target identity information can refer to the identity of the 3D digital human, such as age, gender, and ethnicity. The target emotion information can be the emotions that the 3D digital human can express, such as happiness, anger, and sadness. There are no specific restrictions here, and the settings can be based on the emotions that humans possess.
[0234] It should be noted that the target personalized parameters can be personalized parameters that are determined by the client and sent to the server after the user makes a selection or input operation through the client.
[0235] In one embodiment, the target personalization parameters may include one or more parameters.
[0236] For example, Tables 2 and 3 can be used to represent the target personalization parameters:
[0237] Table 2
[0238] gender age race mood male / / /
[0239] Table 2 shows target personalization parameters with only one parameter, such as gender (male). Other parameters are not limited. Table 2 is just one example. In actual implementation, there may also be only one other parameter, such as age, race, or emotion.
[0240] Table 3
[0241] gender age race mood female 20 yellow race Happy
[0242] Table 3 shows the target personalization parameters with multiple parameters, such as: gender female, age 20, race yellow, and mood happy. Table 3 uses four different parameters to form the target personalization parameters as an example. In actual implementation, there may be two, three, five or more parameters. No specific restrictions are made here.
[0243] It should be noted that the emotions in Tables 2 and 3 can be a fixed emotion, a combination of multiple emotions, or a changing emotion. They can be set according to actual needs, such as an emotion with a happy percentage of 0.5 and a sad percentage of 0.5, or a changing emotion of being sad first and then happy. Furthermore, the duration of sadness and happiness can be set accordingly. There are no specific restrictions here, and the emotion parameter can be set according to actual needs.
[0244] When characterizing the target personalization parameters, each type of parameter (including parameters without a specific type, such as age in Table 2) can be represented as a target personalization parameter in the manner described in Tables 2 and 3 above. For example, the target personalization parameters corresponding to Table 2 can be characterized as: male, age not limited, race not limited, and emotion not limited; or, only parameters with a specific type can be used as target personalization parameters, for example, the target personalization parameters corresponding to Table 2 can be characterized as: male.
[0245] S1020: Determine the target adjustment parameters based on the target's personalized parameters and mapping relationship.
[0246] The mapping relationship can include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters.
[0247] The adjustment parameters are used to indicate the deviation of head parameters between the animated digital humans obtained by inputting the same voice into a personalized digital human generation model and a general digital human generation model.
[0248] Each personalized digital human generation model can correspond to at least one personalized parameter. It should be noted that since the target personalized parameter can be one or more parameters, and these target personalized parameters can be different, multiple personalized digital human generation models can be set on the server side, and each target personalized parameter can correspond to one personalized digital human generation model.
[0249] A personalized digital human generation model can be a model used to generate a corresponding three-dimensional digital human based on at least one personalized parameter. For example, a voice recording can be input into the personalized digital human generation model to obtain a three-dimensional digital human with features having at least one personalized parameter.
[0250] A general digital human generation model can be a model used to generate a non-personalized 3D digital human. For example, it can be a model used to generate a general 3D digital human that is unrelated to gender, age, race, and emotion. Through this model, a general 3D digital human can be obtained, which has no gender bias, no age bias, no race bias, and no emotion bias.
[0251] The same voice input can be fed into a personalized digital human generation model and a general digital human generation model to obtain a 3D digital human with at least one personalized parameter and a general 3D digital human, respectively. The head parameters of the 3D digital human with at least one personalized parameter and the head parameters of the general 3D digital human in 3D space can be determined, and the deviation between the two head parameters can be calculated.
[0252] The head parameter can be the position coordinates of the head key points in three-dimensional space. For example, N points can be selected on the head of a three-dimensional digital human with at least one personalized parameter, and N points can be selected at the corresponding position on the head of a general three-dimensional digital human. The difference between the positions of these N points in the spatial coordinate system of the two three-dimensional digital humans can be compared to obtain the deviation of the head parameter. The deviation of the position of each head key point can be used as one of the adjustment parameters. If there are differences among the above N points, there can be N adjustment parameters corresponding to the at least one personalized parameter.
[0253] The mapping relationship can include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The server can determine the target adjustment parameter corresponding to the target personalized parameter by looking up the mapping relationship.
[0254] In actual implementation, the target adjustment parameters corresponding to the target personalized parameters can be determined from the mapping relationship using the above mapping relationship and the target personalized parameters.
[0255] S1030: Based on the target adjustment parameters and the initial animation digital human, generate the animation target digital human.
[0256] The initial digital human in the animation is obtained based on the first voice input to the general digital human generation model.
[0257] Optionally, the first voice can be voice input by the user through the client, or voice obtained by converting text into speech using a preset text-to-speech tool. The text-to-speech tool can run on the client or on the server. If it runs on the client, the client can obtain the text, convert it, and send the resulting speech to the server. If it runs on the server, the client can send the text to the server, which will then convert it into speech.
[0258] After the server obtains the first voice recording, it can input the first voice recording into the general digital human generation model to obtain the initial digital human for animation.
[0259] After the server determines the initial digital mannequin for the animation, it can adjust the initial digital mannequin using target adjustment parameters to obtain the target digital mannequin for the animation. The target adjustment parameters may record the positional deviations of some head key points. The head parameters of the initial digital mannequin for the animation can be adjusted based on these positional deviations of head key points, that is, the positions of the corresponding head key points are adjusted. After all adjustments are completed, the target digital mannequin for the animation can be obtained.
[0260] In a 3D digital human generation method provided in this application embodiment, target personalized parameters can be obtained, and target adjustment parameters can be determined based on the target personalized parameters and mapping relationships. The mapping relationships include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters indicate the deviation of head parameters between the animated digital humans obtained by inputting the same speech into a personalized digital human generation model and a general digital human generation model. Therefore, an animated target digital human can be generated based on the target adjustment parameters and the initial animated digital human. In the process of generating the 3D digital human, since the aforementioned general digital human generation model is used, a more accurate and stable animated target digital human can be obtained.
[0261] The following explains a feasible implementation process for generating an animated target digital human provided in the embodiments of this application.
[0262] Optionally, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, with each frame adjustment parameter corresponding to one speech frame.
[0263] The first voice can be composed of multiple voice frames, and the initial digital human in the animation can also be composed of multiple frames. Each frame of the initial digital human in the animation can be adjusted by a frame adjustment parameter, and each frame adjustment parameter can correspond to a voice frame. That is to say, each frame of the initial digital human in the animation can also correspond to a voice frame.
[0264] In one embodiment, generating an animated target digital human based on target adjustment parameters and an initial animated digital human specifically includes: adjusting each frame of the initial animated digital human corresponding to each speech frame based on a frame adjustment parameter corresponding to each speech frame to generate the animated target digital human.
[0265] It should be noted that during the adjustment of the initial digital human in the animation, each frame of the initial digital human can be adjusted based on a frame adjustment parameter corresponding to each voice frame. For example, the first voice can be divided into 100 voice frames, and each voice frame can correspond to a frame adjustment parameter. Based on these 100 frame adjustment parameters, 100 frames of the initial digital human in the animation can be adjusted. After the adjustment is completed, the target digital human in the animation can be obtained.
[0266] Among them, the frame adjustment parameters can be used to adjust some head parameters of the initial digital human in the animation. For example, the position of some key points of the head that have been selected can be adjusted to achieve the adjustment of a certain frame. After the initial digital human in the animation is adjusted according to the frame adjustment parameters corresponding to each voice frame, the target digital human in the animation can be obtained.
[0267] In a method for generating a 3D digital human provided in this application embodiment, each frame of the initial animated digital human corresponding to each speech frame can be adjusted based on a frame adjustment parameter corresponding to that speech frame to generate the target animated digital human. Adjusting each frame of the initial animated digital human allows for a more comprehensive adjustment of the 3D digital human. Furthermore, since each frame adjustment parameter corresponds to a speech frame, adjusting each frame of the initial animated digital human based on the frame adjustment parameter corresponding to each speech frame can also yield a more realistic 3D digital human that better matches the content of the first speech.
[0268] The following explains a feasible implementation process for adjusting each frame of the initial digital human animation provided in the embodiments of this application.
[0269] In one embodiment, each frame of the digital human includes multiple grid cells in three-dimensional space, and the adjustment parameters for each frame include: the identifier of the grid cell to be adjusted among the multiple grid cells and the offset value of each grid cell to be adjusted.
[0270] It should be noted that both the initial digital human and the target digital human in the animation can include multiple frames, and each frame can include multiple grid units in three-dimensional space.
[0271] The mesh unit can be a mesh point in a mesh structure that forms a triangular mesh or a quadrangular mesh. These mesh structures are distributed throughout the head of the 3D digital human, thus forming a head skeleton of the 3D digital human. The head parameters of the 3D digital human can be represented by these mesh units. The head parameters may include, for example, head contour parameters, facial parameters, lip parameters, etc.
[0272] Figure 11 This is a schematic diagram of the head mesh of the 3D digital human provided in the embodiments of this application. Please refer to... Figure 11 , Figure 11 The 3D digital human shown can be the initial digital human in the above animation, for example, it can be a digital human output by a general digital human generation model, or it can be a digital human with personalized features output by a personalized digital human generation model, without any specific restrictions.
[0273] It should be noted that the mesh unit to be adjusted can be any one or more mesh units selected from the mesh units of the 3D digital human head. There are no specific restrictions here, and the selection can be made according to actual needs. For example, if it is necessary to adjust facial movements, the mesh units of the 3D digital human face can be selected; if it is necessary to adjust the lip shape, the mesh units of the 3D digital human lips can be selected.
[0274] The identifier of the grid cell to be adjusted can be the corresponding label of the grid cell. For example, it can be represented by a preset number or letter. Here, we make specific restrictions and use this label to represent each grid cell to be adjusted.
[0275] The offset value of the mesh cell to be adjusted can be used to indicate the amount of offset of the mesh cell, for example, the amount of change in the three-dimensional coordinates of the mesh cell in the xyz axis of three-dimensional space.
[0276] In one embodiment, each frame of the initial animated digital human corresponding to the audio frame is adjusted based on the adjustment parameters of each frame to generate the target animated digital human. Specifically, this includes adjusting the three-dimensional spatial coordinates of the grid cells to be adjusted in the initial animated digital human corresponding to the audio frame based on the identifier of the grid cells to be adjusted in the adjustment parameters of each frame and the offset value of each grid cell to be adjusted, thereby generating the target animated digital human corresponding to the audio frame.
[0277] It should be noted that during the adjustment process, the coordinates of the grid cells to be adjusted in the three-dimensional space of the initial digital human of the corresponding frame can be adjusted based on the identifier of the grid cell to be adjusted and the offset value of the grid cell to be adjusted in the adjustment parameters of each frame.
[0278] Example: In actual implementation, the grid cell to be adjusted can be determined from multiple grid cells based on the identifier of the grid cell to be adjusted in the adjustment parameters of each frame. Then, the coordinates of the corresponding grid cell to be adjusted in the initial digital man of the animation can be adjusted based on the offset value of the grid cell to be adjusted. The offset value can indicate the amount of change in the x-axis, y-axis and z-axis. The position of each grid cell to be adjusted in the initial digital man of the animation can be adjusted based on the adjustment method indicated by the offset value, so as to obtain the target digital man of the animation corresponding to the voice frame.
[0279] The initial digital human corresponding to each voice frame can be adjusted based on the above method to obtain the target digital human for animation.
[0280] In the three-dimensional digital human generation method provided in this application embodiment, the three-dimensional spatial coordinates of the grid units to be adjusted in the initial animation digital human corresponding to the voice frame can be adjusted based on the identifier of the grid unit to be adjusted and the offset value of each grid unit to be adjusted in the adjustment parameters of each frame, thereby generating the animation target digital human corresponding to the voice frame. In this way, by using the identifier and offset value of the grid unit to be adjusted, the grid unit to be adjusted and the value to be adjusted for each grid unit can be determined more accurately and quickly, thereby achieving the adjustment of the grid unit more quickly and accurately, and obtaining a more accurate animation target digital human.
[0281] The following explains the implementation steps that can be performed after obtaining the animated target digital human in the three-dimensional digital human generation method provided in the embodiments of this application.
[0282] The method also includes sending the animated target digital figure to the client so that the client can render the animated target digital figure.
[0283] It should be noted that after the server adjusts the initial digital human for the animation in the above way to obtain the target digital human for the animation, it can send the target digital human for the animation to the client. The client can be a client that provides the target personalized parameters to the server, or a client that provides the first voice. There are no specific restrictions here.
[0284] After the client obtains the animated target digital human, it can render the animated target digital human.
[0285] Optionally, rendering can include various different rendering methods. For example, the skeleton and texture files sent by the server can be received via data stream, and then rendering can be performed according to the digital human's configuration file. The configuration file indicates the coordinates of the digital human skeleton in three-dimensional space, the specific position of each texture on the digital human skeleton, and the texture type to be rendered.
[0286] In addition, during the rendering process, lighting and shadows can be used to simulate realistic effects, and the rendering engine can be used to generate a 3D digital human that is displayed on the client's monitor, which is the rendered 3D digital human.
[0287] In the 3D digital human generation method provided in this application embodiment, the animated target digital human can be sent to the client so that the client can render the animated target digital human. By transmitting the target animated digital human, the client can more accurately and quickly render the animated target digital human, and then the client can display the rendered 3D digital human.
[0288] The following explains the implementation process of sending the animated target digital human to the client in the three-dimensional digital human generation method provided in this application embodiment.
[0289] Sending the animated target digital human to the client specifically includes sending the target digital human and target deviation parameters corresponding to the first voice frame in the first voice to the client.
[0290] The target deviation parameter is at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent speech frames in the animated target digital human.
[0291] It should be noted that the target digital human can be the digital human corresponding to the first voice frame in the first speech. The target digital human can be the first frame in the animated target digital human, or it can be another frame. There are no specific restrictions here. It can be set according to the correspondence between the frames of the animated target digital human and the voice frames of the first speech.
[0292] Optionally, the deviation parameter can be the deviation of the head parameters of the animated target digital human in two adjacent frames corresponding to every two adjacent voice frames. The target deviation parameter can be a subset of these head deviation parameters; for example, if the deviation parameters include the positional deviations of 100 head key points, then the target deviation parameter can be the positional deviations of 50 head key points.
[0293] For example, if the first audio frame comprises 200 frames, then there can be a corresponding deviation in the head parameters of the animated target digitizer between every two adjacent frames, meaning there are 199 deviations in the head parameters of the animated target digitizer. The target digitizer corresponding to the first audio frame, along with the remaining 199 target deviation parameters, can be sent to the client.
[0294] It should be noted that the deviation parameter can be recorded as the positional difference of the head key points. For example, if there is a difference of 20 grid units in the position between two adjacent animated target digital figures, the deviation parameter can record the number of these 20 grid units and the deviation value of these 20 grid units.
[0295] Optionally, the client can, based on the target digital human, render each voice frame according to the target deviation parameters of the animated target digital human corresponding to that voice frame and the next voice frame, starting from the second voice frame.
[0296] In the 3D digital human generation method provided in this application embodiment, the target digital human and target deviation parameters corresponding to the first speech frame in the first speech can be sent to the client. The target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent speech frames in the animated target digital human. By determining the target digital human and target deviation parameters, the client can more accurately render the 3D digital human.
[0297] Optionally, the deviation parameters of different frames may not be sent to the client all at once, but rather new deviation parameters are obtained in real time and sent to the client sequentially according to the order of the corresponding voice frames.
[0298] In one embodiment, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital human corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; sending the target digital human and target deviation parameters corresponding to the first speech frame in the first speech to the client includes: sending multiple target deviation parameters to the client sequentially according to the time sequence of each speech frame in the first speech.
[0299] Optionally, the target deviation parameters in the deviation parameters corresponding to each speech frame can be sent to the client sequentially according to the order of the speech frames in the first speech. The deviation parameter corresponding to each speech frame can be the deviation of the head parameters of the animated target digital human corresponding to that speech frame and the next frame.
[0300] Correspondingly, the client can also obtain the corresponding target deviation parameters in the corresponding order, and then perform rendering by the client.
[0301] For the client, in each rendering process, the position of the corresponding mesh unit in the 3D digital human rendered in the previous frame can be adjusted based on the positional difference of the mesh unit in the target deviation parameter.
[0302] In the 3D digital human generation method provided in this application embodiment, multiple target deviation parameters can be sent to the client sequentially according to the chronological order of each speech frame in the first speech. Transmitting multiple target deviation parameters in the order of the speech frames in the first speech allows for more accurate rendering of the 3D digital human by the client, thereby improving the real-time performance and accuracy of the client's 3D digital human rendering.
[0303] During the process of sending deviation parameters, all deviation parameters can be sent to the client, or only a portion of the deviation parameters can be sent to the client. For example, if the deviation parameters record the difference values of 20 grid cells, and the difference of 5 grid cells has a low impact on the overall 3D digital human, then the difference of the remaining 15 grid cells can be sent to the client as the target deviation parameter, and the difference of these 5 grid cells will not be transmitted accordingly.
[0304] In one embodiment, sending the target deviation parameter to the client includes: determining the target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; and sending the target deviation parameter to the client.
[0305] It should be noted that bandwidth configuration information refers to the transmission bandwidth between the client and the server. It can be used to characterize the amount of information that can be transmitted during a single communication. A larger bandwidth allows for the transmission of more information, while a smaller bandwidth allows for the transmission of less information. In this case, if the bandwidth is small, the differences in mesh cells that have a lower impact on the 3D digital human can be deleted from the deviation parameters and not transmitted; that is, only the differences in a subset of mesh cells are transmitted.
[0306] Among them, the grid cells with lower impact can be grid cells located in certain positions, such as grid cells in the forehead and ears of the digital human. The specific grid cells with lower impact can be set by the user and there are no specific restrictions here.
[0307] The client's display requirements can be a personalized configuration on the client. For example, if the number of grid cells that can be displayed by the 3D digital human is reduced due to resolution or other reasons, only a portion of the grid cells can be displayed. In this case, only the differences between the grid cells that can be displayed on the client can be transmitted, and the differences between the grid cells that cannot be displayed are not transmitted.
[0308] The weight of the facial region corresponding to the deviation parameter can refer to the weight of the facial region where the head key point corresponding to the deviation parameter is located. For example, in the facial region of a 3D digital human, the weight of the cheek and mouth positions can be set to be higher, while the weight of other positions can be set to be lower. If the head key point corresponding to a certain deviation parameter is in a position with a higher weight, then the deviation parameter can be transmitted; if the head key point corresponding to a certain deviation parameter is in a position with a lower weight, then the deviation parameter can be not transmitted. It should be noted that the weight can be set according to actual needs and is not a fixed weight.
[0309] In the actual transmission process, the target deviation parameter can be determined from the deviation parameters based on any one or more of the above three conditions, and the determined target deviation parameter can be transmitted to the client.
[0310] In the 3D digital human generation method provided in this application embodiment, a target deviation parameter can be determined from multiple deviation parameters according to preset conditions, and the target deviation parameter can be sent to the client. The preset conditions include at least one of the following: client bandwidth configuration information, client display requirements, and the weight of the facial region corresponding to the deviation parameter. By determining the target deviation parameter from the deviation parameters, the number of mesh cells with discrepancies in the deviation parameters can be reduced, thereby reducing the amount of information transmitted between the client and the server, reducing the computational load on the client during the rendering process, and improving rendering efficiency.
[0311] To more clearly demonstrate the process of rendering a 3D digital human using the deviation parameters provided in this embodiment, the rendering process will be explained below by showing the changes in a 3D digital human between two frames.
[0312] Figure 12 This is a schematic diagram illustrating the interface changes for rendering a 3D digital human based on deviation parameters, as provided in the embodiments of this application. Please refer to... Figure 12 , Figure 12 The display interface on the left is the display interface corresponding to the Nth frame of the 3D digital human, and the display interface on the right is the display interface corresponding to the (N+1)th frame of the 3D digital human. N can be a positive integer greater than or equal to 1.
[0313] according to Figure 12 The 3D digital humans shown in the image show that the movements of the 3D digital humans on the left and the 3D digital humans on the right are somewhat different. This difference can be obtained by rendering the 3D digital humans corresponding to the Nth frame based on the target deviation parameters corresponding to the N+1th frame.
[0314] For example, if N is 1, then Figure 12The 3D digital human displayed on the left side of the interface is the target digital human corresponding to the first voice frame transmitted by the server. The 3D digital human on the right side of the interface is the 3D digital human obtained by rendering the target digital human based on the target deviation parameters of the animated target digital human corresponding to the first and second voice frames.
[0315] Example, Figure 12 The three-dimensional digital figure on the left is showing sadness, while the three-dimensional digital figure on the right is showing fear.
[0316] It should be noted that, in addition to performing the above rendering process, the client can also preload the skeleton, textures, and configuration files of the 3D digital human during the initialization process, so that it can render the 3D digital human to be displayed more quickly after receiving it from the server.
[0317] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0318] Based on the foregoing embodiments, this application provides a three-dimensional digital human rendering device. The modules and units included in the device can be implemented by a processor; of course, they can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.
[0319] Figure 13 This is a schematic diagram of the structure of the training device for the three-dimensional digital human model provided in the embodiments of this application. Please refer to... Figure 13 The training device for a three-dimensional digital human model includes: a video acquisition module 1310, a data acquisition module 1320, a training module 1330, and a normalization module 1340.
[0320] The video acquisition module 1310 is used to acquire sample videos corresponding to different personalized parameters, including at least one of identity information and emotion information.
[0321] The data acquisition module 1320 is used to acquire corresponding sample data from the sample videos corresponding to each personalized parameter. The sample data includes sample speech, sample user's emotion, and sample user's head parameters. The sample user's head parameters are obtained by calibrating the points to be calibrated on the initial digital human based on the target head key points in the calibration digital human established by the sample video.
[0322] Training module 1330 is used to train the audio-visual generation model based on the sample data corresponding to each personalized parameter, so as to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0323] The normalization module 1340 is used to normalize different personalized digital human generation models to obtain a general digital human generation model.
[0324] In one embodiment, the data acquisition module 1320 is specifically used to establish a corresponding training digital human based on the sample video, and determine the head parameters of the sample user based on the head parameters of the training digital human; identify the emotions of the sample user through the facial parameters of the training digital human; determine the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0325] In one embodiment, the sample video includes multiple video frames. The data acquisition module 1320 is specifically used to establish a calibration digital human through the multiple video frames in the sample video. The calibration digital human includes multiple frames. The initial digital human is calibrated according to the target head key points of the calibration digital human to obtain a training digital human.
[0326] In one embodiment, the target head key points include a first head key point and a second head key point. The second head key point is determined based on optical flow information in the sample video. The data acquisition module 1320 is specifically used to perform a first calibration on a first point to be calibrated on an initial digital person based on a first loss function and the position of the first head key point in the calibration digital person, to obtain a first-calibrated digital person; wherein the first head key point corresponds to the first point to be calibrated on the initial digital person. Based on a second loss function and the position of the second head key point in the calibration digital person, a second calibration is performed on a second point to be calibrated on the first-calibrated digital person, to obtain a training digital person; wherein the second head key point corresponds to the second point to be calibrated on the first-calibrated digital person.
[0327] In one embodiment, the device is a digital person obtained when the first calibrated digital person is the first to be calibrated point on the initial digital person and the first loss function is minimized; the training digital person is a digital person obtained when the second loss function is minimized when the second to be calibrated point on the first calibrated digital person is the second to be calibrated point.
[0328] In one embodiment, the data acquisition module 1320 is specifically used to acquire the light density and color in the sample video, determine the head region of the sample user based on the light density and color, and determine a second head key point from the head region including multiple pixels based on the brightness and / or contrast with adjacent pixels of each pixel in the head region, wherein the brightness and / or contrast with adjacent pixels of the second head key point is greater than a preset threshold.
[0329] In one embodiment, the data acquisition module 1320 is specifically used to calculate the similarity between the facial parameters of the training digital human and the facial parameters of each pre-set emotion avatar model; based on the similarity, the emotion label of the training digital human is determined, the emotion label includes the emotion and weight of the target emotion avatar model, the target emotion avatar model is the emotion avatar model corresponding to a similarity greater than or equal to a threshold, and the emotion of the sample user includes the emotion label of the training digital human.
[0330] In one embodiment, the training module 1330 is further configured to determine the weight of the sample data based on the degree of matching between the head parameters of the sample user in the sample video and the sample speech corresponding to the sample video, wherein the weight of the sample data is positively correlated with the degree of matching; and to train the audio-visual generation model based on the sample data whose weights corresponding to each personalized parameter are greater than a preset weight threshold, thereby obtaining a personalized digital human generation model corresponding to each personalized parameter.
[0331] Figure 14 This is a schematic diagram of the structure of the three-dimensional digital human generation device provided in the embodiments of this application. Please refer to... Figure 14 A three-dimensional digital human generation device, the device comprising: an acquisition module 1410, a determination module 1420 and a generation module 1430;
[0332] The acquisition module 1410 is used to acquire target personalized parameters; the target personalized parameters include at least one of target identity information and target emotion information.
[0333] The determination module 1420 is used to determine the target adjustment parameters based on the target personalized parameters and the mapping relationship. The mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters are used to indicate the deviation of the head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model.
[0334] The generation module 1430 is used to generate an animated target digital human based on the target adjustment parameters and the initial animated digital human. The initial animated digital human is generated based on the first voice input to the general digital human generation model.
[0335] In one embodiment, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, with each frame adjustment parameter corresponding to a speech frame; the generation module 1430 is specifically used to adjust each frame of the initial animated digital human corresponding to each speech frame based on a frame adjustment parameter corresponding to each speech frame, thereby generating the animated target digital human.
[0336] In one embodiment, each frame of the digital human includes multiple grid cells in three-dimensional space. The adjustment parameters for each frame include: the identifier of the grid cell to be adjusted in the multiple grid cells and the offset value of each grid cell to be adjusted. The generation module 1430 is specifically used to adjust the three-dimensional space coordinates of the grid cells to be adjusted in the initial animation digital human corresponding to the voice frame based on the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted in the adjustment parameters of each frame, so as to generate the animation target digital human corresponding to the voice frame.
[0337] In one embodiment, the generation module 1430 is further configured to send the animated target digital man to the client so that the client can render the animated target digital man.
[0338] In one embodiment, the generation module 1430 is specifically used to send the target digital person and target deviation parameters corresponding to the first speech frame in the first speech to the client; wherein, the target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital person corresponding to adjacent speech frames in the animated target digital person.
[0339] In one embodiment, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital human corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; the generation module 1430 is specifically used to send multiple target deviation parameters to the client in sequence according to the time order of each speech frame in the first speech.
[0340] In one embodiment, the determining module 1420 is further configured to determine a target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; the generating module 1430 is further configured to send the target deviation parameter to the client.
[0341] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0342] It should be noted that, in the embodiments of this application... Figure 13 as well as Figure 14 The module division of the training device and generation device for the 3D digital human model shown is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit with two or more units. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0343] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0344] Figure 15 This is a schematic diagram of the structure of the computing device provided in the embodiments of this application. Please refer to... Figure 15 This application provides a computing device, which can be the aforementioned server, without specific limitations. Its internal structure diagram can be as follows. Figure 15As shown, the computing device includes a processor 1520, memory, and a network interface 1540 connected via a system bus 1510. The processor 1520 provides computing and control capabilities. The memory includes a non-volatile storage medium 1531 and internal memory 1532. The non-volatile storage medium 1531 stores the operating system, computer programs, and a database. The internal memory 1532 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium 1531. The database is used to store data. The network interface 1540 is used for communication with external terminals via a network connection; for example, communication between a client and a server can be achieved through the network interface 1540. When the computer program is executed by the processor 1520, it implements the above-described methods.
[0345] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.
[0346] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0347] Those skilled in the art will understand that Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computing device on which the present application is applied. The specific computing device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0348] In one embodiment, the training device for the three-dimensional digital human model and the generation device for the three-dimensional digital human provided in this application can be implemented as a computer program, which can be implemented in the form of, for example... Figure 15 The device operates on the computing device shown. The memory of the computing device can store the various program modules that make up the above-described apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the methods in the various embodiments of this application described in this specification.
[0349] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0350] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0351] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0352] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0353] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0354] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0355] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0356] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0357] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0358] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0359] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0360] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0361] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A training method for a three-dimensional digital human model, characterized in that, The method includes: Obtain sample videos corresponding to different personalized parameters, wherein the personalized parameters include at least one of identity information and emotion information; Sample data is obtained from sample videos corresponding to each personalized parameter. The sample data includes sample speech, sample user emotion, and sample user head parameters. The sample user head parameters are obtained by calibrating the points to be calibrated on the initial digital human based on the target head key points in the calibration digital human established by the sample video. The calibration digital human is obtained based on multiple frames of two-dimensional images in the sample video. Based on the sample data corresponding to each personalized parameter, train the audio-visual generation model to obtain a personalized digital human generation model corresponding to each personalized parameter. The different personalized digital human generation models are normalized to obtain a general digital human generation model.
2. The method according to claim 1, characterized in that, The step of obtaining corresponding sample data from sample videos corresponding to each personalized parameter includes: The head parameters of the sample users are determined based on the head parameters of the trained digital human. The emotions of the sample users are identified by using the facial parameters of the trained digital human. The sample speech corresponding to each personalized parameter is determined, as well as the emotion and head parameters of the sample user corresponding to each speech frame of the sample speech; wherein, the training digital human is obtained by calibrating the initial digital human based on the target head key points of the calibration digital human.
3. The method according to claim 2, characterized in that, The target head key points include a first head key point and a second head key point, wherein the second head key point is determined based on optical flow information in the sample video. The step of calibrating the initial digital man based on the target head key points of the calibration digital man to obtain the training digital man includes: Based on the first loss function and the position of the first head key point in the calibration digital man, the first point to be calibrated on the initial digital man is calibrated first to obtain the digital man after the first calibration; wherein the first head key point has a corresponding relationship with the first point to be calibrated on the initial digital man. Based on the second loss function and the position of the second head key point in the calibrated digital man, the second point to be calibrated on the first calibrated digital man is calibrated a second time to obtain the training digital man; wherein the second head key point has a corresponding relationship with the second point to be calibrated on the first calibrated digital man.
4. The method according to claim 3, characterized in that, The first calibrated digital human is the digital human obtained when the first loss function is minimized during the first calibration process of the first calibration point on the initial digital human; The training digital person is the digital person obtained when the second loss function is minimized during the second calibration process of the second calibration point on the first calibrated digital person.
5. The method according to claim 3, characterized in that, The step of determining the second head key points based on the optical flow information in the sample video includes: Obtain the light density and color in the sample video, and determine the head region of the sample user based on the light density and color; Based on the brightness and / or contrast with adjacent pixels of each pixel in the head region, the second head key point is determined from the multiple pixels in the head region, wherein the brightness and / or contrast with adjacent pixels of the second head key point is greater than a preset threshold.
6. The method according to claim 2, characterized in that, The process of identifying the emotions of the sample users using the facial parameters of the trained digital human includes: Calculate the similarity between the facial parameters of the trained digital human and the facial parameters of each pre-set emotional avatar model; Based on the similarity, the emotion label of the training digital human is determined. The emotion label includes the emotion and weight of the target emotion avatar model. The target emotion avatar model is the emotion avatar model corresponding to a similarity greater than or equal to a threshold. The emotion of the sample user includes the emotion label of the training digital human.
7. The method according to claim 1, characterized in that, The method further includes: The weight of the sample data is determined based on the degree of matching between the head parameters of the sample users in the sample video and the sample audio corresponding to the sample video. The weight of the sample data is positively correlated with the degree of matching. The step of training an audio-visual generation model based on sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter includes: The audio-visual generation model is trained based on sample data whose weights for each personalized parameter are greater than a preset weight threshold, thereby obtaining a personalized digital human generation model corresponding to each personalized parameter.
8. A method for generating a three-dimensional digital human, characterized in that, The method includes: Obtain target personalized parameters; the target personalized parameters include at least one of target identity information and target emotion information; Based on the target personalized parameters and mapping relationship, target adjustment parameters are determined; the mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters, and the adjustment parameters are used to indicate the deviation of head parameters between the personalized digital human generation model and the animated digital human obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model as described in any one of claims 1-7. Based on the target adjustment parameters and the initial animation digital human, an animated target digital human is generated. The initial animation digital human is generated based on the first voice input to the general digital human generation model.
9. A computing device, characterized in that, It includes a memory and a processor, the memory storing a computer program that can run on the processor, and the processor executing the program to implement the method as described in claims 1-7 or claim 8.
Citation Information
Patent Citations
Digital human driving method and device, electronic equipment and storage medium
CN116580169A
Techniques for processing reconstructed three-dimensional image data
US20130307848A1