Methods for generating 3D digital humans and training 3D digital human models
By acquiring personalized parameters and mapping relationships to adjust the 3D digital human, the problem of monotonous rendering results in existing technologies is solved, and personalized and efficient digital human generation is achieved.
Patent Information
- Application Number
- CN202411643566.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Current technologies produce monotonous 3D digital human rendering results, resulting in a poor user experience and failing to generate personalized digital humans for different users.
By acquiring the target's personalized parameters, using the mapping relationship to determine the target adjustment parameters, adjusting the initial animation digit man to generate the target animation digit man, combining the grid cell's identifier and offset value for precise adjustment, and transmitting the target deviation parameters to the client for rendering.
It achieves greater diversity and accuracy in the generation of 3D digital humans, improves the real-time performance and efficiency of rendering, and ensures a high degree of matching between the generated digital humans and the speech content.
Smart Images

Figure CN119516060B_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, specifically a method for generating a 3D digital human and a method for training a 3D digital human model. Background Technology
[0002] A 3D digital human is a virtual avatar that can be displayed on a client interface. By inputting voice, the 3D digital human can perform corresponding actions, allowing the audience to visually perceive that the voice is spoken by the 3D digital human.
[0003] In related technologies, the rendering methods used can usually only render a specific digital human, resulting in different users receiving the same digital human. It is evident that the rendering results of digital humans are relatively uniform, leading to a poor user experience. Summary of the Invention
[0004] This application provides a method for generating a 3D digital human and a method for training a 3D digital human model, which can render different digital humans and increase the diversity of rendering results. The method for generating a 3D digital human and a method for training a 3D digital human model provided in this application are implemented as follows:
[0005] One aspect of this application provides a method for generating a three-dimensional digital human, the method comprising:
[0006] Obtain personalized parameters for the target; personalized parameters for the target include at least one of the target's identity information and the target's emotional information.
[0007] Based on the target personalized parameters and mapping relationship, the target adjustment parameters are determined; the mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters, and the adjustment parameters are used to indicate the deviation of head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model.
[0008] Based on the target adjustment parameters and the initial animation digital human, the target animation digital human is generated. The initial animation digital human is generated based on the first voice input to the general digital human generation model.
[0009] In one possible implementation, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, each frame adjustment parameter corresponding to one speech frame; based on the target adjustment parameters and the initial animation digital man, the animation target digital man is generated, specifically including: based on a frame adjustment parameter corresponding to each speech frame, adjusting each frame of the initial animation digital man corresponding to that speech frame to generate the animation target digital man.
[0010] In this scheme, the adjustment of the three-dimensional digital human can be achieved more comprehensively by adjusting each frame of the initial digital human animation. Furthermore, since each frame adjustment parameter corresponds to a voice frame, adjusting each frame of the initial digital human animation based on the frame adjustment parameters corresponding to each voice frame can also result in a more realistic three-dimensional digital human that better matches the content of the first voice.
[0011] In one possible implementation, each frame of the digital human includes multiple grid cells in three-dimensional space. The adjustment parameters for each frame include: the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted. Based on the adjustment parameters for each frame, each frame of the initial animated digital human corresponding to the speech frame is adjusted to generate the target animated digital human. Specifically, this includes: adjusting the three-dimensional spatial coordinates of the grid cells to be adjusted in the initial animated digital human corresponding to the speech frame based on the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted in the adjustment parameters for each frame to generate the target animated digital human corresponding to the speech frame.
[0012] In this scheme, by using the identifier and offset value of the grid cell to be adjusted, the grid cell that needs to be adjusted and the value that needs to be adjusted for each grid cell can be determined more accurately and quickly. This allows for faster and more accurate adjustment of the grid cells, resulting in a more accurate animated target digital human.
[0013] In one possible implementation, the method further includes sending the animated target digital man to the client so that the client renders the animated target digital man.
[0014] In this solution, by transmitting the target animated digital human, the client can more accurately and quickly render the animated target digital human, and then the client can display the rendered 3D digital human.
[0015] In one possible implementation, sending the animated target digital human to the client specifically includes: sending the target digital human and target deviation parameters corresponding to the first voice frame in the first voice to the client; wherein, the target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent voice frames in the animated target digital human.
[0016] In this scheme, by determining the target digital human and the target deviation parameters, the client can more accurately render the 3D digital human.
[0017] In one possible implementation, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital person corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; sending the target digital person and target deviation parameters corresponding to the first speech frame in the first speech to the client includes: sending multiple target deviation parameters to the client sequentially according to the time sequence of each speech frame in the first speech.
[0018] In this scheme, the transmission of multiple target deviation parameters is carried out in the order of the voice frames in the first voice, which can more accurately realize the rendering of the 3D digital human by the client, thereby improving the real-time performance and accuracy of the 3D digital human rendering by the client.
[0019] In one possible implementation, sending the target deviation parameter to the client includes: determining the target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; and sending the target deviation parameter to the client.
[0020] In this scheme, the number of mesh cells with differences in the deviation parameters can be reduced by determining the target deviation parameter from the deviation parameter, thereby reducing the amount of information transmitted between the client and the server, reducing the amount of computation on the client during the rendering process, and improving the rendering efficiency.
[0021] Another aspect of this application embodiment provides a method for training a three-dimensional digital human model, the method comprising:
[0022] Obtain sample videos corresponding to different personalized parameters;
[0023] Obtain corresponding sample data from sample videos corresponding to each personalized parameter. The sample data includes sample audio, sample user's emotion, and sample user's head parameters.
[0024] Based on the sample data corresponding to each personalized parameter, train the audio-visual generation model to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0025] Different personalized digital human generation models are normalized to obtain a general digital human generation model.
[0026] In this scheme, normalization processing is performed on multiple personalized digital human generation models, which makes the resulting general digital human generation model highly versatile and applicable. For different personalized parameters, the corresponding three-dimensional digital human can be generated.
[0027] In one possible implementation, sample data is obtained from sample videos corresponding to each personalized parameter, including: building a corresponding training digital human based on the sample videos, and determining the head parameters of the sample user based on the head parameters of the training digital human; identifying the emotions of the sample user through the facial parameters of the training digital human; determining the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0028] In this scheme, by establishing a training digital human, the head parameters and emotions of the sample users can be obtained more quickly and accurately, thus enabling the acquisition of sample data more quickly and accurately.
[0029] In one possible implementation, the emotions of sample users are identified by training the facial parameters of a digital human, including: calculating the similarity between the facial parameters of the trained digital human and the facial parameters of various pre-set emotion avatar models; determining the emotion label of the trained digital human based on the similarity, wherein the emotion label includes the emotion and weight of the target emotion avatar model, the target emotion avatar model being the emotion avatar model corresponding to a similarity greater than or equal to a threshold, and the emotion of the sample user including the emotion label of the trained digital human.
[0030] In one possible implementation, the sample video includes multiple video frames, the training digital human includes multiple frames, each frame of the training digital human corresponds one-to-one with a video frame, and the sample user's emotion includes the emotion label of the training digital human for each frame.
[0031] In this scheme, by determining emotion labels, the emotions of each training digital human can be determined in more detail. In addition to the emotions themselves, the weights corresponding to the emotions can also be determined, thus obtaining more accurate and detailed emotion labels.
[0032] In one possible implementation, after normalizing different personalized digital human generation models to obtain a general digital human generation model, the method further includes: inputting each segment of P sample speech into the personalized digital human generation model and the general digital human generation model corresponding to M personalized parameters, respectively, to obtain P×M personalized animated digital humans and one general animated digital human; wherein, a segment of speech corresponds to M personalized animated digital humans, and P and M are both positive integers greater than or equal to 2; determining the deviation between the head parameters of each frame of each of the M personalized animated digital humans corresponding to each segment of speech and the head parameters of each frame of the general animated digital human; performing a weighted average on the deviations between the head parameters corresponding to each segment of speech to obtain adjustment parameters, and establishing a mapping relationship between the adjustment parameters and personalized parameters.
[0033] In this scheme, by establishing a mapping relationship, the adjustment parameters can be determined more quickly during the generation of 3D digital humans, thereby obtaining the animated target digital human more quickly.
[0034] In another aspect of this application, a three-dimensional digital human generation device is provided, the device comprising: an acquisition module, a determination module, and a generation module;
[0035] The acquisition module is used to acquire target personalized parameters; the target personalized parameters include at least one of target identity information and target emotion information.
[0036] The determination module is used to determine the target adjustment parameters based on the target personalized parameters and the mapping relationship. The mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters are used to indicate the deviation of the head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model.
[0037] The generation module is used to generate the target animated digital human based on the target adjustment parameters and the initial animated digital human. The initial animated digital human is generated based on the first voice input to the general digital human generation model.
[0038] Another aspect of this application embodiment provides a training device for a three-dimensional digital human model, the device including: a video acquisition module, a data acquisition module, a training module, and a normalization module;
[0039] The video acquisition module is used to acquire sample videos corresponding to different personalized parameters;
[0040] The data acquisition module is used to obtain corresponding sample data from the sample videos corresponding to each personalized parameter. The sample data includes sample audio, sample user's emotion, and sample user's head parameters.
[0041] The training module is used to train the audio-visual generation model based on the sample data corresponding to each personalized parameter, so as to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0042] The normalization module is used to normalize different personalized digital human generation models to obtain a general digital human generation model.
[0043] The computing device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the method of this application.
[0044] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method provided in this application embodiment.
[0045] The 3D digital human generation method and 3D digital human model training method provided in this application embodiment can obtain target personalized parameters and determine target adjustment parameters based on the target personalized parameters and mapping relationships. The mapping relationships include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters indicate the deviation of head parameters between the animated digital humans obtained by inputting the same speech into the personalized digital human generation model and the general digital human generation model. Therefore, an animated target digital human can be generated based on the target adjustment parameters and the initial animated digital human. Since there is a mapping relationship between the target adjustment parameters and the personalized parameters, adjusting the initial animated digital human according to the target adjustment parameters yields an animated target digital human with personalized characteristics. That is, different animated target digital humans can be generated by changing the target personalized parameters, further improving the diversity of animated target digital humans generated by the server. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a schematic diagram showing the interface of the three-dimensional digital human generation method provided in the embodiments of this application;
[0048] Figure 2 This is a schematic diagram illustrating an application scenario of the three-dimensional digital human generation method provided in the embodiments of this application;
[0049] Figure 3 This is a flowchart illustrating the three-dimensional digital human generation method provided in the embodiments of this application;
[0050] Figure 4 This is a schematic diagram of the head mesh of the three-dimensional digital human provided in the embodiments of this application;
[0051] Figure 5 This is a schematic diagram illustrating the interface changes for rendering a 3D digital human based on deviation parameters, as provided in the embodiments of this application.
[0052] Figure 6 This is a flowchart illustrating the training method for the three-dimensional digital human model provided in the embodiments of this application;
[0053] Figure 7 This is a schematic diagram of the process for determining sample data provided in the embodiments of this application;
[0054] Figure 8 This is a schematic diagram of the process for identifying the emotions of sample users provided in the embodiments of this application;
[0055] Figure 9 This is a schematic diagram of the emotion avatar model provided in the embodiments of this application;
[0056] Figure 10 This is a schematic diagram of the process for establishing mapping relationships provided in the embodiments of this application;
[0057] Figure 11 This is a schematic diagram illustrating the head deviation between the personalized animated digital human and the general digital human provided in the embodiments of this application;
[0058] Figure 12 This is a diagram showing the correspondence between the initial digital human and the target digital human in the animation provided in the embodiments of this application;
[0059] Figure 13 This is a schematic diagram illustrating the interaction between the client and the server provided in the embodiments of this application;
[0060] Figure 14 This is a schematic diagram of the structure of the three-dimensional digital human generation device provided in the embodiments of this application;
[0061] Figure 15 This is a schematic diagram of the structure of the training device for the three-dimensional digital human model provided in the embodiments of this application;
[0062] Figure 16 This is a schematic diagram of the structure of the computing device provided in the embodiments of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0065] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0066] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0067] The following explains the three-dimensional digital human provided in the embodiments of this application and its practical application scenarios.
[0068] First, a 3D digital human can be a virtual avatar displayed on a client, such as a virtual anchor or virtual teacher.
[0069] Optionally, a 3D digital human is a virtual character developed based on computer technology and artificial intelligence technology. It features high-quality character image and exquisite visual expression. During use, users can input voice or text, and the 3D digital human will make corresponding actions and expressions based on the voice or the voice converted from text, thus making the audience mistakenly believe that the voice is made by the 3D digital human.
[0070] This technology can be used in various fields to meet corresponding needs. For example, 3D digital humans can be used to simulate real live streamers for live e-commerce, live gaming, or live chatting. Alternatively, 3D digital humans can be used to simulate real teachers for online teaching.
[0071] The 3D digital humans used in the above scenarios are mainly generated in real time. In actual use, it may be necessary to generate them in advance and then render them in non-real time. For example, 3D digital humans are used in movies and TV series to replace actors, or videos are made based on 3D digital humans. There are no specific restrictions here. You can choose to render the 3D digital human in real time or render it in advance according to actual needs.
[0072] In real-time rendering scenarios, taking the aforementioned virtual anchor as an example, the 3D digital human can be displayed on the live broadcast screen, and the 3D digital human can display different actions and expressions based on the input voice, such as happy actions and expressions, angry actions and expressions, etc., without specific limitations.
[0073] In the application of 3D digital humans, in order to reduce the pressure on the client, the server can generate the data needed for the rendering process and send it to the client, which will then perform the rendering.
[0074] For the client, the corresponding 3D digital human can be displayed in the corresponding display interface. The following explains the scenario of the 3D digital human being displayed in the client interface provided in the embodiments of this application.
[0075] Figure 1 This is a schematic diagram illustrating the interface of the 3D digital human generation method provided in this application embodiment. Please refer to... Figure 1 , Figure 1 The interface shown is the live streaming interface, which can display a 3D digital human 110. The appearance of the 3D digital human can be customized by the user, such as simulating a real person or an anime character, etc. There are no specific restrictions here.
[0076] In this live streaming scenario, users can input relevant parameters into the 3D digital human through voice or text input, thereby enabling the 3D digital human to perform corresponding actions.
[0077] For example, if a user inputs a voice message and sets a corresponding emotion, such as an expression of surprise, the corresponding 3D digital human can then express a surprised expression based on the set emotion, and the lips can also make a surprised mouth shape.
[0078] In the application of 3D digital humans, in addition to the client needing to render and display the corresponding 3D digital human, the process of calculating the relevant parameters of each movement of the 3D digital human can be completed by the server.
[0079] In this scenario, the server can be a single server, multiple servers, or a server cluster. The client can be various terminal devices, such as mobile phones, personal computers, tablets, etc., without specific limitations. The client and server are connected via a communication link.
[0080] The following explains the practical application scenarios of the three-dimensional digital human generation method provided in the embodiments of this application, including the client and server sides.
[0081] Figure 2 This is a schematic diagram illustrating the application scenario of the 3D digital human generation method provided in the embodiments of this application. Please refer to... Figure 2 This scenario may include a server 210 and multiple clients 220 corresponding to the server 210.
[0082] Among them, the server 210 can be one end running on a server, such as a cloud server, AI (Artificial Intelligence) server, rack server, etc., which can realize the relevant parameter calculation of three-dimensional digital human. The AI server can be a server used to perform artificial intelligence calculations, and the rack server can be a server installed in a standardized network equipment rack, for example, without specific restrictions.
[0083] Client 220 can be a terminal device used by the user, such as including but not limited to mobile phones, wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), tablets, laptops, in-vehicle terminals, PCs (Personal Computers), etc., without specific restrictions.
[0084] It should be noted that in one application scenario, the server 210 can simultaneously perform data interaction for multiple clients 220. For example, the client sends relevant parameters of the 3D digital human to the server, the server calculates the relevant data required for rendering based on the relevant parameters, and then sends this data to the corresponding client so that the client can complete the rendering work.
[0085] This application provides a method for generating a 3D digital human. The specific implementation process of the 3D digital human generation method provided in this application embodiment is explained below. It should be noted that the execution subject of this method can be the aforementioned server.
[0086] Figure 3 This is a flowchart illustrating the three-dimensional digital human generation method provided in the embodiments of this application. Please refer to... Figure 3 The method includes:
[0087] S310: Obtain target personalized parameters.
[0088] It should be noted that the personalized parameters for the target include at least one of the target's identity information and the target's emotional information.
[0089] Optionally, personalized parameters can be parameters corresponding to the selected preset after the user selects from multiple presets in the client, or parameters determined by the user through input. There are no specific restrictions here, and they can be set according to actual needs.
[0090] The target personalized parameters can be personalized parameters sent from the client to the server, or default personalized parameters in the server. There are no specific restrictions here, and they can be set according to the actual scenario.
[0091] For example, if you need to use pre-set expressions from the server to generate a corresponding 3D digital human, you can use the default personalized parameters in the server; if you need to generate 3D digital humans with different expressions according to the client's needs, you can use the personalized parameters sent from the receiving client.
[0092] It should be noted that if personalized parameters sent by the client are used, multiple personalized parameters can be set for each client, which can be selected in advance, generated in advance, or generated in real time. The personalized parameters corresponding to the 3D digital human that are actively selected by the user or displayed by default by the client can be used as the aforementioned target personalized parameters.
[0093] Among the personalized parameters for the target, the target identity information can refer to the identity of the 3D digital human, such as age, gender, and ethnicity. The target emotion information can be the emotions that the 3D digital human can express, such as happiness, anger, and sadness. There are no specific restrictions here, and the settings can be based on the emotions that humans possess.
[0094] It should be noted that the target personalized parameters can be personalized parameters that are determined by the client and sent to the server after the user makes a selection or input operation through the client.
[0095] In one embodiment, the target personalization parameters may include one or more parameters.
[0096] For example, Table 1 and Table 2 can be used to represent the target personalization parameters:
[0097] Table 1
[0098] gender age mood male / /
[0099] Table 1 shows target personalization parameters with only one parameter, such as gender (male). Other parameters are not limited. Table 1 is just one example. In actual implementation, there may be only one other parameter, such as age or mood.
[0100] Table 2
[0101] gender age mood female 20 Happy
[0102] Table 2 shows the target personalized parameters with multiple parameters, such as: gender female, age 20, and mood happy. Table 2 uses four different parameters to form the target personalized parameters as an example. In actual implementation, there may be two, three, five or more parameters. No specific restrictions are made here.
[0103] It should be noted that the emotions in Tables 1 and 2 can be a fixed emotion, a combination of multiple emotions, or a changing emotion. They can be set according to actual needs, such as an emotion with a happy percentage of 0.5 and a sad percentage of 0.5, or a changing emotion of being sad first and then happy. Furthermore, the duration of sadness and happiness can be set accordingly. There are no specific restrictions here, and the emotion parameter can be set according to actual needs.
[0104] When characterizing the target personalization parameters, each type of parameter (including parameters without a specific type, such as age in Table 1) can be represented as a target personalization parameter in the manner described in Tables 1 and 2 above. For example, the target personalization parameters corresponding to Table 1 can be characterized as: male, age not limited, race not limited, and emotion not limited; or, only parameters with a specific type can be used as target personalization parameters, for example, the target personalization parameter corresponding to Table 1 can be characterized as: male.
[0105] S320: Determine target adjustment parameters based on target personalized parameters and mapping relationships.
[0106] The mapping relationship can include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters.
[0107] The adjustment parameters are used to indicate the deviation of head parameters between the animated digital humans obtained by inputting the same voice into a personalized digital human generation model and a general digital human generation model.
[0108] Each personalized digital human generation model can correspond to at least one personalized parameter. It should be noted that since the target personalized parameter can be one parameter or multiple parameters, and these target personalized parameters can be different, multiple personalized digital human generation models can be set on the server side, and each set of target personalized parameters can correspond to one personalized digital human generation model.
[0109] A personalized digital human generation model can be used to generate a corresponding three-dimensional digital human's audio-visual model based on at least one personalized parameter. For example, a voice segment can be input into the personalized digital human generation model to obtain a three-dimensional digital human with the features of that personalized parameter.
[0110] A general digital human generation model can be a model used to generate a non-personalized 3D digital human. For example, it can be an audio-visual model used to generate a universal 3D digital human that is unrelated to gender, age, race, and emotion. This model can be used to obtain a universal 3D digital human that has no gender bias, no age bias, no race bias, and no emotional bias.
[0111] The same voice input can be fed into a personalized digital human generation model and a general digital human generation model to obtain a 3D digital human with the features of the personalized parameter and a general 3D digital human, respectively. The head parameters of the 3D digital human with the features of the personalized parameter and the general 3D digital human in 3D space can be determined, and the deviation between the two head parameters can be calculated.
[0112] The head parameter can be the position coordinates of the head key points in three-dimensional space. For example, Q points can be selected on the head of a three-dimensional digital human with the features of this personalized parameter, and Q points can be selected at the corresponding positions on the head of a general three-dimensional digital human. The difference between the positions of these Q points in the spatial coordinate systems of the two three-dimensional digital humans can be compared to obtain the deviation of the head parameter. The deviation of the position of each head key point can be used as one of the sub-adjustment parameters. If there are differences among the above N points, then there can be N adjustment parameters corresponding to the at least one personalized parameter.
[0113] The mapping relationship can include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The server can determine the target adjustment parameter corresponding to the target personalized parameter by looking up the mapping relationship.
[0114] The mapping relationships provided in the embodiments of this application are explained more clearly in Table 3 below:
[0115] Table 3
[0116]
[0117]
[0118] Among them, any one of the adjustment parameters δ i Each includes N frame adjustment parameters. Personalization parameters may include at least one of gender, age, and mood.
[0119] In actual implementation, the target adjustment parameters corresponding to the target personalized parameters can be determined from the mapping relationship using the above mapping relationship and the target personalized parameters.
[0120] S330: Based on the target adjustment parameters and the initial animation digital human, generate the animation target digital human.
[0121] The initial digital human in the animation is obtained based on the first voice input to the general digital human generation model.
[0122] Optionally, the first voice can be voice input by the user through the client, or voice obtained by converting text into speech using a preset text-to-speech tool. The text-to-speech tool can run on the client or on the server. If it runs on the client, the client can obtain the text, convert it, and send the resulting speech to the server. If it runs on the server, the client can send the text to the server, which will then convert it into speech.
[0123] After the server obtains the first voice recording, it can input the first voice recording into the general digital human generation model to obtain the initial digital human for animation.
[0124] After the server determines the initial digital human for animation, if the client receives the personalized parameters input by the user, the server can determine the target adjustment parameters corresponding to the personalized parameters through mapping relationships (such as Table 3 or Table 5-1). Then, the initial digital human for animation is adjusted according to the target adjustment parameters to obtain the target digital human for animation. The target adjustment parameters may record the positional deviations of some head key points. The head parameters of the initial digital human for animation can be adjusted according to these positional deviations of head key points, that is, the positions of the corresponding head key points are adjusted. After all adjustments are completed, the target digital human for animation can be obtained.
[0125] In a 3D digital human generation method provided in this application embodiment, target personalized parameters can be obtained, and target adjustment parameters can be determined based on the target personalized parameters and mapping relationships. The mapping relationships include a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters indicate the deviation of head parameters between the animated digital humans obtained by inputting the same speech into a personalized digital human generation model and a general digital human generation model. Therefore, an animated target digital human can be generated based on the target adjustment parameters and the initial animated digital human. Since there is a mapping relationship between the target adjustment parameters and the personalized parameters, adjusting the initial animated digital human according to the target adjustment parameters yields an animated target digital human with personalized characteristics. That is, different animated target digital humans can be generated by changing the target personalized parameters, further improving the diversity of animated target digital humans generated by the server.
[0126] The following explains a feasible implementation process for generating an animated target digital human provided in the embodiments of this application.
[0127] Optionally, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, with each frame adjustment parameter corresponding to one speech frame.
[0128] The first voice can be composed of multiple voice frames, and the initial digital human in the animation can also be composed of multiple frames. Each frame of the initial digital human in the animation can be adjusted by a frame adjustment parameter, and each frame adjustment parameter can correspond to a voice frame. That is to say, each frame of the initial digital human in the animation can also correspond to a voice frame.
[0129] In one embodiment, generating an animated target digital human based on target adjustment parameters and an initial animated digital human specifically includes: adjusting each frame of the initial animated digital human corresponding to each speech frame based on a frame adjustment parameter corresponding to each speech frame to generate the animated target digital human.
[0130] It should be noted that during the adjustment of the initial digital human in the animation, each frame of the initial digital human can be adjusted based on a frame adjustment parameter corresponding to each voice frame. For example, the first voice can be divided into 100 voice frames, and each voice frame can correspond to a frame adjustment parameter. In this case, it can be understood as any δ in Table 3. i Each includes 100 frame adjustment parameters. Based on these 100 frame adjustment parameters, the initial 100 frames of the digital human in the animation can be adjusted. After the adjustment is completed, the target digital human in the animation can be obtained.
[0131] Among them, the frame adjustment parameters can be used to adjust some head parameters of the initial digital human in the animation. For example, the position of some key points of the head that have been selected can be adjusted to achieve the adjustment of a certain frame. After the initial digital human in the animation is adjusted according to the frame adjustment parameters corresponding to each voice frame, the target digital human in the animation can be obtained.
[0132] In the 3D digital human generation method provided in this application embodiment, a mapping relationship table can be queried based on personalized parameters to determine target adjustment parameters. Based on a frame adjustment parameter corresponding to each speech frame included in the target adjustment parameters, each frame of the initial animated digital human corresponding to that speech frame is adjusted to generate the target animated digital human. By adjusting each frame of the initial animated digital human, the adjustment of the 3D digital human can be achieved more comprehensively. Furthermore, since each frame adjustment parameter corresponds to a speech frame, adjusting each frame of the initial animated digital human based on the frame adjustment parameter corresponding to each speech frame can also yield a more realistic 3D digital human that better matches the content of the first speech.
[0133] The following explains a feasible implementation process for adjusting each frame of the initial digital human animation provided in the embodiments of this application.
[0134] In one embodiment, each frame of the digital human comprises multiple grid cells in three-dimensional space, and the adjustment parameters for each frame include: the identifier of the grid cell to be adjusted among the multiple grid cells and the offset value of each grid cell to be adjusted, which is δ in Table 3. i .
[0135] It should be noted that both the initial digital human and the target digital human in the animation can include multiple frames, and each frame can include multiple grid units in three-dimensional space.
[0136] The mesh unit can be a mesh point in a mesh structure that forms a triangular mesh or a quadrangular mesh. These mesh structures are distributed throughout the head of the 3D digital human, thus forming a head skeleton of the 3D digital human. The head parameters of the 3D digital human can be represented by these mesh units. The head parameters may include, for example, head contour parameters, facial parameters, lip parameters, etc.
[0137] Figure 4 This is a schematic diagram of the head mesh of the 3D digital human provided in the embodiments of this application. Please refer to... Figure 4 , Figure 4 The 3D digital human shown can be the initial digital human in the above animation, for example, it can be a digital human output by a general digital human generation model, or it can be a digital human with personalized features output by a personalized digital human generation model, without any specific restrictions.
[0138] It should be noted that the mesh unit to be adjusted can be any one or more key mesh units selected from the mesh units of the 3D digital human head. There are no specific restrictions here, and the selection can be made according to actual needs. For example, if it is necessary to adjust facial movements, the mesh units of the 3D digital human face can be selected; if it is necessary to adjust the lip shape, the mesh units of the 3D digital human lips can be selected.
[0139] The identifier of the grid cell to be adjusted can be the corresponding label of the grid cell. For example, it can be represented by a preset number or letter. Here, we make specific restrictions and use this label to represent each grid cell to be adjusted.
[0140] The offset value of the mesh cell to be adjusted can be used to indicate the amount of offset of the mesh cell, for example, the amount of change in the three-dimensional coordinates of the mesh cell in the xyz axis of three-dimensional space.
[0141] In one embodiment, each frame of the initial animated digital human corresponding to the audio frame is adjusted based on the adjustment parameters of each frame to generate the target animated digital human. Specifically, this includes adjusting the three-dimensional spatial coordinates of the grid cells to be adjusted in the initial animated digital human corresponding to the audio frame based on the identifier of the grid cells to be adjusted in the adjustment parameters of each frame and the offset value of each grid cell to be adjusted, thereby generating the target animated digital human corresponding to the audio frame.
[0142] It should be noted that during the adjustment process, the coordinates of the grid cells to be adjusted in the three-dimensional space of the initial digital human of the corresponding frame can be adjusted based on the identifier of the grid cell to be adjusted and the offset value of the grid cell to be adjusted in the adjustment parameters of each frame.
[0143] Example: In actual implementation, the grid cells to be adjusted can be determined from multiple grid cells based on the identifier of the grid cells to be adjusted in the adjustment parameters of each frame. Then, the coordinates of the corresponding grid cells to be adjusted in the initial digital mannequin of the animation can be adjusted based on the offset value of the grid cells to be adjusted. The position of each grid cell to be adjusted in the initial digital mannequin of the animation can be adjusted according to the adjustment method indicated by the offset value, thereby obtaining the target digital mannequin of the animation corresponding to the voice frame. The offset value is the target adjustment parameter δ in Table 3. i The frame adjustment parameters included in the text represent the amount of change in the three-dimensional coordinates of the cell to be adjusted.
[0144] The initial digital human corresponding to each voice frame can be adjusted based on the above method to obtain the target digital human for animation.
[0145] In the three-dimensional digital human generation method provided in this application embodiment, the three-dimensional spatial coordinates of the grid units to be adjusted in the initial animation digital human corresponding to the voice frame can be adjusted based on the identifier of the grid unit to be adjusted and the offset value of each grid unit to be adjusted in the adjustment parameters of each frame, thereby generating the animation target digital human corresponding to the voice frame. In this way, by using the identifier and offset value of the grid unit to be adjusted, the grid unit to be adjusted and the value to be adjusted for each grid unit can be determined more accurately and quickly, thereby achieving the adjustment of the grid unit more quickly and accurately, and obtaining a more accurate animation target digital human.
[0146] The following explains the implementation steps that can be performed after obtaining the animated target digital human in the three-dimensional digital human generation method provided in the embodiments of this application.
[0147] The method also includes sending the animated target digital figure to the client so that the client can render the animated target digital figure.
[0148] It should be noted that after the server adjusts the initial digital human for the animation in the above way to obtain the target digital human for the animation, it can send the target digital human for the animation to the client. The client can be a client that provides the target personalized parameters to the server, or a client that provides the first voice. There are no specific restrictions here.
[0149] After the client obtains the animated target digital human, it can render the animated target digital human.
[0150] Optionally, rendering can include various rendering methods. For example, the skeleton and texture files sent by the server can be received via data stream, and then rendering can be performed according to the digital human's configuration file. The configuration file indicates the coordinates of the digital human skeleton in three-dimensional space, the specific position of each texture on the digital human skeleton, and the texture type to be rendered.
[0151] In addition, during the rendering process, lighting and shadows can be used to simulate realistic effects, and the rendering engine can be used to generate a 3D digital human that is displayed on the client's monitor, which is the rendered 3D digital human.
[0152] In the 3D digital human generation method provided in this application embodiment, the animated target digital human can be sent to the client so that the client can render the animated target digital human. By transmitting the target animated digital human, the client can more accurately and quickly render the animated target digital human, and then the client can display the rendered 3D digital human.
[0153] The following explains the implementation process of sending the animated target digital human to the client in the three-dimensional digital human generation method provided in this application embodiment.
[0154] Sending the animated target digital human to the client specifically includes sending the target digital human and target deviation parameters corresponding to the first voice frame in the first voice to the client.
[0155] The target deviation parameter is at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent speech frames in the animated target digital human.
[0156] It should be noted that the target digital human can be the digital human corresponding to the first voice frame in the first speech. The target digital human can be the first frame in the animated target digital human, or it can be another frame. There are no specific restrictions here. It can be set according to the correspondence between the frames of the animated target digital human and the voice frames of the first speech.
[0157] Optionally, the deviation parameter can be the deviation of the head parameters of the animated target digital human in two adjacent frames corresponding to every two adjacent voice frames. The target deviation parameter can be a subset of these head deviation parameters; for example, if the deviation parameters include the positional deviations of 100 head key points, then the target deviation parameter can be the positional deviations of 50 head key points.
[0158] For example, if the first audio frame comprises 200 frames, then there can be a corresponding deviation in the head parameters of the animated target digitizer between every two adjacent frames, meaning there are 199 deviations in the head parameters of the animated target digitizer. The target digitizer corresponding to the first audio frame, along with the remaining 199 target deviation parameters, can be sent to the client.
[0159] It should be noted that the deviation parameter can be recorded as the positional difference of the head key points. For example, if there is a difference of 20 grid units in the position between two adjacent animated target digital figures, the deviation parameter can record the number of these 20 grid units and the deviation value of these 20 grid units.
[0160] Optionally, the client can, based on the target digital human, render each voice frame according to the target deviation parameters of the animated target digital human corresponding to that voice frame and the next voice frame, starting from the second voice frame.
[0161] In the 3D digital human generation method provided in this application embodiment, the target digital human and target deviation parameters corresponding to the first speech frame in the first speech can be sent to the client. The target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital human corresponding to adjacent speech frames in the animated target digital human. By determining the target digital human and target deviation parameters, the client can more accurately render the 3D digital human.
[0162] Optionally, the deviation parameters of different frames may not be sent to the client all at once, but rather new deviation parameters are obtained in real time and sent to the client sequentially according to the order of the corresponding voice frames.
[0163] In one embodiment, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital human corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; sending the target digital human and target deviation parameters corresponding to the first speech frame in the first speech to the client includes: sending multiple target deviation parameters to the client sequentially according to the time sequence of each speech frame in the first speech.
[0164] Optionally, the target deviation parameters in the deviation parameters corresponding to each speech frame can be sent to the client sequentially according to the order of the speech frames in the first speech. The deviation parameter corresponding to each speech frame can be the deviation of the head parameters of the animated target digital human corresponding to that speech frame and the next frame.
[0165] Correspondingly, the client can also obtain the corresponding target deviation parameters in the corresponding order, and then perform rendering by the client.
[0166] For the client, in each rendering process, the position of the corresponding mesh unit in the 3D digital human rendered in the previous frame can be adjusted based on the positional difference of the mesh unit in the target deviation parameter.
[0167] In the 3D digital human generation method provided in this application embodiment, multiple target deviation parameters can be sent to the client sequentially according to the chronological order of each speech frame in the first speech. Transmitting multiple target deviation parameters in the order of the speech frames in the first speech allows for more accurate rendering of the 3D digital human by the client, thereby improving the real-time performance and accuracy of the client's 3D digital human rendering.
[0168] During the process of sending deviation parameters, all deviation parameters can be sent to the client, or only a portion of the deviation parameters can be sent to the client. For example, if the deviation parameters record the difference values of 20 grid cells, and the difference of 5 grid cells has a low impact on the overall 3D digital human image, then the difference of the remaining 15 grid cells can be sent to the client as the target deviation parameter, and the difference of these 5 grid cells will not be transmitted accordingly.
[0169] In one embodiment, sending the target deviation parameter to the client includes: determining the target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; and sending the target deviation parameter to the client.
[0170] It should be noted that bandwidth configuration information refers to the transmission bandwidth between the client and the server. It can be used to characterize the amount of information that can be transmitted during a single communication. A larger bandwidth allows for the transmission of more information, while a smaller bandwidth allows for the transmission of less information. In this case, if the bandwidth is small, the differences in mesh cells that have a lower impact on the 3D digital human image can be deleted from the deviation parameters and not transmitted; that is, only the differences in a subset of mesh cells are transmitted.
[0171] Among them, the grid cells with lower impact can be grid cells located in certain positions, such as grid cells in the forehead and ears of the digital human. The specific grid cells with lower impact can be set by the user and there are no specific restrictions here.
[0172] The client's display requirements can be a personalized configuration on the client. For example, if the number of grid cells that can be displayed by the 3D digital human is reduced due to resolution or other reasons, only a portion of the grid cells can be displayed. In this case, only the differences between the grid cells that can be displayed on the client can be transmitted, and the differences between the grid cells that cannot be displayed are not transmitted.
[0173] The weight of the facial region corresponding to the deviation parameter can refer to the weight of the facial region where the head key point corresponding to the deviation parameter is located. For example, in the facial region of a 3D digital human, the weight of the cheek and mouth positions can be set to be higher, while the weight of other positions can be set to be lower. If the head key point corresponding to a certain deviation parameter is in a position with a higher weight, then the deviation parameter can be transmitted; if the head key point corresponding to a certain deviation parameter is in a position with a lower weight, then the deviation parameter can be not transmitted. It should be noted that the weight can be set according to actual needs and is not a fixed weight.
[0174] In the actual transmission process, the target deviation parameter can be determined from the deviation parameters based on any one or more of the above three conditions, and the determined target deviation parameter can be transmitted to the client.
[0175] In the 3D digital human generation method provided in this application embodiment, a target deviation parameter can be determined from multiple deviation parameters according to preset conditions, and the target deviation parameter can be sent to the client. The preset conditions include at least one of the following: client bandwidth configuration information, client display requirements, and the weight of the facial region corresponding to the deviation parameter. By determining the target deviation parameter from the deviation parameters, the number of mesh cells with discrepancies in the deviation parameters can be reduced, thereby reducing the amount of information transmitted between the client and the server, reducing the computational load on the client during the rendering process, and improving rendering efficiency.
[0176] To more clearly demonstrate the process of rendering a 3D digital human using the deviation parameters provided in this embodiment, the rendering process will be explained below by showing the changes in a 3D digital human between two frames.
[0177] Figure 5 This is a schematic diagram illustrating the interface changes for rendering a 3D digital human based on deviation parameters, as provided in the embodiments of this application. Please refer to... Figure 5 , Figure 5 The display interface on the left is the display interface corresponding to the Nth frame of the 3D digital human, and the display interface on the right is the display interface corresponding to the (N+1)th frame of the 3D digital human. N can be a positive integer greater than or equal to 1.
[0178] according to Figure 5The 3D digital humans shown in the image show that the movements of the 3D digital humans on the left and the 3D digital humans on the right are somewhat different. This difference can be obtained by rendering the 3D digital humans corresponding to the Nth frame based on the target deviation parameters corresponding to the N+1th frame.
[0179] For example, if N is 1, then Figure 5 The 3D digital human displayed on the left side of the interface is the target digital human corresponding to the first voice frame transmitted by the server. The 3D digital human on the right side of the interface is the 3D digital human obtained by rendering the target digital human based on the target deviation parameters of the animated target digital human corresponding to the first and second voice frames.
[0180] Example, Figure 5 The three-dimensional digital figure on the left is showing sadness, while the three-dimensional digital figure on the right is showing fear.
[0181] It should be noted that, in addition to performing the above rendering process, the client can also preload the skeleton, textures, and configuration files of the 3D digital human during the initialization process, so that it can render the 3D digital human to be displayed more quickly after receiving it from the server.
[0182] The above process explains the steps performed by the server and client in the actual application of 3D digital human. In actual implementation, the general digital human generation model stored in the server can be obtained after training. The training method of the 3D digital human model provided in the embodiments of this application will be explained below.
[0183] Figure 6 This is a flowchart illustrating the training method for the 3D digital human model provided in this embodiment. Please refer to... Figure 6 The method includes:
[0184] S610: Obtain sample videos corresponding to different personalized parameters.
[0185] Optionally, the sample videos can be pre-acquired 2D videos (two-dimensional videos), and these sample videos can be pre-classified to obtain sample videos corresponding to each personalized parameter.
[0186] It should be noted that corresponding sample videos can be configured according to the personalized parameters required, or the personalized parameters in these sample videos can be determined by collecting sample videos. No specific restrictions are imposed here.
[0187] For example, if there are 100 different sets of personalized parameters, a corresponding sample video can be set for each set of personalized parameters. Alternatively, if 20 sample videos are obtained in advance, the personalized parameters present in these 20 sample videos can be extracted.
[0188] It should be noted that the two-dimensional video can be a video of a real person speaking, such as a video of a teacher giving a lecture. In this sample video, there needs to be a real person's head image, preferably a facial image.
[0189] In any given two-dimensional video, during the actual implementation process, the real person may exhibit different emotions, such as changing from a happy emotion to a sad emotion, or the real person in the video may change, for example, the first half of the video may be filmed as a man, and the second half as a woman, etc.
[0190] Before obtaining sample videos, these two-dimensional videos can be segmented. For example, clustering can be used to identify segments in the video that have different emotions or identities, resulting in multiple segments. Each segment can then be used as a sample video.
[0191] Optionally, by segmenting the two-dimensional video, multiple sample videos can be obtained, and each sample video can correspond to a set of personalized parameters. For example, the personalized parameters corresponding to one sample video are: male, 50 years old, happy; and the personalized parameters corresponding to another sample video are: sad.
[0192] It should be noted that the correspondence between sample videos and personalized parameters can be presented in Table 4.
[0193] Table 4
[0194] Sample Video Personalized parameters Sample Video 1 male Sample Video 2 female Sample Video 3 Male, 20 years old … … Sample video N-1 Male, happy Sample video N Male, 20 years old, happy
[0195] Please refer to Table 4. Each sample video can correspond to a different set of personalized parameters. There can be one or more personalized parameters, and no specific restrictions are imposed here.
[0196] S620: Obtain corresponding sample data from the sample videos corresponding to each personalized parameter.
[0197] The sample data includes sample audio, sample user emotions, and sample user head parameters.
[0198] Optionally, after obtaining the sample video, the corresponding sample data can be determined from the sample video. For the same sample video, multiple sets of sample data can be obtained, where each set of sample data can include sample speech, sample user's emotion, and sample user's head parameters.
[0199] The sample speech can be a part of the speech corresponding to the sample video, the sample user's emotion can be the emotion of a real person in the sample video, and the sample user's head parameters can be the head parameters of a three-dimensional model (i.e., the sample user) built in three-dimensional space based on the head image of a real person in the sample video.
[0200] For each personalized parameter, the sample video can yield the aforementioned multiple sets of sample data.
[0201] It should be noted that the head parameters of the sample user can specifically include the head contour parameters, facial parameters, lip parameters, etc. These parameters can be the positions of the corresponding mesh cells. For example, the head contour parameters can be the positions of the mesh cells that make up the head contour of the entire 3D digital human; the facial parameters can be the positions of the mesh cells that make up the face of the 3D digital human; and the lip parameters can be the positions of the mesh cells that make up the lips of the 3D digital human.
[0202] Optionally, the head parameters and emotions of the sample users can correspond to the sample speech, and each speech frame can correspond to the head parameters and emotions of a sample user.
[0203] For example: If a sample speech in a sample data contains 100 speech frames, then there can be 100 emotions and 100 head parameters of the sample user. The emotion and head parameters of each frame of the sample user correspond to one speech frame of the sample speech.
[0204] S630: Based on the sample data corresponding to each personalized parameter, train the audio-visual generation model to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0205] Optionally, sample data corresponding to each personalized parameter can be obtained through the above method, and then the audio-visual generation model can be trained based on these sample data.
[0206] The audio-visual generation model can be an audio-visual encoder or a neural network structure. Its input can be a sample sound and the emotion of a sample user, and its output can be a three-dimensional digital human. The model can compare the head parameters of the output three-dimensional digital human with the head parameters of the sample user. After extensive training, if the audio-visual generation model converges, for example, if the matching degree between the head parameters of the output three-dimensional digital human and the head parameters of the sample user is greater than a preset similarity threshold, then the audio-visual generation model can be determined to have converged, thus obtaining the personalized digital human generation model corresponding to the personalized parameters.
[0207] It should be noted that the above training process can be performed on a set of personalized parameters corresponding to each sample video to obtain a personalized digital human generation model corresponding to each set of personalized parameters. In other words, multiple personalized digital human generation models are obtained by training the same audio-visual generation model with different sample data.
[0208] S640: Normalize different personalized digital human generation models to obtain a general digital human generation model.
[0209] It should be noted that after obtaining multiple personalized digital human generation models, these personalized digital human generation models can be normalized.
[0210] It should be noted that normalization methods can include various approaches. For example, input feature normalization can be used, meaning the input data is normalized before training these personalized digital human generation models to ensure similar scale across different sample data. Alternatively, inter-layer normalization can be employed. Since each personalized digital human generation model is trained based on the same audio-visual generation model, they possess a certain degree of similarity. For instance, if the neural network has the same number of layers, normalization can be applied after each layer is generated, especially before the activation function. Another approach is weight normalization, where the weights of each layer in the neural network are reparameterized, decoupling the length and direction of the weight vector, thus normalizing these personalized digital human generation models.
[0211] In practice, any of the above-mentioned normalization methods can be used to normalize multiple personalized digital human generation models, thereby obtaining a general digital human generation model.
[0212] A general digital human generation model can be obtained by normalizing all personalized digital human generation models.
[0213] In practical applications, the input parameters of this general digital human generation model can be a first speech, and the output can be a general three-dimensional digital human.
[0214] The training method for the 3D digital human model provided in this application embodiment can acquire sample videos corresponding to different personalized parameters; obtain corresponding sample data from the sample videos corresponding to each personalized parameter; train an audio-visual generation model based on the sample data corresponding to each personalized parameter to obtain a personalized digital human generation model corresponding to each personalized parameter; and normalize the different personalized digital human generation models to obtain a general digital human generation model. Normalizing multiple personalized digital human generation models makes the resulting general digital human generation model highly versatile and applicable, enabling the generation of corresponding 3D digital humans for different personalized parameters.
[0215] The following explains the implementation process of determining sample data based on sample videos provided in the embodiments of this application.
[0216] Figure 7 This is a flowchart illustrating the process of determining sample data provided in the embodiments of this application. Please refer to... Figure 7 For sample videos, the corresponding sample speech in each frame of the sample video can be determined, and a three-dimensional model can be built based on the video image of the real person. Thus, the emotion and head parameters of the three-dimensional model built in three-dimensional space can be determined. The head parameters of the model are used as the head parameters of the sample user, and the facial emotion of the model is used as the emotion of the sample user.
[0217] One approach is to first build a 3D model based on a 2D video, which is essentially training a digital human. Then, based on the head parameters of the training digital human, the head parameters of the sample user are determined. Finally, the emotions of the training digital human are identified to obtain the emotions of the sample user.
[0218] In one embodiment, obtaining corresponding sample data from sample videos corresponding to each personalized parameter includes: establishing a corresponding training digital human based on the sample video, and determining the head parameters of the sample user based on the head parameters of the training digital human; identifying the emotions of the sample user through the facial parameters of the training digital human; determining the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0219] It should be noted that the audio can be extracted from the sample video and used as sample audio in the sample data.
[0220] Optionally, the head region of the real person in the sample video can be determined from the sample video, and then the head parameters of the sample user in the two-dimensional image can be determined from the head region. By using the conversion method from two-dimensional video to three-dimensional modeling, a training digital human can be established, and the head parameters of the training digital human can be obtained. The head parameters of the training digital human can then be used as the head parameters of the sample user in the sample data.
[0221] After obtaining the training digitized human, the emotions of the sample users can be identified based on the facial parameters of the training digitized human. For example, the emotions of the sample users can be determined by comparing the differences between the facial parameters of the 3D digitized human and the facial parameters of a preset 3D digitized human with emotional tendencies.
[0222] For each audio frame, there can be a corresponding training digit. The head parameters of the sample user corresponding to the training digit are the same as the head parameters of the sample user corresponding to the audio frame. Correspondingly, the emotion of the sample user corresponding to the training digit is the same as the emotion of the sample user corresponding to the audio frame.
[0223] The above steps allow us to obtain the head parameters and emotions of the sample user for each audio frame in the sample data.
[0224] It should be noted that the training digital human can be a digital human obtained by transforming two-dimensional graphics into three-dimensional ones. For example, projection modeling can be used to obtain the depth information of two-dimensional images by using different angles of a real person's head in multiple frames of two-dimensional images. Then, the training digital human can be built based on the depth information of the two-dimensional images and the real person's face image displayed in the two-dimensional images.
[0225] In the training method for a 3D digital human model provided in this application embodiment, a corresponding training digital human can be established based on sample videos, and the head parameters of the sample user can be determined based on the head parameters of the training digital human; the emotions of the sample user can be identified through the facial parameters of the training digital human; the sample speech corresponding to each personalized parameter can be determined, as well as the emotion and head parameters of the sample user corresponding to each speech frame of the sample speech. The method of establishing a training digital human allows for faster and more accurate acquisition of the sample user's head parameters and emotions, thus enabling faster and more accurate acquisition of sample data.
[0226] It should be noted that after obtaining the training digital human, the emotions of the training digital human can be determined, and the following methods can be used for determination.
[0227] Figure 8 This is a flowchart illustrating the process of identifying user emotions in sample cases as provided in this application embodiment. Please refer to... Figure 8 By training the facial parameters of digital humans, the system can identify the emotions of sample users, including:
[0228] S810: Calculate the similarity between the facial parameters of the trained digital human and the facial parameters of various pre-set emotional avatar models.
[0229] Optionally, multiple emotion avatar models can be pre-stored on the server side. Each emotion avatar model can be a 3D digital human with an expression tendency. Each 3D digital human with an expression tendency can represent an emotion, such as happiness, sadness, anger, etc. In addition, the facial expression of each emotion model is matched with the corresponding emotion, such as the expression corresponding to the emotion of happiness, the expression corresponding to the emotion of sadness, the expression corresponding to the emotion of anger, etc., without specific restrictions.
[0230] It should be noted that the facial parameters of each emotion model can be the parameters corresponding to the emotion represented by that model.
[0231] For example, different emotion avatar models can be set for different emotions. For instance, human emotions can be divided into 52 different emotions, and an emotion avatar model can be set for each emotion.
[0232] It should be noted that the number and distribution of grid cells in all emotion avatar models can be the same as the number and distribution of grid cells used in training digital humans.
[0233] After obtaining the training digit, the facial parameters of the training digit can be compared with the facial parameters of the emotion avatar model. For example, similarity can be calculated, where the similarity can be determined by comparing the relative positions of the grid cells corresponding to the facial parameters in the training digit and the emotion avatar model.
[0234] The relative position can be the relative position between multiple grid cells in the training digital human and emotion avatar model, such as the positional deviation between any two grid cells.
[0235] It should be noted that the number and distribution of grid cells in the emotion avatar model can be the same as those in the training digital human. In practice, the similarity between the facial parameters of the training digital human and each pre-set emotion avatar model can be determined based on the degree of matching of the relative positions.
[0236] S820: Determine the emotion labels for training digital humans based on similarity.
[0237] The emotion tags include the emotion and weight of the target emotion avatar model, which is the emotion avatar model corresponding to a similarity greater than or equal to a threshold. The emotions of the sample users include the emotion tags used to train the digital human.
[0238] It should be noted that similarity can be represented by weights. For example, if we are comparing the similarity between the training digit and the emotion avatar model corresponding to the emotion of happiness, we can obtain the weight of the emotion of happiness in the training digit. For example, if the similarity is 60%, the weight of the emotion of happiness is 0.6.
[0239] Optionally, the similarity between the trained digital human and each emotion avatar model can be determined in the above manner, thereby obtaining the emotion label of the three-dimensional digital human.
[0240] Among them, emotion labels can be recorded in the form of "emotion + weight". For example, "happy 0.6, sad 0.1, angry 0.1" means that among the expressions corresponding to the emotion label, the weight of the emotion of happiness is 0.6, the weight of the emotion of sadness is 0.1, and the weight of the emotion of anger is 0.1.
[0241] It should be noted that the weight represents the degree of similarity between the trained digital human and the corresponding emotion avatar model. The emotion labels recorded above may indicate that the expression has a similarity of 0.6 to the emotion of happiness, a similarity of 0.1 to the emotion of sadness, and a similarity of 0.1 to the emotion of anger.
[0242] Alternatively, multiple weights can be recorded in the emotion tags in a certain preset order, such as "0.6, 0.1, 0.1, 0.5, ..., 0.2". Each position in the order can correspond to a similarity. Taking a model with 52 emotion avatars as an example, the weights corresponding to 52 emotions can be recorded in the emotion tags.
[0243] It should be noted that in the actual process of determining emotion labels, cases with low similarity can be discarded, and only results with high similarity can be retained. For example, for the label "happy 0.6, sad 0.1, angry 0.1" in the previous example, if the similarity threshold is 0.5, the emotion label can be changed to "happy 0.6".
[0244] In the actual process of representing emotion labels, if there is only one label and the weight is 100%, the emotion corresponding to that label can be used as the label, such as "happy" or "sad". If there is only one label but the weight is not 100%, the emotion and weight corresponding to that label can also be used as the label, such as "happy 0.6" or "sad 0.5". If there are multiple labels and each label has a corresponding weight, they can be represented by a combination, such as "happy 0.6, sad 0.2" or "angry 0.6, sad 0.5, happy 0.1".
[0245] In other words, after determining the similarity, the target emotional avatar model can be identified from the emotional avatar model, that is, the emotional avatar model with a similarity greater than or equal to the threshold, thereby obtaining the corresponding emotional label.
[0246] Optionally, the emotions of the sample users may include emotion labels used to train the digital human. These emotion labels can be used to represent the emotions of the sample users. Alternatively, the emotion labels can be combined with other relevant data to represent the emotions of the sample users. No specific restrictions are imposed here.
[0247] It should be noted that the emotions of sample users can be represented by emotion labels. That is to say, the emotions of sample users in the same frame are not unique, and may include both "happy" and "angry" emotions at the same time. In this case, the emotions of sample users in a certain frame can be represented by emotion labels with weights. For the emotions of sample users in multiple frames, multiple emotion labels can be used to represent them.
[0248] In the training method for the 3D digital human model provided in this application embodiment, the similarity between the facial parameters of the training digital human and each pre-set emotion avatar model can be calculated; based on the similarity, the emotion label of the training digital human is determined. By determining the emotion label, the emotion of each training digital human can be determined in more detail. In addition to the emotion itself, the weight corresponding to the emotion can also be determined, thus obtaining a more accurate and detailed emotion label.
[0249] In one embodiment, the sample video includes multiple video frames, the training digital human includes multiple frames, each frame of the training digital human corresponds one-to-one with a video frame, and the sample user's emotion includes the emotion label of the training digital human for each frame.
[0250] Optionally, the sample speech in the sample video may include multiple speech frames. Each speech frame can correspond to a video frame in the sample video, and correspondingly, it can also correspond to a training digital human frame. The emotions of the training digital human corresponding to each video frame may differ to some extent. For example, the training digital human may have a happy expression in the first video frame and a sad expression in the tenth video frame. In this case, the emotion label of each video frame can be recorded. These multiple emotion labels together constitute the emotion of the sample user, that is, the emotion of the sample user mentioned above.
[0251] It should be noted that during the training of the model, the emotions of the sample users can be trained using one or more specific emotions, or the weights of emotions can be added to the specific emotions for training. For example, the emotions of the sample users in each frame can be input into the audio-visual generation model as an emotion label for training, without any specific restrictions.
[0252] Through the above steps, sample speech, sample user emotion, and sample user head parameters can be obtained, thus obtaining more accurate and detailed sample data. After training the audio-visual generation model based on this sample data, a personalized digital human generation model can be obtained.
[0253] It should be noted that a general digital human generation model can be obtained by normalizing the personalized digital human generation model.
[0254] Correspondingly, the emotional avatar model explained above can also be a three-dimensional digital human with preset expressions, which can be displayed in the following manner.
[0255] Figure 9 This is a schematic diagram of the emotion avatar model provided in the embodiments of this application. Please refer to... Figure 9 , Figure 9 The 3D digital human on the left could be an avatar model corresponding to the emotion of surprise. Figure 9 The 3D digital human on the right could be a model of an emotional avatar corresponding to the emotion of calmness. Figure 9 The example only uses the two emotion avatar models corresponding to these two emotions. In actual implementation, different emotion avatar models can be set for different emotions.
[0256] The following explains a feasible implementation process for establishing the mapping relationship between adjustment parameters and personalized parameters provided in the embodiments of this application.
[0257] Figure 10 This is a schematic diagram of the process for establishing mapping relationships provided in the embodiments of this application. Please refer to... Figure 10 After normalizing different personalized digital human generation models to obtain a general digital human generation model, the method also includes:
[0258] S1010: Input each segment of the P-segment sample speech into the personalized digital human generation model and the general digital human generation model corresponding to the M personalized parameters, respectively, to obtain P×M personalized animated digital humans and one general animated digital human.
[0259] It should be noted that both the personalized digital human generation model and the general digital human generation model can be pre-trained and stored on the server. In actual implementation, each segment of speech in the P-segment sample speech can be input into the personalized digital human generation model and the general digital human generation model corresponding to M personalized parameters. P×M personalized animated digital humans can be obtained through the personalized digital human generation model, and a general animated digital human can be obtained through the general digital human generation model.
[0260] Personalized animated digital humans can be digital humans with personalized characteristics. These personalized characteristics can be features of the personalized parameters corresponding to the personalized digital human generation model, such as gender, age, ethnicity, and emotions, without specific restrictions.
[0261] S1020: Determine the deviation between the head parameters of each frame of each of the M personalized animated digital humans corresponding to each segment of speech in the P-segment speech and the head parameters of each frame of the general animated digital human.
[0262] It should be noted that both personalized animated digital humans and general animated digital humans are digital humans generated based on each speech frame of the same sample speech. The two digital humans have the same number of frames, and each frame corresponds one-to-one. The deviation of head parameters between personalized animated digital humans and general animated digital humans can be compared in each frame. Specifically, it can be the positional deviation of multiple head key points in three-dimensional space.
[0263] For example, the head parameters of each frame of the above P×M personalized animated digital humans can be compared with the head parameters of each frame of the general digital human to obtain the positional deviation of these head key points in three-dimensional space.
[0264] Figure 11 This is a schematic diagram illustrating the head deviation between the personalized animated digital human and the general digital human provided in the embodiments of this application. Please refer to... Figure 11 , Figure 11 In the content shown, the top digital figure is any one of P×M personalized animated digital figures, the bottom digital figure is a general animated digital figure, and the horizontal axis represents the order of the voice frames.
[0265] It should be noted that a speech segment can include N speech frames, which are sequential in time. Inputting the same speech frame into a personalized digital human generation model generates a personalized animated digital human corresponding to that personalized parameter. Each speech frame can correspond to a personalized digital human. See [link to relevant documentation]. Figure 11 N voice frames correspond to N personalized digital humans, and similarly, these N voice frames also correspond to N general digital humans. These N temporally consecutive personalized digital humans constitute a personalized animated digital human corresponding to personalized parameters, and these N temporally consecutive general digital humans constitute a general animated digital human. The Q points in the head mentioned above are all included in... Figure 11 The header parameters of each personalized digital human 1-N and each general digital human an are in the data.
[0266] S1030: Perform weighted averaging on the deviations between the head parameters corresponding to each speech segment to obtain the adjustment parameters, and establish a mapping relationship between the adjustment parameters and the personalized parameters.
[0267] Continue to refer to Figure 11 , Figure 11 The head deviation between the personalized animated digital human and the general animated digital human actually includes the head deviation between the personalized digital human and the general digital human corresponding to N speech frames. In other words, for the same speech, the head deviation between the personalized animated digital human and the general animated digital human obtained after inputting into a personalized digital human generation model and a general digital human generation model includes the head deviation between the personalized digital human and the general digital human corresponding to each speech frame. The head deviation between the personalized digital human and the general digital human corresponding to each speech frame also includes the difference in the three-dimensional spatial coordinates of the Q key points mentioned above that are included in the personalized digital human and the general digital human corresponding to each speech frame.
[0268] For example, see continue. Figure 11 In the first speech frame, the head parameters of both digital human 1 and digital human a include Q key grid cells. The coordinates of the Q key grid cells on digital human 1 are: A1(X1,Y1,Z1), B1(X2,Y2,Z2), ..., N1(XQ,YQ,ZQ); the coordinates of the Q key grid cells on digital human a are: Aa(x1,y1,z1), Ba(x2,y2,z2), ..., Na(xQ,yQ,zQ). Each key grid cell on the personalized digital human and the general digital human is in one-to-one correspondence. Therefore, the head deviation between digital human 1 and digital human a is the difference between the three-dimensional coordinates of grid cells A1-N1 on personalized digital human 1 and grid cells Aa-Na on digital human a, denoted as δ. 111 In this context, any grid cell on a digital human can be understood as a point in a three-dimensional coordinate system.
[0269] Similarly, the calculation method for the head parameter deviation between personalized digital humans 2-N and general digital humans bn can refer to the calculation method between personalized digital humans 1 and general digital humans a, and will not be repeated here. Therefore, N sets of head parameter deviations (δ) can be obtained. 111 δ 112 ……δ 11N ).
[0270] The head parameter deviation between each personalized digital human and each general digital human can be considered as a frame adjustment parameter. This method yields the head parameter deviation between a personalized animated digital human and a general animated digital human. The required adjustment parameters include the N frame adjustment parameters corresponding to these N speech frames. Within the same speech segment, for a personalized animated digital human and a general animated digital human corresponding to a single personalized parameter, the obtained adjustment parameter is (δ... 111 δ 112 ……δ 11N ), where δ 111 Adjust the frame parameters δ corresponding to the first speech frame. 112 δ is the frame adjustment parameter corresponding to the second speech frame. 11N Let N be the frame adjustment parameters corresponding to the Nth speech frame; and so on. Then, for the M personalized animated digital humans corresponding to the M personalized parameters, calculate the N frame adjustment parameters between these M personalized animated digital humans and the general animated digital human, and obtain the first mapping relationship shown in Table 5-1:
[0271] Table 5-1
[0272] Personalized parameters (M items) Adjust parameters male <![CDATA[(δ 111 ,d 112 ……d 11N )]]> Female, 20 years old, of East Asian descent, happy <![CDATA[(δ 121 ,d 122 ……d 12N )]]> Women, happy <![CDATA[(δ 131 ,d 132 ……d 13N )]]> …… …… Male, 40 years old, happiness 0.5 + sadness 0.5 <![CDATA[(δ 1N1 ,d 1N2 ……d 1NN )]]>
[0273] Table 5-1 shows only one example, where δ 111 δ 112 ……δ 11N Each parameter in the table represents a frame adjustment parameter. It should be noted that the first mapping relationship obtained in Table 5-1 is the correspondence between personalized parameters and adjustment parameters for a specific speech segment. However, if different speech segments are input into the personalized digital human generation model and the general digital human generation model corresponding to the speech segments in Table 5-1, the resulting personalized animated digital human and another general animated digital human will differ from the examples above (such as...). Figure 11 The personalized animated digital human and the general animated digital human obtained are different. In other words, Table 5-1 above is actually a mapping relationship corresponding to a speech segment. Different speech segments will result in different mapping relationships.
[0274] Based on this, in one implementation, to improve the accuracy of the adjustment parameters, after any subsequent voice input by the user, the target adjustment parameters corresponding to the personalized parameters input by the user can be obtained based on the same mapping table. When establishing the mapping relationship, P (P is an integer ≥ 1) different voice segments can be input into M personalized digital human generation models and a general digital human generation model, which are the same as those in the embodiment corresponding to Table 5-1 above, to obtain P mapping relationships, as follows: (For example, they can be Table 5-2, Table 5-3, etc., up to Table 5-P.)
[0275] Table 5-2
[0276] Personalized parameters (M items) Adjust parameters male <![CDATA[(δ 211 ,d 212 ……d 21N )]]> Female, 20 years old, of East Asian descent, happy <![CDATA[(δ 221 ,d 222 ……d 22N )]]> Women, happy <![CDATA[(δ 231 ,d 232 ……d 23N )]]> …… …… Male, 40 years old, happiness 0.5 + sadness 0.5 <![CDATA[(δ 2N1 ,d 2N2 ……d 2NN )]]>
[0277] Table 5-3
[0278] ...
[0279] Table 5-P
[0280] Personalized parameters (M items) Adjust parameters male <![CDATA[(δ p11 ,d p12 ……d p1N )]]> Female, 20 years old, of East Asian descent, happy <![CDATA[(δ p21 ,d p22 ……d p2N )]]> Women, happy <![CDATA[(δ p31 ,d p32 ……d p3N )]]> …… …… Male, 40 years old, happiness 0.5 + sadness 0.5 <![CDATA[(δ pN1 ,d pN2 ……d pNN )]]>
[0281] After obtaining the P mapping relationships as shown in Tables 5-1 to 5-P, a weighted average can be calculated for the P adjustment parameters corresponding to each personalized parameter, ultimately yielding an averaged adjustment parameter. The calculation method is as follows:
[0282] For example, for the personalized parameter of male, the weighted average of the corresponding adjusted parameters in Tables 5-1 to 5-P is as follows:
[0283]
[0284] in, Adjust the parameters for the frame corresponding to the first speech frame. Adjust the frame parameters corresponding to the second speech frame. The frame adjustment parameter is the frame adjustment parameter corresponding to the Nth speech frame. The frame adjustment parameter corresponding to any speech frame is the difference in the three-dimensional coordinates of Q key points on the personalized digital human head and the general digital human head.
[0285] Similarly, we can obtain the adjustment parameter δi corresponding to each personalized parameter in Table 5-1 up to Table 5-P, where 1≤i≤M;
[0286]
[0287] in, It can be denoted as δ i1 ;
[0288] It can be denoted as δ i2 ; It can be denoted as δ iN .
[0289] The mapping relationship between the personalized parameters and the adjustment parameters obtained by the above calculation method can be represented by the aforementioned Table 3. The adjustment parameter corresponding to any personalized parameter in Table 3 is calculated by Equation (2).
[0290] For example, referring to Table 3 above, a personalized parameter is: a 20-year-old happy girl. Using the above method, the adjustment parameter corresponding to the 15-year-old happy girl is δ2. Similarly, the adjustment parameters corresponding to other personalized parameters can also be obtained in the above way. In this way, multiple adjustment parameters corresponding to multiple personalized parameters can be obtained, thereby establishing a mapping relationship between adjustment parameters and personalized parameters.
[0291] In the training method for the 3D digital human model provided in this embodiment, a sample speech can be input into a personalized digital human generation model and a general digital human generation model corresponding to the personalized parameters to obtain a personalized animated digital human and a general animated digital human; the deviation between the head parameters of each frame of the personalized animated digital human and the head parameters of each frame of the general animated digital human is determined; the deviation between the head parameters corresponding to each speech segment is weighted and averaged to obtain adjustment parameters, and a mapping relationship between the adjustment parameters and the personalized parameters is established. By establishing the mapping relationship, the adjustment parameters can be determined more quickly during the generation of the 3D digital human, thereby allowing for a faster acquisition of the animated target digital human.
[0292] The execution steps of the above method are mainly explained with the server as the execution subject. The following explanation uses the client as the execution subject to explain the three-dimensional digital human generation method provided in the embodiments of this application.
[0293] Optionally, the client can obtain the target's personalized parameters and the first voice.
[0294] The personalized parameters for the target include the target's identity information and the target's emotional information.
[0295] Optionally, for the client, the target personalization parameters can be actively entered by the user, or the user can select from multiple preset personalization parameters stored in the client; no specific restrictions are imposed here.
[0296] The first voice can be the voice input by the user, or the voice converted from the text input by the user; no specific restrictions are imposed here.
[0297] Optionally, the client sends the target personalization parameters and the first voice to the server.
[0298] Optionally, the client can send the target personalization parameters and the first voice to the server. The target personalization parameters can be sent to the server by the client after initialization, and the first voice can be sent to the server after the user inputs the voice.
[0299] The target personalized parameters and the first voice message can be sent together, or they can be sent sequentially according to the order in which they are obtained; there is no specific restriction here.
[0300] Optionally, the client can receive the animated target digital human sent by the server.
[0301] It should be noted that after the client sends the target personalized parameters and the first voice to the server, the server can obtain the animated target digital human based on these two parameters and the general digital human generation model, and then send the animated target digital human to the client.
[0302] Optionally, in practical application scenarios, one server can correspond to multiple clients. These clients can all perform the above steps, and the same server will complete the generation of the corresponding 3D digital human to be displayed. Then, based on the target personalized parameters and the source of the first voice, the corresponding 3D digital human to be displayed will be sent to the corresponding client.
[0303] Optionally, the client can render based on the target digital human in the animation.
[0304] In one embodiment, the first voice includes multiple voice frames; the method further includes: the client can obtain the target digital person and target deviation parameters sent by the server.
[0305] Optionally, the server can obtain the target digital human and multiple target deviation parameters through the animation target digital human, and can send the target digital human and multiple target deviation parameters to the client, so that the client can render the target digital human according to the target deviation parameters.
[0306] It should be noted that during the display of the 3D digital human, the client can simultaneously output the first voice. This first voice can be output according to its frame correspondence with the rendered 3D digital human. Alternatively, the server can transmit the voice frame of the first voice corresponding to the frame while transmitting the 3D digital human to be displayed and the target deviation parameters. For example, it can transmit the first voice to the client in the form of WebSocket protocol. No specific restrictions are made here.
[0307] It should be noted that the server can first send the target digital human to the client, and then send multiple target deviation parameters to the client in the order of the voice frames.
[0308] For example: After receiving the target digital human, the client can display the target digital human on the client's display interface. After receiving the target deviation parameter, the client can adjust the corresponding mesh cells in the target digital human based on the target deviation parameter to obtain the rendered 3D digital human. Each time the target deviation parameter is received, a rendering can be performed, and each rendering can produce a rendered 3D digital human. These rendered 3D digital humans can be displayed on the client's display interface in the corresponding order.
[0309] In one embodiment, the method further includes: sending target information to a server, the target information being used to indicate preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, display requirements, and the weight of the facial region, so that the server determines a target deviation parameter from a plurality of deviation parameters according to the preset conditions.
[0310] It should be noted that the target information can be sent actively by the client to the server, or it can be sent by the client to the server after the server sends a request; no specific restrictions are imposed here.
[0311] After the client sends the target information indicating the preset conditions to the server, the server can determine the target deviation parameter from multiple deviation parameters and send the target deviation parameter to the client.
[0312] It should be noted that the digital human sent from the server to the client can be the animation target digital human. However, the server does not directly generate the animation target digital human, but rather obtains it by adjusting the initial animation digital human. The following explains an optional implementation process for converting the initial animation digital human into the animation target digital human provided in this application embodiment.
[0313] Figure 12 For a diagram showing the correspondence between the initial digital human and the target digital human in the embodiments of this application, please refer to... Figure 12 , Figure 12 The digital figure on the top is the target digital figure for the animation, and the digital figure on the bottom is the initial digital figure for the animation. Each frame of the initial digital figure can be adjusted by adjusting the corresponding adjustment parameters to obtain the target digital figure for that frame. The deviation of the head parameters between multiple target digital figures is the aforementioned deviation parameter.
[0314] Figure 12 The initial digital person and the target digital person in the animation shown each have N frames.
[0315] Taking the personalized parameter as male as an example, see Table 3. At this time, the target adjustment parameter is δ1, where δ1 includes N frame adjustment parameters. The value of each frame adjustment parameter has been calculated and stored by the server as shown in equation (2) when establishing the mapping relationship. The adjustment parameter of each frame is δ 11 to δ 1N .
[0316] The difference between two adjacent frames of the animated target digital human is a deviation parameter. The deviation parameter between each pair of adjacent frames can be recorded as Ψ. The deviation parameter Ψ between each pair of adjacent frames can include the positions of multiple head key points.
[0317] During the process of sending the animated target digital human from the server to the client, what is sent can be the target digital human of the first frame, and the deviation parameter Ψ for each adjacent two frames thereafter. The deviation parameter Ψ can be sent directly, or the target deviation parameter can be determined from the deviation parameters Ψ of each adjacent two frames.
[0318] The following explanation will focus on another implementation of the three-dimensional digital human generation method provided in this application from the perspective of data interaction between the client and the server.
[0319] Figure 13 This is a schematic diagram illustrating the interaction between the client and server provided in the embodiments of this application. Please refer to... Figure 13 The method includes:
[0320] S1310: The client obtains the target personalized parameters and the first voice.
[0321] S1320: The client sends the target personalization parameters and the first voice to the server.
[0322] S1330: The server determines the target adjustment parameters based on the target personalized parameters and mapping relationship, and obtains the animated target digital human based on the target adjustment parameters and the initial animated digital human generated based on the first voice.
[0323] S1340: The server sends the animated target digital human to the client.
[0324] S1350: The client renders the digital human based on the target of the animation.
[0325] It should be noted that the specific implementation process and principle of the above steps S1310-S1350 are the same as those in the aforementioned embodiments, and will not be repeated here.
[0326] For example, taking a live streaming scenario, the client can initialize itself, such as loading its own personalized configuration and determining the target personalized parameters based on the user's selection. Additionally, if the client has a previously displayed 3D digital human, and the current rendering scene is the same as the previous one, it can directly load the cached 3D digital human, along with its skeleton, textures, and configuration files. Correspondingly, the server can also initialize itself, then build a general digital human generation model based on the received first voice and target personalized parameters, obtaining the 3D digital human to be displayed. Multiple deviation parameters can be determined from the 3D digital human in different frames and sent to the client. The client can continue rendering to complete real-time inference for the digital human until the live stream ends.
[0327] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0328] Based on the foregoing embodiments, this application provides a three-dimensional digital human generation device, which includes various modules and units included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.
[0329] Figure 14 This is a schematic diagram of the structure of the three-dimensional digital human generation device provided in the embodiments of this application. Please refer to... Figure 14 A three-dimensional digital human generation device, the device comprising: an acquisition module 1410, a determination module 1420 and a generation module 1430;
[0330] The acquisition module 1410 is used to acquire target personalized parameters; the target personalized parameters include at least one of target identity information and target emotion information.
[0331] The determination module 1420 is used to determine the target adjustment parameters based on the target personalized parameters and the mapping relationship. The mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters. The adjustment parameters are used to indicate the deviation of the head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model.
[0332] The generation module 1430 is used to generate an animated target digital human based on the target adjustment parameters and the initial animated digital human. The initial animated digital human is generated based on the first voice input to the general digital human generation model.
[0333] In one embodiment, the first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, with each frame adjustment parameter corresponding to a speech frame; the generation module 1430 is specifically used to adjust each frame of the initial animated digital human corresponding to each speech frame based on a frame adjustment parameter corresponding to each speech frame, thereby generating the animated target digital human.
[0334] In one embodiment, each frame of the digital human includes multiple grid cells in three-dimensional space. The adjustment parameters for each frame include: the identifier of the grid cell to be adjusted in the multiple grid cells and the offset value of each grid cell to be adjusted. The generation module 1430 is specifically used to adjust the three-dimensional space coordinates of the grid cells to be adjusted in the initial animation digital human corresponding to the voice frame based on the identifier of the grid cell to be adjusted and the offset value of each grid cell to be adjusted in the adjustment parameters of each frame, so as to generate the animation target digital human corresponding to the voice frame.
[0335] In one embodiment, the generation module 1430 is further configured to send the animated target digital man to the client so that the client can render the animated target digital man.
[0336] In one embodiment, the generation module 1430 is specifically used to send the target digital person and target deviation parameters corresponding to the first speech frame in the first speech to the client; wherein, the target deviation parameters are at least a portion of a plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital person corresponding to adjacent speech frames in the animated target digital person.
[0337] In one embodiment, each deviation parameter includes: the identifier of the grid cell to be adjusted in multiple grid cells of the target digital human corresponding to adjacent speech frames and the offset value of each grid cell to be adjusted; the generation module 1430 is specifically used to send multiple target deviation parameters to the client in sequence according to the time order of each speech frame in the first speech.
[0338] In one embodiment, the determining module 1420 is further configured to determine a target deviation parameter from a plurality of deviation parameters according to preset conditions, the preset conditions including at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter; the generating module 1430 is further configured to send the target deviation parameter to the client.
[0339] Figure 15 This is a schematic diagram of the structure of the training device for the three-dimensional digital human model provided in the embodiments of this application. Please refer to... Figure 15 A training device for a three-dimensional digital human model, the device comprising: a video acquisition module 1510, a data acquisition module 1520, a training module 1530, and a normalization module 1540.
[0340] The video acquisition module 1510 is used to acquire sample videos corresponding to different personalized parameters;
[0341] The data acquisition module 1520 is used to acquire corresponding sample data from the sample videos corresponding to each personalized parameter. The sample data includes sample voice, sample user's emotion, and sample user's head parameters.
[0342] Training module 1530 is used to train the audio-visual generation model based on the sample data corresponding to each personalized parameter, so as to obtain a personalized digital human generation model corresponding to each personalized parameter.
[0343] The normalization module 1540 is used to normalize different personalized digital human generation models to obtain a general digital human generation model.
[0344] In one embodiment, the data acquisition module 1520 is specifically used to establish a corresponding training digital human based on the sample video, and determine the head parameters of the sample user based on the head parameters of the training digital human; identify the emotions of the sample user through the facial parameters of the training digital human; determine the sample speech corresponding to each personalized parameter, and the emotions and head parameters of the sample user corresponding to each speech frame of the sample speech.
[0345] In one embodiment, the data acquisition module 1520 is further configured to calculate the similarity between the facial parameters of the training digital human and the facial parameters of each pre-set emotion avatar model; and determine the emotion label of the training digital human based on the similarity, wherein the emotion label includes the emotion and weight of the target emotion avatar model, the target emotion avatar model being the emotion avatar model corresponding to a similarity greater than or equal to a threshold, and the emotion of the sample user includes the emotion label of the training digital human.
[0346] In one embodiment, the device includes a sample video comprising multiple video frames, a training digital human comprising multiple frames, each frame of the training digital human corresponding one-to-one with a video frame, and the emotions of the sample user comprising the emotion labels of the training digital human for each frame.
[0347] In one embodiment, the normalization module 1540 is further configured to input each segment of the P-segment sample speech into the personalized digital human generation model and the general digital human generation model corresponding to the M personalized parameters, respectively, to obtain P×M personalized animated digital humans and one general animated digital human; wherein, a segment of speech corresponds to M personalized animated digital humans, and P and M are both positive integers greater than or equal to 2; respectively determine the deviation between the head parameters of each frame of each personalized animated digital human and the head parameters of each frame of the general animated digital human in the M personalized animated digital humans corresponding to each segment of speech; perform weighted averaging on the deviation between the head parameters corresponding to each segment of speech to obtain adjustment parameters, and establish a mapping relationship between the adjustment parameters and personalized parameters.
[0348] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0349] It should be noted that, in the embodiments of this application... Figure 14 and Figure 15 The module division of the illustrated 3D digital human generation device and 3D digital human training device is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit with two or more units. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0350] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0351] Figure 16 This is a schematic diagram of the structure of the computing device provided in the embodiments of this application. Please refer to... Figure 16 This application provides a computing device, which can be the aforementioned server, without specific limitations. Its internal structure diagram can be as follows. Figure 16 As shown. The computing device includes a processor 1620, memory, and a network interface 1640 connected via a system bus 1610. The processor 1620 provides computing and control capabilities. The memory includes a non-volatile storage medium 1631 and internal memory 1632. The non-volatile storage medium 1631 stores the operating system, computer programs, and a database. The internal memory 1632 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium 1631. The database is used to store data. The network interface 1640 is used for communication with external terminals via a network connection; for example, communication between a client and a server can be achieved through the network interface 1640. When the computer program is executed by the processor 1620, it implements the above-described methods.
[0352] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.
[0353] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0354] Those skilled in the art will understand that Figure 16The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computing device on which the present application is applied. The specific computing device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0355] In one embodiment, the 3D digital human generation device and the 3D digital human model training device provided in this application can be implemented as a computer program, and the computer program can be implemented in the form of, for example... Figure 16 The device operates on the computing device shown. The memory of the computing device can store the various program modules that make up the above-described apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the methods in the various embodiments of this application described in this specification.
[0356] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0357] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0358] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0359] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0360] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0361] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0362] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0363] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0364] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0365] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0366] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0367] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0368] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating animated digital humans, characterized in that, The method includes: Obtain target personalized parameters; the target personalized parameters include at least one of target identity information and target emotion information; Based on the target personalized parameters and mapping relationship, target adjustment parameters are determined; the mapping relationship includes a one-to-one correspondence between multiple personalized parameters and multiple adjustment parameters, and the adjustment parameters are used to indicate the deviation of head parameters between the animated digital humans obtained by inputting the same voice into the personalized digital human generation model and the general digital human generation model. Based on the target adjustment parameters and the initial animation digital human, an animated target digital human is generated. The initial animation digital human is generated based on the first voice input to the general digital human generation model.
2. The method according to claim 1, characterized in that, The first speech includes multiple speech frames, and the target adjustment parameters include multiple frame adjustment parameters, each of which corresponds to one speech frame; the step of generating the animated target digital human based on the target adjustment parameters and the initial animated digital human specifically includes: Based on a frame adjustment parameter corresponding to each speech frame, each frame of the initial animated digital human corresponding to the first speech is adjusted to generate the target animated digital human.
3. The method according to claim 2, characterized in that, Each frame corresponds to a digital human comprising multiple grid cells in three-dimensional space. Each frame adjustment parameter includes: the identifier of the grid cell to be adjusted among the multiple grid cells and the offset value of each grid cell to be adjusted. The step of adjusting each frame of the initial animated digital human corresponding to the first speech based on the adjustment parameters of each frame to generate the target animated digital human specifically includes: Based on the identifier of the grid cell to be adjusted in the adjustment parameters of each frame and the offset value of each grid cell to be adjusted, the three-dimensional spatial coordinates of the grid cells to be adjusted in the initial animated digital human corresponding to the voice frame are adjusted to generate the animated target digital human corresponding to the voice frame.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: The animated target digital human is sent to the client so that the client can render the animated target digital human.
5. The method according to claim 4, characterized in that: Sending the animated target digital human to the client specifically includes: The target digital person and target deviation parameters corresponding to the first audio frame in the first audio are sent to the client; wherein, the target deviation parameters are at least a portion of the plurality of deviation parameters; each deviation parameter includes the deviation of the head parameters of the animated target digital person corresponding to the adjacent audio frames in the animated target digital person.
6. The method according to claim 5, characterized in that, Each of the deviation parameters includes: the identifier of the grid cell to be adjusted in the plurality of grid cells of the target digital human corresponding to adjacent speech frames, and the offset value of each of the grid cells to be adjusted; The step of sending the target digital person and target deviation parameters corresponding to the first speech frame in the first speech to the client includes: The multiple target deviation parameters are sent to the client sequentially according to the chronological order of each audio frame in the first audio recording.
7. The method according to claim 5 or 6, characterized in that, Sending the target deviation parameter to the client includes: According to preset conditions, a target deviation parameter is determined from the plurality of deviation parameters. The preset conditions include at least one of the client's bandwidth configuration information, the client's display requirements, and the weight of the facial region corresponding to the deviation parameter. The target deviation parameter is sent to the client.
8. A training method for an animated digital human model, characterized in that, The method includes: Obtain sample videos corresponding to different personalized parameters; the personalized parameters include at least one of target identity information and target emotion information. Obtain corresponding sample data from sample videos corresponding to each personalized parameter. The sample data includes sample voice, sample user's emotion, and sample user's head parameters. Based on the sample data corresponding to each personalized parameter, a sound and image generation model is trained to obtain a personalized digital human generation model corresponding to each personalized parameter. The different personalized digital human generation models are normalized to obtain a general digital human generation model. The personalized digital human generation model and the general digital human generation model are applied in the method described in any one of claims 1-7.
9. The method according to claim 8, characterized in that, The step of obtaining corresponding sample data from sample videos corresponding to each personalized parameter includes: A corresponding training digital human is established based on the sample video, and the head parameters of the sample user are determined based on the head parameters of the training digital human; wherein, the training digital human is obtained by transforming a two-dimensional graphic into a three-dimensional graphic. The emotions of the sample users are identified by using the facial parameters of the trained digital human. Determine the sample speech corresponding to each personalized parameter, as well as the emotion and head parameters of the sample user corresponding to each speech frame of the sample speech.
10. The method according to claim 9, characterized in that, The process of identifying the emotions of the sample users using the facial parameters of the trained digital human includes: Calculate the similarity between the facial parameters of the trained digital human and the facial parameters of each pre-set emotional avatar model; Based on the similarity, the emotion label of the training digital human is determined. The emotion label includes the emotion and weight of the target emotion avatar model. The target emotion avatar model is the emotion avatar model corresponding to the similarity being greater than or equal to the threshold. The emotion of the sample user includes the emotion label of the training digital human.
11. The method according to claim 10, characterized in that, The sample video includes multiple video frames, the training digital human includes multiple frames, each frame of the training digital human corresponds one-to-one with the video frame, and the emotions of the sample user include the emotion labels of the training digital human for each frame.
12. The method according to claim 8, characterized in that, After normalizing the different personalized digital human generation models to obtain a general digital human generation model, the method further includes: Each segment of the P-segment sample speech is input into the personalized digital human generation model corresponding to M personalized parameters and the general digital human generation model, respectively, to obtain P×M personalized animated digital humans and one general animated digital human; wherein, a segment of speech corresponds to M personalized animated digital humans, and P and M are both positive integers greater than or equal to 2; In each of the P segments of speech, the head parameters of each frame of each of the M personalized animated digital humans corresponding to each segment of speech are determined to be different from the head parameters of each frame of the general animated digital human. The deviations between the head parameters corresponding to each speech segment are weighted and averaged to obtain adjustment parameters, and a mapping relationship between the adjustment parameters and the personalized parameters is established.
Citation Information
Patent Citations
Personalized face model display method, device and equipment and storage medium
CN110689604A
Speech-based three-dimensional face model driving method and related device
CN116188649A