Training method of deep learning model, virtual image driving method and device
By using multiple sub-models to process virtual images with different topologies in a deep learning model and adjusting parameters through mask loss, the problem of training with a single topology in existing technologies is solved, and the generation of 3D virtual images with multiple topologies and the improvement of model accuracy are realized.
Patent Information
- Application Number
- CN202211660898.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-12-22
AI Technical Summary
Existing speech-driven deep learning models can only train 3D virtual images with a single topological structure, which has a narrow scope of application and insufficient model accuracy due to the small amount of training data.
Multiple sub-models are used to drive virtual images with different topologies. The model parameters are adjusted by calculating the mask loss, so that each sub-model can drive the 3D virtual image with the corresponding topology, thereby improving training efficiency and model generalization ability.
It has enabled the creation of 3D virtual avatars capable of driving various topological structures, improving the accuracy and generalization ability of the models and generating a wider variety of 3D face styles.
Smart Images

Figure CN115906987B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and in particular to the technical field of virtual human, augmented reality, virtual reality, mixed reality, extended reality, metaverse, etc. More specifically, the present disclosure provides a training method of a deep learning model, a virtual image driving method, an apparatus, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of technologies such as the Internet, three-dimensional (3-Dimensional), augmented reality (Augmented Reality), virtual reality (Virtual Reality) and metaverse, virtual images are increasingly widely used in live streaming, virtual socializing, entertainment media and the like. SUMMARY
[0003] The present disclosure provides a training method of a deep learning model, a virtual image driving method, an apparatus, an electronic device and a storage medium.
[0004] According to a first aspect, a training method of a deep learning model is provided, which includes: obtaining a first audio feature of a sample voice, the sample voice having a virtual image label, the virtual image label containing topology structure information; inputting the first audio feature into the deep learning model to obtain a plurality of first driving parameters corresponding to a plurality of topology structures respectively; determining a first target driving parameter from the plurality of first driving parameters according to the topology structure information; and adjusting the deep learning model according to a difference between the topology structure information and the first target driving parameter to obtain a trained deep learning model.
[0005] According to a second aspect, a virtual image driving method is provided, which includes: obtaining a second audio feature of a to-be-processed voice; inputting the second audio feature into the deep learning model to obtain a second driving parameter; and generating a virtual image according to the second driving parameter; wherein the deep learning model is trained according to the training method of the deep learning model described above.
[0006] According to a third aspect, a training apparatus of a deep learning model is provided, which includes: a first obtaining module configured to obtain a first audio feature of a sample voice, the sample voice having a virtual image label, the virtual image label containing topology structure information; a first processing module configured to input the first audio feature into the deep learning model to obtain a plurality of first driving parameters corresponding to a plurality of topology structures respectively; a determining module configured to determine a first target driving parameter from the plurality of first driving parameters according to the topology structure information; and an adjusting module configured to adjust the deep learning model according to a difference between the topology structure information and the first target driving parameter to obtain a trained deep learning model.
[0007] According to a fourth aspect, a virtual figure driving apparatus is provided, the apparatus comprising: a second acquisition module configured to acquire a second audio feature of a voice to be processed; a second processing module configured to input the second audio feature into a deep learning model to obtain a second driving parameter; and a generation module configured to generate a virtual figure according to the second driving parameter; wherein the deep learning model is trained according to the training apparatus of the deep learning model.
[0008] According to a fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.
[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to execute the method provided by the present disclosure.
[0010] According to a seventh aspect, a computer program product is provided, comprising a computer program stored in at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method provided by the present disclosure.
[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:
[0013] Figure 1A is a schematic diagram of a training method of a face driving model driven by voice in the related art;
[0014] Figure 1B is a schematic diagram of a three-dimensional face generation method driven by voice in the related art;
[0015] Figure 2 is a flowchart of a training method of a deep learning model according to an embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram of a training method of a deep learning model according to an embodiment of the present disclosure;
[0017] Figure 4 is a flowchart of a virtual figure driving method according to an embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram of a virtual image driving method according to an embodiment of the present disclosure;
[0019] Figure 6 is a block diagram of a training device of a deep learning model according to an embodiment of the present disclosure;
[0020] Figure 7 is a block diagram of a virtual image driving device according to an embodiment of the present disclosure;
[0021] Figure 8 is a block diagram of an electronic device of a deep learning model training method and / or a virtual image driving method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0023] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.
[0024] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before the user's personal information is obtained or collected.
[0025] Virtual images include, for example, three-dimensional virtual digital humans, which include virtual anchors, virtual customer service, virtual idols, etc. With the development of three-dimensional virtual digital humans, voice-driven three-dimensional face technology has become one of the important research focuses of virtual human interaction.
[0026] Voice-driven three-dimensional virtual image (e.g., three-dimensional face) technology is generally implemented based on deep learning model technology, which uses voice as a driving source, three-dimensional virtual images as a driving target, and uses a deep learning model to generate three-dimensional virtual images.
[0027] Figure 1A is a schematic diagram of a related art voice-driven deep learning model training method.
[0028] As Figure 1AAs shown, the face driving model 110 is a speech driving based deep learning model to be trained. In the training stage, the sample speech is first sent to a feature extraction model for extracting audio features. The feature extraction model can be a deep learning model or a traditional processing module such as a Fourier processing module, etc. The audio features can be frequency features, spectrum features, etc.
[0029] Then, the audio features are input to the face driving model 110, which can convert the audio features into driving parameters, for example, topological structure information containing a three-dimensional face, the topological structure information containing the number and position of key points (vertices), which can be used for the generation of a three-dimensional face. The driving parameters can also be BlendShape (BlendShape) weights, PCA (Principal Component Analysis, Principal Component Analysis Coefficient) coefficients, etc. The BlendShape weights and PCA coefficients can be converted into topological structure for the generation of a three-dimensional face.
[0030] Next, the gap between the driving parameters and the true value (e.g. labeled topological structure) of the sample speech is calculated, and the parameters of the face driving model 110 are adjusted by back propagation. The true value of the sample speech can be specific portrait data collected and labeled in advance for the sample speech.
[0031] Figure 1B is a schematic diagram of a speech driving based three-dimensional face generation method in the related art.
[0032] As shown in Figure 1B , the face driving model 110 can be a deep learning model trained by the training process as shown in Figure 1A . The audio features are obtained by sending the speech to be processed to the feature extraction model, and the driving parameters are obtained by inputting the audio features to the trained face driving model 110. The driving parameters are input into the three-dimensional face generation model to complete the driving to generate a three-dimensional face. The three-dimensional face generation model is, for example, a processing module for performing smoothing, rendering and other post-processing.
[0033] As shown in Figures 1A-1B , the face driving model based on speech driving in the related art is trained for training data with specific topological structure true value, so the trained face driving model can only reconstruct a single specific three-dimensional face with the same topological structure, and the application scope is narrow.
[0034] In other words, to obtain a three-dimensional face with different topological structures, a corresponding face driving model needs to be trained using training data with different topological structure labels, which consumes a lot of time. In addition, the training data with specific topological structure labels generally has a small amount of data, which also leads to insufficient accuracy of the trained face driving model.
[0035] Figure 2 is a flowchart of a training method of a deep learning model according to one embodiment of the present disclosure.
[0036] As shown in Figure 2 The training method 200 of the deep learning model can include operations S210-S240. The deep learning model can be a voice-driven face driving model.
[0037] In operation S210, a first audio feature of a sample voice is obtained.
[0038] For example, the sample voice can be recorded or converted from text. The sample voice has a virtual image label containing topological structure information. The virtual image label of the sample voice can be three-dimensional face data collected and labeled in advance for the sample voice, which has specific topological structure information.
[0039] For example, the sample voice can be input into a feature extraction model to obtain the first audio feature of the sample voice. The feature extraction model can be a deep learning model or a traditional processing module such as a Fourier processing module. The first audio feature can be a frequency feature, a spectrum feature, etc.
[0040] In operation S220, the first audio feature is input into the deep learning model to obtain a plurality of first driving parameters each corresponding to a plurality of topological structures.
[0041] For example, the first audio feature is input into the deep learning model, and the deep learning model can perform virtual image driving processing for the first audio feature in multiple branches, each branch corresponding to a topological structure. The virtual image driving processing can convert the first audio feature into a visual feature (visual parameter), i.e., the first driving parameter.
[0042] For example, the first driving parameter can be topological structure information containing a three-dimensional virtual image, including the number and position of key points (vertices). The first driving parameter can also be a Blend Shape weight, a PCA coefficient, etc., which can be converted into corresponding topological structure information, respectively.
[0043] For example, the deep learning model can obtain first driving parameters corresponding to different topological structures by processing the first audio features through different topological structure branches.
[0044] In operation S230, a first target driving parameter is determined from the plurality of first driving parameters according to the topological structure information.
[0045] For example, the avatar label of each sample voice has specific topological structure information. For the plurality of first driving parameters output by the deep learning model, a first driving parameter corresponding to the topological structure information of the label can be selected as the first target driving parameter. The difference between the topological structure information of the label and the first target driving parameter can be calculated to obtain the loss of the voice sample.
[0046] In operation S240, the deep learning model is adjusted according to the difference between the topological structure information and the first target driving parameter to obtain a trained deep learning model.
[0047] For example, the label of the sample voice contains topological structure information of topological structure A, and a first driving parameter corresponding to topological structure A can be selected from the plurality of first driving parameters as the first target driving parameter. The mean square error or mean absolute error between the topological structure information contained in the label of the sample voice and the first target driving parameter can be calculated as the loss of the sample voice.
[0048] For example, the loss of the sample voice can be used to adjust the parameters of the corresponding branch processing module of the deep learning model, so that the corresponding processing branch module has the ability to drive the avatar of the specific topological structure corresponding to the label of the sample voice. Thus, each branch processing module of the deep learning model can have the ability to drive the avatar of the corresponding specific topological structure.
[0049] The embodiment can obtain a plurality of first driving parameters by performing avatar driving processing for different topological structures on the input audio features, select a first target driving parameter corresponding to the sample voice from the plurality of first driving parameters, calculate the loss of the sample voice with the topological structure information contained in the label of the sample voice, and adjust the parameters of the deep learning model using the loss, so that the deep learning model has the ability to drive multiple avatars.
[0050] According to an embodiment of the present disclosure, the deep learning model comprises a plurality of sub-models each corresponding to a plurality of topological structures; operation S220 comprises inputting the first audio feature into the plurality of sub-models to obtain a plurality of first driving parameters each output by the plurality of sub-models. Operation S230 comprises determining a target sub-model from the plurality of sub-models according to the topological structure information; determining the first driving parameter output by the target sub-model as a first target driving parameter. Operation S240 comprises calculating a mask loss of the target sub-model according to the difference between the topological structure information and the first target driving parameter; adjusting the parameters of the target sub-model according to the mask loss of the target sub-model to obtain a trained deep learning model.
[0051] For example, the plurality of branch processing modules of the deep learning model each corresponding to a plurality of topological structures can be a plurality of sub-models, and the plurality of first driving parameters can be a plurality of first driving parameters each output by the plurality of sub-models.
[0052] For example, the deep learning model comprises a plurality of sub-models each corresponding to a plurality of topological structures, and a sample voice is input into the plurality of sub-models, and the plurality of sub-models respectively perform avatar driving processing on the sample voice to respectively output a plurality of first driving parameters.
[0053] For each sample voice, the topological structure information contained in the label of the sample voice can be used to determine a sub-model corresponding to the sample voice from the plurality of sub-models as a target sub-model. For example, the topological structure information contained in the label of the sample voice is information of topological structure A, and then a sub-model corresponding to topological structure A is selected from the plurality of sub-models as a target sub-model. Correspondingly, the first driving parameter output by the target sub-model is the first target driving parameter.
[0054] The loss of each sample voice can be the difference between the topological structure information contained in the label of the sample voice and the first target driving parameter output by the corresponding target sub-model. Since other first driving parameters do not participate in the calculation of the sample voice, the loss can be called a mask loss. The mask loss of the sample can be used to adjust the parameters of the target sub-model.
[0055] In this way, for each sample, the mask loss of each sample voice only affects the parameters of the sub-model corresponding to itself. Since the sample voices are input in batches, the mask loss is calculated for each sample voice, and therefore, for each sub-model, the mask loss of the corresponding sample voice is used to update the sub-model each time the back propagation is performed. Thus, after a plurality of times of back propagation and sub-model parameter adjustment, after the training is completed, each sub-model can have the ability to drive the corresponding avatar (three-dimensional face).
[0056] The embodiments of this disclosure use multiple sub-models to perform virtual image driving processing on the input audio features for different topological structures to obtain multiple first driving parameters. A first target driving parameter corresponding to the sample speech is selected from the multiple first driving parameters. The mask loss of the sample speech is calculated with the topological structure information contained in the label of the sample speech. The parameters of the corresponding sub-model are adjusted using the mask loss, so that each sub-model has the ability to drive the corresponding virtual image.
[0057] Therefore, compared to the speech-driven deep learning models in related technologies, which can only drive a single specific virtual image with the same topology, the deep learning model provided in this disclosure can drive three-dimensional virtual images with multiple topologies. For example, it can be applied to drive three-dimensional faces with multiple topologies to obtain three-dimensional faces of multiple images.
[0058] Furthermore, compared to related technologies that require training corresponding face-driven models using sample speech with different topological structure labels, the embodiments of this disclosure can train virtual images capable of driving multiple topological structures using sample speech with different topological structure labels, thereby improving training efficiency.
[0059] Furthermore, compared to related technologies that only train deep learning models targeting a single topology using training data with specific topological structure labels, which suffers from insufficient model accuracy due to small sample sizes, the deep learning model in this disclosure can obtain features of sample speech with multiple topological structure labels for each sub-model. Therefore, it can improve the generalization ability of each sub-model and the accuracy of the sub-model output, thereby improving the generalization ability and accuracy of the deep learning model.
[0060] Figure 3 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure.
[0061] like Figure 3 As shown, the three-dimensional virtual image in this embodiment is a three-dimensional human face. The deep learning model 320 includes K face-driven sub-models (sub-model 1, sub-model 2, ..., sub-model K, where K is an integer greater than 2), and each sub-model corresponds to a topological structure.
[0062] The sample speech is input into the feature extraction model 310 to obtain the first audio feature. The first audio feature is then input into K sub-models of the deep learning model 320. Sub-model 1 outputs driving parameter 1, sub-model 2 outputs driving parameter 2, ..., sub-model K outputs driving parameter K.
[0063] Since each sample speech has a virtual image label with a specific topological structure, multiple driving parameters are masked when calculating the loss function. This ensures that the label of each sample speech is used to calculate the loss only with the driving parameters of the corresponding topological structure, while other driving parameters do not participate in the loss calculation of that sample speech. Therefore, the loss of that sample speech can be called mask loss.
[0064] For example, speech samples are processed in batches, and the audio features of the batch-processed speech samples are used as input to the deep learning model 320, which can be denoted as X = {x1, x2, ..., x...}. N}, where N is the number of sample speech words, and the output of the deep learning model 320 is in, This is the i-th driving parameter obtained from the j-th input. The ground truth (label) of the speech sample is Y = {y1, y2, ..., y...}. N}, then the mask loss of the deep learning model can be calculated according to the following formula (1), and the mask loss of each sub-model can be obtained according to the following formula (2).
[0065]
[0066]
[0067] Where N is the number of sample speech, j represents the j-th sample speech, j = 1, ..., N, y j The topological structure information in the virtual avatar label of the j-th sample speech.
[0068] K represents the number of sub-models, and i represents the i-th sub-model, i = 1, ..., K. is the first driving parameter output by the i-th sub-model for the speech output of the j-th sample, where, in the case that the i-th sub-model is the target sub-model. These are the driving parameters for the first objective.
[0069] L1(·) represents the mean absolute error function. Let represent the mask loss of the i-th sub-model, and ML (Masked Loss) represents the mask loss of the deep learning model.
[0070] For example, the mask loss (ML) of a deep learning model includes the mask loss of each sub-model. For each sub-model, the corresponding mask loss is updated during each backpropagation. Thus, after training, each sub-model has the ability to drive its corresponding 3D virtual avatar.
[0071] Figure 4 This is a flowchart of a virtual avatar driving method according to an embodiment of the present disclosure.
[0072] As Figure 4 shown, the virtual image driving method 400 can include operation S410 to operation S430.
[0073] In operation S410, a second audio feature of a to-be-processed voice is acquired.
[0074] In operation S420, the second audio feature is input into a deep learning model to obtain a second driving parameter.
[0075] In operation S430, a virtual image is driven according to the second driving parameter.
[0076] For example, the to-be-processed voice is obtained by recording or by text conversion, and the embodiment is used to generate a three-dimensional virtual image (for example, a three-dimensional face) for the to-be-processed voice using a trained deep learning model based on voice driving.
[0077] For example, the to-be-processed voice is input into a feature extraction model to obtain the second audio feature. The feature extraction model can be a deep learning model or a traditional processing module, such as a Fourier processing module. The second audio feature can be a frequency feature, a spectrum feature, etc.
[0078] For example, the deep learning model is trained according to the training method of the deep learning model described above. The second audio feature is input into the deep learning model to obtain the second driving parameter. The second driving parameter can be topology structure information containing a three-dimensional virtual image, and the topology structure information includes the number and position of key points (vertices). The second driving parameter can also be a Blend Shape weight, a PCA coefficient, etc., which can be converted into corresponding topology structure information, respectively.
[0079] For example, after the second driving parameter is converted into topology structure information, a virtual image can be generated through post-processing operations such as smoothing processing and rendering processing.
[0080] The deep learning model of the embodiment can include sub-models corresponding to a plurality of topology structures, respectively. The second audio feature is input into the deep learning model to obtain a second driving parameter corresponding to each of the plurality of topology structures. Each driving parameter is post-processed to obtain a three-dimensional virtual image with a plurality of topology structures.
[0081] In one example, the to-be-processed voice can have index information indicating a topology structure corresponding to the to-be-processed voice, for example, a topology structure of a three-dimensional virtual image that is intended to be obtained for the voice, that is, a three-dimensional virtual image of which style. Therefore, according to the index information, a sub-model in the deep learning model corresponding to the to-be-processed voice can be determined.
[0082] In this example, the second audio feature is input into the deep learning model, and the sub-model corresponding to the to-be-processed voice in the deep learning model can process the second audio feature to obtain a second driving parameter, which can be used to generate a three-dimensional virtual image with a corresponding topology.
[0083] After the second audio feature is input into the deep learning model, the corresponding sub-model is specified based on the index information, and other sub-models can not process the second audio feature. Therefore, while generating a three-dimensional virtual image with a corresponding topology, the processing efficiency of the deep learning model can be ensured, that is, the generation efficiency of the three-dimensional virtual image.
[0084] In another example, the to-be-processed voice has no index information, that is, the topology of the three-dimensional virtual image that the to-be-processed voice wants to generate is not specified. In this example, the second audio feature is input into the deep learning model, and multiple sub-models in the deep learning model can process the second audio feature respectively to obtain multiple second driving parameters corresponding to multiple topologies respectively. Based on the multiple driving parameters, multiple three-dimensional virtual images with corresponding topologies can be generated.
[0085] Therefore, the embodiment can generate multiple three-dimensional virtual images for the to-be-processed voice, so that the style of the three-dimensional virtual image is more rich.
[0086] Figure 5 is a schematic diagram of a virtual image driving method according to one embodiment of the present disclosure.
[0087] As shown in Figure 5 , the three-dimensional virtual image of the embodiment is a three-dimensional face. The deep learning model 520 includes K face driving sub-models (sub-model 1, sub-model 2, …, sub-model K, K is an integer greater than 2), and each sub-model corresponds to a topology.
[0088] The to-be-processed voice is input into the feature extraction model 510 to obtain a second audio feature, and the second audio feature and the index information of the to-be-processed voice are input into the deep learning model 520. If the index information indicates that the Kth sub-model is processed, then the sub-models 1 to (K-1) can not participate in the processing of the second audio feature, and only the Kth sub-model processes the second audio feature to output a driving parameter K.
[0089] The driving parameter K is input into the three-dimensional face generation model 530 for post-processing operations such as smoothing and rendering, and a three-dimensional face can be generated.
[0090] Figure 6 is a block diagram of a training device of a deep learning model according to one embodiment of the present disclosure.
[0091] As shown in Figure 6 The training apparatus 600 of the deep learning model includes a first acquisition module 601, a first processing module 602, a determination module 603, and an adjustment module 604.
[0092] The first acquisition module 601 is configured to acquire a first audio feature of a sample voice, the sample voice having a virtual image label, the virtual image label containing topology information.
[0093] The first processing module 602 is configured to input the first audio feature into the deep learning model to obtain a plurality of first driving parameters respectively corresponding to the plurality of topology structures.
[0094] The determination module 603 is configured to determine a first target driving parameter from the plurality of first driving parameters according to the topology information.
[0095] The adjustment module 604 is configured to adjust the deep learning model according to a difference between the topology information and the first target driving parameter to obtain a trained deep learning model.
[0096] According to an embodiment of the present disclosure, the deep learning model includes a plurality of sub-models respectively corresponding to the plurality of topology structures.
[0097] The first processing module 602 is configured to input the first audio feature into the plurality of sub-models to obtain a first driving parameter output by each of the plurality of sub-models.
[0098] The determination module 603 includes a first determination unit and a second determination unit.
[0099] The first determination unit is configured to determine a target sub-model from the plurality of sub-models according to the topology information.
[0100] The second determination unit is configured to determine the first driving parameter output by the target sub-model as the first target driving parameter.
[0101] The adjustment module 604 includes a calculation unit and an adjustment unit.
[0102] The calculation unit is configured to calculate a mask loss of the target sub-model according to a difference between the topology information and the first target driving parameter.
[0103] The adjustment unit is configured to adjust a parameter of the target sub-model according to the mask loss of the target sub-model to obtain a trained deep learning model.
[0104] The calculation unit is configured to calculate the mask loss of the target sub-model according to the following formula:
[0105]
[0106] wherein j represents the jth sample voice, j = 1, …, N, N is the number of sample voices, yj is the topology information in the virtual image label of the jth sample voice;
[0107] i represents the ith sub-model, i = 1, …, K, K is the number of sub-models, is the first driving parameter output by the ith sub-model for the jth sample voice, wherein, in the case that the ith sub-model is a target sub-model, is the first target driving parameter;
[0108] L1(·) represents a mean absolute error function, represents the mask loss of the ith sub-model, in the case that the ith sub-model is a target sub-model, represents the mask loss of the target sub-model.
[0109] According to an embodiment of the present disclosure, the topology information includes the number and positions of key points constituting the topology in the virtual image label, and the first driving parameter includes the number and positions of key points of the topology corresponding to the first driving parameter.
[0110] Figure 7 is a block diagram of a virtual image driving device according to an embodiment of the present disclosure.
[0111] As shown in Figure 7 , the virtual image driving device 700 can include a second acquisition module 701, a second processing module 702, and a generation module 703.
[0112] The second acquisition module 701 is configured to acquire a second audio feature of a to-be-processed voice.
[0113] The second processing module 702 is configured to input the second audio feature into a deep learning model to obtain a second driving parameter.
[0114] The generation module 703 is configured to drive a virtual image according to the second driving parameter.
[0115] wherein the deep learning model is trained according to the training device of the deep learning model.
[0116] According to an embodiment of the present disclosure, the deep learning model includes a plurality of sub-models corresponding to a plurality of topologies respectively, and the to-be-processed voice includes index information, the index information being used to indicate a sub-model corresponding to the to-be-processed voice.
[0117] The second processing module 702 is configured to input the second audio feature into a sub-model corresponding to the to-be-processed voice to obtain a second driving parameter.
[0118] The second processing module 702 is configured to input the second audio feature into the plurality of sub-models to obtain second driving parameters output by the plurality of sub-models respectively; and the generating module 703 is configured to drive the plurality of virtual images according to the second driving parameters output by the plurality of sub-models respectively.
[0119] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0120] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0121] As shown in Figure 8 The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0122] Various components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; the storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0123] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 801 performs various methods and processes described above, such as the training method of a deep learning model and / or the avatar driving method. For example, in some embodiments, the training method of a deep learning model and / or the avatar driving method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the training method of a deep learning model and / or the avatar driving method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the training method of a deep learning model and / or the avatar driving method by any other appropriate means, such as by means of firmware.
[0124] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0125] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0126] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0128] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0129] The computer system can include clients and servers. This relationship can be
[0130] It should be understood that the procedures shown above can be re-ordered, added to, or removed from, while still falling within the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation, as long as the desired results of the technology disclosed in the present disclosure are achieved.
[0131] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure. Any alternatives, modifications, equivalents, and the like, along with many apparent variations that would be or become apparent to one of ordinary skill in the art are intended to be embraced by the scope of the present disclosure.
Claims
1. A training method for a voice-driven deep learning model of multiple 3D faces, comprising: The first audio feature of the sample speech is obtained. The sample speech has a virtual avatar label. The virtual avatar label is three-dimensional face data that has been collected and labeled in advance for the sample speech and has specific topological structure information. The first audio feature is input into a deep learning model that includes multiple sub-models corresponding to multiple topologies to obtain multiple first driving parameters for driving the multiple three-dimensional faces, wherein the first driving parameters are output by each of the multiple sub-models and correspond to the corresponding topology. Based on the topology information, a first target driving parameter is determined from the plurality of first driving parameters; and Based on the difference between the topology information and the first target driving parameters, the deep learning model is adjusted to obtain a trained deep learning model. Determining the first target driving parameter from the plurality of first driving parameters based on the topology information includes: - Based on the topological information of the virtual character label of the sample speech, determine the sub-model corresponding to the topological information from the plurality of sub-models as the target sub-model; - The first driving parameter output by the target sub-model is determined as the first target driving parameter, and The process of adjusting the deep learning model based on the difference between the topological information and the first target driving parameters to obtain a trained deep learning model includes: - Calculate the mask loss of the target sub-model based on the difference between the topology information and the first target driving parameters; - Adjust the parameters of the target sub-model based on the mask loss of the target sub-model to obtain a trained deep learning model.
2. The method according to claim 1, wherein, The step of calculating the mask loss of the target sub-model based on the difference between the topology information and the first target driving parameters includes: The mask loss of the target sub-model is calculated using the following formula: Where j represents the j-th sample speech, j=1, ..., N, and N is the number of sample speech. The topological structure information in the virtual avatar label of the j-th sample speech; i represents the i-th sub-model, i=1, ...,K, where K is the number of sub-models. is the first driving parameter of the i-th sub-model for the speech output of the j-th sample, where, in the case that the i-th sub-model is the target sub-model. The first target driving parameters; This represents the mean absolute error function. Let represent the mask loss of the i-th sub-model, where the i-th sub-model is the target sub-model. This represents the mask loss of the target sub-model.
3. The method according to claim 1 or 2, wherein, The topology information includes the number and position of key points in the topology that make up the virtual image tag, and the first driving parameter includes the number and position of key points in the topology corresponding to the first driving parameter.
4. A virtual avatar driving method, comprising: Obtain the second audio features of the speech to be processed; The second audio feature is input into the deep learning model to obtain the second driving parameters; as well as Drive the virtual avatar according to the second driving parameter; The deep learning model is trained using the method described in any one of claims 1 to 3.
5. The method according to claim 4, wherein, The deep learning model includes multiple sub-models corresponding to multiple topologies, and the speech to be processed includes index information, which is used to indicate the sub-model corresponding to the speech to be processed. The step of inputting the second audio feature into the deep learning model to obtain the second driving parameters includes: The second audio feature is input into the sub-model corresponding to the speech to be processed to obtain the second driving parameters.
6. The method according to claim 4, wherein, The deep learning model includes multiple sub-models corresponding to multiple topologies; The step of inputting the second audio feature into the deep learning model to obtain the second driving parameters includes: The second audio feature is input into the plurality of sub-models to obtain the second driving parameters output by each of the plurality of sub-models; The step of driving the virtual image according to the second driving parameter includes: Multiple virtual characters are driven based on the second driving parameters output by each of the multiple sub-models.
7. A training apparatus for a voice-driven deep learning model of multiple 3D faces, comprising: The first acquisition module is used to acquire the first audio features of the sample speech, wherein the sample speech has a virtual image tag, and the virtual image tag is three-dimensional face data that has been collected and labeled in advance for the sample speech and has specific topological structure information. The first processing module is used to input the first audio feature into a deep learning model including multiple sub-models corresponding to multiple topological structures, and obtain multiple first driving parameters for driving the multiple three-dimensional faces, wherein the first driving parameters are output by each of the multiple sub-models and correspond to the corresponding topological structure. The determining module is configured to determine a first target driving parameter from the plurality of first driving parameters based on the topology information; and An adjustment module is used to adjust the deep learning model based on the difference between the topology information and the first target driving parameters, so as to obtain a trained deep learning model. The determining module includes: The first determining unit is configured to determine, based on the topological information of the virtual image tag of the sample speech, a sub-model corresponding to the topological information from the plurality of sub-models as a target sub-model; The second determining unit is used to determine the first driving parameter output by the target sub-model as the first target driving parameter, and The adjustment module includes: A calculation unit is used to calculate the mask loss of the target sub-model based on the difference between the topology information and the first target driving parameters; An adjustment unit is used to adjust the parameters of the target sub-model according to the mask loss of the target sub-model, so as to obtain a trained deep learning model.
8. The apparatus according to claim 7, wherein, The computing unit is used to calculate the mask loss of the target sub-model according to the following formula: Where j represents the j-th sample speech, j=1, ..., N, and N is the number of sample speech. The topological structure information in the virtual avatar label of the j-th sample speech; i represents the i-th sub-model, i=1, ...,K, where K is the number of sub-models. is the first driving parameter of the i-th sub-model for the speech output of the j-th sample, where, in the case that the i-th sub-model is the target sub-model. The first target driving parameters; This represents the mean absolute error function. Let represent the mask loss of the i-th sub-model, where the i-th sub-model is the target sub-model. This represents the mask loss of the target sub-model.
9. The apparatus according to claim 7 or 8, wherein, The topology information includes the number and position of key points in the topology that make up the virtual image tag, and the first driving parameter includes the number and position of key points in the topology corresponding to the first driving parameter.
10. A virtual avatar driving device, comprising: The second acquisition module is used to acquire the second audio features of the speech to be processed. The second processing module is used to input the second audio feature into the deep learning model to obtain the second driving parameters; as well as The generation module is used to drive the virtual image according to the second driving parameters; The deep learning model is trained using the apparatus according to any one of claims 7 to 9.
11. The apparatus according to claim 10, wherein, The deep learning model includes multiple sub-models corresponding to multiple topologies, and the speech to be processed includes index information, which is used to indicate the sub-model corresponding to the speech to be processed; the second processing module is used to input the second audio features into the sub-model corresponding to the speech to be processed to obtain the second driving parameters.
12. The apparatus according to claim 10, wherein, The deep learning model includes multiple sub-models corresponding to multiple topologies; The second processing module is used to input the second audio feature into the plurality of sub-models to obtain the second driving parameters output by each of the plurality of sub-models; The generation module is used to drive multiple virtual images according to the second driving parameters output by each of the multiple sub-models.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 6.
15. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Driving method of virtual digital human and training method of pose acquisition model
CN114419205A
Deep learning model training method, text recognition method, device and equipment
CN114998881A