Speech-based 3D Facial Driving Method, Model Training Method and Device
Through the method of generating and style conversion driving sequences, the problem of mismatch between the facial action style of the virtual image and the speaking object is solved, and the three-dimensional facial model's movements are smoother and more in line with the specific style, improving the user experience.
Patent Information
- Application Number
- CN202311766861.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-12-20
Smart Images

Figure CN117788654B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as computer vision, deep learning, augmented reality, virtual reality, etc., and can be applied to scenarios such as the metaverse, digital humans, and generative artificial intelligence; in particular, it relates to a voice-based three-dimensional face driving method, model training method, and device. Background Art
[0002] Currently, with the continuous development of artificial intelligence technology, a large number of virtual avatars have been created for information transmission. Moreover, in order to ensure that users have a good visual experience, during the display of virtual avatars, it is also necessary to adaptively adjust the facial movements of the virtual avatars, especially the lip movements, to meet the requirements of the voice to be played. Summary of the Invention
[0003] The present disclosure provides a voice-based three-dimensional face driving method, model training method, and device, so that the facial movement changes of the three-dimensional face model are more in line with the facial movement style of the speaker when speaking.
[0004] According to a first aspect of the present disclosure, there is provided a voice-based three-dimensional face driving method, including:
[0005] Determine a to-be-processed driving sequence of the to-be-processed voice; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial movement when a first object outputs the to-be-processed voice;
[0006] Perform style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial movement when a second object outputs the to-be-processed voice; the style conversion model is obtained by training an initial model according to at least one group of first driving sequences and second driving sequences; the first driving sequence is used to indicate the three-dimensional facial movement when the first object outputs a target voice; the second driving sequence is used to indicate the three-dimensional facial movement when the second object outputs the target voice; the style conversion model is used to output a driving sequence that conforms to the facial style of the second object when speaking;
[0007] Drive the three-dimensional face model corresponding to the second object to perform facial movements according to the target driving sequence.
[0008] According to a second aspect of the present disclosure, there is provided a training method for a style conversion model, including:
[0009] Obtain at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial movements when a first object outputs a target speech; the second driving sequence is used to indicate the three-dimensional facial movements when a second object outputs the target speech;
[0010] Train an initial model according to the at least one set of training sets to obtain a style conversion model; wherein, the style conversion model is used to output a driving sequence that conforms to a target style; the target style is the facial style when the second object speaks.
[0011] According to a third aspect of the present disclosure, there is provided a three-dimensional facial driving device based on speech, including:
[0012] A determination unit, configured to determine a to-be-processed driving sequence of the to-be-processed speech; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial movements when a first object outputs the to-be-processed speech;
[0013] A processing unit, configured to perform style conversion processing on the to-be-processed driving sequence according to the style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial movements when a second object outputs the to-be-processed speech; the style conversion model is obtained by training an initial model according to at least one set of first driving sequences and second driving sequences; the first driving sequence is used to indicate the three-dimensional facial movements when a first object outputs a target speech; the second driving sequence is used to indicate the three-dimensional facial movements when a second object outputs the target speech; the style conversion model is used to output a driving sequence that conforms to the facial style when the second object speaks;
[0014] A driving unit, configured to drive the three-dimensional facial model corresponding to the second object to perform facial movements according to the target driving sequence.
[0015] According to a fourth aspect of the present disclosure, there is provided a training device for a style conversion model, including:
[0016] An acquisition unit, configured to acquire at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial movements when a first object outputs a target speech; the second driving sequence is used to indicate the three-dimensional facial movements when a second object outputs the target speech;
[0017] A training unit, configured to train an initial model according to the at least one set of training sets to obtain a style conversion model; wherein, the style conversion model is used to output a driving sequence that conforms to a target style; the target style is the facial style when the second object speaks.
[0018] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0019] at least one processor; and
[0020] a memory communicatively connected to the at least one processor; wherein,
[0021] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the first aspect, or the at least one processor is enabled to execute the method described in the second aspect.
[0022] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are for causing the computer to execute the method described in the first aspect, or the computer instructions are for causing the computer to execute the method described in the second aspect.
[0023] According to a seventh aspect of the present disclosure, there is provided a computer program product, the computer program product including: a computer program stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and when the at least one processor executes the computer program, the electronic device is caused to execute the method described in the first aspect, or when the at least one processor executes the computer program, the electronic device is caused to execute the method described in the second aspect.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0026] Figure 1 is a schematic flowchart of a voice-based three-dimensional face driving method provided by an embodiment of the present disclosure;
[0027] Figure 2 is a schematic flowchart of a second voice-based three-dimensional face driving method provided by an embodiment of the present disclosure;
[0028] Figure 3 is a schematic flowchart of a method for training a style conversion model provided by an embodiment of the present disclosure;
[0029] Figure 4 is a schematic flowchart of a second method for training a style conversion model provided by an embodiment of the present disclosure;
[0030] Figure 5 Schematic diagram of a voice-based three-dimensional face driving device provided by an embodiment of the present disclosure;
[0031] Figure 6 Schematic diagram of a second voice-based three-dimensional face driving device provided by an embodiment of the present disclosure;
[0032] Figure 7 Schematic diagram of a training device for a style conversion model provided by an embodiment of the present disclosure;
[0033] Figure 8 Schematic diagram of a training device for a style conversion model provided by an embodiment of the present disclosure;
[0034] Figure 9 Schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0035] Figure 10 It is a block diagram of an electronic device used to implement the voice-based three-dimensional face driving method or model training method of an embodiment of the present disclosure. Detailed implementation manners
[0036] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0037] Currently, when driving the facial movements corresponding to a virtual avatar according to voice, how to control the facial movements of the three-dimensional face model corresponding to the virtual avatar to ensure that the facial movement style of the virtual avatar is more in line with the facial movement style of the user when speaking is an urgently needed problem to be solved.
[0038] In a possible implementation, a driving sequence corresponding to each phoneme when a user speaks can be manually constructed, where the driving sequence corresponding to the phoneme can be used to indicate the facial movements corresponding to the user's pronunciation based on the above phoneme. Then, when it is necessary to drive the facial movements of the virtual character according to the speech, the phonemes included in the speech can be determined first, and based on the pre-constructed correspondence between the phonemes and the driving sequences, a set of driving sequences corresponding to the current speech can be constructed. Then, according to the driving sequences included in the set of driving sequences, the facial movements of the virtual character are controlled in sequence. It should be noted that the above implementation is likely to cause the facial movements of the virtual character corresponding to different phonemes to be discontinuous, resulting in a poor user experience. It should be noted that the facial movements mentioned in this disclosure include the movements of various parts of the face in the virtual character, for example, movements of different parts such as the lips, cheeks, and forehead.
[0039] In another possible implementation, the facial scan data when a real person speaks can be obtained, and based on the facial scan data, the driving sequence corresponding to the facial scan data can be determined. Then, the speech uttered by the user and the obtained driving sequence are used for end-to-end model training, so that the model can output a driving sequence that conforms to the facial style when the above user speaks based on the input speech in the future. However, the above model training method requires a large amount of data sets, resulting in a high data acquisition cost.
[0040] To avoid at least one of the above technical problems, the inventors of the present disclosure have obtained the inventive concept of the present disclosure through creative labor: when it is necessary to drive the three-dimensional facial model corresponding to the second object, first, a to-be-processed driving sequence that conforms to the facial style when the first object speaks can be generated. Then, according to the style conversion model, the above to-be-processed driving sequence is subjected to style conversion processing to obtain a target driving sequence that conforms to the facial movement style when the second object speaks. Then, the three-dimensional facial model is driven according to the target driving sequence, so that the three-dimensional facial model can perform facial movements more smoothly and be closer to the facial movement style when the second object speaks.
[0041] The present disclosure provides a three-dimensional facial driving method, a model training method, and a device based on speech, which are applied to technical fields such as computer vision, deep learning, augmented reality, and virtual reality in artificial intelligence technology, and can be applied to scenarios such as the metaverse, digital humans, and generative artificial intelligence, so that the movements of the three-dimensional facial model are more smooth and can make the movements of the three-dimensional facial model conform to a specific style.
[0042] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0043] Figure 1 The flowchart of a voice-based three-dimensional facial driving method provided by an embodiment of the present disclosure is shown as Figure 1 follows. The method includes:
[0044] S101. Determine a to-be-processed driving sequence of the to-be-processed voice; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial actions when a first object outputs the to-be-processed voice.
[0045] Exemplarily, the execution subject of this embodiment may be a voice-based three-dimensional facial driving device. The voice-based three-dimensional facial driving device may be a server (such as a local server or a cloud server), or a computer, or a processor, or a chip, etc., which is not limited in this embodiment.
[0046] The first object and the second object in this embodiment may be real people or cartoon animation characters, which are not specifically limited in the present disclosure.
[0047] When it is necessary to perform drive control on the three-dimensional facial model corresponding to the second object according to the to-be-processed voice, first, a to-be-processed driving sequence may be generated according to the to-be-processed voice.
[0048] It should be noted that the to-be-processed driving sequence generated here is a driving sequence that conforms to the facial action style when the first object speaks, and the to-be-processed driving sequence is specifically used to indicate the facial actions corresponding to the first object when outputting the above to-be-processed voice.
[0049] In one example, the above to-be-processed driving sequence may be generated by a model for generating the driving sequence corresponding to the first speaking object. By inputting the above to-be-processed voice into the above model, the above to-be-driven sequence can be obtained. Specifically, the model in this example may be the above end-to-end sequence generation model from voice to driving sequence.
[0050] S102. Perform style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial actions when the second object outputs the to-be-processed voice; the style conversion model is obtained by training an initial model according to at least one group of first driving sequences and second driving sequences; the first driving sequence is used to indicate the three-dimensional facial actions when the first object outputs a target voice; the second driving sequence is used to indicate the three-dimensional facial actions when the second object outputs a target voice; the style conversion model is used to output a driving sequence that conforms to the facial style when the second object speaks.
[0051] Exemplarily, in this embodiment, after obtaining the to-be-processed driving sequence that conforms to the style of the first object, the pre-trained style conversion model can be used to perform style conversion processing on the to-be-processed driving sequence, so as to obtain the target driving sequence that conforms to the facial action style when the second object speaks, and the target driving sequence can specifically be used to indicate the facial actions corresponding to the second object when emitting the to-be-processed voice.
[0052] It should be noted that the style conversion model in this embodiment is used to convert the style of the driving sequence. Among them, the driving sequence can be understood as a series of parameters used to indicate the facial actions of the model. Style conversion can be understood as converting from the facial action style when one object speaks to the facial action style when another object speaks.
[0053] Moreover, when training the above-mentioned style conversion model, at least one training set can be obtained in advance. Among them, the training set includes the first driving sequence corresponding to the first object and the second driving sequence corresponding to the second object for the same voice (i.e., the above-mentioned target voice). Specifically, the first driving sequence is a driving sequence that conforms to the facial style when the first object speaks, and the first driving sequence is specifically used to indicate the facial actions of the first object when outputting the target voice. Similarly, the second driving sequence is a driving sequence that conforms to the facial style when the second object speaks, and the second driving sequence is specifically used to indicate the facial actions of the second object when outputting the target voice. Among them, different training sets can correspond to different target voices. Then, in combination with the above training set, the initial model is trained to obtain the style conversion model that can convert the driving sequence that conforms to the speaking facial style of the first object into the driving sequence that conforms to the speaking facial style of the second object.
[0054] It should be noted that the model architecture of the style conversion model in this embodiment is not specifically limited.
[0055] S103. Drive the three-dimensional facial model corresponding to the second object to perform facial actions according to the target driving sequence.
[0056] Exemplarily, in this embodiment, after obtaining the target driving sequence that conforms to the facial style when the second object speaks, the three-dimensional facial model corresponding to the second object can be controlled for facial actions according to the target driving sequence, so that the facial actions of the three-dimensional facial model can conform to the facial action style when the second object speaks the to-be-processed voice.
[0057] It can be understood that in this embodiment, when driving the three-dimensional facial model of the second object according to the voice, a to-be-processed driving sequence conforming to the facial style of the first object can be generated first, and then the to-be-processed driving sequence can be subjected to style conversion processing in combination with the style conversion model to obtain a target driving sequence conforming to the facial action style of the second object. Compared with the method of generating a target driving sequence based on an end-to-end model from voice to a driving sequence conforming to the facial style of the second object, the style conversion model provided in this embodiment for performing style conversion of the driving sequence has a lower training difficulty and requires a smaller training set, which can improve the efficiency of three-dimensional facial driving.
[0058] Figure 2 FIG. is a schematic flowchart of a second voice-based three-dimensional facial driving method provided by an embodiment of the present disclosure, as Figure 2 shown, the method includes:
[0059] S201. Determine the phoneme information corresponding to the to-be-processed voice; wherein, the phoneme information is a set of phonemes that make up the to-be-processed voice.
[0060] Exemplarily, in this embodiment, when it is necessary to control the three-dimensional facial model corresponding to the second object to accurately simulate the facial actions of the second object when outputting the to-be-processed voice, in this embodiment, the to-be-processed voice is first analyzed to obtain the phoneme information corresponding to the to-be-processed voice. Among them, the phoneme information can be understood as a set composed of the phonemes arranged in sequence obtained after splitting the to-be-processed voice into phonemes.
[0061] S202. Determine the to-be-processed driving sequence according to the mapping relationship and the phoneme information; wherein, the mapping relationship is the corresponding relationship between the phoneme and the preset parameter set; the preset parameter set is used to indicate the three-dimensional facial actions of the first object when uttering the phoneme; the to-be-processed driving sequence includes at least one preset parameter set. The to-be-processed driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting the to-be-processed voice.
[0062] Exemplarily, in this embodiment, after obtaining the phoneme information, the preset parameter sets corresponding to the phonemes included in the phoneme information can be determined according to the pre-set corresponding relationship between the phoneme and the preset parameter set, and the sequence composed of the preset parameter sets corresponding to each phoneme is used as the to-be-processed sequence in this embodiment.
[0063] It should be noted that the above preset parameter set can be understood as the facial parameters corresponding to the facial actions of the first object when the first object utters the phoneme corresponding to the set.
[0064] It can be understood that in this embodiment, the sequence to be driven can be generated by splicing the preset parameter sets corresponding to the phonemes in the speech to be processed. Furthermore, it is possible to avoid the problem of long training time caused by training the above end-to-end model for generating the driving sequence from speech to the speaking facial style of the first object.
[0065] S203. Perform style conversion processing on the sequence to be processed according to the style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to instruct the three-dimensional facial actions of the second object when outputting the speech to be processed; the style conversion model is obtained by training an initial model according to at least one set of first driving sequences and second driving sequences; the first driving sequence is used to instruct the three-dimensional facial actions of the first object when outputting the target speech; the second driving sequence is used to instruct the three-dimensional facial actions of the second object when outputting the target speech; the style conversion model is used to output a driving sequence that conforms to the facial style of the second object when speaking.
[0066] Exemplarily, in this embodiment, the specific principle of step S203 can refer to step S102, which will not be elaborated here. And when the sequence to be processed is obtained in the above S201 - S202 manner, the first driving sequence used by the further style conversion model during training can also be obtained by referring to the above S201 - S202 manner.
[0067] S204. Drive the three-dimensional facial model corresponding to the second object to perform facial actions according to the target driving sequence.
[0068] Exemplarily, the specific principle of step S204 can refer to step S103, which will not be elaborated here.
[0069] It can be understood that in this embodiment, through the style conversion model, performing style conversion processing on the sequence to be processed obtained by splicing the corresponding phonemes, and controlling the facial actions of the three-dimensional facial model corresponding to the second object based on the converted target driving sequence, it is possible to improve the smoothness of the facial actions of the three-dimensional facial model, and make the facial actions of the three-dimensional facial model conform to the facial style of the second object when speaking, so as to improve the user viewing experience.
[0070] In one example, the first driving sequence includes N first parameter sets; the first parameter set is used to instruct the three-dimensional facial actions of the first object at the time frame corresponding to the first parameter set; N is a positive integer greater than 1; the style conversion model is obtained by adjusting the parameters of the initial model according to the third driving sequence and the second driving sequence; wherein, the third driving sequence is the output of the initial model according to N first parameter sets; the third driving sequence includes N third parameter sets. It should be noted that the specific principle here can refer to Figure 4The description in the embodiments will not be repeated here.
[0071] In one example, the style conversion model is obtained by adjusting the parameters of the initial model according to the first facial model and the second facial model; wherein, the first facial model is obtained by adjusting the parameters of the preset facial model according to the third driving sequence; the second facial model is obtained by adjusting the parameters of the preset facial model according to the second driving sequence. It should be noted that the specific principle here can be referred to Figure 4 The description in the embodiments will not be repeated here.
[0072] In one example, the style conversion model is obtained by adjusting the parameters of the initial model according to the first loss function and the second loss function; the first loss function is obtained according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object at the time frame corresponding to the second parameter set; the second loss function is obtained according to the first facial model and the second facial model; wherein, the first facial model is obtained by adjusting the parameters of the preset facial model according to the third driving sequence; the second facial model is obtained by adjusting the parameters of the preset facial model according to the second driving sequence. It should be noted that the specific principle here can be referred to Figure 4 The description in the embodiments will not be repeated here.
[0073] Figure 3 It is a schematic flow chart of a method for training a style conversion model provided by an embodiment of the present disclosure, as Figure 3 shown, the method includes:
[0074] S301. Obtain at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting the target speech; the second driving sequence is used to indicate the three-dimensional facial actions of the second object when outputting the target speech.
[0075] Exemplarily, the execution subject of this embodiment can be a training device for the style conversion model. The training device for the style conversion model can be a server (such as a local server or a cloud server), or a computer, or a processor, or a chip, etc., which is not limited in this embodiment. And, the above-mentioned training device for the style conversion model can be the same device as the above-mentioned three-dimensional facial driving device based on voice, or different devices.
[0076] Among them, the technical principle of step S301 can be referred to step S102, and will not be repeated here.
[0077] S302. Train an initial model based on at least one set of training sets to obtain a style conversion model, where the style conversion model is used to output a driving sequence that conforms to the target style, and the target style is the facial style when the second object speaks.
[0078] In one example, the first driving sequence may include multiple sets of driving parameters, and each set of driving parameters has a corresponding time frame. The set of driving parameters in the first driving sequence can be understood as the set of parameters used to describe the facial movements of the first object at the time frame corresponding to the set during the process of the first object outputting the target speech. Moreover, the second driving sequence may also include multiple sets of driving parameters. Specifically, the set of driving parameters in the second driving sequence can be understood as the set of parameters used to describe the facial movements of the second object at the time frame corresponding to the set during the process of the first object outputting the target speech.
[0079] When training the initial model based on the training set, a set of driving parameters corresponding to a target time frame can be selected from the first driving sequence, and a set of driving parameters corresponding to the above target time frame can be selected from the second driving sequence for model training. That is, the input to the initial model is the set of parameters in the first driving sequence corresponding to a time frame.
[0080] It can be understood that in this embodiment, a style conversion model is trained by using the training set so that subsequently, the driving sequence can be style-converted based on the style conversion model to quickly obtain a driving sequence that conforms to the style of the speaking object to control the three-dimensional facial model corresponding to the speaking object.
[0081] Figure 4 FIG. is a schematic flowchart of the second method for training a style conversion model provided by an embodiment of the present disclosure. As Figure 4 shown, the method includes:
[0082] S401. Obtain at least one set of training sets, where the training set includes a first driving sequence and a second driving sequence. The first driving sequence is used to indicate the three-dimensional facial movements of the first object when outputting the target speech, and the second driving sequence is used to indicate the three-dimensional facial movements of the second object when outputting the target speech.
[0083] Exemplarily, the execution subject of this embodiment may be a training device for a style conversion model. The training device for the style conversion model may be a server (such as a local server or a cloud server), or a computer, or a processor, or a chip, etc., which is not limited in this embodiment. Moreover, the above training device for the style conversion model may be the same device as the above three-dimensional facial driving device based on speech, or may be different devices.
[0084] Among them, the first driving sequence includes N first parameter sets; the first parameter set is used to indicate the three-dimensional facial actions of the first object in the time frame corresponding to the first parameter set; N is a positive integer greater than 1.
[0085] That is, in this embodiment, the first driving sequence is a sequence composed of N first parameter sets, and the first parameter set corresponds to the time frame in the target speech. The first parameter set can be specifically understood as the three-dimensional facial actions corresponding to the first object when outputting the speech in the time frame corresponding to it in the target speech.
[0086] In one example, step S401 includes the following steps: determining the phoneme information corresponding to the target speech; where the phoneme information is a set of phonemes that make up the target speech; determining the first driving sequence according to the mapping relationship and the phoneme information; where the mapping relationship is the corresponding relationship between phonemes and a preset parameter set; the preset parameter set is used to characterize the three-dimensional facial actions when pronouncing phonemes; the first driving sequence includes at least one preset parameter set; obtaining a second driving sequence, and using the second driving sequence and the first driving sequence as a set of training sets.
[0087] Exemplarily, the specific principle of the acquisition method of the first driving sequence in this embodiment is similar to the principles of steps S201 - S202, and will not be elaborated here.
[0088] In addition, the acquisition method of the second driving sequence can be generated by using the end-to-end model of speech-to-driving sequence mentioned in related technologies, which will not be elaborated here; or, it can also be generated by collecting the facial images when the second object occurs.
[0089] It can be understood that in this embodiment, by using the first driving sequence spliced from the preset parameter sets corresponding to phonemes as the sequence conforming to the facial style of the first object's speech to train the style conversion model, not only can the smoothness of the actions of the final three-dimensional facial model be improved, but also the acquisition method of the above training data is relatively simple, avoiding the complexity of model training that requires combining models to generate corresponding driving sequences.
[0090] S402: Input the N first parameter sets into the initial model to obtain a third driving sequence; where the third driving sequence includes N third parameter sets.
[0091] Exemplarily, in this embodiment, when training the initial model, the N first parameter sets corresponding to N time frames are simultaneously input into the initial model, and the initial model performs style conversion processing on the N first parameter sets, so that the initial model can combine the relevance between the facial actions indicated by the first parameter sets of multiple time frames to convert and obtain a driving sequence composed of N third parameter sets.
[0092] S403. Adjust the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model, where the style conversion model is used to output a driving sequence that conforms to the target style, and the target style is the facial style when the second object speaks.
[0093] Exemplarily, in this embodiment, after obtaining the third driving sequence output by the initial model, the second driving sequence in the training set can be combined with the third driving sequence to adjust the parameters of the initial model so that the style conversion model obtained after training meets the preset training stop condition. It should be noted that the stop condition for model training in this embodiment is similar to the setting method in the related art and will not be elaborated here.
[0094] In one example, when adjusting the parameters of the initial model based on the third driving sequence and the second driving sequence, the loss function obtained by combining the third driving sequence and the second driving sequence can be used to adjust the parameters of the initial model. Among them, the process of generating the loss function according to the third driving sequence and the second driving sequence can be obtained by combining various loss function generation methods provided in the related art, such as the L1 norm loss function type, mean square error loss function, cross-entropy loss function, etc., and no specific limitation is made in this embodiment.
[0095] It can be understood that since the changes in the user's facial movements are also temporally continuous when the user speaks, that is, the facial movements corresponding to the subsequent moment are usually affected by the facial movements of the previous moment. Therefore, when training the initial model, the first parameter sets under multiple consecutive time frames are simultaneously input into the initial model, so that when the initial model performs style conversion on the input data, it can also fully combine the correlation between the facial movements under adjacent time frames to perform style conversion on the first parameter set, so as to improve the accuracy of the style conversion result. And it can avoid the problem that when two different objects output the same segment of speech, there is a time difference in the opening and closing of the mouth, resulting in large noise when training frame by frame, making the final facial movements not smooth.
[0096] In one example, step S403 includes the following steps:
[0097] Adjust the parameters of the preset facial model according to the third driving sequence to obtain a first facial model; adjust the parameters of the preset facial model according to the second driving sequence to obtain a second facial model; adjust the parameters of the initial model according to the first facial model and the second facial model to obtain a style conversion model.
[0098] Exemplarily, in this embodiment, when adjusting the initial model parameters according to the third driving sequence and the second driving sequence, the parameters of the same preset facial model can be adjusted based on the third driving sequence and the second driving sequence. That is, the facial actions in the preset facial model are adjusted according to the third driving sequence to obtain the first facial model. And the facial actions in the preset facial model are adjusted according to the fourth driving sequence to obtain the second facial model.
[0099] Among them, the preset facial model can be understood as a three-dimensional facial model constructed according to the face corresponding to the speaking object. Specifically, the preset facial model here can be the three-dimensional facial model corresponding to the second object, or the three-dimensional facial model corresponding to the rest of the speaking objects, which is not specifically limited in this embodiment.
[0100] After obtaining the first facial model and the second facial model, the parameters of the initial model can be adjusted according to the difference between the first facial model and the second facial model. For example, the position information corresponding to at least one key point in the first facial model and the position information corresponding to at least one key point in the second facial model can be extracted to construct a loss function, and then, based on the obtained loss function, the parameters of the initial model are adjusted.
[0101] It can be understood that in this embodiment, the third driving sequence obtained by the model is used to drive the preset facial model, and further facial action verification is performed on the preset facial model to further ensure the accuracy and rationality of the third driving sequence output by the model, avoid the phenomenon of overfitting of the model, and improve the model training efficiency.
[0102] In one example, step S403 includes the following steps:
[0103] Determine the first loss function according to the third driving sequence and the second driving sequence; where the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object in the time frame corresponding to the second parameter set; adjust the parameters of the preset facial model according to the third driving sequence to obtain the first facial model; adjust the parameters of the preset facial model according to the second driving sequence to obtain the second facial model; and determine the second loss function according to the first facial model and the second facial model; adjust the parameters of the initial model according to the first loss function and the second loss function to obtain the style conversion model.
[0104] Exemplarily, in this embodiment, on the basis of the above example, the second driving sequence also includes N second parameter sets corresponding to the time frames one by one. Here, the second parameter set can be understood as the facial parameters of the second object in one time frame during the process of emitting the target voice.
[0105] When the third driving sequence and the second driving sequence are obtained, not only can a loss function be constructed based on the third driving sequence and the second driving sequence to obtain the above first loss function. In addition, a second loss function can also be constructed according to the first facial model and the second facial model obtained by driving the preset facial model actions by the third driving sequence and the second driving sequence respectively. Then, the initial model can be adjusted in parameters by combining the first loss function and the second loss function to train and obtain a style conversion model.
[0106] It can be understood that in this embodiment, the initial model can be trained by combining the first loss function constructed from the differences between the driving sequences and the second loss function constructed from the differences between the facial models, which can improve the efficiency of model training, avoid the phenomenon of overfitting of the model and the unreasonable facial actions during facial driving, so as to improve the accuracy of style conversion of the style conversion model.
[0107] In one example, the style conversion model includes M one-dimensional convolutional layers, where the one-dimensional convolutional layer is used to perform convolutional processing on the input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
[0108] Exemplarily, the style conversion model in this embodiment specifically includes M one-dimensional convolutional layers. Moreover, when each one-dimensional convolutional layer performs a convolutional operation on the input data input to this convolutional layer, the obtained convolutional result is the same as the size of the input data corresponding to this one-dimensional convolutional layer. That is, the convolutional operation of the shift convolutional layer in this embodiment does not change the size of the input data. Furthermore, through the above model construction method, the size consistency of the input and output results of the style conversion model is ensured.
[0109] For example, in practical applications, a fully convolutional network can be used as the model architecture of the above style conversion model, and further, the stride of each convolutional layer in the fully convolutional network can be set to 1 to ensure the consistency of the input and output sizes of each convolutional layer.
[0110] It can be understood that in this embodiment, by setting multiple one-dimensional convolutional layers to perform style conversion processing on the input driving sequence, the model structure is simple, and the consistency of the input and output data sizes can be ensured to ensure the smoothness of facial actions during final facial driving.
[0111] In one example, the style conversion model can also perform multiple downsampling processes and multiple upsampling processes on the input driving sequence by multiple successively connected convolutional layers in sequence, and then obtain a model output result with the same size as the input driving sequence.
[0112] Figure 5The structural schematic diagram of a voice-based three-dimensional facial driving device provided by an embodiment of the present disclosure is shown as Figure 5 follows. The voice-based three-dimensional facial driving device 500 includes:
[0113] A determination unit 501, configured to determine a to-be-processed driving sequence of the to-be-processed voice; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial actions when a first object outputs the to-be-processed voice.
[0114] A processing unit 502, configured to perform style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial actions when a second object outputs the to-be-processed voice; the style conversion model is obtained by training an initial model according to at least one set of first driving sequences and second driving sequences; the first driving sequence is used to indicate the three-dimensional facial actions when the first object outputs a target voice; the second driving sequence is used to indicate the three-dimensional facial actions when the second object outputs a target voice; the style conversion model is used to output a driving sequence that conforms to the facial style when the second object speaks.
[0115] A driving unit 503, configured to drive the three-dimensional facial model corresponding to the second object to perform facial actions according to the target driving sequence.
[0116] The device provided in this embodiment is used to implement the technical solution provided by the above method, and its implementation principle and technical effects are similar, so details are not described herein again.
[0117] Figure 6 The structural schematic diagram of a second voice-based three-dimensional facial driving device provided by an embodiment of the present disclosure is shown as Figure 6 follows. The voice-based three-dimensional facial driving device 600 includes:
[0118] A determination unit 601, configured to determine a to-be-processed driving sequence of the to-be-processed voice; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial actions when a first object outputs the to-be-processed voice.
[0119] A processing unit 602, configured to perform style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial actions when a second object outputs the to-be-processed voice; the style conversion model is obtained by training an initial model according to at least one set of first driving sequences and second driving sequences; the first driving sequence is used to indicate the three-dimensional facial actions when the first object outputs a target voice; the second driving sequence is used to indicate the three-dimensional facial actions when the second object outputs a target voice; the style conversion model is used to output a driving sequence that conforms to the facial style when the second object speaks.
[0120] A driving unit 603, configured to drive a three-dimensional facial model corresponding to a second object to perform facial actions according to a target driving sequence.
[0121] In one example, the determination unit 601 includes:
[0122] A first determination module 6011, configured to determine phoneme information corresponding to the speech to be processed; wherein, the phoneme information is a set of phonemes that make up the speech to be processed;
[0123] A second determination module 6012, configured to determine a driving sequence to be processed according to a mapping relationship and the phoneme information; wherein, the mapping relationship is a correspondence between phonemes and a set of preset parameters; the set of preset parameters is used to indicate three-dimensional facial actions when a first object emits phonemes; the driving sequence to be processed includes at least one set of preset parameters.
[0124] In one example, the first driving sequence includes N sets of first parameters; the set of first parameters is used to indicate three-dimensional facial actions of the first object in a time frame corresponding to the set of first parameters; N is a positive integer greater than 1; the style conversion model is obtained by adjusting parameters of an initial model according to a third driving sequence and a second driving sequence; wherein, the third driving sequence is output by the initial model according to N sets of first parameters; the third driving sequence includes N sets of third parameters.
[0125] In one example, the style conversion model is obtained by adjusting parameters of an initial model according to a first facial model and a second facial model; wherein, the first facial model is obtained by adjusting parameters of a preset facial model according to a third driving sequence; the second facial model is obtained by adjusting parameters of a preset facial model according to a second driving sequence.
[0126] In one example, the style conversion model is obtained by adjusting parameters of an initial model according to a first loss function and a second loss function; the first loss function is obtained according to a third driving sequence and a second driving sequence; wherein, the second driving sequence includes N sets of second parameters; the set of second parameters is used to indicate three-dimensional facial actions of a second object in a time frame corresponding to the set of second parameters; the second loss function is obtained according to a first facial model and a second facial model; wherein, the first facial model is obtained by adjusting parameters of a preset facial model according to a third driving sequence; the second facial model is obtained by adjusting parameters of a preset facial model according to a second driving sequence.
[0127] In one example, the style conversion model includes M one-dimensional convolutional layers, wherein, the one-dimensional convolutional layer is configured to perform convolutional processing on input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
[0128] The device provided in this embodiment is used to implement the technical solution provided by the above method, and its implementation principle and technical effect are similar, so details are not described herein again.
[0129] Figure 7 The following is a schematic structural diagram of a training device for a style conversion model provided by an embodiment of the present disclosure. As Figure 7 shown, the training device 700 for the style conversion model includes:
[0130] An obtaining unit 701, configured to obtain at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting a target voice; the second driving sequence is used to indicate the three-dimensional facial actions of the second object when outputting the target voice.
[0131] A training unit 702, configured to train an initial model according to at least one set of training sets to obtain a style conversion model; wherein, the style conversion model is used to output a driving sequence that conforms to a target style; the target style is the facial style when the second object is speaking.
[0132] The device provided in this embodiment is used to implement the technical solution provided by the above method, and its implementation principle and technical effect are similar, so details are not described herein again.
[0133] Figure 8 The following is a schematic structural diagram of a training device for a style conversion model provided by an embodiment of the present disclosure. As Figure 8 shown, the training device 800 for the style conversion model includes:
[0134] An obtaining unit 801, configured to obtain at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting a target voice; the second driving sequence is used to indicate the three-dimensional facial actions of the second object when outputting the target voice.
[0135] A training unit 802, configured to train an initial model according to at least one set of training sets to obtain a style conversion model; wherein, the style conversion model is used to output a driving sequence that conforms to a target style; the target style is the facial style when the second object is speaking.
[0136] In one example, the first driving sequence includes N first parameter sets; the first parameter set is used to indicate the three-dimensional facial actions of the first object at the time frame corresponding to the first parameter set; N is a positive integer greater than 1;
[0137] The training unit 802 includes:
[0138] The first acquisition module 8021 is configured to input N first parameter sets into the initial model to obtain a third driving sequence; wherein, the third driving sequence includes N third parameter sets.
[0139] The adjustment module 8022 is configured to adjust the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model.
[0140] In one example, the adjustment module 8022 includes:
[0141] The first adjustment sub-module 80221 is configured to adjust the parameters of the preset facial model according to the third driving sequence to obtain a first facial model.
[0142] The second adjustment sub-module 80222 is configured to adjust the parameters of the preset facial model according to the second driving sequence to obtain a second facial model.
[0143] The third adjustment sub-module 80223 is configured to adjust the parameters of the initial model according to the first facial model and the second facial model to obtain a style conversion model.
[0144] In one example, the adjustment module 8022 includes:
[0145] The first determination sub-module is configured to determine a first loss function according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object at the time frame corresponding to the second parameter set.
[0146] The fourth adjustment sub-module is configured to adjust the parameters of the preset facial model according to the third driving sequence to obtain a first facial model.
[0147] The fifth adjustment sub-module is configured to adjust the parameters of the preset facial model according to the second driving sequence to obtain a second facial model.
[0148] The second determination sub-module is configured to determine a second loss function according to the first facial model and the second facial model.
[0149] The sixth adjustment sub-module is configured to adjust the parameters of the initial model according to the first loss function and the second loss function to obtain a style conversion model.
[0150] In one example, the acquisition unit 801 includes:
[0151] The third determination module 8011 is configured to determine the phoneme information corresponding to the target speech; wherein, the phoneme information is a set of phonemes that make up the target speech.
[0152] The fourth determination module 8012 is configured to determine a first driving sequence according to the mapping relationship and the phoneme information; wherein, the mapping relationship is the correspondence between phonemes and a set of preset parameters; the set of preset parameters is used to represent three-dimensional facial movements when phonemes are uttered; the first driving sequence includes at least one set of preset parameters;
[0153] The second acquisition module 8013 is configured to acquire a second driving sequence;
[0154] The fifth determination module 8014 is configured to use the second driving sequence and the first driving sequence as a set of training sets.
[0155] The device provided in this embodiment is used to implement the technical solution provided by the above method, and its implementation principle and technical effect are similar, and will not be elaborated here.
[0156] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0157] The present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by any one of the above embodiments.
[0158] Figure 9 As a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure, as Figure 9 shown, the electronic device 900 in the present disclosure may include: a processor 901 and a memory 902.
[0159] A memory 902 for storing programs; the memory 902 may include volatile memory (e.g., random-access memory, such as static random-access memory (SRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.); the memory may also include non-volatile memory, such as flash memory. The memory 902 is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories 902. And the above computer programs, computer instructions, data, etc. can be called by the processor 901.
[0160] The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories 902. And the above computer programs, computer instructions, data, etc. can be called by the processor 901.
[0161] A processor 901 for executing the computer programs stored in the memory 902 to implement each step in the method involved in the above embodiments.
[0162] For details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0163] The processor 901 and the memory 902 may be independent structures or integrated structures integrated together. When the processor 901 and the memory 902 are independent structures, the memory 902 and the processor 901 can be coupled and connected through a bus 903.
[0164] The electronic device of this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0165] The present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided in any one of the above embodiments.
[0166] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which includes: a computer program stored in a readable storage medium, and at least one processor of the electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the solution provided in any of the above embodiments.
[0167] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0168] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0169] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0170] The computing unit 1001 can be various general-purpose and / or special-purpose processing groups with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the voice-based 3D face driving method and the model training method. For example, in some embodiments, the voice-based 3D face driving method and the model training method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the voice-based 3D face driving method and the model training method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the voice-based 3D face driving method and the model training method by any other suitable means (e.g., by means of firmware).
[0171] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0172] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0173] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0174] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0175] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0176] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0177] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0178] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A three-dimensional face driving method based on voice, including: determining a to-be-processed driving sequence of the to-be-processed voice; wherein, the to-be-processed driving sequence is used to indicate the three-dimensional facial actions of a first object when outputting the to-be-processed voice; performing style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to indicate the three-dimensional facial actions of a second object when outputting the to-be-processed voice; the style conversion model is obtained by adjusting the parameters of an initial model according to a third driving sequence and a second driving sequence, the third driving sequence is obtained by inputting N first parameter sets included in a first driving sequence into the initial model; N is a positive integer greater than 1; the first driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting a target voice; the second driving sequence is used to indicate the three-dimensional facial actions of the second object when outputting the target voice; the style conversion model is used to output a driving sequence conforming to the facial style when the second object speaks; driving a three-dimensional facial model corresponding to the second object to perform facial actions according to the target driving sequence.
2. The method according to claim 1, wherein, determining the to-be-processed driving sequence of the to-be-processed voice includes: determining the phoneme information corresponding to the to-be-processed voice; wherein, the phoneme information is a set of phonemes constituting the to-be-processed voice; determining the to-be-processed driving sequence according to a mapping relationship and the phoneme information; wherein, the mapping relationship is a correspondence between phonemes and a preset parameter set; the preset parameter set is used to indicate the three-dimensional facial actions of the first object when uttering a phoneme; the to-be-processed driving sequence includes at least one preset parameter set.
3. The method according to claim 1 or 2, wherein, the first parameter set is used to indicate the three-dimensional facial actions of the first object in a time frame corresponding to the first parameter set; the third driving sequence includes N third parameter sets.
4. The method according to claim 3, wherein, the style conversion model is obtained by adjusting the parameters of the initial model according to a first facial model and a second facial model; wherein, the first facial model is obtained by adjusting the parameters of a preset facial model according to the third driving sequence; the second facial model is obtained by adjusting the parameters of a preset facial model according to the second driving sequence.
5. The method according to claim 3, wherein, the style conversion model is obtained by adjusting the parameters of the initial model according to a first loss function and a second loss function; the first loss function is obtained according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object in a time frame corresponding to the second parameter set; The second loss function is obtained based on the first facial model and the second facial model; wherein, the first facial model is obtained by adjusting the parameters of a preset facial model according to the third driving sequence; and the second facial model is obtained by adjusting the parameters of the preset facial model according to the second driving sequence.
6. The method according to any one of claims 1-2, 4-5, wherein, The style conversion model includes M one-dimensional convolutional layers, wherein the one-dimensional convolutional layer is used to perform convolutional processing on the input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
7. A method for training a style conversion model, including: Obtaining at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial actions of a first object when outputting a target voice; and the second driving sequence is used to indicate the three-dimensional facial actions of a second object when outputting the target voice; Inputting the N first parameter sets included in the first driving sequence into an initial model to obtain a third driving sequence; and adjusting the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model; wherein, N is a positive integer greater than 1; the style conversion model is used to output a driving sequence conforming to a target style; and the target style is the facial style when the second object speaks.
8. The method according to claim 7, wherein, The first parameter set is used to indicate the three-dimensional facial actions of the first object at the time frame corresponding to the first parameter set; and the third driving sequence includes N third parameter sets.
9. The method according to claim 8, wherein, Adjusting the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model includes: Adjusting the parameters of a preset facial model according to the third driving sequence to obtain a first facial model; Adjusting the parameters of the preset facial model according to the second driving sequence to obtain a second facial model; Adjusting the parameters of the initial model according to the first facial model and the second facial model to obtain a style conversion model.
10. The method according to claim 8, wherein, Adjusting the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model includes: Determining a first loss function according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; and the second parameter set is used to indicate the three-dimensional facial actions of the second object at the time frame corresponding to the second parameter set; Adjusting the parameters of a preset facial model according to the third driving sequence to obtain a first facial model; adjusting the parameters of the preset facial model according to the second driving sequence to obtain a second facial model; and determining a second loss function according to the first facial model and the second facial model. Adjust the parameters of the initial model according to the first loss function and the second loss function to obtain a style conversion model.
11. The method according to any one of claims 7-10, wherein, Obtain at least one set of training sets, including: Determine the phoneme information corresponding to the target speech; wherein, the phoneme information is a set of phonemes that make up the target speech; Determine a first driving sequence according to the mapping relationship and the phoneme information; wherein, the mapping relationship is the correspondence between phonemes and a set of preset parameters; the set of preset parameters is used to characterize the three-dimensional facial movements when pronouncing phonemes; the first driving sequence includes at least one set of preset parameters; Obtain a second driving sequence, and use the second driving sequence and the first driving sequence as a set of training sets.
12. The method according to any one of claims 7-10, wherein, The style conversion model includes M one-dimensional convolutional layers, wherein the one-dimensional convolutional layer is used to perform convolutional processing on the input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
13. A three-dimensional facial driving device based on speech, including: A determination unit for determining a to-be-processed driving sequence of the to-be-processed speech; wherein, the to-be-processed driving sequence is used to instruct the three-dimensional facial movements of the first object when outputting the to-be-processed speech; A processing unit for performing style conversion processing on the to-be-processed driving sequence according to a style conversion model to obtain a target driving sequence; wherein, the target driving sequence is used to instruct the three-dimensional facial movements of the second object when outputting the to-be-processed speech; the style conversion model is obtained by adjusting the parameters of the initial model according to a third driving sequence and a second driving sequence, the third driving sequence is obtained by inputting N first parameter sets included in the first driving sequence into the initial model; N is a positive integer greater than 1; the first driving sequence is used to instruct the three-dimensional facial movements of the first object when outputting the target speech; the second driving sequence is used to instruct the three-dimensional facial movements of the second object when outputting the target speech; the style conversion model is used to output a driving sequence that conforms to the facial style when the second object speaks; A driving unit for driving the three-dimensional facial model corresponding to the second object to perform facial movements according to the target driving sequence.
14. The device according to claim 13, wherein, The determination unit includes: A first determination module for determining the phoneme information corresponding to the to-be-processed speech; wherein, the phoneme information is a set of phonemes that make up the to-be-processed speech; A second determination module for determining the to-be-processed driving sequence according to the mapping relationship and the phoneme information; wherein, the mapping relationship is the correspondence between phonemes and a set of preset parameters; the set of preset parameters is used to instruct the three-dimensional facial movements of the first object when pronouncing phonemes; the to-be-processed driving sequence includes at least one set of preset parameters.
15. The device according to claim 13 or 14, wherein, The first parameter set is used to indicate the three-dimensional facial actions of the first object in the time frame corresponding to the first parameter set; the third driving sequence includes N third parameter sets.
16. The device according to claim 15, wherein, the style conversion model is obtained by adjusting the parameters of the initial model according to the first facial model and the second facial model; wherein, the first facial model is obtained by adjusting the parameters of a preset facial model according to the third driving sequence; the second facial model is obtained by adjusting the parameters of a preset facial model according to the second driving sequence.
17. The device according to claim 15, wherein, the style conversion model is obtained by adjusting the parameters of the initial model according to the first loss function and the second loss function; the first loss function is obtained according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object in the time frame corresponding to the second parameter set; the second loss function is obtained according to the first facial model and the second facial model; wherein, the first facial model is obtained by adjusting the parameters of a preset facial model according to the third driving sequence; the second facial model is obtained by adjusting the parameters of a preset facial model according to the second driving sequence.
18. The device according to any one of claims 13-14, 16-17, wherein, the style conversion model includes M one-dimensional convolutional layers, wherein the one-dimensional convolutional layer is used to perform convolutional processing on the input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
19. A training device for a style conversion model, comprising: an acquisition unit, configured to acquire at least one set of training sets; wherein, the training set includes a first driving sequence and a second driving sequence; the first driving sequence is used to indicate the three-dimensional facial actions of the first object when outputting a target voice; the second driving sequence is used to indicate the three-dimensional facial actions of the second object when outputting the target voice; a training unit, configured to input the N first parameter sets included in the first driving sequence into an initial model to obtain a third driving sequence; and adjust the parameters of the initial model according to the third driving sequence and the second driving sequence to obtain a style conversion model; wherein, N is a positive integer greater than 1; the style conversion model is used to output a driving sequence that conforms to the target style; the target style is the facial style when the second object speaks.
20. The device according to claim 19, wherein, the first parameter set is used to indicate the three-dimensional facial actions of the first object in the time frame corresponding to the first parameter set; the third driving sequence includes N third parameter sets.
21. The device according to claim 20, wherein, the adjustment module includes: a first adjustment sub-module, configured to adjust the parameters of a preset facial model according to the third driving sequence to obtain a first facial model; A second adjustment sub-module, configured to adjust the parameters of the preset facial model according to the second driving sequence to obtain a second facial model; A third adjustment sub-module, configured to adjust the parameters of the initial model according to the first facial model and the second facial model to obtain a style conversion model.
22. The apparatus according to claim 20, wherein, The adjustment module includes: A first determination sub-module, configured to determine a first loss function according to the third driving sequence and the second driving sequence; wherein, the second driving sequence includes N second parameter sets; the second parameter set is used to indicate the three-dimensional facial actions of the second object at the time frame corresponding to the second parameter set; A fourth adjustment sub-module, configured to adjust the parameters of the preset facial model according to the third driving sequence to obtain a first facial model; A fifth adjustment sub-module, configured to adjust the parameters of the preset facial model according to the second driving sequence to obtain a second facial model; A second determination sub-module, configured to determine a second loss function according to the first facial model and the second facial model; A sixth adjustment sub-module, configured to adjust the parameters of the initial model according to the first loss function and the second loss function to obtain a style conversion model.
23. The apparatus according to any one of claims 19-22, wherein, The acquisition unit includes: A third determination module, configured to determine the phoneme information corresponding to the target speech; wherein, the phoneme information is a set of phonemes that make up the target speech; A fourth determination module, configured to determine a first driving sequence according to the mapping relationship and the phoneme information; wherein, the mapping relationship is a correspondence between phonemes and preset parameter sets; the preset parameter set is used to characterize the three-dimensional facial actions when pronouncing phonemes; the first driving sequence includes at least one preset parameter set; A second acquisition module, configured to acquire a second driving sequence; A fifth determination module, configured to use the second driving sequence and the first driving sequence as a set of training sets.
24. The apparatus according to any one of claims 19-22, wherein, The style conversion model includes M one-dimensional convolutional layers, wherein the one-dimensional convolutional layer is configured to perform convolutional processing on the input data to obtain a processing result; and the size of the processing result is the same as the size of the input data; M is a positive integer greater than 1.
25. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-12.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
27. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-12.
Citation Information
Patent Citations
Face video synthesis method and device, equipment and medium
CN112215927A
Virtual image driving and model training method and device, equipment and storage medium
CN117115317A