Training of a virtual avatar mouth driving model and driving method, device and equipment thereof
By training a virtual avatar lip-syncing model, and utilizing pure human voice audio and video images from pure music and synchronized audio-visual videos, combined with a human voice information extraction and lip-syncing coefficient prediction network, the accuracy problem of virtual avatar lip-syncing in noisy environments was solved, achieving accurate lip-syncing in noisy environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-03-24
AI Technical Summary
In existing technologies, virtual avatar lip-syncing driving solutions are difficult to accurately and reliably drive the lip movements of virtual avatars in noisy environments, especially those based on facial expression capture devices and phonetic sequence data, which have errors in real-world scenarios.
By acquiring pure music audio samples and audio-visual synchronized video samples containing pure human voices, a mixed audio sample is synthesized. A lip-syncing model for virtual characters is trained using a human voice information extraction network and a lip-syncing coefficient prediction network. Pure human voice audio and lip-syncing coefficients are used as supervision information to train the model to improve the accuracy of lip-syncing.
In noisy environments, it can accurately and reliably drive the lip movements of virtual avatars based on audio, improving the accuracy and reliability of virtual avatar lip movement driving.
Smart Images

Figure CN115691544B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network live broadcast and artificial intelligence, in particular to a virtual image mouth shape driving model training method, a virtual image driving method and device, electronic equipment and a computer readable storage medium. BACKGROUND
[0002] With the development of network live broadcast technology, virtual image live broadcast has been widely applied in game, e-commerce and other business fields.
[0003] In the current technology, the mouth shape driving of the virtual image mainly comes from the face expression capture device, which collects the face image of the host through the camera and calculates the mouth shape driving coefficient accordingly, but this scheme needs to rely on good lighting environment and collection angle, and it is difficult to accurately and reliably drive the mouth shape of the virtual image. In the current virtual image mouth shape driving technology based on sound, the envelope amplitude of the voice and the phonetic symbol in the sound is analyzed, and the mouth shape is driven through the corresponding preset timing data, but this scheme needs to preset the timing data for the phonetic symbol in the actual scene, and the timing data corresponding to the limited phonetic symbol is also difficult to accurately and reliably drive the mouth shape of the virtual image. SUMMARY
[0004] Therefore, it is necessary to provide a virtual image mouth shape driving model training method, a virtual image driving method, a device, electronic equipment and a computer readable storage medium to solve the above technical problems.
[0005] In a first aspect, the present application provides a virtual image mouth shape driving model training method. The method comprises:
[0006] obtaining a pure music audio sample and an audio-visual synchronous video sample containing pure human voice;
[0007] According to the pure human voice in the audio-visual synchronous video sample and the pure music audio sample, a mixed audio sample is synthesized, and according to the video image corresponding to the pure human voice in the audio-visual synchronous video sample, a mouth shape driving coefficient corresponding to the pure human voice is obtained;
[0008] The mixed audio sample is input into the virtual image mouth shape driving model to be trained, the human voice information extraction network in the virtual image mouth shape driving model extracts the human voice part information in the mixed audio sample according to the mixed audio sample, and provides the human voice part information to the mouth shape coefficient prediction network in the virtual image mouth shape driving model, and the mouth shape coefficient prediction network obtains the corresponding predicted mouth shape driving coefficient according to the human voice part information;
[0009] According to the human voice part information extracted by the human voice information extraction network, a corresponding predicted pure human voice audio is obtained, and a first model loss is obtained according to the consistency of the predicted pure human voice audio and the pure human voice audio;
[0010] According to the consistency of the predicted mouth shape driving coefficient and the mouth shape driving coefficient, a second model loss is obtained;
[0011] According to the first model loss and the second model loss, the virtual image mouth shape driving model to be trained is trained.
[0012] In a second aspect, the present application provides a virtual image driving method. The method comprises:
[0013] Collecting audio of an anchor; inputting the audio into a trained virtual image mouth shape driving model to obtain a predicted mouth shape driving coefficient output by the virtual image mouth shape driving model; wherein the virtual image mouth shape driving model is trained according to the method described above; and driving the mouth shape of a virtual image of the anchor according to the predicted mouth shape driving coefficient.
[0014] In a third aspect, the present application provides a virtual image mouth shape driving model training device. The device comprises:
[0015] A sample acquisition module is configured to acquire a pure music audio sample and acquire a video and audio synchronous video sample containing pure human voice;
[0016] A sample processing module is configured to synthesize a mixed audio sample according to the pure human voice audio in the video and audio synchronous video sample and the pure music audio sample, and acquire a mouth shape driving coefficient corresponding to the pure human voice audio according to a video image corresponding to the pure human voice audio in the video and audio synchronous video sample;
[0017] A sample input module is configured to input the mixed audio sample into a virtual image mouth shape driving model to be trained, extract human voice part information in the mixed audio sample by a human voice information extraction network in the virtual image mouth shape driving model according to the mixed audio sample, and provide the human voice part information to a mouth shape coefficient prediction network in the virtual image mouth shape driving model, so as to obtain a corresponding predicted mouth shape driving coefficient by the mouth shape coefficient prediction network according to the human voice part information;
[0018] A first loss acquisition module is configured to obtain a corresponding predicted pure human voice audio according to human voice part information extracted by the human voice information extraction network, and obtain a first model loss according to the consistency of the predicted pure human voice audio and the pure human voice audio;
[0019] A second loss acquisition module is configured to obtain a second model loss according to the consistency of the predicted mouth shape driving coefficient and the mouth shape driving coefficient.
[0020] a model training module configured to train the virtual image lip driving model to be trained according to the first model loss and the second model loss.
[0021] In a fourth aspect, the present application provides a virtual image driving device. The device comprises:
[0022] an audio acquisition module configured to acquire audio of an anchor;
[0023] an audio input module configured to input the audio into a trained virtual image lip driving model to obtain predicted lip driving coefficients output by the virtual image lip driving model; wherein the virtual image lip driving model is trained by using the device as described above;
[0024] a lip driving module configured to drive a lip of a virtual image of the anchor according to the predicted lip driving coefficients.
[0025] In a fifth aspect, the present application provides an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program: obtaining a pure music audio sample, and obtaining a video sample with pure vocal synchronization; synthesizing a mixed audio sample according to the pure vocal audio in the video sample with pure vocal synchronization and the pure music audio sample, and obtaining a lip driving coefficient corresponding to the pure vocal audio according to a video image corresponding to the pure vocal audio in the video sample with pure vocal synchronization; inputting the mixed audio sample into a virtual image lip driving model to be trained, extracting vocal part information in the mixed audio sample by a vocal information extraction network in the virtual image lip driving model, and providing the vocal part information to a lip coefficient prediction network in the virtual image lip driving model, obtaining corresponding predicted lip driving coefficients by the lip coefficient prediction network according to the vocal part information; obtaining corresponding predicted pure vocal audio according to the vocal part information extracted by the vocal information extraction network, obtaining a first model loss according to consistency of the predicted pure vocal audio and the pure vocal audio; obtaining a second model loss according to consistency of the predicted lip driving coefficients and the lip driving coefficients; training the virtual image lip driving model to be trained according to the first model loss and the second model loss.
[0026] In a sixth aspect, the present application provides an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program: collecting audio of an anchor; inputting the audio into a trained virtual image lip movement driving model to obtain a predicted lip movement driving coefficient output by the virtual image lip movement driving model; wherein the virtual image lip movement driving model is trained according to the method described above; and driving the lip movement of a virtual image of the anchor according to the predicted lip movement driving coefficient.
[0027] In a seventh aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0028] obtaining a pure music audio sample and an audio-visual synchronous video sample containing pure vocal; synthesizing a mixed audio sample according to the pure vocal in the audio-visual synchronous video sample and the pure music audio sample, and obtaining a lip movement driving coefficient corresponding to the pure vocal according to a video image corresponding to the pure vocal in the audio-visual synchronous video sample; inputting the mixed audio sample into a virtual image lip movement driving model to be trained, extracting vocal part information in the mixed audio sample by a vocal information extraction network in the virtual image lip movement driving model, and providing the vocal part information to a lip movement coefficient prediction network in the virtual image lip movement driving model, obtaining a corresponding predicted lip movement driving coefficient by the lip movement coefficient prediction network according to the vocal part information; obtaining a corresponding predicted pure vocal according to the vocal part information extracted by the vocal information extraction network, obtaining a first model loss according to the consistency of the predicted pure vocal and the pure vocal; obtaining a second model loss according to the consistency of the predicted lip movement driving coefficient and the lip movement driving coefficient; and training the virtual image lip movement driving model to be trained according to the first model loss and the second model loss.
[0029] In an eighth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0030] collecting audio of an anchor; inputting the audio into a trained virtual image lip movement driving model to obtain a predicted lip movement driving coefficient output by the virtual image lip movement driving model; wherein the virtual image lip movement driving model is trained according to the method described above; and driving the lip movement of a virtual image of the anchor according to the predicted lip movement driving coefficient.
[0031] The training method of the virtual image mouth shape driving model, the virtual image driving method, the device, the electronic equipment and the computer readable storage medium, the mixed audio sample is synthesized according to the pure vocal audio and the pure music audio sample in the audio-visual synchronous video sample containing the pure vocal, and the corresponding mouth shape driving coefficient is obtained according to the video image corresponding to the pure vocal audio in the audio-visual synchronous video sample; the mixed audio sample is input into the virtual image mouth shape driving model to be trained, the vocal part information in the mixed audio sample is extracted by the vocal information extraction network in the model and provided to the mouth shape coefficient prediction network in the model, the corresponding predicted mouth shape driving coefficient is obtained by the mouth shape coefficient prediction network according to the vocal part information; then the corresponding predicted pure vocal audio is obtained according to the vocal part information extracted by the vocal information extraction network, the first model loss is obtained according to the consistency of the predicted pure vocal audio and the pure vocal audio, and the second model loss is obtained according to the consistency of the predicted mouth shape driving coefficient and the mouth shape driving coefficient; the virtual image mouth shape driving model is trained according to the first and second model losses. In the model training, the mixed audio is obtained by mixing the pure music and the pure vocal audio in the audio-visual synchronous video, the corresponding mouth shape driving coefficient is obtained according to the corresponding video image in the audio-visual synchronous video, the mixed audio is used as the model training input data, and the foregoing pure vocal audio and the mouth shape driving coefficient are used as the model supervision information. On the one hand, the vocal part information provided by the vocal information extraction network in the model is supervised whether it is accurate, and on the other hand, the predicted mouth shape driving coefficient output by the mouth shape coefficient prediction network in the model is supervised whether it is accurate, so as to train the virtual image mouth shape driving model according to the corresponding first and second loss functions. Thus, the virtual image mouth shape driving model which can output the virtual image mouth shape coefficient based on the audio to drive the virtual image mouth shape and can cope with the noisy environment can be trained, so that the model can accurately and reliably drive the virtual image mouth shape based on the audio, and the accuracy and reliability of the virtual image mouth shape driving are improved. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 The application environment diagram of the related method in the embodiments of the present application is shown;
[0033] Figure 2 The flowchart of the training method of the virtual image mouth shape driving model in the embodiments of the present application is shown;
[0034] Figure 3 The schematic diagram of part of the basic mouth shape in the embodiments of the present application is shown;
[0035] Figure 4 The schematic diagram of the virtual image mouth shape driving model to be trained in the embodiments of the present application is shown;
[0036] Figure 5(a) is a schematic diagram of a virtual image mouth shape driving model in an embodiment of the present application;
[0037] Fig. 5(b) is a schematic diagram of another virtual image mouth shape driving model in an embodiment of the present application;
[0038] Figure 6 Fig. 1 is a flowchart of a virtual image driving method in an embodiment of the present application;
[0039] Figure 7 Fig. 4 is a schematic diagram of a virtual image mouth shape in an embodiment of the present application;
[0040] Figure 8 Fig. 7 is a structural block diagram of a virtual image mouth shape driving model training device in an embodiment of the present application;
[0041] Figure 9 Fig. 8 is a structural block diagram of a virtual image driving device in an embodiment of the present application;
[0042] Figure 10 Fig. 10 is an internal structure diagram of an electronic device in an embodiment of the present application;
[0043] Figure 11 Fig. 11 is an internal structure diagram of an electronic device in another embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0045] The virtual image mouth shape driving model training method and the virtual image driving method provided by the present application can be applied in the application environment as shown in Figure 1 The terminal can communicate with the server through the network. Specifically, the virtual image mouth shape driving model training method provided by the present application can be executed by the server, and the virtual image driving method provided by the present application can be executed by the terminal. The server can obtain the trained virtual image mouth shape driving model according to the virtual image mouth shape driving model training method provided by the present application, and then send the trained virtual image mouth shape driving model to the terminal for storage and application. The virtual image mouth shape of the anchor can be driven according to the virtual image driving method provided by the present application. In the application environment as shown in Figure 1 The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones and tablet computers; the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0046] The following is based on the application environment as shown in Figure 1The application environment shown, in combination with the embodiments and corresponding drawings, sequentially describes the training method of the virtual image mouth shape driving model and the driving method of the virtual image provided by the present application.
[0047] In one embodiment, as shown, a training method of a virtual image mouth shape driving model is provided, which can include the following steps: Figure 2
[0048] Step S201, obtaining pure music audio samples and obtaining audio-visual synchronous video samples containing pure human voice.
[0049] Step S202, synthesizing a mixed audio sample according to the pure human voice in the audio-visual synchronous video sample and the pure music audio sample, and obtaining a mouth shape driving coefficient corresponding to the pure human voice according to the video image corresponding to the pure human voice in the audio-visual synchronous video sample.
[0050] In this embodiment, steps S201 and S202 are mainly steps of obtaining and processing related samples. Specifically, in step S201, pure music audio samples and audio-visual synchronous video samples containing pure human voice are obtained. The pure music audio samples are used to mix the background sound of pure human voice with the pure human voice, and the mixed audio sample will be used for model training, so that the model obtained by training can cope with noisy environments, such as when the host is playing music, the host's voice and music will be collected together, which will interfere with mouth shape recognition. Therefore, in the training stage, the present application embodiment obtains pure music audio samples for mixing the background sound of pure human voice with the pure human voice and applies them to the synthesis of mixed audio samples and subsequent model training. For pure music audio samples, specifically, they can be various types of musical instruments, various styles of pure music audio samples, etc. In addition, in step S201, audio-visual synchronous video samples containing pure human voice are also obtained, which can be audio-visual synchronous video of pure human voice of a single person. The audio-visual synchronous video needs to include pure human voice and video images containing the portrait of the individual who makes the sound. The mouth shape of the portrait in the video image is synchronized with the pure human voice. For example, the audio-visual synchronous video sample containing pure human voice can use video samples from relevant news, knowledge lectures, etc.
[0051] Thus, in step S202, on the one hand, the mixed audio samples for model training input data are synthesized, and on the other hand, the mouth shape driving coefficients for model training supervision information are obtained. Among them, for the synthesis of the mixed audio sample, according to the pure human audio and pure music audio samples in the audio-visual synchronous video sample, the mixed audio sample is synthesized, specifically, the audio sequence can be extracted from the audio-visual synchronous video sample in units of, for example, 150 milliseconds or 200 milliseconds, and each piece of audio contained in the audio sequence is basically pure human voice, that is, pure human audio, so that the pure human audio in the audio-visual synchronous video sample can be obtained. Then, the pure human audio and the pure music audio sample are mixed to obtain the mixed audio sample, so that a pair of mixed audio sample and pure human audio can be obtained. Among them, for the acquisition of the mouth shape driving coefficient, according to the video image corresponding to the aforementioned pure human audio in the audio-visual synchronous video sample, the mouth shape driving coefficient corresponding to the pure human audio is obtained, specifically, the mouth shape in the video image corresponding to the pure human audio in the audio-visual synchronous video sample can be recognized by using an existing face expression capture model, so that the mouth shape driving coefficient corresponding to the pure human audio can be obtained.
[0052] Among them, for the mouth shape driving coefficient It can contain 28 components, the sum of each component can be set to 1, j represents the serial number of the component, and the mouth shape B of the virtual image can be obtained by fusing the different basic mouth shapes blendshape{B0, B1, …, B Figure 3 as shown in the following formula: 27} by fusing the mouth shape driving coefficient to obtain:
[0053] Since the pure human audio and the mixed audio sample are paired, the mixed audio sample and the mouth shape driving coefficient can be obtained by step S202, so that the mixed audio sample can correspond to a pure human audio and a mouth shape driving coefficient. The mixed audio sample will be used as model training input data, and the corresponding pure human audio and mouth shape driving coefficient will be used as two supervision information of model training.
[0054] In step S203, the mixed audio sample is input into the virtual image mouth shape driving model to be trained. The human voice information extraction network in the virtual image mouth shape driving model extracts the human voice part information in the mixed audio sample according to the mixed audio sample, and provides the human voice part information to the mouth shape coefficient prediction network in the virtual image mouth shape driving model. The mouth shape coefficient prediction network obtains the corresponding predicted mouth shape driving coefficient according to the human voice part information.
[0055] Specifically, as mentioned above, in a noisy environment, for example, when the host is playing music, the host's voice and the music will enter the microphone together at this time, and the music will interfere with the recognition of the mouth shape. Therefore, in combination with Figure 4The virtual avatar lip-syncing model to be trained in this application includes a voice information extraction network and a lip-syncing coefficient prediction network. During training, the voice information extraction network extracts voice information from mixed audio samples. Then, the lip-syncing coefficient prediction network predicts the corresponding lip-syncing driving coefficients (denoted as predicted lip-syncing driving coefficients) based on the voice information extracted from the mixed audio samples. In other words, the voice information extraction network extracts voice-related information from the input audio data and then feeds this information to the lip-syncing coefficient prediction network to predict the predicted lip-syncing driving coefficients, thereby eliminating interference from non-voice information on lip-syncing recognition. Specifically, as shown... Figure 4 As shown, mixed audio samples are input into the virtual avatar lip-syncing model to be trained. The human voice information extraction network in the model extracts the human voice portion information from the mixed audio samples and provides this human voice portion information to the lip-syncing coefficient prediction network in the model. The lip-syncing coefficient prediction network then obtains the corresponding predicted lip-syncing driving coefficients based on the human voice portion information. In specific implementations of the virtual avatar lip-syncing model, the human voice information extraction network can be implemented based on a U-shaped network with a deep separable convolutional structure, and the lip-syncing coefficient prediction network can be implemented based on a neural network with a deep separable convolutional structure, including but not limited to MobileNet.
[0056] Furthermore, the voice information extracted by the voice information extraction network from the mixed audio samples can be pure human voice audio from the mixed audio samples, or it can be the time spectrum corresponding to the pure human voice audio from the mixed audio samples. Depending on the different extracted voice information, in some embodiments, virtual avatar lip-syncing models with different specific structures can be used to process it.
[0057] In one embodiment, step S203, in which the human voice information extraction network in the virtual avatar lip-sync driving model extracts human voice information from the mixed audio samples based on the mixed audio samples and provides this human voice information to the lip-sync coefficient prediction network in the virtual avatar lip-sync driving model, and the lip-sync coefficient prediction network obtains the corresponding predicted lip-sync driving coefficients based on the human voice information, may include:
[0058] The voice information extraction network extracts pure human voice audio from the mixed audio samples and provides it as human voice part information to the lip shape coefficient prediction network. The lip shape coefficient prediction network obtains the corresponding time spectrum from the pure human voice audio in the mixed audio samples and obtains the corresponding predicted lip shape driving coefficient based on the time spectrum.
[0059] In this embodiment, the voice part information extracted by the voice information extraction network from the mixed audio sample is the pure voice audio in the mixed audio sample. Specifically, in combination with FIG. 5(a), first, the voice information extraction network extracts the pure voice audio in the mixed audio sample from the mixed audio sample, and then the voice information extraction network provides the extracted pure voice audio in the mixed audio sample as the voice part information to the lip coefficient prediction network. The lip coefficient prediction network can include a short-time Fourier transform unit and a lip coefficient prediction unit, the lip coefficient prediction unit can be implemented by using a neural network with a depth separable convolution structure such as a mobilenet, the lip coefficient prediction network can calculate and obtain the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample based on the short-time Fourier transform unit, and then the lip coefficient prediction unit in the lip coefficient prediction network predicts the corresponding predicted lip driving coefficient according to the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample.
[0060] In another embodiment, the voice information extraction network in the virtual image lip driving model extracts the voice part information in the mixed audio sample from the mixed audio sample in step S203, and provides the voice part information to the lip coefficient prediction network in the virtual image lip driving model, and the lip coefficient prediction network obtains the corresponding predicted lip driving coefficient according to the voice part information, which can include:
[0061] The voice information extraction network obtains the time-frequency spectrum corresponding to the mixed audio sample from the mixed audio sample, extracts the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample from the time-frequency spectrum corresponding to the mixed audio sample, and provides the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample as the voice part information to the lip coefficient prediction network; the lip coefficient prediction network obtains the corresponding predicted lip driving coefficient from the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample.
[0062] In this embodiment, the voice part information extracted by the voice information extraction network from the mixed audio sample is the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample. Specifically, in combination with FIG. 5(b), the voice information extraction network can include a short-time Fourier transform unit and a voice information extraction unit, the voice information extraction unit can be implemented by using a U-shaped network with a depth separable convolution structure, the voice information extraction network can calculate and obtain the time-frequency spectrum corresponding to the mixed audio sample based on the short-time Fourier transform unit, then the voice information extraction unit in the voice information extraction network extracts the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample from the time-frequency spectrum corresponding to the mixed audio sample, provides the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample as the voice part information to the lip coefficient prediction network, and then the lip coefficient prediction network predicts the corresponding predicted lip driving coefficient from the time-frequency spectrum corresponding to the pure voice audio in the mixed audio sample.
[0063] Step S204, according to the human voice information extraction network extracted human voice part information to obtain the corresponding prediction of pure human voice frequency, according to the consistency of the prediction of pure human voice frequency and pure human voice frequency to obtain the first model loss.
[0064] Step S205, according to the consistency of the predicted mouth shape driving coefficient and the mouth shape driving coefficient, the second model loss is obtained.
[0065] Step S204 and step S205 are the related steps for obtaining the corresponding first and second model losses by using the aforementioned pure human voice frequency corresponding to the mixed audio sample and the mouth shape driving coefficient corresponding to the mixed audio sample to train the two models. Combined with Figure 4 , using the pure human voice frequency in the audio-visual synchronous video sample as the output of the human voice information extraction network, the human voice information extraction network can accurately extract the human voice part information in the mixed audio sample, and using the mouth shape driving coefficient obtained based on the video image corresponding to the pure human voice frequency in the audio-visual synchronous video sample as the output of the mouth shape coefficient prediction network, the mouth shape coefficient prediction network can accurately extract the predicted mouth shape driving coefficient.
[0066] Specifically, in step S204, the corresponding prediction of pure human voice frequency needs to be obtained according to the human voice part information extracted by the human voice information extraction network, and then the first model loss is obtained according to the consistency of the prediction of pure human voice frequency V i and pure human voice frequency V o , through the back propagation algorithm, the prediction of pure human voice frequency can be close to the pure human voice frequency, so that the human voice information extraction network can accurately extract the human voice part information in the mixed audio sample, and the calculation method of the first model loss L1 can be used as follows: L1 = ||V i -V o || 2 .
[0067] For step S204, different methods can be used to obtain the corresponding predicted pure human voice audio based on the human voice information extracted by the human voice information extraction network, depending on the different extracted human voice information. In one embodiment, obtaining the corresponding predicted pure human voice audio based on the human voice information extracted by the human voice information extraction network in step S204 can include: using the pure human voice audio in the mixed audio sample extracted by the human voice information extraction network as the corresponding predicted pure human voice audio. That is, referring to Figure 5(a), when the human voice information extracted by the human voice information extraction network from the mixed audio sample is the pure human voice audio in the mixed audio sample, the pure human voice audio in the mixed audio sample extracted by the human voice information extraction network can be directly used as the corresponding predicted pure human voice audio, and then the first model loss L1 is obtained based on the predicted pure human voice audio and the pure human voice audio in the audio-visual synchronized video sample.
[0068] In another embodiment, step S204, obtaining the corresponding predicted pure human voice audio based on the human voice portion information extracted by the human voice information extraction network, may include: obtaining the corresponding predicted pure human voice audio based on the time spectrum corresponding to the pure human voice audio in the mixed audio samples extracted by the human voice information extraction network. In this embodiment, referring to Figure 5(b), when the human voice portion information extracted by the human voice information extraction network based on the mixed audio samples is the time spectrum corresponding to the pure human voice audio in the mixed audio samples, the inverse short-time Fourier transform unit can be used to calculate the corresponding predicted pure human voice audio based on the time spectrum corresponding to the pure human voice audio in the mixed audio samples. Then, the first model loss L1 is obtained based on the predicted pure human voice audio and the pure human voice audio in the audio-visual synchronized video samples.
[0069] For step S205, the lip-sync driving coefficients obtained from the video images corresponding to pure human voice audio in the audio-visual synchronized video samples are used as supervision for the output of the lip-sync coefficient prediction network, enabling the lip-sync coefficient prediction network to accurately extract and predict the lip-sync driving coefficients. Specifically, as follows... Figure 4 to Figure 5(b) As shown, the predicted lip shape driving coefficients can be predicted based on the lip shape coefficients of the network output. With mouth shape driving coefficient The consistency of the second model loss L2 is used to obtain the predicted lip shape driving coefficients, making the predicted lip shape driving coefficients close to the predicted lip shape driving coefficients. This allows the lip shape coefficient prediction network to accurately extract the predicted lip shape driving coefficients. For example, the second model loss L2 can be calculated as follows:
[0070] Step S206: Train the virtual avatar lip-syncing driven model to be trained based on the first model loss and the second model loss.
[0071] In this step, specifically, the overall model loss L of the virtual image mouth driving model to be trained can be calculated according to the first model loss L1 and the second model loss L2. The voice information extraction network and the mouth shape coefficient prediction network in the virtual image mouth driving model to be trained are updated based on the overall model loss L, so as to train the virtual image mouth driving model. As an implementation manner, the training of the virtual image mouth driving model to be trained can be determined to be completed when the overall model loss L is less than or equal to a preset model loss threshold, and a trained virtual image mouth driving model is obtained.
[0072] In the training method of the virtual image mouth driving model, the mixed audio sample is synthesized according to the pure voice audio and the pure music audio sample in the audio-visual synchronous video sample containing pure voice, and the corresponding mouth shape driving coefficient is obtained according to the video image corresponding to the pure voice audio in the audio-visual synchronous video sample; the mixed audio sample is input into the virtual image mouth driving model to be trained, the voice part information in the mixed audio sample is extracted by the voice information extraction network in the model and provided to the mouth shape coefficient prediction network in the model, and the corresponding predicted mouth shape driving coefficient is obtained by the mouth shape coefficient prediction network according to the voice part information; then the corresponding predicted pure voice audio is obtained according to the voice part information extracted by the voice information extraction network, the first model loss is obtained according to the consistency of the predicted pure voice audio and the pure voice audio, and the second model loss is obtained according to the consistency of the predicted mouth shape driving coefficient and the mouth shape driving coefficient; and the virtual image mouth driving model is trained according to the first and second model losses. In the model training, the mixed audio is obtained by mixing the pure music and the pure voice audio in the audio-visual synchronous video, the corresponding mouth shape driving coefficient is obtained according to the corresponding video image in the audio-visual synchronous video, the mixed audio is used as the input data for model training, and the foregoing pure voice audio and the mouth shape driving coefficient are used as the supervision information of the model. On the one hand, the voice part information provided by the voice information extraction network in the model is supervised to be accurate, and on the other hand, the predicted mouth shape driving coefficient output by the mouth shape coefficient prediction network in the model is supervised to be accurate, so that the virtual image mouth driving model is trained according to the corresponding first and second loss functions. Thus, the virtual image mouth driving model that can output the virtual image mouth shape coefficient based on the audio to drive the virtual image mouth shape and can cope with the noisy environment can be trained, so that the model can accurately and reliably drive the virtual image mouth shape based on the audio, and the accuracy and reliability of the virtual image mouth driving are improved.
[0073] In some embodiments, the obtaining the pure music audio sample in step S201 can include: obtaining multiple types of pure music audio samples. Specifically, pure music audio samples of different types such as various musical instruments, various music styles, etc. can be obtained. The synthesizing the mixed audio sample according to the pure human voice audio and the pure music audio sample in the audio-visual synchronous video sample in step S202 further includes: mixing at least two types of pure music audio samples in the multiple types of pure music audio samples with the pure human voice audio in the audio-visual synchronous video sample according to a mixing ratio adapted to an audio collection scene to obtain the mixed audio sample.
[0074] Specifically, after obtaining multiple types of pure music audio samples, the multiple types of pure music audio samples can be mixed with the pure human voice audio in the audio-visual synchronous video sample to obtain the mixed audio sample. In this embodiment, in order to adapt to the actual audio collection scene in the model application stage, a mixing ratio adapted to the audio collection scene can be obtained first. The audio collection scene refers to a specific scene in which the terminal collects audio through the microphone in the application stage of the virtual image lip driving model, such as a live broadcast scene, etc. The mixing ratio refers to a ratio used when mixing different pure music audio samples as background audio with the pure human voice audio in the audio-visual synchronous video sample. The mixing ratio adapted to the audio collection scene can be set by relevant personnel according to experience, can be determined by analyzing the audio collected in the actual scene by using a related algorithm, etc. This embodiment does not limit this. Then, at least two types of pure music audio samples can be selected from the multiple types of pure music audio samples according to the mixing ratio to mix with the pure human voice audio to obtain the mixed audio sample. For example, the mixing ratio is X1: X2: Y, two types of pure music audio samples can be selected to mix with the pure human voice audio according to the amplitude ratio of X1: X2: Y. For another example, the mixing ratio is X1: X2: X3: Y, three types of pure music audio samples can be selected to mix with the pure human voice audio according to the amplitude ratio of X1: X2: X3: Y, etc. Thus, the scheme of this embodiment can generate a mixed audio sample more adapted to the actual audio collection scene for model training, so that the mixed audio sample can have better performance in driving the lip of the virtual image in the audio collection scene such as live broadcast.
[0075] In some embodiments, the obtaining the pure human voice audio corresponding to the lip driving coefficient according to the video image corresponding to the pure human voice audio in the audio-visual synchronous video sample in step S202 can include:
[0076] According to a time period corresponding to the pure human voice in the audio-visual synchronous video sample, a corresponding video image sequence is obtained; according to the video image sequence, a video image used for extracting the lip driving coefficient is obtained; the video image is input into the facial expression capturing model, and a facial expression coefficient corresponding to the video image output by the facial expression capturing model is obtained; and according to the facial expression coefficient, the lip driving coefficient corresponding to the pure human voice is obtained.
[0077] In the embodiment, after the pure human voice is extracted from the audio-visual synchronous video sample, according to a time period corresponding to the pure human voice in the audio-visual synchronous video sample, a video image sequence corresponding to the time period is obtained. For example, if a pure human voice of 200 milliseconds is extracted from the audio-visual synchronous video sample, a corresponding video image sequence can be extracted from the audio-visual synchronous video sample according to the time period. The video image sequence can include multiple video images, and the last video image in the video image sequence can be extracted as a video image used for extracting the lip driving coefficient. Since the pure human voice is usually extracted from the audio-visual synchronous video sample according to a time length corresponding to a pronunciation unit of a person, the last video image in the video image sequence can represent the pronunciation unit when the pronunciation unit is completed in the video image sequence corresponding to the time period. Then, the video image is input into the existing facial expression capturing model, and the facial expression coefficient corresponding to the video image is output by the facial expression capturing model according to the face in the video image. The facial expression coefficient can include coefficients of eyebrows, eyes, and lips, and the like. Thus, the lip driving coefficient corresponding to the pure human voice can be obtained according to the coefficient of the lip part in the facial expression coefficient provided by the facial expression capturing model. The embodiment extracts the lip driving coefficient synchronized with the pure human voice from the video image by using the facial expression capturing model, avoids the complex design of the lip driving coefficient by manual operation, improves the model training efficiency, and takes into account the accuracy of the lip recognition.
[0078] In some embodiments, the obtaining of the audio-visual synchronous video sample containing the pure human voice in step S201 can include: obtaining a video sample collected from a pure human voice broadcasting scene; performing face tracking on the video image in the video sample; and obtaining the audio-visual synchronous video sample containing the pure human voice from the video sample according to the face tracking result.
[0079] Specifically, in the acquisition stage of audio-visual synchronous video samples, it is necessary to acquire as much as possible audio-visual synchronous video samples of pure human voice of a single person, and the audio-visual synchronous video samples need to contain pure human voice and the portrait of the individual who makes the sound, and the mouth shape of the portrait is synchronized with the human voice. In this embodiment, video samples collected from a pure human voice broadcasting scene can be acquired. The pure human voice broadcasting scene can be a news broadcast, a legal knowledge explanation scene, etc. These scenes usually have a specific person continuously making sound and the picture is less switched, so the individual making the sound can usually be guaranteed to be continuously in the picture. In order to more accurately acquire audio-visual synchronous video samples containing pure human voice based on this, in this embodiment, after acquiring the video samples collected from the pure human voice broadcasting scene, the video images in the video samples are subjected to face tracking to obtain a face tracking result. The face tracking result can indicate whether there is an individual making sound in the video images continuously. The video images in which the face tracking is lost in the video samples can be discarded, and the video images in which the face tracking result is an individual making sound are retained, thereby obtaining audio-visual synchronous video samples containing pure human voice.
[0080] In one embodiment, as shown in Figure 6 a driving method of a virtual image is provided, which can be executed by a terminal of an anchor, and can include the following steps:
[0081] In step S601, audio of an anchor is collected.
[0082] In this step, the terminal of the anchor can collect the audio of the anchor through a microphone after the anchor starts broadcasting.
[0083] In step S602, the audio is input into a trained virtual image mouth shape driving model to obtain a predicted mouth shape driving coefficient output by the virtual image mouth shape driving model.
[0084] The virtual image mouth shape driving model is trained according to the training method of the virtual image mouth shape driving model described in the above embodiments of the present application. Specifically, as an embodiment, in combination with FIG. 5(b), the trained virtual image mouth shape driving model can calculate a time-frequency graph of the audio of the anchor using a short-time Fourier transform unit, then extract a time-frequency graph of pure human voice audio in the audio of the anchor according to the time-frequency graph of the audio of the anchor using a human voice information extraction unit, and then input the time-frequency graph of the pure human voice audio in the audio of the anchor to a mouth shape coefficient prediction network to output a predicted mouth shape driving coefficient according to the time-frequency graph. That is, in the application stage or test stage of the virtual image mouth shape driving model, no related calculation is needed using an inverse short-time Fourier transform unit.
[0085] In step S603, the mouth shape of the virtual image of the anchor is driven according to the predicted mouth shape driving coefficient.
[0086] In this step, the mouth shape of the virtual image of the anchor is driven according to the predicted mouth shape driving coefficient He Ru Figure 3 The different basic mouth shapes shown are blendshape{B0, B1, ..., B 27}, to obtain and drive the lip movements of the streamer's virtual avatar Among them, such as Figure 7 This shows the lip movements of a virtual avatar of a broadcaster driven by this.
[0087] The solution in this embodiment can apply the virtual avatar lip-sync driving model trained by the training method of the virtual avatar lip-sync driving model provided in this application to drive the lip-sync of the virtual avatar of the anchor in a live streaming scenario. It can accurately drive the lip-sync of the virtual avatar of the anchor in noisy environments based on audio, even when the anchor is not facing the camera or the lighting conditions are relatively dark. This overcomes the problem that the current camera-based virtual avatar lip-sync driving scheme performs poorly in dark lighting conditions. It also avoids the problem that the traditional virtual avatar lip-sync driving scheme cannot be used normally in noisy environments. It improves the accuracy and reliability of driving the lip-sync of the anchor's virtual avatar. Moreover, the virtual avatar lip-sync driving model in this application adopts an end-to-end design, which simultaneously realizes the extraction of human voice information and lip-sync driving. The audio sequence used for inference is short, and it has the characteristics of low latency and real-time performance, thus improving the virtual anchor technology.
[0088] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0089] Based on the same inventive concept, this application also provides a related apparatus for implementing the aforementioned related methods. The solution provided by this apparatus is similar to the implementation scheme described in the above methods; therefore, the specific limitations in one or more related apparatus embodiments provided below can be found in the limitations of the related methods described above, and will not be repeated here.
[0090] In one embodiment, such as Figure 8 As shown, a training device for a virtual character lip-syncing driven model is provided. The device 800 may include:
[0091] The sample obtaining module 801 is configured to obtain a pure music audio sample and obtain a video sample containing pure human voice;
[0092] The sample processing module 802 is configured to synthesize a mixed audio sample according to the pure human voice in the video sample and the pure music audio sample, and obtain a lip driving coefficient corresponding to the pure human voice according to a video image corresponding to the pure human voice in the video sample;
[0093] The sample input module 803 is configured to input the mixed audio sample into a virtual image lip driving model to be trained, extract human voice part information in the mixed audio sample by a human voice information extraction network in the virtual image lip driving model, and provide the human voice part information to a lip coefficient prediction network in the virtual image lip driving model to obtain a corresponding predicted lip driving coefficient by the lip coefficient prediction network according to the human voice part information.
[0094] The first loss obtaining module 804 is configured to obtain a corresponding predicted pure human voice according to the human voice part information extracted by the human voice information extraction network, and obtain a first model loss according to consistency of the predicted pure human voice and the pure human voice.
[0095] The second loss obtaining module 805 is configured to obtain a second model loss according to consistency of the predicted lip driving coefficient and the lip driving coefficient.
[0096] The model training module 806 is configured to train the virtual image lip driving model to be trained according to the first model loss and the second model loss.
[0097] In one embodiment, the sample input module 803 is configured to extract pure human voice in the mixed audio sample by the human voice information extraction network according to the mixed audio sample, and provide the pure human voice in the mixed audio sample as the human voice part information to the lip coefficient prediction network; obtain a corresponding time-frequency spectrum according to the pure human voice in the mixed audio sample by the lip coefficient prediction network, and obtain a corresponding predicted lip driving coefficient according to the time-frequency spectrum; and the first loss obtaining module 804 is configured to take the pure human voice in the mixed audio sample extracted by the human voice information extraction network as the corresponding predicted pure human voice.
[0098] In one embodiment, the sample input module 803 is configured to acquire, by the human voice information extraction network, a time-frequency spectrum corresponding to the mixed audio sample according to the mixed audio sample, extract a time-frequency spectrum corresponding to pure human audio in the mixed audio sample according to the time-frequency spectrum corresponding to the mixed audio sample, and provide the time-frequency spectrum corresponding to the pure human audio in the mixed audio sample as the human voice part information to the lip coefficient prediction network; the lip coefficient prediction network acquires a corresponding predicted lip driving coefficient according to the time-frequency spectrum corresponding to the pure human audio in the mixed audio sample; the first loss acquisition module 804 is configured to acquire a corresponding predicted pure human audio according to the time-frequency spectrum corresponding to the pure human audio in the mixed audio sample extracted by the human voice information extraction network.
[0099] In one embodiment, the sample acquisition module 801 is configured to acquire a plurality of types of pure music audio samples; the sample processing module 802 is configured to mix at least two types of pure music audio samples in the plurality of types of pure music audio samples with pure human audio in the audio-visual synchronous video sample according to a mixing ratio suitable for an audio acquisition scene, to obtain the mixed audio sample.
[0100] In one embodiment, the sample processing module 802 is configured to acquire a corresponding video image sequence according to a time period corresponding to the pure human audio in the audio-visual synchronous video sample, acquire a video image used for extracting a lip driving coefficient according to the video image sequence, input the video image into a facial expression capture model to obtain a facial expression coefficient output by the facial expression capture model corresponding to the video image, and acquire a lip driving coefficient corresponding to the pure human audio according to the facial expression coefficient.
[0101] In one embodiment, the sample acquisition module 801 is configured to acquire a video sample collected from a pure human voice broadcast scene, perform facial tracking on a video image in the video sample, and acquire the audio-visual synchronous video sample containing pure human voice from the video sample according to a facial tracking result.
[0102] In one embodiment, as shown in Figure 9 the driving device of a virtual image is provided, and the device 900 can include:
[0103] The audio acquisition module 901 is configured to acquire audio of an anchor;
[0104] The audio input module 902 is configured to input the audio to a trained virtual image lip driving model to obtain a predicted lip driving coefficient output by the virtual image lip driving model; wherein the virtual image lip driving model is trained by using the training device of the virtual image lip driving model as described above;
[0105] a mouth shape driving module 903, configured to drive the mouth shape of the virtual image of the host according to the predicted mouth shape driving coefficient.
[0106] Each of the above modules in the device can be implemented wholly or partially by software, hardware, and combinations thereof. The above modules can be embedded in or independent of a processor in the electronic device in hardware form, or stored in a memory in the electronic device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above modules.
[0107] In one embodiment, an electronic device, which can be a server, has an internal structure diagram as shown in Figure 10 The electronic device includes a processor, a memory, and a network interface connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is configured to store related sample data. The network interface of the electronic device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement a virtual image mouth shape driving model training method.
[0108] In one embodiment, an electronic device, which can be a terminal, has an internal structure diagram as shown in Figure 11 The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is configured to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved by WIFI, mobile cellular network, NFC (near field communication), or other technologies. The computer program is executed by the processor to implement a virtual image driving method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0109] Those skilled in the art can understand that Figure 10 and Figure 11The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0110] In one embodiment, an electronic device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above-mentioned method embodiments when executing the computer program.
[0111] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and the computer program is executed by a processor to implement the steps in the above-mentioned method embodiments.
[0112] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium and can include the processes of the above-mentioned embodiments when executed. Any reference to a memory, database or other medium used in the embodiments provided by the present application can include at least one of a non-volatile and volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory, an optical storage, a high-density embedded non-volatile memory, a resistive memory (ReRAM), a magnetoresistive random access memory (MRAM), a ferroelectric random access memory (FRAM), a phase change memory (PCM), a graphene memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory, etc. As an illustration but not as a limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0113] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.
[0114] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.
[0115] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A training method for a virtual character lip-syncing driven model, characterized in that, The method includes: Obtain pure music audio samples and obtain audio-visual synchronized video samples containing pure human voices; Based on the pure human voice audio and the pure music audio sample in the audio-visual synchronized video sample, a mixed audio sample is synthesized, and based on the video image corresponding to the pure human voice audio in the audio-visual synchronized video sample, the lip-sync driving coefficient corresponding to the pure human voice audio is obtained. The mixed audio samples are input into the virtual avatar lip-sync driving model to be trained. The human voice information extraction network in the virtual avatar lip-sync driving model extracts the human voice information from the mixed audio samples and provides the human voice information to the lip-sync coefficient prediction network in the virtual avatar lip-sync driving model. The lip-sync coefficient prediction network obtains the corresponding predicted lip-sync driving coefficients based on the human voice information. Based on the human voice information extracted by the human voice information extraction network, the corresponding predicted pure human voice audio is obtained, and the first model loss is obtained based on the consistency between the predicted pure human voice audio and the pure human voice audio. Based on the consistency between the predicted mouth shape driving coefficient and the mouth shape driving coefficient, the second model loss is obtained; The virtual avatar lip-syncing model to be trained is trained based on the first model loss and the second model loss.
2. The method according to claim 1, characterized in that, The process involves the human voice information extraction network in the virtual avatar lip-sync driving model extracting human voice information from the mixed audio samples and providing this information to the lip-sync coefficient prediction network in the virtual avatar lip-sync driving model. The lip-sync coefficient prediction network then obtains the corresponding predicted lip-sync driving coefficients based on the human voice information. This includes: The human voice information extraction network extracts pure human voice audio from the mixed audio samples based on the mixed audio samples, and provides the pure human voice audio from the mixed audio samples as the human voice part information to the lip shape coefficient prediction network; The lip-sync coefficient prediction network obtains the corresponding temporal spectrum based on the pure human voice audio in the mixed audio samples, and obtains the corresponding predicted lip-sync driving coefficients based on the temporal spectrum. The step of obtaining the corresponding predicted pure human voice audio based on the human voice portion information extracted by the human voice information extraction network includes: The pure human voice audio extracted from the mixed audio samples by the human voice information extraction network is used as the corresponding predicted pure human voice audio.
3. The method according to claim 1, characterized in that, The process involves the human voice information extraction network in the virtual avatar lip-sync driving model extracting human voice information from the mixed audio samples and providing this information to the lip-sync coefficient prediction network in the virtual avatar lip-sync driving model. The lip-sync coefficient prediction network then obtains the corresponding predicted lip-sync driving coefficients based on the human voice information. This includes: The human voice information extraction network obtains the time spectrum corresponding to the mixed audio sample based on the mixed audio sample, extracts the time spectrum corresponding to the pure human voice audio in the mixed audio sample based on the time spectrum corresponding to the mixed audio sample, and provides the time spectrum corresponding to the pure human voice audio in the mixed audio sample as the human voice part information to the lip shape coefficient prediction network. The lip-sync coefficient prediction network obtains the corresponding predicted lip-sync driving coefficients based on the time spectrum of the pure human voice audio in the mixed audio samples; The step of obtaining the corresponding predicted pure human voice audio based on the human voice portion information extracted by the human voice information extraction network includes: Based on the time spectrum of the pure human voice audio in the mixed audio samples extracted by the human voice information extraction network, the corresponding predicted pure human voice audio is obtained.
4. The method according to any one of claims 1 to 3, characterized in that, The acquisition of pure music audio samples includes: Obtain various types of pure music audio samples; The process of synthesizing a mixed audio sample based on the pure human voice audio and the pure music audio sample in the audio-visual synchronized video sample includes: Based on a mixing ratio adapted to the audio acquisition scenario, at least two types of pure music audio samples from the various types of pure music audio samples are mixed with pure human voice audio from the audio-visual synchronized video samples to obtain the mixed audio sample.
5. The method according to any one of claims 1 to 3, characterized in that, The step of obtaining the lip-sync driving coefficients corresponding to the pure human voice audio from the video image corresponding to the pure human voice audio in the audio-visual synchronized video sample includes: Based on the time period corresponding to the pure human voice audio in the audio-visual synchronized video sample, obtain the corresponding video image sequence; Based on the video image sequence, a video image for extracting lip-sync driving coefficients is obtained; The video image is input into the facial expression capture model to obtain the facial expression coefficients corresponding to the video image output by the facial expression capture model. Based on the facial expression coefficients, the lip-syncing coefficients corresponding to the pure human voice audio are obtained.
6. The method according to any one of claims 1 to 3, characterized in that, The acquisition of audio-visual synchronized video samples containing only human voices includes: Acquire video samples from pure human voice broadcasting scenarios; Face tracking is performed on the video images in the video samples; Based on the face tracking results, the audio-visual synchronized video sample containing only human voice is obtained from the video sample.
7. A method for driving a virtual avatar, characterized in that, The method includes: Collect the broadcaster's audio; The audio is input to a trained virtual avatar lip-syncing model to obtain the predicted lip-syncing driving coefficients output by the virtual avatar lip-syncing model; wherein the virtual avatar lip-syncing model is trained according to the method described in any one of claims 1 to 6. The predicted lip-syncing driving coefficient is used to drive the lip-syncing of the virtual avatar of the broadcaster.
8. A training device for a virtual character lip-syncing driven model, characterized in that, The device includes: The sample acquisition module is used to acquire pure music audio samples and audio-visual synchronized video samples containing pure human voices. The sample processing module is used to synthesize a mixed audio sample based on the pure human voice audio and the pure music audio sample in the audio-visual synchronized video sample, and to obtain the lip-sync driving coefficient corresponding to the pure human voice audio based on the video image corresponding to the pure human voice audio in the audio-visual synchronized video sample. The sample input module is used to input the mixed audio samples into the virtual avatar lip-sync driving model to be trained. The human voice information extraction network in the virtual avatar lip-sync driving model extracts the human voice information from the mixed audio samples and provides the human voice information to the lip-sync coefficient prediction network in the virtual avatar lip-sync driving model. The lip-sync coefficient prediction network obtains the corresponding predicted lip-sync driving coefficients based on the human voice information. The first loss acquisition module is used to obtain the corresponding predicted pure human voice audio based on the human voice information extracted by the human voice information extraction network, and to obtain the first model loss based on the consistency between the predicted pure human voice audio and the pure human voice audio. The second loss acquisition module is used to acquire the second model loss based on the consistency between the predicted mouth shape driving coefficient and the mouth shape driving coefficient. The model training module is used to train the virtual avatar lip-syncing driven model to be trained based on the first model loss and the second model loss.
9. A driving device for a virtual avatar, characterized in that, The device includes: The audio acquisition module is used to capture the broadcaster's audio. An audio input module is used to input the audio to a trained virtual avatar lip-syncing driving model to obtain the predicted lip-syncing driving coefficients output by the virtual avatar lip-syncing driving model; wherein the virtual avatar lip-syncing driving model is trained using the device described in claim 8. The lip-syncing module is used to drive the lip movements of the virtual avatar of the broadcaster according to the predicted lip-syncing driving coefficient.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6 or claim 7.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 6 or claim 7.
Citation Information
Patent Citations
Mouth action driving model training method and assembly based on ASR acoustic model
CN113111813A
Lip shape model training method and device, and voice animation synthesis method and device
CN113314094A