Speech synthesis and lip-driven method, device, equipment and storage medium

By obtaining phoneme sequence characteristics and generating relevant audio feature information, and directly determining the lip feature parameters, the problem of low efficiency in audio and lip animation generation in the prior art is solved, and efficient and direct audio and lip animation generation is achieved.

CN116246609BActive Publication Date: 2025-05-06CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310162639.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-05-06
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

In the virtual avatar display, it is necessary to create audio and then extract features from the audio to generate lip animations, resulting in low generation efficiency and delays.

Method used

By acquiring phoneme sequence features, audio PPG feature information is generated, and pitch and energy feature information is generated using a pre-trained prediction model, thereby generating superimposed audio feature information. Based on these feature information, the lip feature parameters are directly determined and corresponding audio and lip animations are generated.

Benefits of technology

This method does not need to extract features from audio, and directly generates audio and lip animations, avoids delays, simplifies the process, and improves generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246609B_ABST
    Figure CN116246609B_ABST
Patent Text Reader

Abstract

The present application provides a speech synthesis and lip-sync driving method, apparatus, device and storage medium, which obtains phoneme sequence features, then generates audio PPG feature information based on the phoneme sequence features, generates pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model, generates superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information, determines lip-sync feature parameters according to the superimposed audio feature information, generates corresponding audio according to the superimposed audio feature information, determines corresponding lip-sync feature parameters based on the lip-sync feature parameters, and plays the lip-sync animation and audio. Since the corresponding audio and the corresponding lip-sync feature parameters can be directly generated according to the superimposed audio feature information, there is no need to extract features from the audio, which avoids delayed generation of lip-sync animation, simplifies the process of generating audio and corresponding lip-sync animation, and improves generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to a speech synthesis and lip-driven method, device, equipment and storage medium. Background Art

[0002] In the current field of virtual image display, it is usually necessary to synthesize the audio emitted by the virtual image and the corresponding lip animation. At present, the corresponding audio is usually generated first, and then features are extracted from the audio. The corresponding lip animation is determined based on the extracted features. In other words, it is necessary to generate audio first, and then generate the corresponding lip animation based on the generated audio. The lip animation generation is delayed and the generation efficiency is low. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a speech synthesis and lip-driven method, device, equipment and storage medium to solve the above-mentioned technical problems.

[0004] On the one hand, a speech synthesis and lip-activated method is provided, the method comprising:

[0005] Obtain phoneme sequence features;

[0006] Generate audio PPG feature information based on the phoneme sequence feature;

[0007] Generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model;

[0008] Generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information;

[0009] Determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information;

[0010] Determine the corresponding lip animation based on the lip feature parameters;

[0011] Play the lip animation and the audio.

[0012] In one embodiment, the obtaining of phoneme sequence features includes:

[0013] Determine the text information to be broadcasted;

[0014] Generate a corresponding phoneme sequence according to the text information;

[0015] The phoneme sequence is encoded to obtain corresponding phoneme sequence features.

[0016] In one embodiment, generating audio PPG feature information based on the phoneme sequence feature includes:

[0017] Determining voiceprint information for audio synthesis;

[0018] The voiceprint information and the phoneme sequence features are input into a pre-trained PPG prediction model to obtain corresponding audio PPG feature information; the PPG prediction model is a model trained based on audio training samples and corresponding audio text sequence training samples.

[0019] In one of the embodiments, the PPG prediction model is a model obtained by training with the error between the phoneme duration prediction feature and the PPG prediction feature as an additional loss, the phoneme duration prediction feature is a feature obtained by predicting the phoneme duration based on the voiceprint feature of the audio training sample and the phoneme sequence feature of the corresponding audio text sequence training sample, and the PPG prediction feature is a feature obtained by performing speech recognition processing on the audio training sample.

[0020] In one embodiment, determining voiceprint information for audio synthesis includes:

[0021] Collecting the target voice through an audio collection device, and extracting voiceprint information from the target voice;

[0022] or,

[0023] A timbre selection instruction is received, and corresponding voiceprint information is selected from a preset voiceprint information library according to the timbre selection instruction.

[0024] In one embodiment, generating pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model includes:

[0025] The voiceprint information and the audio PPG feature information are respectively input into a pre-trained pitch prediction model and a pre-trained energy prediction model to obtain pitch feature information and energy feature information.

[0026] In one embodiment, generating corresponding audio according to the superimposed audio feature information includes:

[0027] Generate audio of corresponding timbre according to the voiceprint information and the superimposed audio feature information.

[0028] On the other hand, a speech synthesis and lip-actuating device is provided, the device comprising:

[0029] An acquisition module, used for acquiring phoneme sequence features;

[0030] A first generating module, configured to generate audio PPG feature information based on the phoneme sequence feature;

[0031] A second generating module is used to generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model;

[0032] A third generating module, configured to generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information;

[0033] A fourth generating module, used to determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information;

[0034] A determination module, used for determining a corresponding lip animation based on the lip feature parameters;

[0035] A playing module is used to play the lip animation and the audio.

[0036] On the other hand, the present application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the above-mentioned methods.

[0037] On the other hand, the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, it implements any of the above-mentioned methods.

[0038] The speech synthesis and lip-sync driving method, apparatus, device and storage medium provided in the present application obtain phoneme sequence features, then generate audio PPG feature information based on the phoneme sequence features, generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model, generate superimposed audio feature information according to the audio PPG feature information, pitch feature information and energy feature information, determine lip feature parameters according to the superimposed audio feature information, generate corresponding audio according to the superimposed audio feature information, determine corresponding lip animation based on the lip feature parameters, play the lip animation and audio, and since the superimposed audio feature information is a highly abstract audio feature, it can be directly used as a prediction of lip animation parameters. Compared with the existing solution that needs to extract features from the audio to generate the corresponding lip animation, the solution provided in the present application can directly generate the corresponding audio and the corresponding lip feature parameters according to the superimposed audio feature information, that is, there is no need to extract features from the audio, which avoids delayed generation of lip animation, simplifies the process of generating audio and corresponding lip animation, and improves generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A flowchart of the speech synthesis and lip-activated method provided in Example 1 of the present application;

[0040] Figure 2 A schematic diagram of a process for obtaining phoneme sequence features provided in Embodiment 1 of the present application;

[0041] Figure 3 A schematic diagram of a process for generating audio PPG feature information provided in Example 1 of the present application;

[0042] Figure 4 A flowchart of the speech synthesis and lip-activated method provided in Example 1 of the present application;

[0043] Figure 5 A flowchart of the training process of the PPG prediction model provided in Example 1 of the present application;

[0044] Figure 6 A flowchart of the training process of the pitch prediction model and the energy prediction model provided in Example 1 of the present application;

[0045] Figure 7 A flowchart of the training process of the audio decoder provided in Embodiment 1 of the present application;

[0046] Figure 8 A flow chart of the training process of the lip decoder provided in the first embodiment of the present application;

[0047] Fig. 9 A schematic diagram of the structure of the speech synthesis and lip-actuating device provided in Example 2 of the present application;

[0048] Fig.10 This is a schematic diagram of the structure of an electronic device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] Embodiment 1:

[0051] This application embodiment provides a speech synthesis and lip-driven method, see Figure 1 As shown, the following steps may be included:

[0052] S11: Obtain phoneme sequence features.

[0053] S12: Generate audio PPG feature information based on the phoneme sequence features.

[0054] S13: Generate pitch feature information and energy feature information based on the audio PPG feature information and the pre-trained prediction model.

[0055] S14: Generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information.

[0056] S15: Determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information.

[0057] S16: Determine the corresponding lip animation based on the lip feature parameters.

[0058] S17: Play lip animation and audio.

[0059] The specific process of the above steps is described in detail below.

[0060] See also Figure 2 As shown, step S11 may include the following sub-steps:

[0061] S111: Determine the text information to be broadcast.

[0062] S112: Generate a corresponding phoneme sequence according to the text information.

[0063] S113: Encode the phoneme sequence to obtain corresponding phoneme sequence features.

[0064] It is understandable that the text information in step S111 may be input by the user, and the user may input the text content of the voice that the virtual image expects to emit. Of course, the text information may also be obtained by collecting the voice emitted by the user through an audio collection device and performing text recognition on the collected voice.

[0065] In sub-step S112, the text information may be parsed to obtain a phoneme sequence, such as a consonant, a vowel, and / or a phonetic symbol sequence.

[0066] In sub-step S113, the phoneme sequence may be encoded to obtain a phoneme vector, and then the phoneme vector may be encoded to obtain a phoneme sequence feature, that is, to obtain an implicit feature z1.

[0067] The audio PPG feature information in the embodiment of the present application is independent of the pronunciation timbre, volume and pitch of the sound, and is an implicit feature representing the probability of the pronunciation phoneme, see Figure 3 As shown, step S12 may include the following sub-steps:

[0068] S121: Determine voiceprint information for audio synthesis.

[0069] S122: Input the voiceprint information and phoneme sequence features into a pre-trained PPG prediction model to obtain corresponding audio PPG feature information; the PPG prediction model is a model trained based on audio training samples and corresponding audio text sequence training samples.

[0070] In sub-step S121, the target voice can be collected by an audio collection device and voiceprint information can be extracted from the target voice; or a timbre selection instruction can be received and corresponding voiceprint information can be selected from a preset voiceprint information library according to the timbre selection instruction, that is, the user can select the desired speaker.

[0071] The logic block diagram of the speech synthesis and lip-driven method in the embodiment of the present application can be found in Figure 4 As shown in the figure, the PPG prediction model can predict the phoneme sequence according to the voiceprint information and the phoneme sequence features, and expand the implicit feature z1 into the implicit feature z2 according to the phoneme sequence. The implicit feature z2 here is the audio PPG feature information predicted by the PPG prediction model.

[0072] In step S13, the voiceprint information and the audio PPG feature information may be input into a pre-trained pitch prediction model to obtain pitch feature information, and the voiceprint information and the audio PPG feature information may be input into a pre-trained energy prediction model to obtain energy feature information. It is understandable that the voiceprint information and the audio PPG feature information may also be input into a pre-trained pitch-volume prediction model, and the pitch feature information and the energy feature information may be directly output from this model.

[0073] The implicit feature z2 is superimposed with the pitch feature information and the energy feature information to obtain the implicit feature z3, that is, the superimposed audio feature information.

[0074] The pre-trained audio decoder decodes the implicit feature z3 and the corresponding voiceprint information to obtain the audio of the corresponding timbre, and the pre-trained lip shape decoder decodes the implicit feature z3 to obtain the corresponding lip shape feature parameters, and determines the corresponding lip shape animation according to the lip shape feature parameters.

[0075] In step S17, the lip animation and the audio may be played synchronously.

[0076] Since different people have different timbres, the PPG prediction model, pitch prediction model, energy prediction model and audio decoder in the embodiments of the present application are all embedded with voiceprint information, so that the corresponding PPG features, pitch features and energy features can be predicted for different people, and then audio with different timbres can be generated.

[0077] It should be noted that in other embodiments, the above-mentioned module may not embed voiceprint information, that is, it may directly predict the audio PPG feature information based on the phoneme sequence features, directly predict the pitch feature information and energy feature information based on the audio PPG feature information, and directly generate the corresponding audio based on the superimposed audio feature information.

[0078] For ease of understanding, the training process of the PPG prediction model is explained below.

[0079] See also Figure 5 As shown, the PPG prediction model is a model obtained by training with the error between the phoneme duration prediction feature and the PPG prediction feature as an additional loss, the phoneme duration prediction feature is a feature obtained by predicting the phoneme duration based on the voiceprint feature of the audio training sample and the phoneme sequence feature of the corresponding audio text sequence training sample, and the PPG prediction feature is a feature obtained by performing speech recognition processing on the audio training sample. That is, based on the FastSpeech2 model, the voiceprint feature of the audio training sample and the phoneme sequence feature of the corresponding audio text sequence training sample can be used to predict the phoneme duration to obtain the implicit feature z2, and the mean square error between the implicit feature z2 here and the feature obtained by performing speech recognition processing on the audio training sample is calculated as the loss of model training, so that the implicit feature z2 predicted by the model is constrained to be PPG information.

[0080] Similarly, the training process of the pitch prediction model and energy prediction model can be found in Figure 6 As shown, the training process of the audio decoder can be found in Figure 7 As shown, the training process of the lip decoder can be found in Figure 8 As shown, the specific process will not be repeated here.

[0081] It should be understood that although Figure 1-3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-3 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0082] The speech synthesis and lip-sync driving method provided in the embodiment of the present application jointly designs an audio-driven lip-sync model based on deep learning and a speech synthesis model, thereby avoiding the need to extract audio features from the audio when the audio drives the lip-sync, and the operating efficiency is higher; in addition, the audio is generated based on the voiceprint information, and the speech synthesizer replication of the audio timbre that drives the lip-sync is realized; the timbre consistent with the training set is used to drive the lip-sync, and the effect is more natural; the PPG prediction model is used to extract audio PPG feature information as an implicit feature, avoiding text annotation of lip-sync driven audio, and also realizing the separation of timbre and rhythm, that is, selecting the rhythm of a fully trained speaker and using lip-sync driven timbre for speech broadcasting.

[0083] Embodiment 2:

[0084] Based on the same inventive concept, the present application embodiment provides a speech synthesis and lip-actuating device, see Fig. 9 As shown, it should be understood that the specific functions of the speech synthesis and lip-actuating device can be found in the above description. To avoid repetition, the detailed description is appropriately omitted here.

[0085] The speech synthesis and lip-activated device includes at least one software functional unit that can be stored in a memory in the form of software or firmware or fixed in the operating system of the device. Specifically, the speech synthesis and lip-activated device includes:

[0086] An acquisition module 901 is used to acquire phoneme sequence features;

[0087] A first generating module 902, configured to generate audio PPG feature information based on the phoneme sequence feature;

[0088] A second generating module 903 is used to generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model;

[0089] A third generating module 904 is used to generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information;

[0090] A fourth generating module 905, configured to determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information;

[0091] A determination module 906, configured to determine a corresponding lip animation based on the lip feature parameters;

[0092] The playing module 907 is used to play the lip animation and the audio.

[0093] It should be noted that, for the sake of brevity, the contents described in the above embodiments will not be repeated in this embodiment.

[0094] Embodiment three:

[0095] This embodiment provides an electronic device. Fig.10 As shown, the electronic device includes a processor 1001 and a memory 1002, the memory 1002 stores a computer program, the processor 1001 and the memory 1002 communicate via a communication bus, and the processor 1001 executes the computer program to implement the steps of the power saving function control method in the above-mentioned embodiment 1, which will not be repeated here. It can be understood that Fig.10 The structure shown is for illustration only. The electronic device may also include Fig.10 More or fewer components as shown, or with Fig.10 It should be noted that the electronic device in the embodiment of the present application can be arranged on a car, for example, it can be an ECU on a car.

[0096] The processor 1001 may be an integrated circuit chip with signal processing capabilities. The processor 1001 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. It may implement or execute various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0097] The memory 1002 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc.

[0098] This embodiment also provides a computer-readable storage medium, such as a floppy disk, a CD, a hard disk, a flash memory, a U disk, an SD card, an MMC card, etc., in which one or more programs for implementing the above steps are stored. This one or more programs can be executed by one or more processors 301 to implement the steps of the power saving function control method in the above embodiment 1, which will not be repeated here.

[0099] It should be noted that the diagram provided in the present embodiment only illustrates the basic concept of the present invention in a schematic manner, so the diagram only shows the components related to the present invention rather than drawing according to the number, shape and size of the components during actual implementation. The type, quantity and ratio of each component during actual implementation can be a random change, and the component layout type may also be more complicated. The structure, ratio, size, etc. illustrated in the drawings of the present specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the limiting conditions that the present invention can implement, so they have no technical substantive significance. Any modification of the structure, change of the proportional relationship or adjustment of the size should still fall within the scope of the technical content disclosed by the present invention without affecting the effect that the present invention can produce and the purpose that can be achieved. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of narration, and are not used to limit the scope of the present invention. The change or adjustment of its relative relationship should also be regarded as the scope of the present invention without substantially changing the technical content.

[0100] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0101] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A speech synthesis and lip-activated method, characterized in that: include: Obtain phoneme sequence features; Generate audio PPG feature information based on the phoneme sequence feature; Generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model; Generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information; Determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information; Determine the corresponding lip animation based on the lip feature parameters; Play the lip animation and the audio.

2. The speech synthesis and lip-activated method according to claim 1, wherein: The obtaining of phoneme sequence features comprises: Determine the text information to be broadcasted; Generate a corresponding phoneme sequence according to the text information; The phoneme sequence is encoded to obtain corresponding phoneme sequence features.

3. The speech synthesis and lip-activated method according to claim 1, wherein: The generating audio PPG feature information based on the phoneme sequence feature includes: Determining voiceprint information for audio synthesis; The voiceprint information and the phoneme sequence features are input into a pre-trained PPG prediction model to obtain corresponding audio PPG feature information; the PPG prediction model is a model trained based on audio training samples and corresponding audio text sequence training samples.

4. The speech synthesis and lip-activated method according to claim 3, wherein: The PPG prediction model is a model obtained by training with the error between the phoneme duration prediction feature and the PPG prediction feature as an additional loss, the phoneme duration prediction feature is a feature obtained by predicting the phoneme duration based on the voiceprint feature of the audio training sample and the phoneme sequence feature of the corresponding audio text sequence training sample, and the PPG prediction feature is a feature obtained by performing speech recognition processing on the audio training sample.

5. The speech synthesis and lip-activated method according to claim 3, wherein: The determining of voiceprint information for audio synthesis includes: Collecting the target voice through an audio collection device, and extracting voiceprint information from the target voice; or, A timbre selection instruction is received, and corresponding voiceprint information is selected from a preset voiceprint information library according to the timbre selection instruction.

6. The speech synthesis and lip-activated method according to claim 3, wherein: The generating pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model includes: The voiceprint information and the audio PPG feature information are respectively input into a pre-trained pitch prediction model and a pre-trained energy prediction model to obtain pitch feature information and energy feature information.

7. The speech synthesis and lip-activated method according to claim 3, wherein: The generating corresponding audio according to the superimposed audio feature information includes: Generate audio of corresponding timbre according to the voiceprint information and the superimposed audio feature information.

8. A speech synthesis and lip-operating device, characterized in that: include: An acquisition module, used for acquiring phoneme sequence features; A first generating module, configured to generate audio PPG feature information based on the phoneme sequence feature; A second generating module is used to generate pitch feature information and energy feature information based on the audio PPG feature information and a pre-trained prediction model; A third generating module, configured to generate superimposed audio feature information according to the audio PPG feature information, the pitch feature information and the energy feature information; A fourth generating module, used to determine lip shape feature parameters according to the superimposed audio feature information, and generate corresponding audio according to the superimposed audio feature information; A determination module, used for determining a corresponding lip animation based on the lip feature parameters; A playing module is used to play the lip animation and the audio.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by at least one processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and apparatus for converting words into animation

    CN101482975A

  • Speech synthesis method, neural network model training method, and speech synthesis model

    CN114464162A