A method, system and storage medium for unconstrained lip-reading to speech synthesis

Through the combination of non-autoregressive end-to-end architecture and adversarial generation network, the direct synthesis of high-quality voice on unconstrained videos solves the problems of complex deployment and low quality in existing methods, achieving faster inference speed and higher voice quality.

CN114974206BActive Publication Date: 2025-05-16ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210677656.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-16
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

Existing unconstrained lip-to-speech synthesis methods have complex model deployment steps, speech quality degradation, inference speed, and audio quality limitations.

Method used

The visual encoder and acoustic encoder with a non-autoregressive end-to-end architecture are used, combined with a GAN-based vocoder for adversarial training, and voice is directly synthesized on unconstrained video.

Benefits of technology

Higher quality speech synthesis is achieved and significantly improved in Mel spectrum and audio inference speeds, 9.14 times faster and 19.76 times faster than the current state-of-the-art models on 3-second video-long datasets, respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974206B_ABST
    Figure CN114974206B_ABST
Patent Text Reader

Abstract

The present invention discloses an unconstrained lip reading to speech synthesis method, system and storage medium, belonging to the field of speech synthesis. A visual feature vector is extracted and encoded from a lip reading video sequence by a visual encoder; the length of the visual feature vector is adjusted to the length of the corresponding audio content to obtain a visual feature vector aligned with the corresponding audio content; the aligned visual feature vector is converted into a corresponding acoustic feature vector by an acoustic encoder; a corresponding Mel spectrum is generated according to the acoustic feature vector, and the visual encoder and the acoustic encoder are trained in combination with the real Mel spectrum; the parameters of the visual encoder and the acoustic encoder are fixed, and an audio generator is trained, and the acoustic feature vector is synthesized into an audio waveform using the trained audio generator, and converted into predicted audio. The present invention can synthesize higher quality speech directly on unconstrained videos at a faster reasoning speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech, and in particular to a method, system and storage medium for unconstrained lip language to speech synthesis. Background Art

[0002] The unconstrained lip-to-speech synthesis task aims to synthesize corresponding speech audio from a silent video of a speaker without head pose or vocabulary constraints. Current approaches mainly use sequence-to-sequence models to solve this problem, whether in autoregressive architectures or flow-based non-autoregressive architectures. However, these models have the following disadvantages:

[0003] (1) These models do not generate audio directly, but generate audio in two steps, first generating the Mel spectrum and then synthesizing audio from the Mel spectrum. This leads to complex model deployment steps and speech quality degradation due to error propagation.

[0004] (2) The audio reconstruction algorithms used by these models limit the inference speed and audio quality, and neural vocoders cannot be used for these models because their output spectrograms on unconstrained inputs are not accurate enough;

[0005] (3) Models based on autoregressive architectures have high inference latency, while models based on streaming architectures have high memory usage, which are not efficient in terms of time and memory usage. Summary of the invention

[0006] In view of the above problems, the present invention provides an unconstrained lip reading to speech synthesis method, system and storage medium, which can synthesize higher quality speech directly on unconstrained video at a faster reasoning speed.

[0007] To this end, the technical solution adopted in the present invention is as follows:

[0008] In a first aspect, the present invention provides a method for unconstrained lip reading to speech synthesis, comprising the following steps:

[0009] S1: extract and encode the visual feature vector from the lip reading video sequence through the visual encoder;

[0010] S2: adjusting the length of the visual feature vector obtained in step S1 to the length of the corresponding audio content, to obtain a visual feature vector aligned with the corresponding audio content;

[0011] S3: converting the aligned visual feature vector obtained in step S2 into a corresponding acoustic feature vector through an acoustic encoder;

[0012] S4: Generate a corresponding Mel spectrum according to the acoustic feature vector obtained in step S3, and train the visual encoder and the acoustic encoder in combination with the real Mel spectrum;

[0013] S5: Fix the parameters of the visual encoder and the acoustic encoder, train the audio generator, and use the trained audio generator to synthesize the acoustic feature vector obtained in step S3 into an audio waveform and convert it into predicted audio.

[0014] In a second aspect, the present invention provides an unconstrained lip-reading to speech synthesis system for implementing the above-mentioned unconstrained lip-reading to speech synthesis method.

[0015] In a third aspect, the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, is used to implement the above-mentioned unconstrained lip reading to speech synthesis method.

[0016] Compared with the prior art, the advantages of the present invention are:

[0017] The present invention proposes an end-to-end model for unconstrained lip synthesis, which adopts a non-autoregressive end-to-end architecture to effectively reduce the computing delay, and establishes a method for improving the audio quality by using a GAN-based vocoder for adversarial training. The results show that the speech synthesized by the model proposed in the present invention is of higher quality, and the Mel spectrum inference speed and audio inference speed are 9.14 times and 19.76 times faster than the current most advanced model on a dataset with a video duration of 3 seconds, respectively, achieving the goal of directly synthesizing higher quality speech with lower inference delay and smaller model size under unconstrained conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of the overall architecture of an unconstrained lip reading to speech synthesis method proposed according to an exemplary embodiment;

[0019] Figure 2 A comparison diagram of Mel spectrum inference speed according to an exemplary embodiment;

[0020] Figure 3 is a comparison diagram of audio reasoning speed proposed according to an exemplary embodiment;

[0021] Figure 4 The figure is a schematic diagram of a device terminal with data processing capability according to an exemplary embodiment. DETAILED DESCRIPTION

[0022] The present invention is further described below in conjunction with the accompanying drawings and embodiments. The accompanying drawings are only schematic diagrams of the present invention, and some of the block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities, and these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0023] This paper proposes for the first time a transformer-based visual encoder for encoding lip movements in the lip-reading to speech synthesis task, which has high performance on both large unconstrained datasets and small datasets; it proposes a non-autoregressive end-to-end acoustic encoder and an audio generator based on a generative adversarial network, which are used to directly synthesize higher quality speech with lower inference latency and smaller model size under unconstrained conditions.

[0024] The problem of unconstrained lip synthesis can be expressed as follows: Assume there is a lip reading video sequence V = {v1, v2, ..., v n}, where n represents the length of the video sequence, v i represents the i-th frame in the video sequence, v i and v i There may be a big difference, that is, the position of the speaker's head is not constrained. The task of lip synthesis is to generate the corresponding speech audio A = {a1, a2, ..., aL}, where L represents the length of the speech, a j represents the jth word in speech, without being restricted by a finite vocabulary.

[0025] T·sr=L·FPS

[0026] Among them, sr is the audio sampling rate and FPS is the video frame rate.

[0027] like Figure 1 As shown, the model used in the present invention is recorded as FastLTS, which is mainly composed of three parts: a visual encoder, an acoustic encoder module and an audio generator. The visual encoder is used to extract and encode visual features from the input lip reading video sequence, the acoustic encoder is used to convert the visual features into corresponding acoustic features, and the audio generator is used to synthesize audio waveforms based on the acoustic features. During the training phase, the FastLTS model also introduces an auxiliary Mel spectrum layer after the output layer of the acoustic encoder to pre-train the visual encoder and acoustic encoder modules.

[0028] Combination Figure 1 As shown, in a specific implementation of the present invention, the unconstrained lip reading to speech synthesis method mainly includes the following steps:

[0029] S1: A visual feature vector is extracted and encoded from an input video sequence through a visual encoding module; the visual encoder comprises a visual labeling layer, a spatial transformer, and a temporal transformer.

[0030] In this step, the visual marker layer is used to preliminarily extract the local features of the input video sequence and generate spatiotemporal markers to obtain the visual marker sequence T. After adding position embedding to the visual marker sequence T, it is used as the input of the spatial transformer; the spatial transformer is used to model the correlation between adjacent visual markers, and the linear approximate multi-head attention layer is used to reduce the computational burden of attention. For the feedforward part of the spatial transformer, a local enhanced feedforward network is used to increase the local feature modeling capability to obtain the spatially encoded visual marker sequence T′, which is linearly mapped to a low dimension and then positionally encoded to obtain the initial visual feature vector F′; the temporal transformer is used to temporally encode the visual feature vector F′ to obtain the final visual feature vector F.

[0031] Specifically, the implementation process of step S1 is:

[0032] S1-1: Input video sequence V = {v1, v2, ..., v n}, where v i represents the i-th frame in the video sequence, and n represents the length of the video sequence; the visual marker layer comprises a 3D convolution layer, a normalization layer, and a maximum pooling layer, which are used to preliminarily extract the local features of the video sequence and generate visual markers containing spatiotemporal information, and positionally encode the visual markers to obtain a visual marker sequence T = {t1, t2, ..., t n}, where t i Represents the visual label of the i-th frame in the video sequence;

[0033] S1-2: The spatial correlation between adjacent visual tags is encoded on the visual tag sequence T obtained in step S1-1 through a spatial transformer, where the spatial transformer includes two normalization layers, a linear similarity multi-head attention layer, and a local enhancement feedforward network to obtain a visual tag sequence T′;

[0034] S1-3: linearly map multiple hidden layers with the same temporal index in the visual tag sequence T′ obtained in step S1-2 into a low-dimensional single hidden layer, and perform position encoding to obtain a visual feature vector sequence F′;

[0035] S1-4: The visual feature vector sequence F′ obtained in step S1-3 is encoded with temporal correlation between hidden layers through a temporal transformer, which includes two residual connections and normalization layers, a multi-head self-attention layer, and a feedforward neural network to obtain the final visual feature vector F;

[0036] S2: adjusting the length of the final visual feature vector F obtained in step S1 to the length of the corresponding audio content through a length adjuster, so as to obtain a visual feature vector matching the corresponding audio content;

[0037] In this step, the main purpose is to align the length of the visual features to match the length of the acoustic features. Specifically, the implementation process of step S2 is as follows:

[0038] S2-1: Based on the length L of the audio feature sequence per second aud And the video frames per second FPS, calculate the adjustment factor d, the calculation formula is as follows:

[0039]

[0040] S2-2: If the adjustment factor d is an integer, copy the visual features of each video frame in the final visual feature vector F obtained in step S1 d times; if the adjustment factor d is not an integer, take L aud , the greatest common divisor of FPS is K, and the final visual feature vector is divided into K groups, each group The adjustment factor sequence of each group of visual feature vectors is d i represents the number of visual feature replications corresponding to the i-th video frame in each group, that is, each adjustment factor in the adjustment factor sequence corresponds to the visual feature of a video frame in the group, and the number of visual feature replications of the video frame corresponds to the value of the adjustment factor;

[0041] After the above adjustments, the aligned visual feature vector F is finally obtained.

[0042] In this embodiment, the adjustment factor sequence δ satisfies the following two conditions:

[0043] max(δ)-min(δ)≤1

[0044]

[0045] Wherein, max(δ) represents the maximum value in the adjustment factor sequence δ, min(δ) represents the minimum value in the adjustment factor sequence δ, and ∑δ represents the number of adjustment factors in the adjustment factor sequence.

[0046] For example, if the frame rate is 30fps and the length of the audio feature is 80 per second, the final visual feature vector can be divided into 10 groups, each group includes feature vectors of 3 video frames, and the adjustment factor sequence of each group of visual feature vectors is δ={3, 3, 2}.

[0047] S3: Convert the visual feature vector F obtained in step S2 into the corresponding acoustic feature vector through an acoustic encoder. The acoustic encoder contains two residual connections and normalization, a multi-head attention layer and a 1D convolution layer.

[0048] In this step, the acoustic encoder is used to convert the aligned visual feature vector F into an acoustic feature vector.

[0049] Since it is difficult to directly use the original audio as a monitoring signal to train the entire model in an end-to-end manner, the present invention proposes a two-stage training method.

[0050] In the first stage of training, the audio generator is not involved, and only the auxiliary Mel spectrum layer is used to train the visual encoder and acoustic encoder;

[0051] In this step, the visual encoder and acoustic encoder are optimized using SSIM loss and L1 loss according to the mel-spectrogram output by the auxiliary mel-spectrogram layer.

[0052] In the second stage of training, the auxiliary Mel spectrum layer is removed, and a linear projection layer is used to transform the acoustic feature vector obtained in step S3 into an acoustic feature vector of the same dimension as the Mel spectrum. The adversarial training is performed using the adversarial generative network. During the training process, the parameters of the visual encoder and the acoustic encoder are fixed, and only the audio generator parameters are optimized. The training objectives of the second stage include three parts: adversarial loss, Mel spectrum loss, and feature matching loss.

[0053] The two-stage training method includes the following steps S4 and S5.

[0054] S4: The acoustic feature vector obtained in step S3 is converted into a corresponding Mel spectrum through an auxiliary Mel spectrum layer to complete the training of the visual encoding module and the acoustic encoder;

[0055] In this step, the auxiliary Mel spectrum layer is only used for the pre-training process of the visual coding module and the acoustic encoder. Specifically, the implementation process of step S4 is as follows: S4-1: generating a Mel spectrum through the auxiliary Mel spectrum layer according to the visual feature vector matched with the corresponding audio content obtained in step S3;

[0056] S4-2: Continuously iteratively update the SSIM loss function during training And L1 loss function After completing the training of the visual encoder and acoustic encoder, the total loss function is The calculation formula is as follows:

[0057]

[0058]

[0059]

[0060] Among them, L mel Represents the length of the Mel spectrum; y i represents the true Mel spectrum of the i-th frame, represents the predicted Mel spectrum of the i-th frame, SSIM(.,.) represents the calculation of the structural similarity index between two vectors, λ SSIM , L1 is a hyperparameter, and ||.||1 represents the L1 norm.

[0061] S5: Remove the auxiliary Mel spectrum layer and replace it with an audio generator. The audio generator generates the corresponding speech from the acoustic feature vector obtained in step S3. The audio generator includes a linear projection layer and a generative adversarial network.

[0062] In this step, the linear projection layer and the auxiliary Mel spectrum layer have the same dimension, and are used to project the acoustic feature vector obtained in step S3 and convert it into an acoustic feature vector of the same dimension as the Mel spectrum; the generator in the adversarial generative network is used to synthesize an audio waveform according to the converted acoustic feature vector.

[0063] Specifically, the implementation process of step S5 is:

[0064] S5-1: After completing the training of the visual encoder and the acoustic encoder, remove the auxiliary Mel-spectrogram layer and replace it with the audio generator;

[0065] S5-2: transform the acoustic feature vector obtained in step S3 into an acoustic feature with the same dimension as the Mel spectrum through a linear projection layer to obtain an acoustic feature vector A;

[0066] S5-3: The acoustic feature vector A obtained in step S5-2 is used to synthesize speech through a generative adversarial network, where the generative adversarial network includes a discriminator D and a generator G.

[0067] The adversarial loss calculation formula of the discriminator D and the generator G during training is as follows:

[0068]

[0069]

[0070] Among them, s is the predicted audio, x is the real audio, represents the expectation, D(.) represents the discriminator, G(.) represents the generator, represents the adversarial training loss of the discriminator D, represents the adversarial training loss of the generator G.

[0071] The calculation formula of Mel spectrum loss is as follows:

[0072]

[0073] Here, φ(·) represents the function for converting audio into Mel spectrum.

[0074] The feature matching loss calculation formula is as follows:

[0075]

[0076] Among them, T represents the number of layers of the discriminator D, D i represents the features of the i-th layer of the discriminator, N i Represents the number of features in the i-th layer of the discriminator.

[0077] The total loss during training is:

[0078]

[0079]

[0080] Among them, λ a , m , f is a hyperparameter, represents the total generator G loss, represents the total discriminator D loss.

[0081] In a specific implementation of the present invention, it is found during the training process that calculating the loss of all audio clips will take up a lot of computing resources and memory capacity. Therefore, a window sampling mechanism is used in the second stage audio generator training process to sample a continuous subsequence from the acoustic feature vector obtained in step S3 and use the corresponding real audio clip for training. This method has a good effect on small data sets, but poor effect on large unconstrained data sets. To address this problem, the present invention first completes the training of the visual encoder and the acoustic encoder in step S4, and only trains the audio generator in step S5. The advantage of doing this is that it can improve the performance of the model for small data sets and large unconstrained data sets.

[0082] In this embodiment, an unconstrained lip-reading to speech synthesis system is also provided, which is used to implement the above embodiment. The terms "module", "unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, it is also possible to implement hardware, or a combination of software and hardware.

[0083] The system comprises:

[0084] A visual encoding module, which is used to extract and encode visual feature vectors from lip reading video sequences;

[0085] A length adjustment module, which is used to adjust the length of the visual feature vector to the length of the corresponding audio content, so as to obtain a visual feature vector aligned with the corresponding audio content;

[0086] An acoustic encoding module, which is used to convert the aligned visual feature vector into a corresponding acoustic feature vector;

[0087] An audio generation module, which is used to synthesize an audio waveform according to the acoustic feature vector and convert it into a predicted audio output;

[0088] An auxiliary training module, which is used to generate a corresponding Mel spectrum according to the acoustic feature vector, and train the visual encoder and the acoustic encoder in combination with the real Mel spectrum;

[0089] The secondary training module is used to fix the parameters of the visual encoder and the acoustic encoder and train the audio generator.

[0090] The implementation process of the functions and effects of each module in the above system is specifically described in the implementation process of the corresponding steps in the above method. For example, a specific process of the above system may be:

[0091] (1) extracting and encoding a visual feature vector from the lip reading video sequence through a visual encoder;

[0092] (2) adjusting the length of the visual feature vector to the length of the corresponding audio content to obtain a visual feature vector aligned with the corresponding audio content;

[0093] (3) Convert the aligned visual feature vector into the corresponding acoustic feature vector through an acoustic encoder;

[0094] (4) generating a corresponding Mel spectrum according to the acoustic feature vector, and training the visual encoder and the acoustic encoder in combination with the real Mel spectrum;

[0095] (5) Fix the parameters of the visual encoder and the acoustic encoder, train the audio generator, and use the trained audio generator to synthesize the acoustic feature vector into an audio waveform and convert it into predicted audio.

[0096] For the system embodiment, since it basically corresponds to the method embodiment, the specific implementation of each step can refer to the description of the method part, which will not be repeated here. The system embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0097] The embodiments of the unconstrained lip-reading to speech synthesis system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the processor of any device with data processing capabilities reads the corresponding computer program instructions in the non-volatile memory into the internal memory and runs them. From the hardware level, if Figure 4 As shown, it is a hardware structure diagram provided by this embodiment, except Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the system in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0098] The embodiment of the present invention further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned unconstrained lip reading to speech synthesis method is implemented.

[0099] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0100] The technical effect of the present invention is verified by experiments below.

[0101] In this embodiment, the Lip2Wav dataset and the GRID dataset are used. The audio sampling rate is 16000 Hz, the video duration is 3 seconds, the sampling window size is 800, the contract length is 200, and the Mel spectrum length is 80. The video frame size is 96×96.

[0102] This embodiment adopts the model configuration: the convolution kernel size of the 3D convolution layer of the tokenization layer is 5×5×5. The token dimension is 32, the number of spatial transformers is 4, the hidden layer dimension is 36, and the number of attention heads is 6. The number of temporal transformers is 4, the number of attention heads is 8, the hidden layer dimension is 384 when processing the Lip2Wav dataset, and the hidden layer dimension is 160 when processing the GRID dataset. The acoustic encoder is configured with a temporal transformer. The adam optimizer is used for optimization training, the learning rate of step S4 is 0.002, and the learning rate of step S5 is 0.0002.

[0103] This embodiment uses the subjective evaluation method MOS and the objective evaluation method PESQ to evaluate the performance of the FastLTS model of the present invention. The experimental results are shown in the following table.

[0104] Table 1 Statistics of MOS scores on the Lip2Wav dataset

[0105]

[0106] Table 2 MOS score statistics on the GRID dataset

[0107]

[0108] It can be seen from Table 1 and Table 2 that the FastLTS model proposed in the present invention has excellent performance in speech generation quality, clarity and naturalness, whether on the Lip2Wav dataset or the GRID dataset, which also proves the superiority of the visual encoder, acoustic encoder and audio generator proposed in the present invention.

[0109] Table 3 PESQ score statistics on the GRID dataset

[0110]

[0111] From Table 3, we can see that the PESQ score of the FastLTS model proposed in the present invention is only 0.067 lower than that of the most advanced VCA-GAN. It can be considered that the FastLTS model is one of the most advanced models in the lip-reading to speech synthesis task.

[0112] from Figure 2 , Figure 3 It can be seen that as the video length increases, the Mel spectrum and audio inference speed of the Lip2Wav model increases dramatically, which shows that the Lip2Wav model is not suitable for processing unconstrained large data sets. The FastLTS model proposed in the present invention has better performance when the video is long. When the video length is 3 seconds, the Mel spectrum inference speed and audio inference speed are 9.14 times and 19.76 times faster than the Lip2Wav model. This is because the FastLTS model uses a non-autoregressive end-to-end architecture that can perform parallel predictions.

[0113] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.

Claims

1. A method for unconstrained lip reading to speech synthesis, characterized in that: The steps include: S1: extract and encode the visual feature vector from the lip reading video sequence through the visual encoder; The visual encoder includes a visual marking layer, a spatial transformer and a temporal transformer; the step S1 includes: S1-1: Obtain lip reading video sequence V = {v1, v2, ..., v n }, where v i represents the i-th frame in the video sequence, and n represents the length of the video sequence. The local features of the lip reading video sequence V are extracted using the visual marker layer, and visual markers containing spatiotemporal information are generated. The visual markers are positionally encoded to obtain a visual marker sequence T = {t1, t2, ..., t n }, where t i Represents the visual label of the i-th frame in the video sequence; S1-2: Use the spatial transformer to encode the spatial correlation between adjacent visual markers in the visual marker sequence T obtained in step S1-1 to obtain the spatially encoded visual marker sequence T ′ ; S1-3: The spatially encoded visual marker sequence T obtained in step S1-2 is ′ Multiple hidden layers with the same temporal index in are linearly mapped into a low-dimensional single hidden layer, and position encoding is performed to obtain the visual feature vector F ′ ; S1-4: The visual feature vector F obtained in step S1-3 is transformed through the temporal transformer ′ Perform temporal correlation coding, and use the visual feature vector after temporal coding as the final visual feature vector F; S2: adjusting the length of the visual feature vector obtained in step S1 to the length of the corresponding audio content, to obtain a visual feature vector aligned with the corresponding audio content; S3: converting the aligned visual feature vector obtained in step S2 into a corresponding acoustic feature vector through an acoustic encoder; S4: Generate a corresponding Mel spectrum according to the acoustic feature vector obtained in step S3, and train the visual encoder and the acoustic encoder in combination with the real Mel spectrum; S5: Fix the parameters of the visual encoder and the acoustic encoder, train the audio generator, and use the trained audio generator to synthesize the acoustic feature vector obtained in step S3 into an audio waveform and convert it into predicted audio.

2. The unconstrained lip reading to speech synthesis method according to claim 1, characterized in that: The step S2 comprises: S2-1: Based on the length L of the audio feature sequence per second aud And the video frames per second FPS, calculate the adjustment factor d, the calculation formula is as follows: S2-2: If the adjustment factor d is an integer, the visual features of each video frame in the final visual feature vector F obtained in step S1 are copied d times, that is, the adjustment factor sequence is δ = {d, d, …, d}; If the adjustment factor d is not an integer, then take L aud , the greatest common divisor of FPS is K, and the final visual feature vector is divided into K groups, each group The adjustment factor sequence of each group of visual feature vectors is d i represents the number of visual feature replications corresponding to each group of i-th video frames; S2-3: Adjust the length of the visual feature vector obtained in step S1 according to the adjustment factor sequence to align with the corresponding audio content, and finally obtain the aligned visual feature vector 3. The unconstrained lip-reading to speech synthesis method according to claim 2, characterized in that: If the adjustment factor d is not an integer, the adjustment factor sequence should satisfy: max(δ)-min(δ)≤1 Wherein, max(δ) represents the maximum value in the adjustment factor sequence δ, min(δ) represents the minimum value in the adjustment factor sequence δ, and ∑δ represents the number of adjustment factors in the adjustment factor sequence.

4. The unconstrained lip reading to speech synthesis method according to claim 1, characterized in that: The step S4 comprises: S4-1: Generate a mel spectrum through an auxiliary mel spectrum layer according to the visual feature vector matched with the corresponding audio content obtained in step S3; S4-2: Combine the generated Mel spectrum and the real Mel spectrum, iteratively update the structural similarity SSIM loss function and the L1 loss function, and complete the training of the visual encoder and the acoustic encoder; the total loss function is the weighted sum of the SSIM loss function and the L1 loss function.

5. The unconstrained lip reading to speech synthesis method according to claim 1, characterized in that: The audio generator includes a linear projection layer and a generative adversarial network; The step S5 comprises: S5-1: Fix the parameters of the trained visual encoder and acoustic encoder, and replace the auxiliary Mel-spectrogram layer with the audio generator; S5-2: transform the acoustic feature vector obtained in step S3 into an acoustic feature with the same dimension as the Mel spectrum through a linear projection layer to obtain an acoustic feature vector A; S5-3: The acoustic feature vector A obtained in step S5-2 is used to synthesize speech through a generative adversarial network, wherein the generative adversarial network includes a discriminator D and a generator G; during the training process, the training loss includes a weighted sum of three parts: adversarial loss, mel-spectrogram loss, and feature matching loss.

6. The unconstrained lip reading to speech synthesis method according to claim 1, characterized in that: A window sampling mechanism is used in the audio generator training process to sample a continuous subsequence from the acoustic feature vector obtained in step S3 and use the corresponding real audio segment for training.

7. An unconstrained lip reading to speech synthesis system, used to implement the unconstrained lip reading to speech synthesis method according to any one of claims 1 to 6, characterized in that: The system comprises: A visual encoding module, which is used to extract and encode visual feature vectors from lip reading video sequences; The visual encoder includes a visual labeling layer, a spatial transformer and a temporal transformer; the calculation process of the visual encoding module includes: Get the lip reading video sequence V = {v1, v2, ..., v n }, where v i represents the i-th frame in the video sequence, and n represents the length of the video sequence. The local features of the lip reading video sequence V are extracted using the visual marker layer, and visual markers containing spatiotemporal information are generated. The visual markers are positionally encoded to obtain a visual marker sequence T = {t1, t2, ..., t n }, where t i Represents the visual label of the i-th frame in the video sequence; The spatial correlation between adjacent visual markers is encoded by the spatial transformer to obtain the spatially encoded visual marker sequence T. ′ ; The obtained spatially encoded visual tag sequence T ′ Multiple hidden layers with the same temporal index in are linearly mapped into a low-dimensional single hidden layer, and position encoding is performed to obtain the visual feature vector F ′ ; Through the temporal transformer, the visual feature vector F ′ Perform temporal correlation coding, and use the visual feature vector after temporal coding as the final visual feature vector F; A length adjustment module, which is used to adjust the length of the visual feature vector to the length of the corresponding audio content, so as to obtain a visual feature vector aligned with the corresponding audio content; An acoustic encoding module, which is used to convert the aligned visual feature vector into a corresponding acoustic feature vector; An audio generation module, which is used to synthesize an audio waveform according to the acoustic feature vector and convert it into a predicted audio output; An auxiliary training module, which is used to generate a corresponding Mel spectrum according to the acoustic feature vector, and train the visual encoder and the acoustic encoder in combination with the real Mel spectrum; The secondary training module is used to fix the parameters of the visual encoder and the acoustic encoder and train the audio generator.

8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, it is used to implement the unconstrained lip reading to speech synthesis method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech synthesis and feature extraction model training method and device, medium and equipment

    CN111883107A

  • Rapid lip movement-voice alignment method based on parallel flow model

    CN113852851A