Voice synthesis method and device fusing visual information, electronic equipment and medium
Through the speech synthesis method that integrates visual information, high-quality emotional speech is generated using video information, which solves the problem of difficulty in collecting data during modeling training of specific characters in the prior art, and achieves the efficiency and quality of emotional speech synthesis.
Patent Information
- Application Number
- CN202510579001.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When training on modeling and training specific characters, existing speech synthesis technology faces the problems of difficulty in collecting data and low data quality, which leads to major challenges in emotional speech synthesis technology.
By fusing visual information, the information in the video is used to generate speech. The specific steps include extracting text features based on the text encoder, performing feature extraction, vector quantization encoding and cross-attention processing on the video information, combining text and video features for cross-attention processing, inputting it into a pre-trained speech generation model for modeling and decoding, and generating synthetic speech corresponding to the target text information.
This method not only solves the shortcomings of the prior art in emotional control, but also uses a small amount of sample data to achieve high-quality speech synthesis, and the generated emotional characteristics are consistent with the emotional characteristics of the target video information.
Smart Images

Figure CN120089122A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech synthesis, and in particular to a speech synthesis method, device, electronic device and medium that integrate visual information. Background Technique
[0002] At the current stage, under the condition of collecting a large amount of data for training, the speech synthesis technology can already approach the level of real people. However, when training for a specific person's model, there are often problems such as difficulty in collecting the data of this person and low data quality, and there is no other emotional speech in the collected specific person's data. In such a situation, there are still great challenges in emotional speech synthesis technology. The existing speech synthesis technology relies on a large amount of high-quality data recorded in a recording studio, and different people need to record a large amount of data of the corresponding person. In addition, if you want to achieve emotional speech synthesis of a specific person, you need to record a large amount of speech data of the specific person under different emotions in advance, so it will consume a lot of energy. Since the types of emotions synthesized are limited by the recorded data, in the case of only a very small amount of sample data of a specific person, it is difficult for the existing speech to model this person, and it is difficult to form a high-quality emotion-controlled speech synthesis ability. Summary of the Invention
[0003] In view of this, the purpose of the present application is to provide a speech synthesis method, device, electronic device and medium that integrate visual information, and generate more expressive speech by integrating the information in the video. This method not only solves the deficiencies of the existing technology in terms of emotion control, but also can achieve high-quality speech synthesis using a small amount of sample data.
[0004] The embodiment of the present application provides a speech synthesis method that integrates visual information. The speech synthesis method includes: Extracting the text features of the target text information based on a text encoder; Performing feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine the video features in the target video information; Performing cross-attention processing on the text features and the video features to determine the joint features; Inputting the joint features into a pre-trained speech generation model, performing modeling processing on the joint features to generate video text features, and then performing random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate the synthetic speech corresponding to the target text information; wherein, the emotion feature of the synthetic speech is consistent with the emotion feature of the target video information.
[0005] In a possible implementation manner, the performing feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine the video features in the target video information includes: Mark the target video information based on a vector quantization encoder, and add random noise to the marked target video information; Input the target video information with added noise into a fusion diffusion model for sampling, and extract the first visual feature map of each video frame; Extract features of the target video information based on a multimodal neural network model to obtain the visual feature map of each video frame, and perform multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map with the same dimension as the first visual feature map; Combine and perform position encoding processing on the first visual feature map and the corresponding second visual feature map of each video frame to determine the initial video feature of each video frame; Perform mean processing on multiple frames of the initial video features to generate the video features.
[0006] In a possible implementation manner, the modeling processing of the joint feature to generate video text features, and then the random prosody prediction processing, feature augmentation processing, and decoding processing of the video text features to generate the synthesized speech corresponding to the target text information include: Capture the cross-modal relationship between the text feature and the video feature by performing modeling processing on the joint feature based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder of the speech generation model to generate video text features; Input the video text features into the random prosody predictor of the speech generation model, perform random prosody prediction processing and feature augmentation processing on the video text features to generate the augmented video text features; Input the duration of phonemes, the target person's speech embedding vector, and the augmented video text features in the video text features into the decoder of the speech generation model for speech synthesis processing to generate the synthesized speech corresponding to the target text information.
[0007] In a possible implementation manner, the capturing of the cross-modal relationship between the text feature and the video feature by performing modeling processing on the joint feature based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder to generate video text features includes: Process the joint feature based on the self-attention mechanism to generate the joint feature after self-attention processing; Perform positional encoding processing on the temporal order information of video frames and the positional information of target text information among the jointly processed features after self-attention processing, to generate the jointly processed features after positional encoding processing; Based on the multi-head attention mechanism, perform multi-angle modeling processing on the jointly processed features after positional encoding processing to generate video-text features.
[0008] In a possible implementation manner, input the video-text features into the random prosody predictor of the speech generation model, perform random prosody prediction processing and high-dimensional feature expansion processing on the video-text features, to generate the expanded video-text features, including: Based on the emotion features and context relationships included in the video features in the video-text features, predict the duration of each phoneme in the video-text features; Based on the features corresponding to each phoneme in the video-text features, perform feature interpolation processing according to the corresponding duration to generate the expanded video-text features.
[0009] In a possible implementation manner, determine the speech generation model through the following steps: Replace the text encoder in the original VITS model with a video-text joint encoder to obtain an initial speech generation model; Input the sample joint features determined based on the sample video information and sample text information into the initial speech generation model for speech synthesis processing to generate the predicted synthesized speech corresponding to the sample text information; Based on the predicted synthesized speech and the actual synthesized speech of the sample text information, determine the loss value of the initial speech generation model; Based on the loss value and training samples, perform iterative training on the initial speech generation model to generate the speech generation model.
[0010] The embodiments of the present application further provide a speech synthesis device integrating visual information, and the speech synthesis device includes: A text feature extraction module, configured to extract text features of target text information based on a text encoder; A video feature extraction module, configured to perform feature extraction processing, vector quantization encoding processing, and cross-attention processing on target video information to determine video features in the target video information; A joint module, configured to perform cross-attention processing on the text features and the video features to determine joint features; A speech generation module, configured to input the joint features into a pre-trained speech generation model, perform modeling processing on the joint features to generate video text features, and then perform random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate a synthetic speech corresponding to the target text information; wherein, the emotion feature of the synthetic speech is consistent with the emotion feature of the target video information.
[0011] In a possible implementation manner, when the video feature extraction module is configured to perform feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine the video features in the target video information, the video feature extraction module is further configured to: Mark the target video information based on a vector quantization encoder, and add random noise to the marked target video information; Input the target video information with added noise into a diffusion model for sampling, and extract a first visual feature map of each video frame; Perform feature extraction processing on the target video information based on a multi-modal neural network model to obtain a visual feature map of each video frame, and perform multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map with the same dimension as the first visual feature map; Perform combination processing and position encoding processing on the first visual feature map of each video frame and the corresponding second visual feature map to determine the initial video feature of each video frame; Perform mean processing on multiple frames of the initial video features to generate the video features.
[0012] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the speech synthesis method for fusing visual information as described above are executed.
[0013] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the speech synthesis method for fusing visual information as described above are executed.
[0014] The voice synthesis method, device, electronic device and medium integrating visual information provided by the embodiments of the present application, the voice synthesis method includes: extracting text features of target text information based on a text encoder; performing feature extraction processing, vector quantization coding processing and cross-attention processing on target video information to determine video features in the target video information; performing cross-attention processing on the text features and the video features to determine joint features; inputting the joint features into a pre-trained voice generation model, performing modeling processing on the joint features to generate video text features, and then performing random prosody prediction processing, feature augmentation processing and decoding processing on the video text features to generate a synthesized voice corresponding to the target text information; wherein, the emotion feature of the synthesized voice is consistent with the emotion feature of the target video information. By integrating the information in the video, a more expressive voice is generated. This method not only solves the deficiencies of the prior art in terms of emotion control, but also can achieve high-quality voice synthesis using a small amount of sample data.
[0015] To make the above objects, features and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0017] Figure 1 One of the flowcharts of a voice synthesis method integrating visual information provided by the embodiments of the present application; Figure 2 Another flowchart of a voice synthesis method integrating visual information provided by the embodiments of the present application; Figure 3 One of the structural diagrams of a voice synthesis device integrating visual information provided by the embodiments of the present application; Figure 4 Another structural diagram of a voice synthesis device integrating visual information provided by the embodiments of the present application; Figure 5 The structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without creative efforts belongs to the scope of protection of this application.
[0019] First, the applicable application scenarios of this application will be introduced. This application can be applied to the field of speech synthesis technology.
[0020] Through research, it is found that at the current stage, speech synthesis technology can already approach the level of real people under the condition of collecting a large amount of data for training. However, when training for a specific person's model, there are often problems such as difficulty in collecting the data of this person and low data quality, and there is no other emotional speech in the collected specific person's data. In such a situation, there are still great challenges in emotional speech synthesis technology. Existing speech synthesis technology relies on a large amount of high-quality data recorded in a recording studio, and different people need to record a large amount of data of the corresponding person. In addition, to achieve emotional speech synthesis of a specific person, a large amount of speech data of the specific person in different emotions needs to be recorded in advance, so it will consume a lot of energy. Since the types of synthesized emotions are limited by the recorded data, in the case of only a very small amount of sample data of a specific person, it is difficult for existing speech to model this person and it is difficult to form a high-quality emotion-controllable speech synthesis ability.
[0021] Based on this, the embodiments of this application provide a speech synthesis method that fuses visual information, which generates more expressive speech by fusing the information in the video. This method not only solves the deficiencies of the existing technology in terms of emotion control, but also can achieve high-quality speech synthesis using a small amount of sample data.
[0022] Please refer to Figure 1 , Figure 1 which is one of the flowcharts of a speech synthesis method that fuses visual information provided by the embodiments of this application. As shown in Figure 1 , the speech synthesis method provided by the embodiments of this application includes: S101: Extract the text features of the target text information based on a text encoder.
[0023] In this step, first convert the target text information into phoneme information, and use a text editor to perform text feature encoding processing on the phoneme information to obtain the text features of the target text information.
[0024] S102: Perform feature extraction processing, vector quantization coding processing, and cross-attention processing on the target video information to determine the video features in the target video information.
[0025] In this step, perform feature extraction processing, vector quantization coding processing, and cross-attention processing on the target video information to determine the video features in the target video information. In a possible implementation manner, the performing feature extraction processing, vector quantization coding processing, and cross-attention processing on the target video information to determine the video features in the target video information includes: A: Based on the vector quantization encoder, mark the target video information, and add random noise to the marked target video information.
[0026] Here, perform quantization marking on the target video information according to the vector quantization encoder to obtain discrete vectors, and add random noise to the target video information after quantization marking.
[0027] B: Input the target video information after adding noise into the fusion diffusion model for sampling, and extract the first visual feature map of each video frame.
[0028] Here, input the target video information after adding noise into the fusion diffusion model for sampling, and extract the first visual feature map of each video frame from the second upsampling module of the fusion diffusion model.
[0029] Among them, the fusion diffusion model defines a Markov chain of diffusion steps, gradually adds random noise to the data, and then learns the inverse diffusion process to construct the required data samples from the noise. The fusion diffusion model is a latent variable model, which uses a fixed Markov chain to map to the latent space and can generate corresponding video features from audio features through the diffusion process. This system designs a fusion diffusion model to model the relationship between audio modality information and video modality information to achieve the purpose of generating corresponding video information from audio information.
[0030] C: Perform feature extraction processing on the target video information based on the multimodal neural network model to obtain the visual feature map of each video frame, and perform multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map with the same dimension as the first visual feature map.
[0031] Here, perform feature extraction processing on the target video information according to the multimodal neural network model to obtain the visual feature map of each video frame, and perform multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map with the same dimension as the first visual feature map.
[0032] Among them, the multi-modal neural network model is the CLIP model.
[0033] D: Combine and process the first visual feature map of each video frame and the corresponding second visual feature map, and perform position encoding processing to determine the initial video feature of each video frame.
[0034] Here, the first visual feature map of each video frame and the corresponding second visual feature map are combined and processed to obtain a combined feature map. A set of learnable position encodings are added to the combined feature map to further enhance the localization ability, and the initial video feature of each video frame is obtained.
[0035] E: Perform mean processing on multiple frames of the initial video features to generate the video features.
[0036] In a specific embodiment, Step 1: Train a fusion diffusion model based on a diffusion model, and retain and freeze the weights of the U-Net therein; Step 2: Use the pre-trained CLIP model and retain and freeze its parameters; Step 3: Tokenize the target video information through a vector quantization (VQ) encoder, add random noise, and input it into the U-Net model saved in Step 1. Extract the visual feature map from the second upsampling block in the U-Net; Step 4: Use the CLIP model in Step 2 to extract visual features, project them through a multi-layer perceptron, and keep the dimension consistent with the dimension of the visual features extracted in Step 3. Finally, the CLIP visual feature map and the feature map extracted by the U-Net model are combined together and a set of learnable position encodings are added to further enhance the localization ability, and finally form the video features. Step 5: Extract visual features for each video frame image of the input video through Steps 3 and 4, and finally perform averaging to obtain the video features of the input video.
[0037] S103: Perform cross-attention processing on the text features and the video features to determine the joint features.
[0038] In this step, cross-attention processing is performed on the text features and the video features to determine the joint features.
[0039] Here, the joint features of the fused video features have high expressiveness and can generate emotion features such as sadness, happiness, and joy.
[0040] S104: Input the joint features into a pre-trained speech generation model, perform modeling processing on the joint features to generate video text features, and then perform random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate the synthetic speech corresponding to the target text information.
[0041] In this step, the joint features are input into a speech generation model, which models the joint features to generate video text features. Then, random prosody prediction processing, feature augmentation processing, and decoding processing are performed on the video text features to generate the synthetic speech corresponding to the target text information.
[0042] Among them, the text information of the synthetic speech is the target text information, and the emotional feature of the synthetic speech is the emotional feature of the target video information. Input the text to be inferred and the corresponding video information, and output the audio corresponding to the text to be inferred. This audio has the information presented by the video, such as happy emotions and sad emotions, etc.
[0043] Here, during the process of inputting the joint features into the pre-trained speech generation model, the speaker embedding vector of the synthetic speech needs to be determined, so as to generate the speech synthesis of the target text information according to the voice feature of the speaker embedding vector carrying the emotional feature of the target video information. Among them, the speaker embedding vector can be screened in the speaker database.
[0044] In a possible implementation manner, the modeling the joint features to generate video text features, and then performing random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate the synthetic speech corresponding to the target text information includes: (1): Based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder of the speech generation model, model the joint features to capture the cross-modal relationship between the text features and the video features, and generate video text features.
[0045] Here, based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder of the speech generation model, model the joint features to capture the cross-modal relationship between the text features and the video features, and generate video text features.
[0046] Among them, the video text joint encoder can be a transformer network layer.
[0047] In a possible implementation manner, the modeling the joint features based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder to capture the cross-modal relationship between the text features and the video features and generate video text features includes: a: Process the joint features based on the self-attention mechanism to generate the joint features after self-attention processing.
[0048] Here, the combined features are processed according to the self-attention mechanism to generate the combined features after self-attention processing.
[0049] b: Perform position encoding processing on the position information of the temporal order information of the video frames and the target text information among the combined features after self-attention processing to generate the combined features after position encoding processing.
[0050] Here, perform position encoding processing on the position information of the temporal order information of the video frames and the target text information among the combined features after self-attention processing to generate the combined features after position encoding processing.
[0051] c: Perform multi-angle modeling processing on the combined features after position encoding processing based on the multi-head attention mechanism to generate video-text features.
[0052] Here, use the multi-head attention mechanism to perform multi-angle modeling processing on the combined features after position encoding processing to generate video-text features.
[0053] (2): Input the video-text features into the random rhythm predictor of the speech generation model to perform random rhythm prediction processing and feature augmentation processing on the video-text features to generate the augmented video-text features.
[0054] Here, input the video-text features into the random rhythm predictor of the speech generation model to perform random rhythm prediction processing and feature augmentation processing on the video-text features to generate the augmented video-text features.
[0055] In a possible implementation manner, inputting the video-text features into the random rhythm predictor of the speech generation model to perform random rhythm prediction processing and high-dimensional feature augmentation processing on the video-text features to generate the augmented video-text features includes: i: Predict the duration of each phoneme in the video-text features based on the emotion features and context relationships included in the video features in the video-text features.
[0056] Here, predict the duration of each phoneme in the video-text features according to the emotion features and context relationships included in the video features.
[0057] Here, the core of the random prosody predictor is a regression model used to predict the duration of each phoneme. The specific steps are as follows: The video text features are non-linearly transformed through a multi-layer perceptron (MLP) to obtain intermediate features, and a probability distribution is modeled for the duration of each phoneme in the intermediate features. Assume that the duration of each phoneme follows a normal distribution, and a specific duration is randomly sampled from the normal distribution. This process introduces randomness, making the generated speech have diverse rhythms. The random prosody predictor not only depends on the text features but also combines the emotional features in the video features. The specific impact is as follows: The emotional information in the video features will affect the output of the random duration predictor. For example, if the person in the video shows "excitement", the predicted rhythm may be faster, resulting in a shorter duration of phonemes. If the person in the video shows "sadness", the predicted rhythm may be slower, resulting in a longer duration of phonemes.
[0058] ii: Perform feature interpolation processing on the features corresponding to each phoneme in the video text features according to the corresponding duration to generate the augmented video text features.
[0059] Here, feature interpolation processing is performed on the features corresponding to each phoneme in the video text features according to the corresponding duration to generate the augmented video text features.
[0060] (3): Input the duration of phonemes, the target person's speech embedding vector, and the augmented video text features in the video text features into the decoder of the speech generation model for speech synthesis processing to generate the synthetic speech corresponding to the target text information.
[0061] Here, the duration of phonemes, the target person's speech embedding vector, and the augmented video text features in the video text features are input into the decoder of the speech generation model for decoding processing of discrete vectors to generate the synthetic speech corresponding to the target text information.
[0062] Among the specific embodiments, the video-text joint encoder plays a key role in this application. It fuses video features and text features to form a "joint feature" that combines the information of both, thereby guiding the subsequent speech synthesis process. The following is a detailed analysis of its specific working principle: The video features are extracted from the input video and contain expressive information such as emotions and tones. These features are extracted through pre-trained CLIP models and diffusion models, which can capture the emotional states (such as happiness and sadness) and tone characteristics in the video. The text features are extracted from the input text and represent the content and structural information of the text. First, the text is converted into a phoneme sequence, and then a high-dimensional vector representation is generated through a text encoder to describe the meaning of the text. To effectively combine the video features and text features, a cross-attention module is used to obtain a joint feature that contains both the content information of the text and the emotional and tone information of the video. The obtained joint feature is further input into a multi-layer perceptron (MLP) for non-linear transformation to enhance the expressive power of the feature. The feature processed by the MLP is then input into the video-text joint encoder part of the speech generation model for processing to obtain the video-text feature. In the speech generation model, the duration of each phoneme is predicted based on the video-text feature, thereby controlling the rhythm and intonation of the speech. The video-text feature is expanded into a higher-dimensional representation to match the resolution of the audio. The expanded feature is input into the decoder to generate the final audio waveform. In this application, the video-text joint encoder effectively fuses the video features and text features through the cross-attention mechanism to form a comprehensive joint feature. This joint feature guides the subsequent speech synthesis process, making the generated speech not only conform to the text content but also reflect the emotional and tone characteristics in the video. This method significantly improves the expressiveness and flexibility of speech synthesis.
[0063] In one possible implementation, the speech generation model is determined through the following steps: I: Replace the text encoder in the original VITS model with a video-text joint encoder to obtain an initial speech generation model.
[0064] Here, VITS is a parallel end-to-end TTS method. In this application, to achieve speech synthesis that integrates visual information, the text encoder in the original VITS model is replaced with a video-text joint encoder to obtain an initial speech generation model.
[0065] II: Input the sample joint feature determined based on the sample video information and sample text information into the initial speech generation model for speech synthesis processing to generate the predicted synthetic speech corresponding to the sample text information.
[0066] Here, the speech synthesis process in the initial speech generation model is consistent with the steps of the above speech generation model for generating synthesized speech, and this part will not be elaborated further.
[0067] III: Based on the predicted synthesized speech and the actual synthesized speech of the sample text information, determine the loss value of the initial speech generation model.
[0068] Here, calculate the predicted synthesized speech and the actual synthesized speech according to the loss function calculation formula to determine the loss value of the initial speech generation model.
[0069] IV: Based on the loss value and the training samples, perform iterative training on the initial speech generation model to generate the speech generation model.
[0070] Here, if the loss value is less than or equal to the preset threshold, use the initial speech generation model as the speech generation model. If the loss value is greater than the preset threshold, change the network parameters of the initial speech generation model, and continue to use the training samples to perform model training on the initial speech generation model with the changed network parameters.
[0071] Further, please refer to Figure 2 , Figure 2 which is the second flowchart of a speech synthesis method integrating visual information provided by an embodiment of the present application. As Figure 2 shown, input the joint feature constructed by the text feature and the audio feature into the video text joint encoder of the speech generation model for processing to obtain the video text feature, input the video text feature into the random prosody predictor of the speech generation model to obtain the duration of phonemes in the video text feature and the expanded video text feature, and input the duration of phonemes in the video text feature, the target person's speech embedding vector, and the expanded video text feature into the decoder of the speech generation model for speech synthesis processing to generate the synthesized speech corresponding to the target text information.
[0072] In this application, the video is processed by a pre-trained video feature extractor to obtain video features, which include features such as emotions and tones, and are used to supervise the expressiveness of the generated audio. The text is first converted into phonemes and then passed through a text encoder to obtain text encoding features. The text encoding features and the video features pass through a cross-attention module to obtain joint video-text features. Finally, through a multi-layer perceptron (MLP), the features are input into the video-text joint encoder of the speech generation model to further model the features of the video and the text. Then, the modeled video-text features are processed to obtain the synthesized speech. By fusing the information in the video, more expressive speech is generated. This method not only solves the deficiencies of the prior art in terms of emotion control but also enables high-quality speech synthesis using a small amount of sample data.
[0073] A speech synthesis method for fusing visual information provided by an embodiment of this application. The speech synthesis method includes: extracting text features of target text information based on a text encoder; performing feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine video features in the target video information; performing cross-attention processing on the text features and the video features to determine joint features; inputting the joint features into a pre-trained speech generation model to perform modeling processing on the joint features to generate video-text features, and then performing random prosody prediction processing, feature augmentation processing, and decoding processing on the video-text features to generate the synthesized speech corresponding to the target text information; wherein, the emotion features of the synthesized speech are consistent with the emotion features of the target video information. By fusing the information in the video, more expressive speech is generated. This method not only solves the deficiencies of the prior art in terms of emotion control but also enables high-quality speech synthesis using a small amount of sample data.
[0074] Please refer to Figure 3 、 Figure 4 , Figure 3 which is one of the structural schematic diagrams of a speech synthesis device for fusing visual information provided by an embodiment of this application; Figure 4 which is the second structural schematic diagram of a speech synthesis device for fusing visual information provided by an embodiment of this application. As Figure 3 shown in The text feature extraction module 310 is used to extract text features of target text information based on a text encoder; The video feature extraction module 320 is used to perform feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine video features in the target video information; A joint module 330 is used to perform cross-attention processing on the text features and the video features to determine joint features; A speech generation module 340 is used to input the joint features into a pre-trained speech generation model, perform modeling processing on the joint features to generate video text features, and then perform random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate a synthetic speech corresponding to the target text information; wherein, the emotional feature of the synthetic speech is consistent with the emotional feature of the target video information.
[0075] Further, when the video feature extraction module 320 is used to perform feature extraction processing, vector quantization encoding processing, and cross-attention processing on the target video information to determine the video features in the target video information, the video feature extraction module 320 is further used for: Mark the target video information based on a vector quantization encoder, and add random noise to the marked target video information; Input the target video information with added noise into a fusion diffusion model for sampling, and extract the first visual feature map of each video frame; Perform feature extraction processing on the target video information based on a multi-modal neural network model to obtain the visual feature map of each video frame, and perform multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map with the same dimension as the first visual feature map; Perform combination processing and position encoding processing on the first visual feature map of each video frame and the corresponding second visual feature map to determine the initial video feature of each video frame; Perform mean processing on multiple frames of the initial video features to generate the video features.
[0076] Further, when the speech generation module 340 is used to perform modeling processing on the joint features based on the self-attention mechanism, position encoding mechanism, and multi-head attention mechanism in the video text joint encoder to capture the cross-modal relationship between the text features and the video features and generate video text features, the speech generation module 340 specifically is used for: Process the joint features based on the self-attention mechanism to generate the joint features after self-attention processing; Perform position encoding processing on the time sequence information of the video frames and the position information of the target text information in the joint features after self-attention processing to generate the joint features after position encoding processing; Perform multi-angle modeling processing on the joint features after position encoding processing based on the multi-head attention mechanism to generate video text features.
[0077] Further, when the speech generation module 340 is used to input the video text features into the random prosody predictor of the speech generation model, perform random prosody prediction processing and high-dimensional feature augmentation processing on the video text features, and generate the augmented video text features, the speech generation module 340 is specifically configured to: Predict the duration of each phoneme in the video text features based on the emotion features and context relationships included in the video features in the video text features; Perform feature interpolation processing on the features corresponding to each phoneme in the video text features according to the corresponding duration to generate the augmented video text features.
[0078] Further, as Figure 4 shown, the speech synthesis device 300 that fuses visual information further includes a model training module 350. The model training module 350 determines the speech generation model through the following steps: Replace the text encoder in the original VITS model with a video text joint encoder to obtain an initial speech generation model; Input the sample joint features determined based on the sample video information and the sample text information into the initial speech generation model for speech synthesis processing to generate the predicted synthetic speech corresponding to the sample text information; Determine the loss value of the initial speech generation model based on the predicted synthetic speech and the actual synthetic speech of the sample text information; Iteratively train the initial speech generation model based on the loss value and training samples to generate the speech generation model.
[0079] A speech synthesis device integrating visual information provided by an embodiment of the present application, the speech synthesis device includes: a text feature extraction module, configured to extract text features of target text information based on a text encoder; a video feature extraction module, configured to perform feature extraction processing, vector quantization encoding processing, and cross-attention processing on target video information to determine video features in the target video information; a joint module, configured to perform cross-attention processing on the text features and the video features to determine joint features; a speech generation module, configured to input the joint features into a pre-trained speech generation model, perform modeling processing on the joint features to generate video text features, and then perform random prosody prediction processing, feature augmentation processing, and decoding processing on the video text features to generate a synthesized speech corresponding to the target text information; wherein, an emotional feature of the synthesized speech is consistent with an emotional feature of the target video information. By integrating information in the video, more expressive speech is generated. This method not only solves the deficiencies of the prior art in terms of emotion control, but also can achieve high-quality speech synthesis using a small amount of sample data.
[0080] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown in, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.
[0081] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, the steps of the speech synthesis method integrating visual information in the method embodiments as described above Figure 1 and Figure 2 shown can be executed. The specific implementation manner can refer to the method embodiments and will not be elaborated here.
[0082] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the speech synthesis method integrating visual information in the method embodiments as described above Figure 1 and Figure 2 shown can be executed. The specific implementation manner can refer to the method embodiments and will not be elaborated here.
[0083] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0084] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0085] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0086] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0087] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0088] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A speech synthesis method integrating visual information, characterized in that: The speech synthesis method comprises: Extract text features of target text information based on the text encoder; Performing feature extraction processing, vector quantization encoding processing, and cross attention processing on the target video information to determine the video features in the target video information; Performing cross-attention processing on the text features and the video features to determine joint features; The joint features are input into a pre-trained speech generation model, the joint features are modeled to generate video text features, and then the video text features are subjected to random rhythm prediction, feature expansion and decoding to generate synthetic speech corresponding to the target text information; wherein the emotional features of the synthetic speech are consistent with the emotional features of the target video information.
2. The speech synthesis method according to claim 1, characterized in that: The step of performing feature extraction processing, vector quantization encoding processing, and cross attention processing on the target video information to determine the video features in the target video information includes: Marking the target video information based on a vector quantization encoder, and adding random noise to the marked target video information; Inputting the target video information after adding noise into the fusion diffusion model for sampling, and extracting the first visual feature map of each video frame; Performing feature extraction processing on the target video information based on a multimodal neural network model to obtain a visual feature map of each video frame, and performing multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map having the same dimension as the first visual feature map; Combining and position-coding the first visual feature map and the corresponding second visual feature map of each video frame to determine initial video features of each video frame; Performing mean processing on the initial video features of multiple frames to generate the video features.
3. The speech synthesis method according to claim 1, characterized in that: The step of modeling the joint features to generate video text features, and then performing random prosody prediction, feature expansion, and decoding on the video text features to generate synthetic speech corresponding to the target text information includes: Based on the self-attention mechanism, position encoding mechanism and multi-head attention mechanism in the video-text joint encoder of the speech generation model, the joint features are modeled to capture the cross-modal relationship between the text features and the video features, and generate video-text features; Inputting the video text features into the random prosody predictor of the speech generation model, performing random prosody prediction processing and feature expansion processing on the video text features, and generating the expanded video text features; The duration of the phonemes in the video text features, the target person's speech embedding vector and the expanded video text features are input into the decoder of the speech generation model for speech synthesis processing to generate synthesized speech corresponding to the target text information.
4. The speech synthesis method according to claim 3, characterized in that: The method of modeling the joint features based on the self-attention mechanism, the position encoding mechanism and the multi-head attention mechanism in the video-text joint encoder to capture the cross-modal relationship between the text features and the video features and generate video-text features includes: Processing the joint features based on the self-attention mechanism to generate the joint features after self-attention processing; Performing position coding processing on the time sequence information of the video frames and the position information of the target text information in the joint features after the self-attention processing to generate the joint features after the position coding processing; Based on the multi-head attention mechanism, the joint features after position encoding are modeled from multiple angles to generate video text features.
5. The speech synthesis method according to claim 4, characterized in that: Inputting the video text features into the random prosody predictor of the speech generation model, performing random prosody prediction processing and high-dimensional feature expansion processing on the video text features, and generating the expanded video text features, including: Based on the emotional features and contextual relationships contained in the video features in the video text features, predict the duration of each phoneme in the video text features; Based on the feature corresponding to each of the phonemes in the video text feature, feature interpolation processing is performed according to the corresponding duration to generate the expanded video text feature.
6. The speech synthesis method according to claim 1, characterized in that: The speech generation model is determined by the following steps: Replace the text encoder in the original VITS model with a video-text joint encoder to obtain the initial speech generation model; Inputting the sample joint features determined based on the sample video information and the sample text information into the initial speech generation model for speech synthesis processing to generate predicted synthesized speech corresponding to the sample text information; Determining a loss value of the initial speech generation model based on the predicted synthesized speech and the actual synthesized speech of the sample text information; The initial speech generation model is iteratively trained based on the loss value and the training samples to generate the speech generation model.
7. A speech synthesis device integrating visual information, characterized in that: The speech synthesis device comprises: A text feature extraction module, used to extract text features of target text information based on a text encoder; A video feature extraction module, used to perform feature extraction processing, vector quantization encoding processing and cross attention processing on target video information to determine video features in the target video information; A joint module, used for performing cross-attention processing on the text features and the video features to determine a joint feature; The speech generation module is used to input the joint features into a pre-trained speech generation model, perform modeling processing on the joint features to generate video text features, and then perform random rhythm prediction processing, feature expansion processing and decoding processing on the video text features to generate a synthetic speech corresponding to the target text information; wherein the emotional features of the synthetic speech are consistent with the emotional features of the target video information.
8. The speech synthesis device according to claim 7, characterized in that: When the video feature extraction module is used to perform feature extraction processing, vector quantization encoding processing and cross attention processing on the target video information to determine the video features in the target video information, the video feature extraction module is also used to: Marking the target video information based on a vector quantization encoder, and adding random noise to the marked target video information; Inputting the target video information after adding noise into the fusion diffusion model for sampling, and extracting the first visual feature map of each video frame; Performing feature extraction processing on the target video information based on a multimodal neural network model to obtain a visual feature map of each video frame, and performing multi-layer perceptron projection processing on the visual feature map of each video frame to generate a second visual feature map having the same dimension as the first visual feature map; Combining and position-coding the first visual feature map and the corresponding second visual feature map of each video frame to determine initial video features of each video frame; Performing mean processing on the initial video features of multiple frames to generate the video features.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the method for speech synthesis that integrates visual information as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for speech synthesis integrating visual information as claimed in any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Personalized speech synthesis method, electronic equipment, server and storage medium
CN118098199A
Speech synthesis method and device, computer equipment and storage medium
CN118280341A
Companion audio generation method, related device and medium
CN118737121A
Low-resource speech synthesis system based on BERT feature and style coding
CN119207373A
Cards-based chinese character learning board game
KR1020260131812A
Cited By
Large model-based audio and video data generation method, training method and intelligent agent
CN121306089A