Speech synthesis method and device

By constructing a multimodal dataset and a three-stage training speech synthesis model, the problem of lack of fine-grained speech style in existing technologies is solved, and high-quality conversion of static images to dynamic speech is achieved. The generated speech style is highly consistent with the visual information.

CN119152837BActive Publication Date: 2025-09-19TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411000066.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-09-19
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively utilize multiple clues in images when generating video content, resulting in a lack of fine-grained generated speech style and an inability to be harmonious with the visual information.

Method used

By constructing a multimodal dataset and using an image encoder, a speech decoder, and a query converter Q-former, a speech synthesis model is trained in three stages to combine visual information and speech features to generate more fine-grained speech styles.

Benefits of technology

High-quality conversion of static images to dynamic speech is achieved, and the generated speech style is highly consistent with the visual information, which improves the granularity and vividness of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152837B_ABST
    Figure CN119152837B_ABST
Patent Text Reader

Abstract

The present invention provides a speech synthesis method and device, relating to the field of speech processing technology. The method comprises: obtaining a target image and a speech manuscript, and inputting the target image and speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech; wherein the target image contains multiple visual information, and the target synthesized speech contains multiple acoustic features, with each visual information corresponding to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset. The method provided by the present invention modally enhances an existing speech dataset to construct a multimodal dataset, thereby addressing the problem of dataset scarcity; based on the one-to-one correspondence between visual information in a static image and acoustic features in speech audio, and based on the speech synthesis model trained using the multimodal dataset, the synthesized target synthesized speech has a more fine-grained speech style.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a speech synthesis method and device. Background Art

[0002] In the era of AI-generated content (AIGC), AI technology is generating an increasing amount of multimedia content. For example, video generation aims to transform static, human-centered images into dynamic, talking animations. To enhance the liveliness of generated videos, it is crucial to ensure that the visual information in the input image is in harmony with the speech characteristics in the audio.

[0003] Existing studies mainly focus on leveraging facial information to infer basic speaker characteristics such as gender, age, and emotion, but they often ignore the abundant additional clues present in images and fail to generate audio with fine-grained speech style.

[0004] How to simulate and synthesize audio with finer-grained speech style through a given image is a technical problem that needs to be solved. Summary of the Invention

[0005] The present invention provides a speech synthesis method and device to solve the defects in the prior art.

[0006] The present invention provides a speech synthesis method, comprising the following steps:

[0007] Acquire a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0008] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0009] According to a speech synthesis method provided by the present invention, the speech synthesis model includes an image encoder, a speech decoder and a query converter Q-former;

[0010] The image encoder is used to extract multiple visual information from the target image;

[0011] The speech decoder is used to generate a plurality of acoustic features in the target synthesized speech;

[0012] The query converter Q-former is used for modality interaction training.

[0013] According to a speech synthesis method provided by the present invention, the training process of the speech synthesis model includes:

[0014] Performing a first-stage training on the speech decoder based on the multimodal dataset and obtaining a training loss value of the speech decoder; wherein the first-stage training is unsupervised speech style learning training;

[0015] Performing a second-stage training on the query converter Q-former based on the multimodal dataset; wherein the second-stage training is a visual representation learning training related to speech style;

[0016] The speech decoder after the first stage training and the query converter Q-former after the second stage training are connected, and the connected speech decoder and query converter Q-former are trained in the third stage based on the training loss value of the speech decoder; wherein the third stage training is speech style control training under visual conditions.

[0017] According to a speech synthesis method provided by the present invention, acquiring a target image and a speech manuscript, and inputting the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech, includes:

[0018] Acquire a target image and a speech manuscript, and extract a plurality of key frames of the target image; wherein the key frames represent an image sequence centered on the speaker in the target image;

[0019] Inputting the multiple key frames of the target image and the speech manuscript into a pre-trained speech synthesis model to obtain multiple speech audios;

[0020] The target synthesized speech is synthesized based on the multiple speech audios.

[0021] According to a speech synthesis method provided by the present invention, before acquiring the target image and the speech manuscript, the method further includes:

[0022] A target data set is obtained, and modality enhancement is performed on the target data set to obtain the multimodal data set; wherein the target data set is a natural language prompt speech data set.

[0023] According to a speech synthesis method provided by the present invention, the target data set includes a plurality of speech segments and a speech description corresponding to each speech segment;

[0024] The performing modality enhancement on the target dataset to obtain the multimodal dataset includes:

[0025] Inputting the speech description in the target dataset into a pre-trained text modality conversion model to obtain a visual description corresponding to the speech description; wherein the speech description represents a text description corresponding to the speech feature of the target dataset, and the visual description represents a picture description corresponding to the speech scene of the target dataset;

[0026] The visual description is input into a pre-trained image generation model to obtain the corresponding target image.

[0027] The present invention also provides a speech synthesis device, comprising the following modules:

[0028] A speech synthesis module is used to obtain a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0029] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0030] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described speech synthesis methods when executing the program.

[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned speech synthesis methods when executed by a processor.

[0032] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned speech synthesis methods.

[0033] The present invention provides a speech synthesis method and device, which obtains a target synthesized speech by acquiring a target image and a speech manuscript, and inputting the target image and the speech manuscript into a pre-trained speech synthesis model. The target image contains multiple visual information, and the target synthesized speech contains multiple acoustic features, with one visual information corresponding to at least one acoustic feature. The speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset. Therefore, it can be seen that the present invention solves the problem of dataset scarcity by modally enhancing the existing speech dataset to construct a multimodal dataset. Based on the one-to-one correspondence between the visual information in the static image and the acoustic features in the speech audio, and based on the speech synthesis model trained using the multimodal dataset, the synthesized target synthesized speech has a more fine-grained speech style. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 It is a flow chart of the speech synthesis method provided by the present invention.

[0036] Figure 2 This is the overall architecture diagram of the speech synthesis model provided by the present invention.

[0037] Figure 3 It is a schematic diagram of the intrinsic connection between visual information and acoustic features provided by the present invention.

[0038] Figure 4 It is a schematic diagram of the correlation strength between visual information and acoustic features provided by the present invention.

[0039] Figure 5 This is a schematic diagram of the speech synthesis model training process provided by the present invention.

[0040] Figure 6 It is a schematic diagram of the multimodal dataset construction process provided by the present invention.

[0041] Figure 7 It is a structural diagram of the speech synthesis device provided by the present invention.

[0042] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0044] The following combination Figures 1-8 A speech synthesis method and apparatus of the present invention are described.

[0045] Figure 1 Schematic diagram of the speech synthesis method provided by the present invention, such as Figure 1 As shown, the method includes the following:

[0046] Step 100: Obtain a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0047] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0048] It should be noted that TTS, short for Text To Speech, is a technology that converts text information into natural, fluent speech output. This embodiment of the present invention proposes a visually enhanced TTS model that can predict the speaker's voice and control the speaking style at the utterance level based on the corresponding image sequence (keyframes).

[0049] Specifically, Figure 2 This is the overall architecture diagram of the speech synthesis model provided by the present invention, such as Figure 2 As shown in the figure, a speech synthesis model with global visual insights, located between the two data streams, synthesizes audio sequences using keyframes and corresponding sentences, thereby organizing speech at the discourse level. The model implicitly extracts visual information from the image listed on the far left, such as appearance and expression, and then compares it with acoustic features shown on the far right, such as pitch and volume.

[0050] It should be noted that, in order to more comprehensively understand vision-based speech style control, the embodiment of the present invention systematically analyzes the relationship between visual information and acoustic features.

[0051] Specifically, this example studied the relationship between personal appearance and voice quality and found a significant negative correlation between audio quality and measurements of body shape and weight. Facial bone measurements also showed significant correlations with F0 and habitual frequency. Experiments have shown that voice provides identity information comparable to facial information, and that people can match unfamiliar voices with static facial images. From a psychological perspective, speech and other body movements are outward manifestations of internal activities (such as emotions), and therefore the two naturally form a close connection. Different directions and intentions of body sway convey different emotional states. Similarly, individuals with positive or negative emotions exhibit opposite head movements. Both emotional and cognitive functions are affected in different communication environments. Furthermore, different tones can trigger a range of emotional responses. Previous research has primarily focused on one or two aspects of the visual-vocal relationship. A more comprehensive understanding of visual-based voice style control requires a systematic analysis of the visual-auditory correlations in human behavior based on human perception, which is currently lacking and urgently needed.

[0052] Specifically, Figure 3 is a schematic diagram of the intrinsic connection between visual information and acoustic features provided by the present invention, Figure 4 This is a schematic diagram of the correlation strength between visual information and acoustic features provided by the present invention, combined with Figure 3 and Figure 4 This paper describes the research on the correlation between sound and scene images using the experimental psychology method provided in this embodiment. Based on the brief blueprint of the correlation map and 20 human-centered conversation video samples, key frame images and corresponding audio clips were extracted, and a five-level correlation mean opinion score (RMOS) (N=10) was calculated for each pair of visual features. Figure 4 For each visual-sound pairing, the average RMOS of all samples is taken as the correlation value of the pair. For each visual-sound combination, the average RMOS of all samples is taken as the correlation value of the pair. Then, the correlation matrix is ​​visualized as a Sankey diagram, which can be seen in Figure 3 , where only the 50% with the highest correlation are retained.

[0053] Specifically, from the perspective of human perception, we discover the systematic relationship between visual information and acoustic features. Figure 3Visual information related to speech includes the speaker's appearance (hair, face, clothing, etc.), the speaker's facial expressions (eyebrows, eyes, lips, etc.), the speaker's posture (head movements, body posture, and gestures), the speaking scene (background), and visual tones (artistic tones, image exposure, contrast, saturation, etc.). Meanwhile, acoustic features related to vision include the speaker's timbre, speaking rhythm (speed and pace), speaking volume, speaking emotions (emotion and intonation), and speaking topics. In addition to facial features, non-facial visual features such as the speaker's posture also have relatively strong correlations with various speech attributes such as speaking rhythm, speaking volume, and speaking emotions. It is worth noting that visual intonation is also relatively strongly related to speaking emotions, while the speaking scene context is mainly related to the speaking topic.

[0054] Furthermore, the overall cross-modal TTS framework provided by the embodiment of the present invention is described.

[0055] Specifically, the speech synthesis model includes an image encoder, a speech decoder and a query converter Q-former;

[0056] The image encoder is used to extract multiple visual information from the target image;

[0057] The speech decoder is used to generate a plurality of acoustic features in the target synthesized speech;

[0058] The query converter Q-former is used for modality interaction training.

[0059] The training process of the speech synthesis model includes:

[0060] Performing a first-stage training on the speech decoder based on the multimodal dataset and obtaining a training loss value of the speech decoder; wherein the first-stage training is unsupervised speech style learning training;

[0061] Performing a second-stage training on the query converter Q-former based on the multimodal dataset; wherein the second-stage training is a visual representation learning training related to speech style;

[0062] The speech decoder after the first stage training and the query converter Q-former after the second stage training are connected, and the connected speech decoder and query converter Q-former are trained in the third stage based on the training loss value of the speech decoder; wherein the third stage training is speech style control training under visual conditions.

[0063] In one embodiment, see Figure 2The TTS framework consists of a CLIP image encoder, a query converter Q-former as a bridge network, and a VITS-based speech backbone network. The VITS-based speech backbone network is an end-to-end variational autoencoder, including a reference speech encoder and a speech decoder. Due to its end-to-end design, it shows strong capabilities in multi-speaker speech synthesis. The image encoder extracts global visual embeddings from the input image, while the Q-former extracts visual reproduction related to the speech style and provides conditions for the generation of the speech decoder. Since end-to-end image-audio alignment relies on a large amount of strongly correlated image-audio data, this embodiment adopts a pre-training-fine-tuning training method and designs a three-stage learning framework. Figure 5 This is a flow chart of the speech synthesis model training process provided by the present invention. Figure 5 As shown in Figure 2, the three-stage learning includes: (1) unsupervised speech style learning; (2) learning of visual representations related to speech style; and (3) speech style control under visual conditions.

[0064] (1) Phase 1: Unsupervised speech style learning.

[0065] In order to give the VITS-based backbone system powerful speech modeling capabilities, in the first phase, the VITS-based speech backbone system was unsupervisedly trained on large-scale emotional speech data. The speech encoder consists of a convolutional neural network layer and a gated recurrent unit (GRU) layer, and learns a global style embedding from the melt spectrogram of real audio. After normalization, the global style embedding is passed to the generator to supervise the training of various modules such as the post-encoder, pre-encoder, and duration predictor. This embodiment uses multiple discriminators for adversarial training. For simplicity, the training loss of the backbone is annotated as L vits .

[0066] (2) The second stage: learning visual representations related to speech style.

[0067] To effectively utilize high-quality, relevant image and audio data, in the second stage, the Q-former learns visual representations from images, guided by live audio and text descriptions that describe basic speech properties (including phonetics). Live audio provides rich, fine-grained style information but is more difficult to model. Text descriptions provide coarser style information but are easier to model. To leverage the style information in audio and descriptions, this embodiment employs multiple cross-modal contrastive learning methods as training targets.

[0068] It should be noted that the query transformer (Q-former) is initialized from a pre-trained BERT model, along with a set of randomly initialized trainable queries and a set of cross-attention layers injected into the transformer encoder layer. The Q-former plays three roles in the second stage: 1) The image encoder inputs the style query, which interacts with the visual features through the cross-attention layer to obtain the image embedding E i ; 2) The text encoder inputs the text description, extracts the rough speech style information, and obtains the text embedding E t 3) The bidirectional encoder takes as input the merged style query and text description and outputs a fused representation. The style query and description can attend to each other during the forward pass of the transformation layer. In other words, two unimodal encoders and one bimodal encoder share a set of parameters in the self-attention and feed-forward layers. This mechanism not only facilitates the learning of coarse-to-fine information from text descriptions and audio for the style query but also provides a path for visual-textual information fusion. During the inference phase, the style query can replace the text description and audio information.

[0069] 1. Image-audio contrastive learning: The goal of this stage is to align the unimodal image embedding E from the Q-former i and the speech embedding E from the frozen speech encoder a In this embodiment, for every N pairs of images I K and Audio A K , the matching image-audio pairs are labeled as positive samples, and the mismatched image-audio pairs are labeled as negative samples. The dot product of the image embedding and the speech embedding is used as the classification vector for binary classification and as the matching score. Specifically, for each image, the softmax function and the cross entropy loss function are applied to the score to obtain the image-audio contrast loss L iac-i , see formula (1).

[0070]

[0071] Similarly, for each audio, the same steps are applied to the scores to obtain the image-audio contrastive loss L iac-a , see formula (2). The total contrast loss should be L iac = L iac-i + L iac-a .

[0072]

[0073] In order to obtain more precise image and audio alignment, this embodiment uses an additional classifier to fuse the image embedding E i and speech embedding E a, and pass an output linear layer to obtain two categories (positive / negative samples) classification vectors. The image-audio matching loss is L iam .

[0074] 2. Image-Text Contrastive Learning: Text descriptions provide pure voice attribute information in the second stage, which helps with image-audio alignment. In other words, the involvement of text descriptions allows audio-visual alignment to be performed in text mode, thereby reducing the modeling complexity of the Q-former. Similar to image-audio contrastive learning, cross-modal contrastive loss and matching loss are used as training objects. Contrastive loss L itc Single model output from Q-former, E i and E i , because in this case, the style query is prohibited from focusing on the text description. The matching loss L itm Bidirectional output from the Q-former, in this case, the style query and text description are concatenated and allowed to follow each other.

[0075] 3. Text-audio contrastive learning: In the second stage, the text description provides pure speech attribute information, which helps image-audio alignment. In order to simplify the text-audio contrastive learning, this embodiment only uses the contrast loss L atc . L atc The form of L iac resemblance.

[0076] The total training loss of Q-former in the second stage is L, which can be seen in formula (3).

[0077]

[0078] Considering the coarse-to-fine supervision of text description and speech audio, an annealing function is set for λia, and a fixed λia value will degrade the performance. ia The weight of gradually increases and finally reaches 1.

[0079] (3) The third stage: speech style control under visual conditions.

[0080] Despite efforts to align visual and speech representations, a domain gap still exists between the two modalities. To address this domain gap, a unified fine-tuning strategy is adopted in the third stage. The Q-former is directly connected to the VITS-based speech decoder and L vitsAs training objects for image-audio pairs. It is worth noting that at this stage, all parameters of the Q-former and speech decoder are trainable.

[0081] It should be noted that the speech backbone adopts the design of VITS2, but replaces the random duration predictor with a normal duration predictor, and enables the gradient of the duration predictor with respect to the global speech embedding during backpropagation. To improve prosody, phonemes and text features from the pre-trained BERT model are used as inputs to the speech decoder, and multiple discriminators are employed for the flow model and duration predictor. The open-source CLIP-VIT-H-14 is used for the CLIP image encoder. The Q-former is initialized with the XLM-Roberta-base and 32 style queries with a hidden size of 768 are set. A hard negative mining strategy is used to select negative pairs for the cross-modal matching loss. Hard negative mining is an important strategy in machine learning, particularly for improving model performance in classification problems. The core of this strategy is to enhance the model's discriminative and generalization capabilities by identifying and focusing on negative examples that are misclassified by the model (i.e., hard negative examples).

[0082] Based on the above-trained speech synthesis model, the target image and speech manuscript are obtained, and the target image and speech manuscript are input into the pre-trained speech synthesis model to obtain the target synthesized speech, including:

[0083] Acquire a target image and a speech manuscript, and extract a plurality of key frames of the target image; wherein the key frames represent an image sequence centered on the speaker in the target image;

[0084] Inputting the multiple key frames of the target image and the speech manuscript into a pre-trained speech synthesis model to obtain multiple speech audios;

[0085] The target synthesized speech is synthesized based on the multiple speech audios.

[0086] The above is a description of the steps of the speech synthesis method provided by the present invention. From the description of the above steps, it can be seen that according to the speech synthesis method provided by the present invention, a target synthesized speech is obtained by acquiring a target image and a speech manuscript, and inputting the target image and the speech manuscript into a pre-trained speech synthesis model; wherein, the target image contains multiple visual information, and the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is obtained by training based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset. It can be seen from this that the present invention solves the problem of dataset scarcity by constructing a multimodal dataset by modally enhancing the existing speech dataset; according to the one-to-one correspondence between the visual information in the static image and the acoustic features in the speech audio, based on the speech synthesis model obtained by training the multimodal dataset, the synthesized target synthesized speech has a more fine-grained speech style.

[0087] Based on the above embodiment, in this embodiment, before step 100 of acquiring the target image and the speech manuscript, the method further includes:

[0088] A target data set is obtained, and modality enhancement is performed on the target data set to obtain the multimodal data set; wherein the target data set is a natural language prompt speech data set.

[0089] The target data set includes a plurality of speech segments and a speech description corresponding to each speech segment;

[0090] The performing modality enhancement on the target dataset to obtain the multimodal dataset includes:

[0091] Inputting the speech description in the target dataset into a pre-trained text modality conversion model to obtain a visual description corresponding to the speech description; wherein the speech description represents a text description corresponding to the speech feature of the target dataset, and the visual description represents a picture description corresponding to the speech scene of the target dataset;

[0092] The visual description is input into a pre-trained image generation model to obtain the corresponding target image.

[0093] It should be noted that, based on the above examples, body posture and speaking context significantly impact speech timbre. This places significant demands on training speech datasets: person-centric images should contain visual information, including not only facial expressions but also head posture, gestures, and speech context. Existing audio-visual datasets are unsuitable for this task for the following reasons: First, audio-visual datasets are generally designed for generating talking heads. They often emphasize the accuracy of visual motion to capture the synchronization between audio signals and visual motion, but overlook the expressiveness of speech, which is crucial for TTS tasks. Second, in traditional audio-visual dataset construction, there is an inevitable trade-off between audio quality and scene diversity. Audio-visual datasets typically use studio or online videos to extract image-audio pairs. Studio videos have good speech quality but poor scene diversity. Online videos have diverse speaking scenes, but due to challenges such as the cocktail party effect, significant data cleaning efforts are required. The cocktail party effect is a unique phenomenon in the human auditory system, which refers to the ability of people, in a noisy environment, to focus on a single person's conversation while ignoring other conversations or noise in the background.

[0094] Specifically, emotional speech datasets and natural language prompt speech datasets typically contain high-quality speech waveforms and detailed textual descriptions of the speech, including gender, age, emotion, pitch, and volume. These rich textual descriptions allow humans to naturally imagine the speaking scene, which can then be directly converted into visual components using mature AIGC technology. Given that modality conversion should be consistent with human perception, we chose to expand the modality of existing speech datasets based on the above examples. Figure 6 It is a schematic diagram of the multimodal dataset construction process provided by the present invention. As shown in FIG6 , a novel text-audio-vision dataset is constructed, which is tailored for text-to-speech after visual improvement.

[0095] 1) Text-to-Modal Conversion: A prompt template is pre-set, and LLM converts the speech description (a textual description of speech features) into a visual description (a picture caption describing the speech scene). Modal conversion is performed in the text domain.

[0096] 2) Visualization: Based on the visual description, the text-to-image model synthesizes images that capture the essence of the audio context and speaking style, completing the visual part of the text-audio-visual speech dataset.

[0097] In this embodiment, GPT-3.5 turbo is selected as the text modality converter. DALL-E 3 is selected for image generation. This embodiment does not impose any special restrictions on this.

[0098] Furthermore, we obtained a bilingual text-visual-audio dataset based on the internal natural language prompt speech dataset. It contains 28,929 image-audio pairs, of which 28,063 samples are used for training segmentation and 866 samples are used for testing segmentation.

[0099] The speech synthesis method provided in this embodiment solves the problem of dataset scarcity by performing modal enhancement on an existing speech dataset to construct a multimodal dataset.

[0100] In summary, this embodiment of the present invention systematically studies the relationship between human visual and auditory perception through psychological experiments. It further proposes a multimodal contrastive learning framework that demonstrates profound voice style control capabilities. Furthermore, it leverages generative artificial intelligence and modality augmentation techniques to construct a multimodal dataset centered around speech. The model trained on this multimodal dataset demonstrates generalization capabilities to real-world photos, including emojis.

[0101] The speech synthesis device provided by the present invention is described below. The speech synthesis device described below and the speech synthesis method described above can be referenced to each other.

[0102] Figure 7 Schematic diagram of the structure of the speech synthesis device provided by the present invention. Figure 7 As shown, the speech synthesis device provided by the present invention includes:

[0103] The speech synthesis module 701 is used to obtain a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0104] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0105] The speech synthesis device provided by the present invention obtains a target synthesized speech by acquiring a target image and a speech manuscript, and inputting the target image and the speech manuscript into a pre-trained speech synthesis model. The target image contains multiple visual information, and the target synthesized speech contains multiple acoustic features, with one visual information corresponding to at least one acoustic feature. The speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset. Therefore, it can be seen that the present invention solves the problem of dataset scarcity by modally enhancing the existing speech dataset to construct a multimodal dataset. Based on the one-to-one correspondence between the visual information in the static image and the acoustic features in the speech audio, and based on the speech synthesis model trained using the multimodal dataset, the synthesized target synthesized speech has a more fine-grained speech style.

[0106] Based on the above embodiment, in this embodiment, the speech synthesis model includes an image encoder, a speech decoder and a query converter Q-former;

[0107] The image encoder is used to extract multiple visual information from the target image;

[0108] The speech decoder is used to generate a plurality of acoustic features in the target synthesized speech;

[0109] The query converter Q-former is used for modality interaction training.

[0110] Based on the above embodiment, in this embodiment, the device further includes a training module, which is specifically configured to:

[0111] Performing a first-stage training on the speech decoder based on the multimodal dataset and obtaining a training loss value of the speech decoder; wherein the first-stage training is unsupervised speech style learning training;

[0112] Performing a second-stage training on the query converter Q-former based on the multimodal dataset; wherein the second-stage training is a visual representation learning training related to speech style;

[0113] The speech decoder after the first stage training and the query converter Q-former after the second stage training are connected, and the connected speech decoder and query converter Q-former are trained in the third stage based on the training loss value of the speech decoder; wherein the third stage training is speech style control training under visual conditions.

[0114] Based on the above embodiment, in this embodiment, the speech synthesis module 701 is specifically configured to:

[0115] Acquire a target image and a speech manuscript, and extract a plurality of key frames of the target image; wherein the key frames represent an image sequence centered on the speaker in the target image;

[0116] Inputting the multiple key frames of the target image and the speech manuscript into a pre-trained speech synthesis model to obtain multiple speech audios;

[0117] The target synthesized speech is synthesized based on the multiple speech audios.

[0118] Based on the above embodiment, in this embodiment, the device further includes a modality enhancement module, which is specifically configured to:

[0119] Before acquiring the target image and speech manuscript, a target data set is acquired, and modality enhancement is performed on the target data set to obtain the multimodal data set; wherein the target data set is a natural language prompt speech data set.

[0120] Based on the above embodiment, in this embodiment, the target data set includes multiple voice segments and a voice description corresponding to each voice segment;

[0121] The modal enhancement module is specifically used to:

[0122] Inputting the speech description in the target dataset into a pre-trained text modality conversion model to obtain a visual description corresponding to the speech description; wherein the speech description represents a text description corresponding to the speech feature of the target dataset, and the visual description represents a picture description corresponding to the speech scene of the target dataset;

[0123] The visual description is input into a pre-trained image generation model to obtain the corresponding target image.

[0124] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the speech synthesis method, which includes:

[0125] Acquire a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0126] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0127] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0128] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the speech synthesis method provided by each of the above methods, which includes:

[0129] Acquire a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0130] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0131] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the speech synthesis method provided by the above methods, the method comprising:

[0132] Acquire a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech;

[0133] Among them, the target image contains multiple visual information, the target synthesized speech contains multiple acoustic features, and one visual information corresponds to at least one acoustic feature; the speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing the target dataset.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0135] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech synthesis method, characterized in that: include: Acquire a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech; The target image includes multiple visual information, the target synthesized speech includes multiple acoustic features, and one visual information corresponds to at least one acoustic feature; The speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing a target dataset.

2. The speech synthesis method according to claim 1, wherein: The speech synthesis model includes an image encoder, a speech decoder and a query converter Q-former; The image encoder is used to extract multiple visual information from the target image; The speech decoder is used to generate a plurality of acoustic features in the target synthesized speech; The query converter Q-former is used for modality interaction training.

3. The speech synthesis method according to claim 2, wherein: The training process of the speech synthesis model includes: Performing a first-stage training on the speech decoder based on the multimodal dataset and obtaining a training loss value of the speech decoder; wherein the first-stage training is unsupervised speech style learning training; Performing a second-stage training on the query converter Q-former based on the multimodal dataset; wherein the second-stage training is a visual representation learning training related to speech style; The speech decoder after the first stage training and the query converter Q-former after the second stage training are connected, and the connected speech decoder and query converter Q-former are trained in the third stage based on the training loss value of the speech decoder; wherein the third stage training is speech style control training under visual conditions.

4. The speech synthesis method according to claim 1, wherein: The step of acquiring a target image and a speech manuscript, and inputting the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech includes: Acquire a target image and a speech manuscript, and extract a plurality of key frames of the target image; wherein the key frames represent an image sequence centered on the speaker in the target image; Inputting the multiple key frames of the target image and the speech manuscript into a pre-trained speech synthesis model to obtain multiple speech audios; The target synthesized speech is synthesized based on the multiple speech audios.

5. The speech synthesis method according to claim 1, wherein: Before acquiring the target image and the speech manuscript, the method further includes: A target data set is obtained, and modality enhancement is performed on the target data set to obtain the multimodal data set; wherein the target data set is a natural language prompt speech data set.

6. The speech synthesis method according to claim 5, characterized in that: The target data set includes a plurality of speech segments and a speech description corresponding to each speech segment; The performing modality enhancement on the target dataset to obtain the multimodal dataset includes: Inputting the speech description in the target dataset into a pre-trained text modality conversion model to obtain a visual description corresponding to the speech description; wherein the speech description represents a text description corresponding to the speech feature of the target dataset, and the visual description represents a picture description corresponding to the speech scene of the target dataset; The visual description is input into a pre-trained image generation model to obtain the corresponding target image.

7. A speech synthesis device, characterized in that: include: A speech synthesis module is used to obtain a target image and a speech manuscript, and input the target image and the speech manuscript into a pre-trained speech synthesis model to obtain a target synthesized speech; The target image includes multiple visual information, the target synthesized speech includes multiple acoustic features, and one visual information corresponds to at least one acoustic feature; The speech synthesis model is trained based on a multimodal dataset, and the multimodal dataset is obtained by modally enhancing a target dataset.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the speech synthesis method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Knowledge fusion multi-modal interaction method and device based on improved alignment method

    CN117113270A

  • Voice data processing method and device, equipment and storage medium

    CN117219049A