Speech synthesis method and device based on multi-modal style embedding, equipment and medium
By adopting a multimodal style embedded speech synthesis method in zero-sample text to speech system, the problem that speech synthesis is difficult to reproduce the voice characteristics of the speaker and control the speech style is solved, and a richer and more natural style diversity is achieved.
Patent Information
- Application Number
- CN202510221661.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
AI Technical Summary
In the zero-sample text to the voice system, it is difficult to accurately reproduce the voice characteristics of the speaker and control the voice style. Especially in application scenarios such as medical intelligent voice Q&A and financial intelligent customer service, the lack of appropriate style audio samples makes it difficult to customize the speech synthesis style.
The speech synthesis method based on multimodal style embedding is adopted, by obtaining multimodal feature data and phoneme data, the first encoder extracts and fuses multimodal features to generate speech embedding data; the second encoder extracts phoneme features to generate text feature data; and then, based on the feature fusion model, the speech embedding data and text feature data are feature-fused to generate stylized speech data.
It achieves the improvement of style diversity and naturalness of speech synthesis under zero sample conditions, and can more accurately capture and retain the style and emotional characteristics in the original speech signal, which is suitable for a variety of application scenarios.
Smart Images

Figure CN119993114A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and is applicable to the medical and financial fields, and in particular to a speech synthesis method, apparatus, device and medium based on multimodal style embedding. Background Art
[0002] In zero-shot text-to-speech systems, a major challenge is to accurately reproduce the speaker's voice characteristics and be able to control the style of the voice, such as emotion, accent, or speed and volume, while maintaining a high degree of editing flexibility. This is especially true in application scenarios such as medical intelligent voice question and answer and financial intelligent customer service.
[0003] For example, in the medical intelligent voice Q&A application scenario, patients ask medical questions through voice, such as asking about disease-related knowledge, how to use medicines, etc. After receiving the voice, voice recognition is first performed to convert the voice into text information. Then, the text information is analyzed and answers are generated based on the knowledge base. Since it is difficult to obtain audio samples with specific emotional styles, when synthesizing voice answers, only relatively neutral and general voice styles can be used.
[0004] For example, in the application scenario of financial intelligent customer service, customers call the customer service hotline to inquire about financial services, such as account balance inquiries and financial product consultations. After the customer's voice input, it is also converted into text through voice recognition, and the text content of the answer is generated according to business rules and database information. When the voice synthesis output answers, there is also the problem of style customization. Because the financial field has high requirements for the stability of voice style, it must be both professional and clear, but due to the lack of suitable style audio samples, it is difficult to flexibly adjust the voice style according to different customer emotions or scenarios.
[0005] A common practice is to use audio samples as style references, which provides a high degree of customization. However, it is difficult to obtain such emotional audio samples, and because the sound characteristics are closely intertwined with style or other acoustic information, it is impossible to stably guide style generation. Although style manipulation can be performed using label control, this approach limits the variability of audio references. There are also methods for directly generating speech through prompts, but this method sacrifices the ability to accurately replicate the sound characteristics from the speaker's audio.
[0006] Therefore, how to improve the style diversity of zero-sample speech synthesis has become a technical problem that needs to be solved urgently. Summary of the invention
[0007] The present application provides a speech synthesis method, apparatus, device and storage medium based on multimodal style embedding, aiming to improve the style diversity of zero-sample speech synthesis.
[0008] In a first aspect, the present application provides a speech synthesis method based on multimodal style embedding, and the speech synthesis method based on multimodal style embedding comprises the following steps:
[0009] Acquire multimodal feature data and phoneme data;
[0010] Based on the first encoder, perform feature extraction and feature fusion on the multimodal feature data to generate speech embedding data;
[0011] Based on the second encoder, feature extraction is performed on the phoneme data to generate text feature data;
[0012] Based on the feature fusion model, the speech embedding data and the text feature data are fused to generate stylized speech data.
[0013] In a second aspect, the present application further provides a speech synthesis device based on multimodal style embedding, the speech synthesis device based on multimodal style embedding comprising:
[0014] A data acquisition module, used to acquire multimodal feature data and phoneme data;
[0015] An embedded data encoding module, used for performing feature extraction and feature fusion on the multimodal feature data based on the first encoder to generate speech embedded data;
[0016] A text feature extraction module, used for extracting features from the phoneme data based on a second encoder to generate text feature data;
[0017] The feature fusion module is used to perform feature fusion on the speech embedding data and the text feature data based on a feature fusion model to generate stylized speech data.
[0018] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the speech synthesis method based on multimodal style embedding as described above are implemented.
[0019] The present application provides a speech synthesis method, device, computer equipment and storage medium based on multimodal style embedding. The method of the present application includes obtaining multimodal feature data and phoneme data; based on a first encoder, extracting features and fusing features of the multimodal feature data to generate speech embedding data; based on a second encoder, extracting features of the phoneme data to generate text feature data; based on a feature fusion model, fusing features of the speech embedding data and the text feature data to generate stylized speech data. In the above manner, the present application extracts and fuses multimodal features through the first encoder to generate speech embedding data, which helps to capture and retain the style and emotional features in the original speech signal. The second encoder extracts features of the phoneme data to generate text feature data, which helps to understand the text content and convert it into a synthesizable speech signal. The speech embedding data and the text feature data are combined through the feature fusion model to generate stylized speech data, and the naturalness of the speech and the style of the text are considered simultaneously when synthesizing speech, thereby achieving richer and more natural style diversity in zero-sample speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of a first embodiment of a speech synthesis method based on multimodal style embedding provided by the present application;
[0022] Figure 2 A schematic diagram of the data processing flow of a speech synthesis model for implementing a speech synthesis method based on multimodal style embedding provided in this application;
[0023] Figure 3 A schematic diagram of a data encoding process of a first encoder provided in this application;
[0024] Figure 4 It is a structural schematic diagram of a first embodiment of a speech synthesis device based on multimodal style embedding provided by the present application;
[0025] Figure 5 It is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application.
[0026] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0028] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0029] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0030] Please refer to Figure 1 , Figure 1 A flowchart of a first embodiment of a speech synthesis method based on multimodal style embedding provided in the present application.
[0031] like Figure 1 As shown, the speech synthesis method based on multimodal style embedding includes steps S101 to S104.
[0032] S101, obtaining multimodal feature data and phoneme data;
[0033] In one embodiment, the multimodal feature data may include style hint text, style voice, speaker voice, etc.
[0034] In one embodiment, the multimodal feature data may be obtained from a variety of data sources, such as social media, news, movies, music, etc., and may also be obtained by recording a specific subject.
[0035] For example, in the field of medical intelligent voice question and answer, style hint texts can be collected from medical professional literature, authoritative medical websites, and examples of doctors' communication with patients. For example, collect gentle and encouraging language texts used by doctors to soothe patients' emotions, as well as easy-to-understand, rigorous and professional text content used when explaining the condition to patients. These texts can be used as style references to generate voice answers with corresponding styles in subsequent speech synthesis. By recording the voices of professional doctors in actual medical question and answer scenarios, style voice samples are obtained. For example, recording the voices of doctors in different scenarios such as asking patients about symptoms, informing test results, and conducting health education. These voices contain the doctor's tone, speed, pauses and other style features, which can provide style guidance for speech synthesis. Collect the voices of specific doctors as speaker reference audio to accurately clone the doctor's unique voice timbre. This can be achieved by having doctors read some common medical terms, question and answer sentences and other text content in a quiet environment to ensure that clear and high-quality voice samples are obtained so that the doctor's voice characteristics can be accurately restored in speech synthesis.
[0036] For example, in the application field of financial intelligent customer service, style prompt texts can be collected from news reports related to the financial industry, financial product manuals, customer service and customer communication scripts, and other channels. For example, collect professional, rigorous, and objective text descriptions used by customer service when introducing financial products to customers, as well as patient, sincere, and soothing text content used when handling customer complaints, as style reference texts to generate voice answers with corresponding styles. Record the voices of financial customer service personnel in actual customer service scenarios, including answering customer consultation calls, answering customer questions, and promoting financial products, to obtain style voice samples. These voice samples contain the style characteristics of customer service personnel such as intonation, speaking speed, and tone, which can provide style guidance for speech synthesis, so that the synthesized speech is more in line with the needs of financial customer service scenarios. Collect the voices of specific customer service personnel as speaker reference audio to accurately clone the unique voice timbre of customer service personnel. High-quality voice samples can be obtained by asking customer service personnel to read some financial terms, common question and answer sentences, and other text content in a quiet environment to ensure that the voice characteristics of customer service personnel can be accurately restored in speech synthesis, thereby improving customers' recognition and trust in customer service voice.
[0037] Generally, the collected multimodal feature data can be cleaned and preprocessed, including deduplication, scaling, normalization, format conversion, and labeling to ensure the quality of the multimodal feature data.
[0038] In one embodiment, a phoneme is the smallest unit in speech and can be identified by extracting acoustic features such as Mel-Frequency Cepstral Coefficients (MFCC). Generally, in order to accurately identify a phoneme, the boundary of a phoneme can be determined by analyzing whether the initial consonant or the final vowel is added when the posterior probability of each frame is the largest.
[0039] For example, a deep learning framework such as TensorFlow or PyTorch can be used to build a phoneme recognition model, which usually involves the use of convolutional neural networks (CNN) and recurrent neural networks (RNN). The trained model is applied to real-time phoneme recognition, which involves real-time audio signal processing and model prediction.
[0040] Specifically, professional audio processing software and speech analysis tools are used to pre-process the collected medical speech samples, including removing background noise, speech enhancement and other operations to improve the quality of speech signals. Acoustic feature extraction methods such as Mel-Frequency Cepstral Coefficient (MFCC) are used to analyze the processed speech signals and identify the phonemes therein.
[0041] For example, when analyzing the speech of a doctor asking a patient "Have you felt unwell recently?", the MFCC features of each frame of the speech signal are calculated and combined with the time domain and frequency domain information of the speech to determine the boundary of each phoneme, thereby accurately identifying the start and end times of each phoneme in the sentence, such as "you", "most", "near", etc.
[0042] For example, when analyzing the voice of a customer service staff member introducing a financial product to a customer, "This financial product has high returns and low risks," the MFCC features of each frame of the speech signal are calculated, and combined with the time domain and frequency domain information of the speech, the boundaries of each phoneme are determined, thereby accurately identifying the start and end times of each phoneme in the sentence, such as "this," "money," "management," and "wealth."
[0043] In one embodiment, the method of the present application allows control of style and speaker characteristics through text prompts and / or audio references, with the goal of improving the editability and naturalness of speech synthesis in existing research. The method of the present application can combine three input modalities: text prompts for natural, interactive conversations, style reference audio for precise style customization, and speaker reference audio for accurate zero-shot speaker identity cloning. This triple control input method enhances precise control of style elements and the speaker's unique voice timbre, fully realizing zero-shot capabilities in a multimodal environment.
[0044] S102, based on the first encoder, performing feature extraction and feature fusion on the multimodal feature data to generate speech embedding data;
[0045] In one embodiment, the present application can efficiently process multimodal inputs including text prompts, audio references, and speaker timbre through a universal front-end encoder, and generate independent style and speaker control embeddings. In addition, a hierarchical Conformer structure is used to fuse these control embeddings, in order to achieve the best feature fusion effect in the current advanced TTS system.
[0046] For example, for different modal data, appropriate techniques can be used to extract features. For example, a convolutional neural network (CNN) can be used to extract visual features from images, a recurrent neural network (RNN) can be used to extract speech features from audio, and a Transformer can be used to extract semantic features from text.
[0047] In one embodiment, the first encoder includes a style encoder and a speaker speech encoder; and the speech embedding data includes style embedding data and speaker speech embedding data.
[0048] Furthermore, based on the style encoder, the style prompt text and the style speech are encoded to generate the style embedding data; based on the speaker speech encoder, the speaker speech is encoded to generate the speaker speech embedding data.
[0049] In one embodiment, the purpose of the style encoder is to extract style features, which can be obtained from style hint text and style speech. The style encoder can learn highly expressive style representations through an inverse autoregressive flow (IAF) method. In addition, the style encoder also includes a reference encoder, an IAF flow, and a style classifier, and constrains the style representation of the source sentence to be close to the representation of the target style through a style difference loss.
[0050] For example, when you need to generate a voice response with a "gentle and soothing" style, the style prompt text can be "Please don't worry too much, we will have a suitable treatment plan", and the style voice is the voice sample of the doctor when he comforts the patient in the past. When you need to generate a voice response with a "professional and enthusiastic" style, the style prompt text can be "We are very happy to introduce the advantages of this financial product to you", and the style voice is the voice sample of the customer service staff when they enthusiastically introduced the product in the past.
[0051] The style encoder learns highly expressive style representations of these data through the inverse autoregressive flow (IAF) method, where the reference encoder extracts basic style features from style hint text and style speech, the IAF flow further enhances the expressiveness of style features, and the style classifier ensures that the extracted style features accurately correspond to the target style. Through the constraints of style difference loss, the style representation of the source sentence (such as the generated preliminary voice answer) can be close to the representation of the target style (such as the gentle and soothing style), thereby generating style embedding data that meets the requirements.
[0052] In one embodiment, the style embedding data is the output of the style encoder, which contains style features extracted from the style hint text and the style speech. These embedding data can be used to control the speech style in the speech synthesis system so that the generated speech has specific style features.
[0053] For example, when answering patients’ questions about more serious conditions, style embedding data can be used to generate a gentle and patient voice style, with a smooth tone and moderate speaking speed, so that patients can better accept information and reduce anxiety. When explaining treatment plans to patients, style embedding data can make the voice style professional and clear, ensuring that patients accurately understand the treatment steps and precautions. Or, when introducing financial products to customers, style embedding data can be used to generate a warm and professional voice style, with an upward tone, moderate speaking speed and infectiousness, to attract customers’ attention and increase their interest in the product. When answering customers’ questions about account security, style embedding data can make the voice style rigorous and reliable, reassuring customers.
[0054] In one embodiment, the speaker speech encoder is used to extract speaker features, which can be obtained from the speaker's speech. Generally, the speaker encoding module can be mainly composed of a three-layer LSTM and a speaker classifier, which is jointly trained with the rest of the network modules to learn discriminative speaker representations.
[0055] In one embodiment, the speaker speech embedding data is the output of the speaker speech encoder, which contains the speaker features extracted from the speaker's speech. These embedding data can be used for tasks such as speech recognition and speech synthesis to ensure that the generated speech retains the unique characteristics of the speaker.
[0056] For example, when a speech sample of a senior doctor (or an excellent customer service representative) is recorded, the speaker speech encoder can extract the unique voice features of the doctor, such as timbre, pronunciation habits, etc., and generate speaker speech embedding data. This allows the voice characteristics of the doctor (or customer service representative) to be maintained during speech synthesis, even if the answers are different, increasing the patient's trust in the doctor (or customer service representative) as if the doctor (or customer service representative) himself is communicating with him.
[0057] Exemplarily, the two submodules (style encoder and speaker voice encoder) in the first encoder perform feature extraction on the input multimodal data (style cue text, style voice and speaker voice) respectively. These feature extraction processes can be implemented by deep learning models such as CNN, RNN or Transformer, which can learn deep feature representations in the data.
[0058] In one embodiment, after extracting the style features and speaker features, the next step is to fuse these features together to generate speech embedding data. Specifically, it is implemented through an encoder-decoder framework, in which the outputs of the style encoding module and the speaker encoding module are simultaneously used as conditions for the generation module to perform speech style transfer. The fused features will form speech embedding data, including style embedding data and speaker speech embedding data, which are used for downstream tasks such as speech synthesis. They provide a compact and information-rich speech representation that can capture the style and speaker characteristics of the speech.
[0059] The style embedding data output by the style encoding module and the speaker voice embedding data output by the speaker encoding module are used as conditions for the generation module at the same time. For example, when answering patients' questions about how to use drugs, the style embedding data carries gentle and clear style features, and the speaker voice embedding data carries the voice features of a specific doctor. The two are fused in the generation module to form voice embedding data. The decoder generates a voice answer based on the voice embedding data, which has a specific style and maintains the unique voice features of the speaker, realizing voice style transfer and personalized voice synthesis, and providing patients with high-quality voice interaction services.
[0060] This embodiment generates style embedding data and speaker speech embedding data by extracting style features from style cue text and style speech, and extracting speaker features from speaker speech, thereby providing rich feature support for speech processing tasks.
[0061] For example, Figure 2 As shown, Figure 2 A data processing flow diagram of a speech synthesis model for implementing a speech synthesis method based on multimodal style embedding provided in this application. The speech synthesis model can represent the speaker's voice timbre and speech emotional style in a multimodal and zero-sample manner through a front-end universal style fusion module; through an enhanced hierarchical Conformer dual-branch style control module, effective feature fusion of zero-sample TTS is ensured.
[0062] The embodiment of the present application improves the decoupling of speaker and style modeling through a universal style fusion encoder for encoding and decoupling multiple control embeddings of speaker identity and emotion. This module facilitates the seamless integration of multimodal inputs, including text style cues, audio style references, and speaker voice features or timbre references, all in a completely zero-sample manner. In addition, by integrating a novel style control fusion module into the most advanced conditional CVAE-based VITS TTS model, optimal feature fusion is ensured and the high naturalness of speech synthesis is maintained.
[0063] By incorporating multimodal references into the VITS architecture of the speech synthesis model, optimal naturalness and enhanced control capabilities are provided for speaker and style information. The speech synthesis model provided in the embodiment of the present application adopts an architecture based on a flow-conditional variational autoencoder to generate audio x under the condition of input text t and control style and timbre embedding. The conversion of text to acoustic features is modeled as a normalized stream f(z), which maps acoustic features z to text features. The optimization process of the model includes maximizing the evidence lower bound (ELBO), which is consistent with the implementation method of the original VITS.
[0064] In one embodiment, model training data is obtained; based on the model training data, the first encoder is trained to obtain a model training prediction result; based on a preset classifier, the model training prediction result is classified and distinguished, and based on the classification and distinction result, the first encoder is iteratively optimized until the prediction accuracy of the model training prediction result is greater than or equal to the preset accuracy.
[0065] In one embodiment, a data set is collected for training the first encoder, and the data set should contain enough samples to cover various features that the model needs to learn. The model training data may include speech signals, text transcriptions, speaker information, and speech style labels.
[0066] Specifically, if the model is applied in the scenario of medical intelligent question-answering, the model training data includes voice signals, such as voice records of communication between doctors and patients, covering voice in various scenarios such as asking about the condition, explaining the diagnosis results, and informing treatment plans; text transcription, that is, accurate text transcription of the above voice records to form corresponding text data, so that the model can learn the correspondence between voice and text; speaker information, including voice feature information of different doctors, such as timbre, intonation, speaking speed, etc., which is used to train the speaker speech encoder so that it can accurately extract speaker features; voice style labels, for example, marking the style adopted by doctors when communicating with patients, such as gentle and soothing style, professional and rigorous style, etc., which are used to train the style encoder so that it can extract corresponding style features.
[0067] Specifically, if the model is applied in a financial intelligent customer service scenario, the model training data includes voice signals, such as voice records of customer service personnel communicating with customers, involving voice in various scenarios such as financial product consultation, account inquiry, and business processing guidance; text transcription, that is, accurate text transcription of the above voice records to form corresponding text data to help the model learn the mapping relationship between voice and text; speaker information, including voice feature information of different customer service personnel, such as timbre, intonation, speaking speed, etc., which is used to train the speaker speech encoder so that it can accurately extract speaker features; voice style labels, for example, marking the style adopted by customer service personnel when communicating with customers, such as enthusiastic and professional style, patient answering style, etc., which are used to train the style encoder so that it can extract corresponding style features.
[0068] In one embodiment, the first encoder is trained using the acquired training data. The encoder may be a deep learning model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer, which can learn feature representations from input data. During the training process, the model generates prediction results through forward propagation and adjusts model parameters through back propagation to minimize the loss function.
[0069] During the training process, the model will produce predictions. These results need to be evaluated to determine the performance of the model. The model training predictions are classified using a preset classifier. This classifier may be an independent model that verifies whether the output of the first encoder meets the expected classification or label.
[0070] In one embodiment, the performance of the first encoder is evaluated based on the classification discrimination result. If the prediction accuracy does not reach the preset accuracy, the first encoder needs to be adjusted and optimized, including adjusting the model architecture, changing hyperparameters, using different optimization algorithms or loss functions, etc. The model is trained for multiple iterations, and the model is evaluated using a preset classifier after each iteration, and optimized based on the evaluation results. The training and evaluation process is repeated until the prediction accuracy of the model training prediction result is greater than or equal to the preset accuracy. This preset accuracy is a performance target determined in advance to ensure that the model has sufficient reliability and effectiveness in practical applications.
[0071] In one embodiment, at least one classification loss is obtained based on the classification discrimination result; based on each of the classification losses, the first encoder is iteratively optimized; wherein the classification losses include text style loss, speech style loss, speaker loss and gradient reversal loss.
[0072] During the training process, an end-to-end approach can be used to combine the HiFi-GAN vocoder and its discriminator for overall optimization.
[0073] For example, Figure 3 The general style fusion encoder component shown is used to accurately generate style and speaker embeddings. These embeddings can effectively control the style and speaker-related properties of the base model and separate these two types of information from the multimodal input.
[0074] In one embodiment, since the multimodal input is entangled with multiple speech information, a gradient reversal layer (GRL) can be used to decouple the speaker and style embedding space, so that the representation of the speaker and style are separated. These components constitute the loss function of the front-end encoder loss and construct the loss of the front-end encoder.
[0075] Among them, text modality data provides a more stable and rough guidance, while audio serves as a supplementary reference for style customization. In order to promote these two style control inputs, a dropout mechanism is used during training to generate the final style embedding. Therefore, the dropout mechanism combines the audio emotion embedding and the prompt emotion embedding during training. The main learning goal is to train the front-end encoder, and the total loss in the whole optimization process is the combination of the front-end encoder loss and the original speech synthesis loss.
[0076] S103, based on the second encoder, extracting features from the phoneme data to generate text feature data;
[0077] In one embodiment, the second encoder can be a deep learning model, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM), for learning a deeper representation from the extracted phoneme features. The second encoder maps the features extracted from the phoneme data to a fixed-length vector for characterizing the acoustic and linguistic features of the phoneme data.
[0078] In one embodiment, the second encoder performs feature extraction on the phoneme data, and the generated text feature data may include acoustic information of the phonemes and may also include semantic information.
[0079] In one embodiment, the input text may be processed by a CLIP text encoder.
[0080] In one embodiment, if Figure 3 As shown, in order to enhance the text prompts of the style labels in the training dataset, the reference text training data can be enhanced using OpenAI's Large Language Model (LLM). This includes generating synonyms for the styles in the dataset to generate multiple keywords, using these keywords to create instructions, and guiding the LLM to directly generate sentences describing each style in the dataset to further enrich the data diversity. As for the reference audio, including audio for speaker cloning and style guidance, the linear spectrum of the audio can be extracted as a front-end feature.
[0081] S104: Based on a feature fusion model, perform feature fusion on the speech embedding data and the text feature data to generate stylized speech data.
[0082] This paper proposes a zero-shot TTS system based on prompts and / or audio references, controllable style and speakers. The optimal integration of speaker and style control embedding is achieved by innovatively adopting a universal style fusion module. The introduction of a universal front-end encoder promotes the effective use of multimodal inputs, including prompts and reference audio, and improves the separation effect. These proposed improvements significantly broaden the application scope and flexibility of TTS technology while maintaining naturalness.
[0083] In one embodiment, in order to further improve the speech naturalness of the TTS model, a dual-branch style control module can be used as an advanced style or speaker control fusion method. And further, the dual control embedding of decoupled speaker and style embedding is processed by a hierarchical Conformer module. This fusion method fuses the control embedding generated by the front-end module to the optimal position of the backbone network. The method hierarchically fuses the input style vector and speaker vector with the frame-level feature input suitable for all positions of the backbone network.
[0084] In one embodiment, the dual-branch style control module combines multi-head self-attention, gated recurrent units, and convolutional networks to achieve local and whole sentence focus. As an effective fusion strategy, speaker information is processed in a hierarchical manner first, followed by style information, treating style as a form of intra-class variation in each speaker modeling. The purpose is to accurately control speaker identity and treat style as a flexible variation for each speaker without losing the similarity of speaker timbre, thereby achieving better naturalness.
[0085] In one embodiment, the hierarchical Conformer module is integrated into the VITS backbone network to provide optimized style and speaker control functions. This function is implemented in three core modules of the backbone network: text content encoder, duration predictor, and text-to-acoustic flow.
[0086] Among them, the text content encoder is used to synthesize the required text content with different speakers and emotional styles.
[0087] The duration predictor is used to predict the best match between the text content frame and the acoustic frame, which is closely related to the content, style and speaker. Therefore, the duration input can be set as a combination of content, style and speaker. The specific process is as follows (taking the inference stage as an example):
[0088] D=f(C,S,P)+∈
[0089] Where D represents the predicted duration, and f is a function that combines content C, style S, and speaker features P to predict the duration. C represents content features, which can be text feature data extracted by a text encoder. S represents style features, which are extracted by a style encoder. P represents speaker features, which are extracted by a speaker encoder. ∈ is extracted from random noise to increase the diversity of prediction.
[0090] Here, the function f can be a deep learning model, such as a two-layer 1D convolutional network with ReLU activations, followed by layer normalization and dropout layers, and an additional linear layer to project the hidden state into the output sequence.
[0091] Then, after control processing, it is fed into the duration flow model to accurately predict the predicted duration of each text content frame under the condition of multiple control embeddings.
[0092] The text-to-acoustic normalization stream is a module designed to convert text to speech at the feature level, tightly coupled to the speaker and style, aiming to produce optimal prosodic and timbre vocoders for decoding into the final waveform.
[0093] Furthermore, based on the text content encoder, text matching is performed on the speech embedding data and the text feature data to determine the correspondence between the text content, speaker identity and speech style; based on the duration prediction model and the correspondence, duration prediction is performed on the text content in the text feature data to determine the predicted duration of each text content; based on the correspondence and the predicted duration, feature matching and fusion are performed on the speech embedding data and the text feature data to generate the stylized speech data.
[0094] In one embodiment, this can be achieved by comparing the similarity of the embedded data, for example, using a metric such as cosine similarity to match the speech embedding data (including style embedding data and speaker speech embedding data) with the text feature data, thereby determining the correspondence between the text content, speaker identity, and speech style. The duration of each text unit (such as a phoneme or word) is estimated to determine their duration in the final speech synthesis.
[0095] Furthermore, based on the corresponding relationship, the speaker identity and voice style corresponding to each of the text contents are determined; based on the predicted duration, the speaker identity and the voice style corresponding to each of the text contents, the audio configuration parameters of each of the text contents are determined; based on the voice conversion algorithm, each of the text contents is converted into audio according to the corresponding audio configuration parameters to generate the stylized voice data.
[0096] In one embodiment, style and speaker features are combined with features of the text content to generate stylized speech data. Fusion can be achieved in a variety of ways, such as by weighted averaging, feature interpolation, or using a deep learning model to learn the best combination of features. Through the above feature fusion process, stylized speech data is generated. These data not only contain the content information of the original text, but also incorporate the specific speaker identity and voice style, so that the synthesized speech not only conveys the correct information, but also has the required expressiveness and personality.
[0097] Generally, duration prediction models can use machine learning methods, such as recurrent neural networks, to learn the relationship between text features and duration.
[0098] Furthermore, based on the audio configuration parameters, a mapping relationship between each of the text contents and the acoustic features is determined; based on the mapping relationship, each of the text contents is converted into the corresponding acoustic features; based on the speech conversion algorithm, the acoustic features corresponding to each of the text contents are converted into audio data to generate the stylized speech data.
[0099] Generally speaking, the main function of the normalized stream module of text-to-acoustic is to convert text to speech at the feature level. It is closely related to the speaker and style, and can generate a voice code with optimal prosody and sound quality for decoding into the final waveform.
[0100] In one embodiment, the normalization flow is a reversible transformation model that can map text features to acoustic feature space through a series of transformations. These transformations are usually parameterized and can be trained to learn the mapping relationship between text and acoustic features.
[0101] During the training process, the model will learn the rhythm and sound quality characteristics of different speakers and styles. For example, for different speakers, the model will learn their respective voice characteristics, such as high pitch, low pitch, brightness of timbre, etc. For different styles, the model will learn the corresponding rhythm changes. For example, in an exciting style, the tone and rhythm of the voice will change significantly.
[0102] In the inference phase, given the text features and related information such as speaker and style, the normalized text-to-acoustic stream can generate the corresponding acoustic features through these reversible transformations, and then the decoder converts the acoustic features into the final speech waveform.
[0103] Specifically, each text content can be converted into audio according to the corresponding audio configuration parameters through a speech synthesis model. For example, voice cloning can be performed using an end-to-end speech synthesis model such as FastSpeech2, which can identify the speaker and generate speech with the desired style. In addition, a speech style conversion technology based on StarGAN-VC can be used, which achieves efficient speech conversion through feature extraction and fundamental frequency conversion methods. Different voices with natural rhythms are synthesized from reference speech utterances, and through self-supervised learning of speaking styles, voices with the same rhythm and emotional tone as any given reference speech are synthesized.
[0104] This embodiment provides a speech synthesis method based on multimodal style embedding. The method of the present application extracts and fuses multimodal features through a first encoder to generate speech embedding data, which helps to capture and retain the style and emotional features in the original speech signal. The second encoder extracts features from phoneme data to generate text feature data, which helps to understand the text content and convert it into a synthesizable speech signal. The speech embedding data and text feature data are combined through a feature fusion model to generate styled speech data. When synthesizing speech, both the naturalness of the speech and the style of the text are considered, thereby achieving richer and more natural style diversity in zero-sample speech synthesis.
[0105] See also Figure 4 , Figure 4 It is a structural schematic diagram of a first embodiment of a speech synthesis device based on multimodal style embedding provided in the present application. The speech synthesis device based on multimodal style embedding is used to execute the aforementioned speech synthesis method based on multimodal style embedding.
[0106] like Figure 4 As shown, the speech synthesis device 200 based on multimodal style embedding includes: a data acquisition module 201, an embedded data encoding module 202, a text feature extraction module 203 and a feature fusion module 204.
[0107] A data acquisition module 201 is used to acquire multimodal feature data and phoneme data;
[0108] The embedded data encoding module 202 is used to extract and fuse the multimodal feature data based on the first encoder to generate speech embedded data;
[0109] A text feature extraction module 203, configured to extract features from the phoneme data based on a second encoder to generate text feature data;
[0110] The feature fusion module 204 is used to perform feature fusion on the speech embedding data and the text feature data based on a feature fusion model to generate stylized speech data.
[0111] In one embodiment, the multimodal feature data includes style cue text, style speech and speaker speech; the first encoder includes a style encoder and a speaker speech encoder; and the speech embedding data includes style embedding data and speaker speech embedding data.
[0112] The embedded data encoding module 202 includes:
[0113] A style encoding unit, configured to encode the style prompt text and the style speech based on the style encoder to generate the style embedding data;
[0114] A speaker speech encoding unit is used to encode the speaker speech based on the speaker speech encoder to generate the speaker speech embedding data.
[0115] In one embodiment, the feature fusion module 204 includes:
[0116] A correspondence determination submodule, configured to perform text matching on the speech embedding data and the text feature data based on a text content encoder to determine a correspondence between text content, speaker identity, and speech style;
[0117] A predicted duration determination submodule, configured to predict the duration of the text content in the text feature data based on the duration prediction model and the corresponding relationship, and determine the predicted duration of each of the text contents;
[0118] A feature fusion submodule is used to perform feature matching and fusion on the speech embedding data and the text feature data based on the corresponding relationship and the predicted duration to generate the stylized speech data.
[0119] In one embodiment, the feature fusion submodule includes:
[0120] A voice style determination unit, configured to determine the speaker identity and voice style corresponding to each of the text contents based on the corresponding relationship;
[0121] An audio configuration parameter determination unit, configured to determine the audio configuration parameters of each of the text contents based on the predicted duration, the speaker identity, and the voice style corresponding to each of the text contents;
[0122] The audio conversion unit is used to perform audio conversion on each of the text contents according to the corresponding audio configuration parameters based on a speech conversion algorithm to generate the stylized speech data.
[0123] In one embodiment, the audio conversion unit includes:
[0124] A mapping relationship determination subunit, used to determine the mapping relationship between each of the text contents and the acoustic features based on the audio configuration parameters;
[0125] A data conversion subunit, configured to convert each of the text contents into the corresponding acoustic features based on the mapping relationship;
[0126] The stylized speech data generating subunit is used to convert the acoustic features corresponding to each of the text contents into audio data based on the speech conversion algorithm to generate the stylized speech data.
[0127] In one embodiment, the speech synthesis apparatus 200 based on multimodal style embedding further includes a model training module, including:
[0128] A training data acquisition unit, used to acquire model training data;
[0129] A model training result obtaining unit, used to train the first encoder based on the model training data to obtain a model training prediction result;
[0130] The model optimization unit is used to classify and distinguish the model training prediction results based on a preset classifier, and iteratively optimize the first encoder based on the classification and distinction results until the prediction accuracy of the model training prediction results is greater than or equal to a preset accuracy.
[0131] In one embodiment, the model optimization unit includes:
[0132] A model loss obtaining subunit, used to obtain at least one classification loss based on the classification discrimination result;
[0133] A first encoder optimization subunit, configured to iteratively optimize the first encoder based on each of the classification losses;
[0134] The classification loss includes text style loss, speech style loss, speaker loss and gradient reversal loss.
[0135] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the above-described device and each module can refer to the corresponding process in the aforementioned embodiment of the speech synthesis method based on multimodal style embedding, and will not be repeated here.
[0136] The apparatus provided in the above embodiment may be implemented in the form of a computer program. The computer program may be Figure 5 Runs on the computer device shown.
[0137] See also Figure 5 , Figure 5 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.
[0138] See also Figure 5 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0139] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any speech synthesis method based on multimodal style embedding.
[0140] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0141] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any speech synthesis method based on multimodal style embedding.
[0142] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0143] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0144] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0145] Acquire multimodal feature data and phoneme data;
[0146] Based on the first encoder, perform feature extraction and feature fusion on the multimodal feature data to generate speech embedding data;
[0147] Based on the second encoder, feature extraction is performed on the phoneme data to generate text feature data;
[0148] Based on the feature fusion model, the speech embedding data and the text feature data are fused to generate stylized speech data.
[0149] In one embodiment, the multimodal feature data includes style hint text, style speech, and speaker speech; the first encoder includes a style encoder and a speaker speech encoder; the speech embedding data includes style embedding data and speaker speech embedding data;
[0150] When the processor extracts and fuses the multimodal feature data based on the first encoder to generate speech embedding data, the processor is used to implement:
[0151] Based on the style encoder, encoding the style prompt text and the style speech to generate the style embedding data;
[0152] Based on the speaker speech encoder, the speaker speech is encoded to generate the speaker speech embedding data.
[0153] In one embodiment, when the processor implements the feature fusion model based on which the speech embedding data and the text feature data are subjected to feature fusion to generate stylized speech data, it is configured to implement:
[0154] Based on a text content encoder, text matching is performed on the speech embedding data and the text feature data to determine the correspondence between text content, speaker identity, and speech style;
[0155] Based on the duration prediction model and the corresponding relationship, predict the duration of the text content in the text feature data to determine the predicted duration of each text content;
[0156] Based on the corresponding relationship and the predicted duration, the speech embedding data and the text feature data are subjected to feature matching and fusion to generate the stylized speech data.
[0157] In one embodiment, when the processor performs feature matching and fusion of the speech embedding data and the text feature data based on the corresponding relationship and the predicted duration to generate the stylized speech data, the processor is used to implement:
[0158] Based on the corresponding relationship, determining the speaker identity and voice style corresponding to each of the text contents;
[0159] Determining audio configuration parameters of each of the text contents based on the predicted duration, the speaker identity, and the voice style corresponding to each of the text contents;
[0160] Based on the speech conversion algorithm, each of the text contents is subjected to audio conversion according to the corresponding audio configuration parameters to generate the stylized speech data.
[0161] In one embodiment, when the processor implements the speech conversion algorithm to perform audio conversion on each of the text contents according to the corresponding audio configuration parameters to generate the stylized speech data, it is used to implement:
[0162] Based on the audio configuration parameters, determining a mapping relationship between each of the text contents and the acoustic features;
[0163] Based on the mapping relationship, convert each of the text contents into the corresponding acoustic features;
[0164] Based on the speech conversion algorithm, the acoustic features corresponding to each of the text contents are converted into audio data to generate the stylized speech data.
[0165] In one embodiment, before implementing the step of extracting features and fusing features on the multimodal feature data based on the first encoder to generate speech embedding data, the processor is further configured to implement:
[0166] Get model training data;
[0167] Based on the model training data, training the first encoder to obtain a model training prediction result;
[0168] Based on a preset classifier, the model training prediction result is classified and judged, and the first encoder is iteratively optimized based on the classification and judgment result until the prediction accuracy of the model training prediction result is greater than or equal to the preset accuracy.
[0169] In one embodiment, when implementing the iterative optimization of the first encoder based on the classification discrimination result, the processor is used to implement:
[0170] Based on the classification and discrimination result, obtaining at least one classification loss;
[0171] Iteratively optimizing the first encoder based on each of the classification losses;
[0172] The classification loss includes text style loss, speech style loss, speaker loss and gradient reversal loss.
[0173] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement any one of the speech synthesis methods based on multimodal style embedding provided in the embodiments of the present application.
[0174] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.
[0175] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A speech synthesis method based on multimodal style embedding, characterized in that: The method comprises: Acquire multimodal feature data and phoneme data; Based on the first encoder, perform feature extraction and feature fusion on the multimodal feature data to generate speech embedding data; Based on the second encoder, feature extraction is performed on the phoneme data to generate text feature data; Based on the feature fusion model, the speech embedding data and the text feature data are fused to generate stylized speech data.
2. The method for speech synthesis based on multimodal style embedding according to claim 1, characterized in that: The multimodal feature data includes style hint text, style speech and speaker speech; the first encoder includes a style encoder and a speaker speech encoder; the speech embedding data includes style embedding data and speaker speech embedding data; The step of extracting and fusing features of the multimodal feature data based on the first encoder to generate speech embedding data includes: Based on the style encoder, encoding the style prompt text and the style speech to generate the style embedding data; Based on the speaker speech encoder, the speaker speech is encoded to generate the speaker speech embedding data.
3. The method for speech synthesis based on multimodal style embedding according to claim 1, characterized in that: The step of performing feature fusion on the speech embedding data and the text feature data based on the feature fusion model to generate stylized speech data includes: Based on a text content encoder, text matching is performed on the speech embedding data and the text feature data to determine the correspondence between text content, speaker identity, and speech style; Based on the duration prediction model and the corresponding relationship, predict the duration of the text content in the text feature data to determine the predicted duration of each text content; Based on the corresponding relationship and the predicted duration, the speech embedding data and the text feature data are subjected to feature matching and fusion to generate the stylized speech data.
4. The method for speech synthesis based on multimodal style embedding according to claim 3, characterized in that: The step of performing feature matching and fusing the speech embedding data with the text feature data based on the corresponding relationship and the predicted duration to generate the stylized speech data includes: Based on the corresponding relationship, determining the speaker identity and voice style corresponding to each of the text contents; Determining audio configuration parameters of each of the text contents based on the predicted duration, the speaker identity, and the voice style corresponding to each of the text contents; Based on the speech conversion algorithm, each of the text contents is subjected to audio conversion according to the corresponding audio configuration parameters to generate the stylized speech data.
5. The method for speech synthesis based on multimodal style embedding according to claim 4, characterized in that: The method of converting each of the text contents into audio according to the corresponding audio configuration parameters based on the voice conversion algorithm to generate the stylized voice data includes: Based on the audio configuration parameters, determining a mapping relationship between each of the text contents and the acoustic features; Based on the mapping relationship, convert each of the text contents into the corresponding acoustic features; Based on the speech conversion algorithm, the acoustic features corresponding to each of the text contents are converted into audio data to generate the stylized speech data.
6. The method for speech synthesis based on multimodal style embedding according to claim 1, characterized in that: Before the method extracts and fuses the multimodal feature data based on the first encoder to generate speech embedding data, the method further includes: Get model training data; Based on the model training data, training the first encoder to obtain a model training prediction result; Based on a preset classifier, the model training prediction result is classified and judged, and the first encoder is iteratively optimized based on the classification and judgment result until the prediction accuracy of the model training prediction result is greater than or equal to the preset accuracy.
7. The method for speech synthesis based on multimodal style embedding according to claim 6, characterized in that: The iterative optimization of the first encoder based on the classification and discrimination result includes: Based on the classification and discrimination result, obtaining at least one classification loss; Iteratively optimizing the first encoder based on each of the classification losses; The classification loss includes text style loss, speech style loss, speaker loss and gradient reversal loss.
8. A speech synthesis device based on multimodal style embedding, characterized in that: The speech synthesis device based on multimodal style embedding includes: A data acquisition module, used to acquire multimodal feature data and phoneme data; An embedded data encoding module, used for performing feature extraction and feature fusion on the multimodal feature data based on the first encoder to generate speech embedded data; A text feature extraction module, used for extracting features from the phoneme data based on a second encoder to generate text feature data; The feature fusion module is used to perform feature fusion on the speech embedding data and the text feature data based on a feature fusion model to generate stylized speech data.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the speech synthesis method based on multimodal style embedding as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the speech synthesis method based on multimodal style embedding as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Digital human speech synthesis method and system based on multi-modal speech feature fusion
CN120833777A
Speech synthesis method, related device and computer program product
CN121306090A