Voice generation method and device, equipment and medium
By constructing a multilingual speech synthesis model and optimization parameters, the problem of low-resource language speech synthesis relying on high-quality data is solved, and natural and smooth speech is generated under low data conditions, improving the quality of speech synthesis.
Patent Information
- Application Number
- CN202510366622.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art highly relies on high-quality paired speech and text data in speech synthesis in low-resource languages, resulting in a decrease in the quality of speech synthesis when the number of available data is small, making it difficult to adapt to multilingual scenarios.
Build a multilingual speech synthesis model containing language-aware embedding layers, encoders and decoders. By obtaining the plain text data set of the target language and paired speech text data, an extended vocabulary is built, and model parameters are optimized through masking language models and supervising training to generate natural and smooth speech.
It improves the speech generation ability of low-resource languages, enhances the accuracy of input text conversion, ensures that the model can still adapt in a low data environment, improves the naturalness and fluency of generated speech, and improves the speech synthesis quality of low-resource languages.
Smart Images

Figure CN120148474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a voice generation method, device, equipment and storage medium. Background Art
[0002] In recent years, neural text-to-speech (TTS) systems have made significant breakthroughs in the development of speech synthesis technology and can generate near-natural speech. However, these deep learning-based TTS systems rely on large-scale and high-quality speech data, which limits their application scope. Currently, TTS systems mainly target resource-rich languages, while low-resource languages globally still face technical bottlenecks and it is difficult to build high-quality TTS models. This limitation brings many challenges in fields such as artificial intelligence, healthcare, and finance.
[0003] There are still multiple technical deficiencies in existing TTS systems in the field of speech synthesis. TTS systems perform excellently on resource-rich languages but are limited on low-resource languages. Existing systems highly rely on large-scale and high-quality paired speech-text data, enabling a few resource-rich languages to benefit, while a large number of low-resource languages are difficult to build TTS models due to the lack of standardized corpora, resulting in a decline in speech generation quality. In recent years, multilingual TTS has been used to make up for the shortage of low-resource language data, but this method relies on existing paired data of the target language. When the paired data of the target language is limited or the quality of the speech data is insufficient, the synthesis effect significantly decreases, manifested as problems such as discontinuous speech and unstable timbre.
[0004] In the field of healthcare, TTS technology is widely used in automatic reading of medical reports, medical Q&A systems, assistive communication devices, etc. However, the application of existing TTS systems in this field still has deficiencies. Medical terms are complex and involve multiple languages. Existing TTS systems are often optimized for specific languages, resulting in a high pronunciation error rate when it is necessary to support multilingual medical terms (such as Latin terms or cross-language professional vocabulary), affecting the accurate transmission of medical reports. Some medical scenarios (such as speech synthesis for language-impaired patients) require generating synthetic speech based on the patient's personalized speech data. However, existing TTS systems rely on a large amount of data for training and are difficult to generate highly personalized speech models under low-data conditions. Telemedicine and intelligent medical devices increasingly rely on voice interaction, but existing TTS systems are difficult to adapt to low-resource languages, resulting in patients in many regions being unable to enjoy voice-based medical services.
[0005] In the financial field, TTS technology is applied to intelligent customer service, financial report reading, transaction reminders, etc. However, the applicability of existing TTS systems in this field is still limited. The financial industry contains a large number of specific terms and numerical combinations (such as currency units, stock codes, investment terms, etc.). Currently, when existing TTS systems process the conversion of financial terms in different languages, there may be situations such as interrupted speech, incorrect pronunciation, or unnatural intonation, affecting the accuracy of financial information. Global financial operations involve multiple languages, but there is insufficient support for low-resource languages in the TTS field, making it difficult for certain markets (such as emerging economies) to build high-quality financial voice assistants and reducing the availability of automated financial services. The financial industry has extremely high requirements for data security and compliance. When existing TTS systems process financial sensitive information, it is difficult to meet strict security standards, and there are problems such as voice synthesis information leakage and voice forgery fraud risks.
[0006] Neural TTS technology still faces many challenges in the support of low-resource languages, the adaptation of terms in the medical and health fields, and the language adaptability and security in the financial field. These technical bottlenecks limit the wide application of TTS systems in multiple industries. Especially in application scenarios with high-precision requirements such as medical and health and finance, how to achieve high-quality speech synthesis under low-data conditions remains a difficult point that current technologies are hard to break through. Summary of the Invention
[0007] The main purpose of the present invention is to provide a speech generation method, device, equipment, and storage medium, aiming to solve the technical problem that in the speech synthesis of low-resource languages in the prior art, it highly depends on high-quality paired speech and text data. When the number of available paired data is small, the speech synthesis quality drops significantly, making it difficult to adapt to multi-language scenarios and resulting in the inability of the speech synthesis system for low-resource languages to generate natural and fluent speech.
[0008] To achieve the above object, the present invention provides a speech generation method, including:
[0009] Construct a multi-language speech synthesis model including a language perception embedding layer, an encoder, and a decoder;
[0010] Obtain a pure text data set of the target language and the corresponding paired speech text data;
[0011] Construct an extended vocabulary of the target language;
[0012] Use the pure text data set to update the parameters of the language perception embedding layer through a masked language model;
[0013] Use the paired speech text data to update the parameters of the multi-language speech synthesis model;
[0014] Convert the target input text into an input token sequence according to the extended vocabulary;
[0015] Convert the input token sequence into an initial vector representation through the language-aware embedding layer;
[0016] Extract the context semantic features of the initial vector representation through the encoder;
[0017] When the extended vocabulary contains phonetic symbols, extract the pronunciation rule features from the phonetic symbols;
[0018] Generate an acoustic feature sequence by fusing the context semantic features and the pronunciation rule features through the decoder;
[0019] Convert the acoustic feature sequence into target speech data.
[0020] Furthermore, to achieve the above object, the present invention provides a speech generation device, including:
[0021] A model construction module for constructing a multilingual speech synthesis model including a language-aware embedding layer, an encoder, and a decoder;
[0022] A data acquisition module for acquiring a pure text data set of the target language and the corresponding paired speech text data;
[0023] A vocabulary construction module for constructing an extended vocabulary of the target language;
[0024] An unsupervised training module for updating the parameters of the language-aware embedding layer using the pure text data set through a masked language model;
[0025] A supervised training module for updating the parameters of the multilingual speech synthesis model using the paired speech text data;
[0026] A text processing module for converting the target input text into an input token sequence according to the extended vocabulary;
[0027] An embedding representation module for converting the input token sequence into an initial vector representation through the language-aware embedding layer;
[0028] A feature extraction module for extracting the context semantic features of the initial vector representation through the encoder;
[0029] A phoneme feature extraction module for extracting pronunciation rule features from the phonetic symbols when the extended vocabulary contains phonetic symbols;
[0030] A feature fusion module for generating an acoustic feature sequence by fusing the context semantic features and the pronunciation rule features through the decoder;
[0031] A voice generation module, configured to convert the acoustic feature sequence into target voice data.
[0032] Furthermore, to achieve the above object, the present invention further provides a computer device, which includes a memory, a processor, and a voice generation program stored on the memory and executable on the processor. When the voice generation program is executed by the processor, the steps of the voice generation method as described above are implemented.
[0033] Furthermore, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a voice generation program is stored. When the voice generation program is executed by a processor, the steps of the voice generation method as described above are implemented.
[0034] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. A voice generation method is disclosed, including: constructing a multilingual speech synthesis model, obtaining a pure text dataset and paired speech text data in a target language, and constructing an extended vocabulary; updating the parameters of the language-aware embedding layer using the pure text data, and updating the model parameters using the paired speech text data. Converting the target input text into an input token sequence according to the extended vocabulary, obtaining an initial vector representation through the language-aware embedding layer, and the encoder extracting context semantic features; when the extended vocabulary contains phonetic symbols, extracting pronunciation rule features; the decoder fusing the context semantic features and the pronunciation rule features to generate an acoustic feature sequence and converting it into target voice data. The present invention combines a multilingual speech synthesis model with a language-aware embedding layer to improve the voice generation ability of low-resource languages; the construction method of the extended vocabulary enhances the accuracy of input text conversion; the unsupervised training of the masked language model enables the model to still learn the features of the target language when there is a lack of paired data; combining limited paired speech text data for supervised training improves the adaptability of the model in a low-data environment; fusing context semantic features and pronunciation rule features enhances the naturalness and fluency of the generated speech, improving the speech synthesis quality of low-resource languages. Description of the Drawings
[0035] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0036] Figure 1 It is a schematic diagram of an application environment of the voice generation method in an embodiment of the present invention;
[0037] Figure 2 It is a schematic flowchart of an embodiment of the voice generation method of the present invention;
[0038] Figure 3Schematic diagram of functional modules of a preferred embodiment of the voice generation device of the present invention;
[0039] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0040] Figure 5 Schematic diagram of another structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] The voice generation method provided by the embodiment of the present invention can be applied to an application environment such as Figure 1 where the client communicates with the server through a network. The server can build a multilingual speech synthesis model through the client, obtain a pure text dataset and paired speech text data in the target language, and build an extended vocabulary; use the pure text data to update the parameters of the language-aware embedding layer, and use the paired speech text data to update the model parameters. Convert the target input text into an input token sequence according to the extended vocabulary, obtain an initial vector representation through the language-aware embedding layer, and the encoder extracts context semantic features; when the extended vocabulary contains phonetic symbols, extract pronunciation rule features; the decoder fuses the context semantic features and pronunciation rule features to generate an acoustic feature sequence and convert it into target speech data. The present invention combines a multilingual speech synthesis model with a language-aware embedding layer to improve the voice generation ability of low-resource languages; the construction method of the extended vocabulary enhances the accuracy of input text conversion; the unsupervised training of the masked language model enables the model to still learn the features of the target language when there is a lack of paired data; combining limited paired speech text data for supervised training improves the adaptability of the model in a low-data environment; fusing context semantic features and pronunciation rule features enhances the naturalness and fluency of the generated speech, and improves the quality of speech synthesis of low-resource languages. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0043] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the voice generation method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0044] As Figure 2 shown, the voice generation method proposed by the present invention includes the following steps:
[0045] S10, construct a multilingual speech synthesis model including a language-aware embedding layer, an encoder, and a decoder;
[0046] In this embodiment, in the multilingual speech synthesis task, a unified model that can adapt to different language characteristics is required to improve the speech generation ability of low-resource languages. To this end, the core of model construction lies in the design of the language-aware embedding layer, the encoder, and the decoder. These parts work together to enable the system to generate natural and fluent speech in the absence of large-scale paired data.
[0047] The core of the multilingual speech synthesis model consists of a language-aware embedding layer, an encoder, and a decoder. The role of this structure is to efficiently model the mapping relationship from multilingual text to speech, reduce the dependence on large-scale paired data, and enhance the adaptability to low-resource languages.
[0048] The design of the language-aware embedding layer is to unify the text representations of different languages, enabling low-resource languages to learn from the speech data of other languages. This embedding layer can be optimized by unsupervised pre-training, that is, using large-scale multilingual text data to learn cross-lingual lexical representations based on the Masked Language Model (MLM). In the specific implementation process, this embedding layer can adopt a Transformer-based text encoder, using the Self-Attention mechanism to capture the shared information between different languages, making the representations of semantically similar words in different languages closer, thereby enhancing the cross-lingual generalization ability of the model. In addition, the output of this embedding layer needs to match the input dimension of the encoder to ensure that information can be transmitted to subsequent calculation modules without loss.
[0049] The role of the encoder is to extract context semantic features from the input text token sequence to provide high-quality feature representations for the decoder to generate speech. The encoder can adopt a stacked Transformer architecture, where each encoding layer contains a Multi-Head Self-Attention mechanism and a Feed-Forward Network. The Multi-Head Self-Attention mechanism enables the model to focus on the key information of the input text in different feature dimensions, enhancing the ability to model complex language dependencies. During training, the encoder can adopt a shared parameter method, that is, the encoding layers of low-resource languages can share some parameters with resource-rich languages, enabling low-resource languages to improve their modeling ability through transfer learning.
[0050] The task of the decoder is to generate the corresponding acoustic feature sequence based on the context semantic features output by the encoder and combined with the pronunciation rule features of the target language. The decoder can adopt a multi-layer Transformer structure, where each layer includes a Masked Multi-Head Self-Attention mechanism, a Cross-Attention mechanism, and a feed-forward network. The Masked Multi-Head Self-Attention mechanism ensures that the decoder can only rely on the current and previous features when generating speech, thus avoiding training biases caused by data leakage. The Cross-Attention mechanism is used to associate the output of the encoder, enabling the decoder to focus on different parts of the input text and improving the accuracy of speech synthesis. In addition, the decoder can perform inference in an autoregressive or non-autoregressive manner during the generation process. The autoregressive manner is suitable for high-precision speech synthesis, while the non-autoregressive manner can improve the inference speed and is suitable for real-time speech synthesis tasks.
[0051] In practical applications, different technical means can be adopted to construct a multilingual speech synthesis model to adapt to different application scenarios and computing resource environments.
[0052] In one implementation, the language-aware embedding layer can be initialized with a pre-trained language model (such as mBERT or XLM-R), enabling the embedding layer to learn cross-lingual general semantic representations using large-scale multilingual data. This approach is suitable for large-scale server-side training and can effectively improve the model's adaptability to low-resource languages.
[0053] In another implementation, the structure of the encoder can be adjusted according to the limitations of computing resources. For example, when computing resources are sufficient, a full Transformer structure can be adopted to learn the deep features of the input text using the global attention mechanism. In scenarios with limited computing resources, a hybrid architecture can be used, that is, a convolutional neural network (CNN) is used for local feature extraction in some layers, while a Transformer structure is used in higher layers to reduce the computational complexity while retaining the global feature extraction ability.
[0054] For example:
[0055]
[0056] This formula represents the set of model initialization parameters, which come from different training stages. Among them, θ initial is initialized with the model parameters of the multilingual pre-training stage; represents the decoder parameters in the multi-task training (mt) stage; represents the parameters of the language-aware embedding layer in the multi-task training stage; Represents the parameters of the language perception embedding layer in the multilingual training stage.
[0057] By constructing a multilingual speech synthesis model that includes a language perception embedding layer, an encoder, and a decoder, the speech generation ability of low-resource languages is improved. The language perception embedding layer can utilize multilingual unsupervised learning to capture the features of the target language in the absence of large-scale paired data, thereby reducing the dependence of low-resource languages on high-quality paired data. The encoder combines cross-lingual transfer learning strategies, enabling low-resource languages to learn the general text-to-speech mapping relationship from resource-rich languages and improving the synthesis quality. The decoder achieves efficient speech synthesis through the cross-attention mechanism and autoregressive or non-autoregressive inference methods, taking into account the adaptability of computational resources while ensuring the naturalness of speech.
[0058] S20, obtain the plain text dataset of the target language and the corresponding paired speech text data;
[0059] In this embodiment, in the speech synthesis task of low-resource languages, the plain text data and paired speech text data of the target language are the core basis for building the model. The quality and quantity of this part of the data directly determine the effect of speech synthesis. However, due to the lack of large-scale paired data in low-resource languages, data acquisition needs to combine multiple data sources and data augmentation strategies to ensure that the speech synthesis system has good adaptability and generalization ability.
[0060] The plain text data of the target language is mainly used for unsupervised pre-training. This dataset can be from sources such as public corpora, news texts, social media data, books, or academic literature. During data acquisition, it is necessary to clean and standardize the text data, including removing special symbols, normalizing punctuation, and handling text encoding problems, etc., to ensure data quality. To enhance the cross-lingual ability of the model, multilingual text alignment technology can also be used, that is, semantically matching the text of high-resource languages with the text of the target language, thereby helping the model learn richer language representations.
[0061] The paired speech text data is the key data source for training the speech synthesis model. Paired data refers to the text of the target language and its corresponding real speech recordings. This data usually comes from existing speech databases, speech transcription data, manually recorded data, or TTS synthesis enhanced data. Since the paired data of low-resource languages is limited, data augmentation techniques can be used to expand the dataset, such as:
[0062] Speech time stretching: Scale the existing speech in time to simulate speech data at different speeds;
[0063] Pitch transformation: Adjust the pitch of the audio to enable the model to better adapt to speech variations;
[0064] Speech splicing: Combine multiple short speech samples into longer speech segments to improve the model's learning ability for long speech.
[0065] Data migration: Utilize paired speech-text data in a resource-rich language and perform transfer training in combination with the speech features of the target language.
[0066] During the data acquisition process, a speech alignment algorithm can also be combined to ensure the alignment accuracy between text and speech, thereby improving the effectiveness of the training data. For example, dynamic time warping (DTW) or an HMM (Hidden Markov Model)-based method can be used for speech alignment to ensure that the model learns the correct speech-text mapping relationship.
[0067] In one implementation, existing speech databases can be directly used, such as open speech datasets (such as Mozilla CommonVoice, LJSpeech, LibriTTS, etc.), to extract the pure text data and paired speech-text data of the target language. For target languages without publicly available datasets, pure text data can be obtained through web crawling from news websites, government documents, books, etc., and an automatic speech recognition (ASR) system can be used for speech transcription to obtain paired speech data.
[0068] In another implementation, paired speech data can be obtained through manual recording. For example, native speakers can be invited to read the specified text, and high-quality recording equipment can be used for collection. To ensure data quality, the recording environment needs to be controlled, such as avoiding background noise and maintaining a consistent speaking speed and pitch. In addition, to further expand the speech data of low-resource languages, a TTS-synthesis enhancement method can also be adopted, that is, using the TTS system of a high-resource language to synthesize the speech of the target language and combining it with real speech for training to improve the generalization ability of the model.
[0069] In another implementation, a data migration strategy can be adopted, that is, utilize the existing speech data of a multilingual TTS system and combine it with the fine-tuning method to enable the model to better adapt to the target language. For example, a high-resource language with language characteristics similar to the target language (such as a language with a similar phonological structure) can be selected, and the speech features can be gradually adjusted to make it more in line with the phonological characteristics of the target language.
[0070] For example:
[0071]
[0072] This formula describes how the training data for low-resource languages is constituted, indicating that the dataset for the target language can be composed of letter-level labeled data and phonetic-level labeled data. When there is a lack of character-to-phoneme conversion tools for the target language, only letter-level labeled data is used for training. This formula shows that the dataset D of the target language l is composed of a dataset of letter sequences and a dataset of phonetic labels . If there is no available character-to-phoneme conversion tool for the target language, then is empty. Where:
[0073] D l (Dataset of the target language): The dataset used to train the speech synthesis model of the target language.
[0074] (Letter-level dataset): Represents text input using the letters of the target language, applicable in cases where there is no phoneme conversion tool.
[0075] (Phonetic-level dataset): Represents text input using the phonemes of the target language to improve the quality of speech synthesis.
[0076] Example illustration: In the field of medical and health, obtaining pure text data and paired speech data in the target language is crucial for applications such as automatic reading of medical reports and intelligent medical Q&A systems. For example, in remote areas, medical institutions may need to provide voice consultation services to patients in low-resource languages, but existing TTS systems cannot accurately synthesize speech due to the lack of target language data. Text and paired speech data in the target language can be extracted from medical literature databases and hospital medical record voice recordings, and combined with data augmentation techniques to generate higher-quality training data, thereby improving the application ability of TTS systems in the medical field.
[0077] In the financial field, accurate text and speech data are crucial for reading financial reports, voice announcements of transaction reminders, and financial intelligent customer service. For example, in emerging markets, financial institutions need to support voice announcement functions in local languages, but due to insufficient data in low-resource languages, the synthesis quality of existing TTS systems is low. Target language data can be extracted from financial news data and historical transaction voice recordings, and voice enhancement techniques and transfer learning can be adopted to enable TTS systems to achieve high-quality speech generation under low-data conditions and improve the intelligent level of financial voice services.
[0078] By constructing a method for obtaining target language text data and paired speech data from multiple sources and multiple strategies, the training ability of a low-resource language speech synthesis system is improved. The diverse collection of pure text data enables the language perception embedding layer to learn richer semantic information and enhances the model's adaptability to the target language. The enhancement strategy for paired speech data reduces the training bottleneck caused by insufficient data in low-resource languages, enabling the model to still generate natural and fluent speech under limited data conditions. Combining speech alignment optimization and transfer learning strategies further improves the accuracy of the text-to-speech mapping and optimizes the quality of speech synthesis.
[0079] S30, construct an extended vocabulary for the target language;
[0080] In this embodiment, in the speech synthesis task of low-resource languages, the construction of an extended vocabulary is an important step in improving the text input processing ability. The extended vocabulary is used to store characters, phonetic symbols, or other language symbols of the target language, enabling the input text to be mapped into a token sequence that can be processed by the model. Due to the significant differences in writing systems, pronunciation rules, and corpus resources of low-resource languages, the construction of the vocabulary needs to take into account the adaptability of character symbols and phonetic symbols to enhance the text-to-speech conversion ability.
[0081] The construction of the extended vocabulary first requires extracting character symbols from the text dataset of the target language. Character symbols are the basic units of text and may include different types of writing systems such as letters, syllables, Chinese characters, Arabic letters, etc. In the implementation process, a character statistics method can be used to extract all independent characters in the target language text and remove unnecessary punctuation marks, numbers, and other non-language characters to ensure the integrity and adaptability of the vocabulary.
[0082] After extracting the character symbols, it is necessary to detect whether the target language has an effective character-to-phonetic conversion tool. A character-to-phonetic conversion tool is a tool used to map text characters to International Phonetic Alphabet (IPA) or other phoneme representations, usually including a rule-based conversion system or a data-driven phonetic prediction model. During the detection process, it can be checked whether there is an existing pronunciation dictionary for the target language, such as the CMU Pronouncing Dictionary, or statistical learning methods can be used to extract the correspondence between characters and phonetic symbols from the corpus. If there is no effective conversion tool for the target language, a phoneme prediction model trained based on speech data or a cross-language mapping method based on language similarity needs to be used to map the phonetic symbols of high-resource languages to the target language.
[0083] When there is an effective character-to-phonetic symbol conversion tool for the target language, the character symbols can be converted into corresponding phonetic symbols through the conversion tool, and an extended vocabulary can be generated based on the character symbols and phonetic symbols. This vocabulary contains two types of symbols: character symbols are used for the input of the original text, while phonetic symbols are used for the pronunciation rule learning of the model. In actual implementation, a dual-tag vocabulary can be adopted, that is, a unique index is assigned to each character, and its corresponding phonetic representation is stored at the same time, enabling the model to utilize both character and phonetic information and improving the text representation ability.
[0084] If there is no effective character-to-phonetic symbol conversion tool for the target language, an extended vocabulary containing only character symbols needs to be constructed. In this case, sub-word segmentation methods (such as Byte Pair Encoding, BPE or Unigram Language Model) can be used to perform sub-word processing on the character symbols to reduce the impact of unknown vocabulary on speech synthesis. Statistical analysis methods can also be combined to determine the most common character combinations based on the corpus of the target language and use them as the basic units of the extended vocabulary to enhance the adaptability to low-resource languages.
[0085] In one implementation, the construction of the extended vocabulary can rely on an existing pronunciation dictionary. If there is a publicly available pronunciation dictionary for the target language, the character symbols can be directly extracted, and a character-phonetic joint vocabulary can be constructed through the phonetic mapping relationship provided by the dictionary. This method is applicable to low-resource languages with relatively mature language resources, such as some South Asian and African languages.
[0086] In another implementation, if there is no ready-made character-to-phonetic symbol conversion tool for the target language, statistical learning methods can be used to automatically generate phonetic mappings. Phonetic alignment data between the target language and a high-resource language can be extracted from bilingual parallel corpora or multilingual TTS systems, and a phonetic prediction model can be trained to enable character-to-phonetic conversion on the target language. This method is applicable to languages with scarce phonetic data.
[0087] In another implementation, if the writing system of the target language does not contain phonetic symbols or its writing and pronunciation are highly irregular (such as English and French), a deep learning-based phoneme prediction model can be adopted. For example, a character-to-phoneme (G2P) model can be trained based on LSTM or Transformer to enable automatic prediction of phoneme-level pronunciation representations when text is input. This method is applicable to low-resource languages with relatively rich speech data but complex character-phonetic mappings.
[0088] For example:
[0089]
[0090] This formula describes how to construct the vocabulary of the target language. By expanding the vocabulary to include graphemes and International Phonetic Alphabet (IPA) symbols of different languages, the adaptability of the model to low-resource languages is enhanced. The expanded vocabulary can support a wider range of language inputs, enabling the speech synthesis system to maintain high quality in low-resource language environments.
[0091] Among them, V represents the overall vocabulary, which contains the vocabulary sets of all languages; represents the grapheme vocabulary set of a certain language, that is, the vocabulary based on letters; represents the International Phonetic Alphabet (IPA) vocabulary set of a certain language, that is, the vocabulary based on phonetic notations.
[0092] Example: In the field of healthcare, constructing an extended vocabulary can be used for speech synthesis of medical terms. Medical terms often contain a large number of cross-language loanwords, such as Latin, Greek, or foreign words, and the pronunciation of these terms may vary significantly from their character spelling. By constructing a character-phonetic mapping relationship through an extended vocabulary, the speech synthesis quality of medical texts can be improved, enabling medical voice assistants or medical report reading systems to pronounce correctly and enhancing the accuracy of information transmission in the medical industry.
[0093] In the financial field, the construction of an extended vocabulary can enhance the speech synthesis ability of financial terms and numerical expressions. There are a large number of numerical values and proper nouns in the financial industry, such as stock codes, currency units, portfolio names, etc. The pronunciation of these terms may be different in different languages, especially in the financial markets of low-resource languages, where standardized pronunciation is crucial. By expanding the vocabulary, it can be ensured that the financial speech synthesis system can accurately read financial reports, market data, and investment advice, improving the professionalism and reliability of intelligent financial assistants.
[0094] By constructing an extended vocabulary that combines character symbols and phonetic symbols, the text processing ability of low-resource languages is enhanced. When there is a character-to-phonetic conversion tool for the target language, the phonetic symbols are used to enhance the vocabulary, improving the pronunciation rule modeling ability of the speech synthesis system and making the synthesized speech more accurate and natural. When the conversion tool is not available, a character-only vocabulary is constructed through subword segmentation methods and statistical learning, enabling the speech synthesis system to still be effectively trained in the absence of prior pronunciation knowledge. This enhances the adaptability of the TTS system for low-resource languages to different languages and improves the accuracy of text-to-speech conversion.
[0095] S40, updating the parameters of the language perception embedding layer using the pure text dataset through a masked language model;
[0096] In this embodiment, in the multi - language speech synthesis task, the core role of the language perception embedding layer is to capture the semantic features of the target language, enabling the model to learn language characteristics under low - resource conditions and improving the accuracy and naturalness of speech synthesis. Since low - resource languages usually lack sufficient paired speech data, relying solely on supervised learning to train the model may lead to overfitting or unstable speech synthesis effects. Therefore, a pure - text dataset can be used to perform unsupervised pre - training through a Masked Language Model (MLM) to optimize the parameters of the language perception embedding layer, enabling it to learn the features of the target language in the absence of paired data.
[0097] The Masked Language Model is an unsupervised pre - training method. Its basic idea is to randomly mask a part of the words or characters in the text data, and then train the model to predict the masked part based on the context, thereby learning the deep semantic structure of the language. In the present invention, this method can be used to optimize the language perception embedding layer, enabling it to be trained on large - scale pure - text data without relying on paired speech data.
[0098] The specific implementation method may include the following steps:
[0099] Text data pre - processing: Extract sentences from the pure - text dataset of the target language and perform normalization processing, including removing punctuation marks, unifying character encoding, word segmentation, etc. For ideographic languages (such as Chinese), a sub - word segmentation method (such as the BPE or Unigram model) can be used for pre - processing to improve the effect of mask prediction.
[0100] Random masking strategy: In the input text sequence, randomly select a certain proportion of words or characters for masking (for example, replace 15% of the characters with special tokens <mask>) and let the model learn to recover the masked parts from the context. This strategy can promote the model to capture context dependencies and improve language understanding ability.
[0101] Model training: Use Transformer-based language models (such as BERT, XLM-R, mBERT, etc.) for unsupervised training, so that the language-aware embedding layer can learn deep semantic representations of the target language. During the training process, cross-lingual contrastive learning strategies can be adopted to enable the model to be pre-trained on multilingual data and enhance its adaptability to low-resource languages.
[0102] Parameter optimization: During the training process, use self-supervised learning objectives (such as cross-entropy loss) to optimize the parameters of the language-aware embedding layer, enabling it to better represent the text features of the target language and improve the performance of subsequent speech synthesis tasks.
[0103] After masked language model training, the language-aware embedding layer can have stronger context understanding ability. Even if there is less speech data in the target language, rich text data can be used for pre-training to improve the speech synthesis quality.
[0104] Through the unsupervised pre-training of the masked language model, the parameters of the language-aware embedding layer are optimized, enabling low-resource languages to still learn the semantic features of the target language in the absence of paired speech data. Through the random masking strategy, the context learning ability of the model is enhanced, and the text understanding ability of the speech synthesis system under low-resource conditions is improved. Through multilingual pre-training and contrastive learning, the cross-lingual adaptability of the embedding layer is enhanced, and the speech synthesis quality of the model in a multilingual environment is improved.
[0105] S50, update the parameters of the multilingual speech synthesis model using the paired speech text data;
[0106] In this embodiment, in order to improve the adaptability of the speech synthesis model to low-resource languages, it is necessary to use paired speech text data for training to optimize the model parameters so that it can learn the mapping relationship from text to speech. During the training process, it is necessary to extract text features and speech features from the paired data and align the two to ensure the accuracy and fluency of speech synthesis.
[0107] The text data is first processed through the language-aware embedding layer, converting the character or phonetic symbol representation into a vector representation and inputting it into the encoder to extract context semantic features. At the same time, the speech data is preprocessed to extract Mel spectrogram or other acoustic features to keep them aligned with the text features. During the training process, the model continuously adjusts the parameters of the encoder and decoder based on the supervised learning method, enabling the text features to be accurately mapped to the corresponding acoustic features.
[0108] The optimization goal of the encoder is to enhance the text representation ability so that it can fully capture sentence structures, lexical information, and context relationships. To improve the generalization ability of low-resource languages, the encoder can adopt a cross-lingual parameter sharing strategy, that is, let multiple languages share some encoding layers, so that low-resource languages can learn common language patterns from the training data of high-resource languages. In addition, data augmentation techniques can be combined, such as through a multilingual alignment mechanism, to semantically match texts of similar language families, further enhancing the learning ability of the encoder.
[0109] The training goal of the decoder is to predict the corresponding acoustic features based on the output of the encoder and optimize the naturalness and clarity of the speech. The decoder can adopt an autoregressive structure, that is, predict the acoustic features of the current frame at each time step and use it as the input to predict the next frame to ensure the fluency of the synthesized speech. To improve the stability of training, an alignment mechanism can be used, such as an attention-based alignment method or dynamic time warping (DTW), to ensure a more accurate time step correspondence between text features and speech features. In addition, the optimization of the decoder can adopt multi-task training, that is, while predicting the Mel spectrogram, learn auxiliary tasks such as pitch and prosody to improve the naturalness of the speech.
[0110] During the training process, the optimization of model parameters adopts the method of minimizing the loss function. For example, the mean squared error (MSE) loss is used for Mel spectrogram prediction, and the cross-entropy loss is used for text token prediction, etc. The model can be jointly trained with high-resource languages, enabling the training data of low-resource languages to obtain more sufficient learning in the multilingual model and improving the speech synthesis effect.
[0111] In one implementation, an end-to-end Transformer TTS model can be used for training. After the text input is processed by the encoder, the decoder directly predicts the Mel spectrogram features, and a pre-trained vocoder is used to convert the Mel spectrogram into a speech waveform. This method is suitable for scenarios with sufficient computing resources and can provide high-quality speech synthesis effects.
[0112] In another implementation, a two-stage training scheme of Tacotron and WaveNet can be adopted. First, use Tacotron to train the mapping from text to acoustic features, and then use WaveNet or HiFi-GAN to train the conversion from acoustic features to speech waveforms. This method is suitable for scenarios with limited data but the desire to improve speech quality and can enhance the controllability and stability of the model.
[0113] In another implementation, a multi - language joint training method can be utilized to enable low - resource languages to share some model parameters with high - resource languages, optimizing the speech synthesis ability. This approach is particularly suitable for situations where there is extremely little data for low - resource languages and can effectively improve the fluency and naturalness of speech synthesis.
[0114] For example:
[0115]
[0116] This formula is used to calculate the mean squared error (MSE) during the training process. As the optimization objective function, it minimizes the error between the predicted speech features and the true speech features. MSE is one of the commonly used loss functions in neural network training. In the speech synthesis task, it is used to measure the deviation between the generated speech features and the true speech features. Minimizing this loss function can improve the accuracy of speech generation, making the synthesized speech closer to the true speech. Here, L represents the mean squared error (MSE) loss function, which is used to measure the speech features predicted by the model and the true speech features y t The deviation between them. This function is used to optimize the model parameters during the training process, making the generated speech features closer to the true speech features. The mean squared error calculates the squared error at all time steps and takes the mean, ensuring the stability of error measurement and giving higher weights to larger errors to guide the model to learn speech features more precisely. T represents the number of time steps in the speech sequence, and each time step corresponds to a speech frame, reflecting the length of the speech data. y t is the true speech feature in the training data, usually obtained by acoustic processing (such as Mel - spectrum extraction) after human reading of the text, while is the predicted speech feature generated by the model, which is compared with the true speech feature to evaluate the synthesis quality.
[0117] θ = {θ E , θ lae , θ D}
[0118]
[0119] This formula describes the parameter optimization process of the speech synthesis model, which involves key elements such as the model parameter set, the gradient of the loss function, and the learning rate. The core idea of the formula is to optimize the model parameters through the gradient descent method, making the value of the loss function L gradually decrease, thereby enhancing the speech synthesis ability of the model. Here, θ represents the model parameter set, which includes the following three parts: θ E : Encoder parameters, used to extract context semantic features from the token sequence of the input text; θ lae : Language-aware embedding layer parameters, used to model the pronunciation rules and text representations between languages, providing cross-lingual knowledge transfer ability; θ D : Decoder parameters, used to receive the semantic features output by the encoder and generate acoustic features. L represents the value of the loss function, which measures the speech features predicted by the model and the true speech features y t The mean squared error (MSE) between them is used to guide parameter optimization; represents the gradient of the loss function with respect to the model parameters, indicating the direction and degree of the impact of the current model parameters on the loss. The calculation of the gradient is usually completed using the Backpropagation algorithm; η represents the learning rate, which is used to control the step size of model parameter updates and determine the amplitude of parameter adjustment. A too large learning rate may lead to instability in the optimization process, while a too small learning rate may result in a too slow convergence speed.
[0120] By using paired speech-text data to optimize the parameters of the speech synthesis model, its performance on low-resource languages is improved. The encoder enhances the ability to model text features through shared parameters and transfer learning, enabling low-resource languages to learn the text-to-speech mapping relationship from high-resource languages. The decoder optimizes speech fluency through autoregressive generation and combines an alignment mechanism to improve the accuracy of the time-step matching between text and speech.
[0121] S60, convert the target input text into an input token sequence according to the extended vocabulary;
[0122] In this embodiment, the process of converting text into an input token sequence determines how the speech synthesis model parses the input text and generates a representation suitable for speech synthesis. Since the vocabulary structures of different languages vary, the construction method of the extended vocabulary directly affects the processing flow of the input text. The target input text needs to be normalized, encoded, and generate a sequence representation suitable for model processing according to the extended vocabulary to ensure the accuracy and stability of the speech synthesis process.
[0123] The conversion of text first requires preprocessing, including operations such as removing irrelevant characters, standardizing punctuation marks, and handling case conversion. For languages with phonemes as the basic unit, a character-level vocabulary can be directly used, while for languages with words or subwords as the unit, word segmentation or subwording processing is required to reduce the impact of out-of-vocabulary words. The extended vocabulary may contain character symbols, phonetic symbols, or other language-specific markers, and the target input text needs to be compared to ensure that the characters or phonetic symbols in the text can find corresponding mappings in the vocabulary.
[0124] Tokenization of text usually adopts sub - word segmentation methods, such as BPE (Byte Pair Encoding) or Unigram Language Model. These methods can reduce the impact of out - of - vocabulary (OOV) words in low - resource language environments and improve the adaptability of the model. During the conversion process, each word or sub - word needs to be checked whether it exists in the extended vocabulary. If it exists, it is directly mapped to the index value in the vocabulary; if not, it needs to be split into multiple in - vocabulary sub - words and mapped to the indices in the vocabulary respectively.
[0125] After the conversion is completed, all the token indices are arranged in the original order of the text to generate the input token sequence. This sequence is used as the input for the subsequent speech synthesis model, enabling the encoder to accurately extract the semantic and pronunciation information of the text and improving the accuracy and fluency of speech synthesis.
[0126] In one implementation, a rule - based text pre - processing method can be used to perform character - level normalization on the target input text and directly map it to the input token sequence according to the character symbols in the extended vocabulary. This method is applicable to languages with relatively regular character writing systems, such as alphabetic scripts of the Indo - European language family.
[0127] In another implementation, a dynamic vocabulary matching method based on sub - word segmentation can be adopted. First, the target input text is segmented into sub - words and gradually matched with the extended vocabulary, preferentially selecting the longest - matching sub - word units. This method is applicable to languages with rich morphology, such as languages of the Uralic or Altaic language families, and can effectively reduce the problems caused by out - of - vocabulary words and improve the controllability of speech synthesis.
[0128] In another implementation, multi - layer vocabulary processing can be combined, that is, extended vocabularies at the character level, sub - word level, and phonetic symbol level are constructed simultaneously, and these vocabularies are comprehensively used during the conversion process to enhance the speech adaptability of the model. For example, for the same input text, first perform sub - word segmentation. If the corresponding sub - word cannot be matched, then fallback to character - level mapping, and apply phonetic symbol mapping in specific cases to ensure that all input characters can find corresponding token indices. This method is applicable to languages with complex writing systems and lack of standardization, such as languages of the Austro - Asiatic family or some African languages.
[0129] Example illustration: In healthcare applications in low-resource languages, the voice synthesis quality requirements for medical report reading aloud, electronic health record (EHR) transcription, and voice interaction of intelligent healthcare assistants are relatively high. However, due to the lack of a complete medical terminology database for many low-resource languages, voice synthesis systems are prone to misreading or failing to recognize professional medical terms. For example, in some African countries or Southeast Asian regions, medical literature is mainly written in languages such as English, French, or Portuguese, while local doctors and patients prefer to communicate in local languages. Traditional text conversion methods often cannot correctly process mixed-language texts when parsing medical texts, resulting in inaccurate reading aloud results and affecting the effectiveness of clinical communication. By constructing a medical-specific extended vocabulary and combining subword segmentation and phonetic matching methods, medical texts in low-resource languages can be correctly converted into input token sequences. For example, in the absence of a character-to-phoneme conversion tool, the mapping relationship between characters and phonemes can be inferred from known medical corpora through statistical learning methods, or cross-language matching can be performed using a phonological structure similar to that of high-resource languages, enabling the system to generate more accurate medical voices. In this way, even if some medical terms are not in the extended vocabulary, they can still be inferred based on the existing subword structure, improving the voice synthesis quality. In addition, to address the issue of scarce data in the medical field, the voice synthesis system can combine an automatic expansion method for medical record texts, automatically transcribe medical records through rule-based or deep learning methods, and combine patient consultation Q&A data to further enrich the extended vocabulary of low-resource languages, enabling the voice synthesis system to maintain a relatively high voice synthesis quality with limited corpus resources. This is of great value for intelligent healthcare assistants, medical learning platforms, and cross-language medical communication in remote areas.
[0130] In financial applications of low-resource languages, scenarios such as financial news broadcasts, financial regulatory announcements, and reading of investment analysis reports pose higher requirements for the professionalism and accuracy of speech synthesis. The text data in the financial industry usually contains a large number of proprietary terms and numerical expressions. However, the financial terms in low-resource languages often lack a standardized vocabulary, making it difficult for speech synthesis systems to correctly convert such texts. For example, in some Southeast Asian countries or South American regions, financial news reports are mainly written in English or Spanish, while local investors may be more accustomed to obtaining market information in their native languages. If the speech synthesis system cannot correctly parse foreign financial terms, it may affect the dissemination and understanding of financial information. A financial-specific extended vocabulary can ensure that financial terms are correctly processed during text conversion. For example, when processing a financial term like "Gross Domestic Product (GDP)", if there is no direct translation in the target language, the system can automatically match the alternative expression in the extended vocabulary and combine it with the subword segmentation method to generate a suitable input token sequence, making the result of speech synthesis more natural and understandable. In addition, due to the limited financial data in low-resource languages, the construction of the extended vocabulary can combine data migration and cross-lingual alignment techniques, that is, using the financial data of high-resource languages to train low-resource languages. For example, existing financial report corpora can be used to enable the extended vocabulary to automatically identify and adapt to the financial terms of low-resource languages through cross-lingual embedding learning, thereby improving the accuracy of speech synthesis and enabling the generation of professional financial voices even in the absence of a large amount of training data.
[0131] Through the text conversion method based on the extended vocabulary, the mapping accuracy from the target input text to the input token sequence is improved. The problem of out-of-vocabulary words is reduced through subword segmentation and character matching, enabling the model to more stably parse different types of target language texts. By combining character-level, subword-level, and phonetic-level vocabularies, the flexibility of text parsing is improved, making the speech synthesis of low-resource languages more accurate.
[0132] S70, converting the input token sequence into an initial vector representation through the language-aware embedding layer;
[0133] In this embodiment, after being converted by the extended vocabulary, the input token sequence is still a discrete symbol sequence and cannot be directly used for training the speech synthesis model. To enable the model to learn the potential patterns of speech generation, it is necessary to convert the input token sequence into a continuous vector representation through the language-aware embedding layer so that it can be processed by the neural network and passed to the subsequent encoder for deep feature extraction.
[0134] The core function of the language-aware embedding layer is to capture the lexical features of the target language and construct a high-dimensional representation that enables the model to learn the correspondence between text and speech. In this process, the input token sequence is first mapped to a high-dimensional dense vector space, where each token corresponds to a vector of a fixed dimension that contains the semantic and pronunciation information of the token.
[0135] To optimize the performance of the embedding layer, various embedding methods can be adopted:
[0136] Single-layer embedding: Directly use a randomly initialized embedding matrix to map the input tokens to vectors of a fixed dimension. This method is suitable for languages with a large amount of data.
[0137] Pre-trained embedding: Use word vectors trained on large-scale multilingual data (such as FastText, mBERT, etc.). Through unsupervised learning methods, the embedding layer can capture richer semantic information.
[0138] Subword embedding: For low-resource languages, subword-level embedding can be adopted, that is, mapping the subword units after BPE or Unigram segmentation to the embedding space to reduce the impact of out-of-vocabulary words on the model.
[0139] Language-specific embedding: In multilingual speech synthesis tasks, to enhance the modeling ability of low-resource languages, language-specific embedding vectors can be combined to enable the model to adapt to the characteristics of different languages. For example, when converting the input token sequence, a language token can be added to the embedding layer so that the model can learn the pronunciation characteristics of different languages during training.
[0140] During the conversion process, each input token sequence forms a series of vector representations after embedding mapping, and these vectors can be further passed to the encoder for feature extraction. To enhance the speech modeling ability, the language-aware embedding layer can combine positional encoding information so that different position information in the sequence can be better perceived by the model, ensuring that context information can be effectively modeled.
[0141] Through methods such as subword embedding, pre-trained embedding, and language-specific embedding, the adaptability of the model to low-resource languages is enhanced, and the quality of speech synthesis is improved. Combining positional encoding information enables the context information of the input text to be effectively modeled, making the speech generation more coherent and natural. The usability of the speech synthesis system under low-resource conditions is improved, so that even when the paired data is limited, high-quality speech synthesis can still be achieved.
[0142] S80, extracting the context semantic features of the initial vector representation through the encoder;
[0143] In this embodiment, the initial vector represents the word vector that remains static after the mapping in the language perception embedding layer and cannot directly capture context information. Therefore, it is necessary to further extract context semantic features through the encoder so that the model can learn the structure, context dependence of the input text, and semantic connections across words and sentences to improve the accuracy and coherence of speech synthesis.
[0144] The main role of the encoder is to construct dynamic context-aware features based on the initial vector representation of the input token sequence, enabling the subsequent decoder to make full use of the text structure and semantic information when generating speech. In the implementation process, the encoder usually adopts a multi-layer stacked neural network structure to improve the feature extraction ability.
[0145] To ensure that the model can efficiently extract context information, the following encoding methods can be adopted:
[0146] RNN-based encoder: Use bidirectional LSTM or GRU networks, enabling the model to extract the context information of the text from both the forward and backward directions. This method is suitable for speech synthesis tasks of shorter texts and can better capture local dependencies.
[0147] Transformer-based encoder: Adopt the self-attention mechanism, enabling the encoder to simultaneously focus on all tokens in the entire input sequence and improve the modeling ability of long-distance dependencies. This method is particularly suitable for speech synthesis tasks that require high naturalness and long texts, such as news broadcasts or literary readings.
[0148] Hybrid model-based encoder: Combine the advantages of RNN and Transformer, use RNN to extract short-distance dependency features at the local level, and use Transformer to extract long-distance dependency features at the global level, thereby enhancing the context understanding ability of speech synthesis.
[0149] During the encoding process, the initial vector representation of the input token sequence undergoes multiple layers of encoding operations, so that the representation of each token not only contains its own feature information but also integrates semantic information related to the context. For example, in the text of the sentence "The market index has risen", the meaning of "index" depends on the context of the word "market". The encoder can capture this semantic relationship through the self-attention mechanism or recurrent neural network and adjust the representation of the word vector to make it more in line with the understanding of the overall context.
[0150] The output of the encoder is a context-enhanced feature sequence, which will be used as the input for the subsequent decoder to generate speech, enabling the speech synthesis system to generate a more natural and coherent speech waveform based on more complete context information.
[0151] By extracting the context semantic features of the input token sequence, the speech synthesis system can make full use of the text structure and context information to improve the fluency and naturalness of speech synthesis. Combining the encoding methods of bidirectional LSTM, Transformer or hybrid models enables the system to adapt to different types of text inputs and improve the speech synthesis quality of low-resource languages. After the encoder is optimized, the speech synthesis not only conforms more to the grammar rules of the target language but also can capture the implicit information in the text, making the finally generated speech closer to the natural pronunciation of humans.
[0152] S90, when the extended vocabulary contains phonetic symbols, extract the pronunciation rule features from the phonetic symbols;
[0153] In this embodiment, when the extended vocabulary contains phonetic symbols, extracting the pronunciation rule features is a key step to ensure that the speech synthesis system can correctly generate the pronunciation of the target language. Phonetic symbols are an explicit speech annotation information, which can accurately describe the pronunciation method at the phoneme level, enabling the speech synthesis system to generate a more natural and clear speech output. By extracting features from phonetic symbols, the phonological features of the target language can be effectively captured, making the speech synthesis of low-resource languages more accurate.
[0154] During the processing, the input phonetic symbols first need to be normalized to ensure that all phonetic symbols conform to the speech rules of the target language, such as IPA (International Phonetic Alphabet) or the phonetic symbol system of a specific language. The normalized phonetic symbols can be mapped to a fixed set of phonemes, enabling the unified conversion of different phonetic symbol representations and reducing the pronunciation errors in speech synthesis.
[0155] Multiple methods can be used to extract the pronunciation rule features:
[0156] Statistical-based method: By analyzing the phonetic symbol patterns in a large-scale corpus, learn the common phoneme sequences, sound change rules and co-occurrence relationships in the target language. For example, in some languages, specific consonants may be weakened or linked together before and after certain vowels. Through statistical analysis, a rule library for phoneme conversion can be established.
[0157] Linguistics-rule-based method: Combine linguistic knowledge to manually define the pronunciation rules of the target language, such as the linking rules between vowels and consonants, the weakening, insertion and deletion rules of phonemes, etc. For example, in French speech synthesis, some word-final consonants may require special treatment when linked in the middle of words.
[0158] Deep learning-based method: Using neural network models such as Transformer or BiLSTM to learn the mapping relationship from phonetic symbols to acoustic features, enabling the system to automatically recognize phonetic patterns and generate pronunciation features suitable for the target language. This method is applicable to multilingual environments and can better adapt to the pronunciation characteristics of different languages.
[0159] After the phonetic features are extracted, the system can generate pronunciation rule features for speech synthesis, including phoneme duration, intonation pattern, liaison information, consonant weakening, etc. These features will be further used in the subsequent acoustic modeling process for generating speech waveforms.
[0160] By extracting pronunciation rule features from phonetic symbols, the speech synthesis system can more accurately simulate the phonological characteristics of the target language, improving the naturalness and clarity of speech. The phonetic feature extraction methods based on rules, statistics, or deep learning enable the system to adapt to the pronunciation patterns of different languages and reduce pronunciation errors caused by insufficient data in low-resource languages.
[0161] S100, fusing the context semantic features and the pronunciation rule features through the decoder to generate an acoustic feature sequence;
[0162] In this embodiment, the core task of the decoder is to generate an acoustic feature sequence that can represent the target speech based on the context semantic features extracted by the encoder and the pronunciation rule features extracted from phonetic symbols. This process determines the naturalness, clarity, and intonation coherence of speech synthesis. Therefore, the decoder needs to have a strong feature fusion ability to enable the speech generation to accurately match the speech characteristics of the target language.
[0163] In the implementation process, the decoder receives the context semantic features from the encoder, which contain the overall semantic information, context relationship, and syntactic structure of the text. At the same time, the decoder also receives the pronunciation rule features, which are mainly used to guide the system to generate speech that conforms to the phonological characteristics of the target language, such as liaison rules, phoneme length changes, consonant weakening, etc.
[0164] The decoder can be implemented using different neural network structures to ensure the full fusion of these two types of features:
[0165] Autoregressive decoder: Adopting an RNN (such as LSTM or GRU) structure to sequentially generate acoustic feature frames, and the generation of each frame depends on the output of the previous frame. This method can maintain the coherence of the speech sequence and is suitable for languages with smooth speech flows.
[0166] Transformer-based decoder: It adopts the self-attention mechanism, enabling the decoder to focus on the entire input sequence and generate multiple acoustic feature frames simultaneously. This approach is suitable for long-text speech synthesis, such as news broadcasts or legal text readings, improving the naturalness of the speech.
[0167] Variational Autoencoder (VAE)-based decoder: By learning the distribution of speech data, it performs latent variable modeling on context semantic features and pronunciation rule features, making the generated acoustic features smoother and more natural.
[0168] During the decoding process, the decoder needs to perform feature fusion, mainly including the following methods:
[0169] Channel dimension concatenation: Concatenate the context semantic features and pronunciation rule features along the channel dimension and map them to a unified feature space through a fully connected layer, enabling the features to complement each other and improving the stability of speech synthesis.
[0170] Attention weighted fusion: Use the attention mechanism to weight the pronunciation rule features, enabling the decoder to dynamically adjust the weights of phoneme conversion at different time steps and optimizing the clarity of the speech.
[0171] Hierarchical feature fusion: In different decoding layers, perform different-level fusions on the context semantic features and pronunciation rule features respectively, enabling the decoder to gradually adjust the pronunciation details of the speech and improving the naturalness and fluency.
[0172] Finally, the output acoustic feature sequence of the decoder includes Mel spectrogram, pitch, energy curve, duration parameters, etc. These features will be used as the input for the subsequent vocoder to generate the final speech waveform.
[0173] For example, it can be represented by the following formula:
[0174]
[0175] This formula describes the core calculation process of the Text-to-Speech (TTS) model, from text input to the generation of acoustic features.
[0176] Among them, X represents the token representation of the input text sequence; θ E represents the encoder parameters; θ lae represents the language perception embedding layer parameters; θ D represents the decoder parameters; Encoder() represents the encoder, which is used to extract the context features of the text; Decoder() represents the decoder, which is used to generate acoustic features; represents the output acoustic feature sequence.
[0177] By fusing context semantic features and pronunciation rule features through a decoder, the speech synthesis system can generate more natural, clear, and coherent speech output. Through different feature fusion strategies, the system can dynamically adjust the pronunciation pattern and optimize the speech quality. Combining decoding structures such as LSTM, Transformer, or VAE makes speech generation more flexible and applicable to various language environments. Especially in low-resource language speech synthesis tasks, it can significantly improve the naturalness and fluency of speech.
[0178] S110, convert the acoustic feature sequence into target speech data.
[0179] In this embodiment, the acoustic feature sequence is an intermediate representation output by the decoder, including Mel spectrogram, pitch information, energy distribution, duration parameters, etc. These features cannot be directly converted into speech waveform data. Therefore, further processing is required to map them into real audio signals. The core objective of this conversion process is to convert the high-dimensional acoustic features generated by the decoder into playable speech waveforms and ensure the naturalness, clarity, and coherence of speech synthesis.
[0180] The conversion of acoustic features usually relies on a neural vocoder or traditional parametric synthesis methods, and can be carried out in the following ways:
[0181] Method based on neural vocoder: Use a trained deep learning model (such as WaveNet, HiFi-GAN, WaveGlow) to map acoustic features into high-quality speech waveforms. This method can generate more natural and fluent speech.
[0182] Parametric synthesis method based on filters: Use traditional audio synthesis techniques (such as LPC, WORLD, STRAIGHT) to convert acoustic features into speech waveforms. This method has a lower computational cost, but the naturalness of the synthesized speech is relatively low and is suitable for devices with low computational resources.
[0183] The conversion process includes the following key steps:
[0184] Time step adjustment: Since the acoustic feature sequence output by the decoder may have problems with uneven frame lengths, time axis alignment is required so that each feature frame can correctly correspond to the time step of the final speech. Linear interpolation or a duration prediction model can be used for alignment.
[0185] Conversion from Mel spectrogram to linear spectrogram: If the neural vocoder requires a linear spectrogram as input, the Mel spectrogram needs to be converted into a linear spectrogram and a Fourier transform is performed for subsequent speech waveform reconstruction.
[0186] Neural vocoder inference: Use a pre-trained neural vocoder to process acoustic features frame by frame to generate speech waveform data. The neural vocoder can adopt autoregressive (such as WaveNet) or non-autoregressive (such as HiFi-GAN, WaveGlow) methods to improve the real-time performance and stability of speech synthesis.
[0187] Post-processing optimization: Include denoising, dynamic range compression, volume normalization, etc., to make the synthesized speech more conform to human auditory perception and improve speech quality.
[0188] Finally, the speech data generated by the system can be used in various application scenarios, such as intelligent voice assistants, news broadcasts, medical voice interactions, etc.
[0189] By converting the acoustic feature sequence, the speech synthesis system can generate high-quality speech waveforms. Using neural vocoders (such as WaveNet, HiFi-GAN) can improve the naturalness and clarity of speech, enabling the speech synthesis system to achieve high-quality speech output even in low-resource language environments. In addition, combining time step adjustment, spectral conversion, and post-processing optimization makes the synthesized speech more stable and reliable in different application scenarios.
[0190] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a speech generation method, including: constructing a multilingual speech synthesis model, obtaining pure text data and paired speech text data, and constructing an extended vocabulary; updating the language-aware embedding layer and model parameters to convert the input text into a token sequence; the encoder extracts context semantic features, extracts pronunciation rule features, the decoder fuses the features to generate an acoustic feature sequence, and converts it into target speech data. The present invention improves the speech generation ability of low-resource languages through the multilingual speech synthesis model combined with the language-aware embedding layer; the extended vocabulary improves the accuracy of text conversion, unsupervised training enhances the target language learning ability, supervised training optimizes the adaptability to low-data environments, and feature fusion improves the naturalness and fluency of speech.
[0191] In one embodiment, the above S10 includes:
[0192] S101, constructing a language-aware embedding layer and initializing the parameters of the language-aware embedding layer;
[0193] S102, constructing an encoder composed of multiple stacked Transformer encoding layers, each encoding layer containing a multi-head self-attention mechanism and a feed-forward network;
[0194] S103, constructing a decoder composed of multiple stacked Transformer decoding layers, each decoding layer containing a masked multi-head self-attention mechanism, a cross-attention mechanism, and a feed-forward network;
[0195] S104, configure the output dimension of the language-aware embedding layer to be the same as the input dimension of the encoder;
[0196] S105, configure the output dimension of the encoder to be the same as the input dimension of the decoder.
[0197] In this embodiment, constructing a multilingual speech synthesis model requires including a language-aware embedding layer, an encoder, and a decoder, enabling the system to adapt to text inputs in different languages and generate high-quality speech outputs. This model adopts the Transformer architecture, with powerful feature extraction and generation capabilities, making the speech synthesis of low-resource languages more stable and fluent.
[0198] The language-aware embedding layer is used to convert the input text token sequence into a high-dimensional vector representation and encode it in combination with the characteristics of the target language, enabling the model to learn semantic information in different languages. In the initialization stage, the parameters of the embedding layer need to be set, including the vocabulary size, embedding dimension, normalization method, etc. To adapt to the characteristics of low-resource languages, random initialization embedding, pre-trained embedding, or sub-word level embedding can be used for optimization. Random initialization embedding directly randomly generates the embedding matrix and optimizes it during the training process, suitable for languages with less data; pre-trained embedding is unsupervised trained based on large-scale multilingual text data (such as mBERT, FastText, etc.), enabling the embedding layer to learn cross-lingual general semantic information and improve the generalization ability of the speech synthesis system; sub-word level embedding is suitable for languages with rich morphological variations, such as Arabic, Finnish, etc., which can reduce the impact of out-of-vocabulary words and improve the adaptability of the system. To ensure the consistency of data transmission, the output dimension of the embedding layer should be the same as the input dimension of the encoder.
[0199] The encoder consists of multiple stacked Transformer encoding layers, and each encoding layer contains a multi-head self-attention mechanism and a feed-forward network, mainly used to capture the context information of the input sequence and generate a feature representation with global dependencies. The multi-head self-attention mechanism can perform global dependency modeling on all tokens in the input sequence, enabling the model to identify long-distance semantic relationships and improve the accuracy of text parsing. The feed-forward network performs a non-linear transformation on the encoded features to enhance the expressive ability of the model and ensure that the features can adapt to subsequent decoding tasks. Residual connections and layer normalization are applied to each Transformer layer to ensure the stability of the gradient and accelerate model training. To ensure the correct progress of the decoding process, the output dimension of the encoder needs to be the same as the input dimension of the decoder.
[0200] The decoder consists of multiple stacked Transformer decoder layers. Each decoder layer contains a masked multi-head self-attention mechanism, a cross-attention mechanism, and a feed-forward network, and is mainly used to generate an acoustic feature sequence for speech synthesis based on the output of the encoder and in combination with the speech features of the target language. The masked multi-head self-attention mechanism ensures that each time step can only access the previously generated features during the decoding process to avoid data leakage (teacher forcing). The cross-attention mechanism performs attention calculation on the output of the encoder, enabling the decoder to make full use of text information for speech feature generation. The feed-forward network further processes the decoded features so that the output can meet the requirements of acoustic modeling, and finally generates an acoustic feature sequence for speech synthesis.
[0201] In this embodiment, by constructing a multi-language speech synthesis model including a language-aware embedding layer, an encoder, and a decoder, the system can effectively perform semantic modeling for text inputs in different languages and generate high-quality speech waveforms. By adopting the Transformer structure, the encoder can capture the long-range dependencies of the input text, improve the ability to understand complex syntactic structures, and at the same time the decoder can dynamically fuse text features and the phonological features of the target language, improving the fluency and naturalness of speech synthesis. Combining optimization strategies such as static word vectors, grouped attention, and phoneme-level time alignment, the speech synthesis quality of low-resource languages has been significantly improved. Even when the training data is limited, it can still generate stable and clear speech outputs.
[0202] In one embodiment, the above S30 includes:
[0203] S301, extracting character symbols from the pure text dataset of the target language;
[0204] S302, detecting whether there is an effective character-to-phoneme conversion tool for the target language;
[0205] S303, when it is detected that there is an effective character-to-phoneme conversion tool for the target language, converting the character symbols into corresponding phoneme symbols through the character-to-phoneme conversion tool, and generating an extended vocabulary based on the character symbols and phoneme symbols;
[0206] S304, when it is detected that there is no effective character-to-phoneme conversion tool for the target language, constructing an extended vocabulary containing only the character symbols.
[0207] In this embodiment, constructing an extended vocabulary for the target language is a fundamental task for a low-resource language speech synthesis system, aiming to provide sufficient language units to support optimizing the character-to-speech conversion effect and improving the accuracy of speech synthesis. In the speech synthesis task of low-resource languages, due to limited corpora, the correspondence between character symbols and pronunciation symbols is imperfect. Therefore, it is necessary to construct an extended vocabulary to cover as many target language words as possible and improve the model's ability to handle out-of-vocabulary (OOV) words.
[0208] The first step in constructing the extended vocabulary is to extract character symbols from the plain text dataset of the target language. This process usually includes preprocessing the corpus to remove non-language symbols (such as punctuation marks, numbers, special characters, etc.) and extracting the basic character set of the target language. In a low-resource language environment, the extraction of character symbols needs to combine language characteristics to ensure coverage of all possible character variants. For example, in some African languages or Southeast Asian languages, the same letter may have multiple variants due to tones or spelling rules. Therefore, when extracting character symbols, different variants need to be merged or recorded separately to ensure the integrity of the vocabulary.
[0209] After extracting the character symbols, it is necessary to detect whether there is an effective grapheme-to-phoneme conversion tool for the target language. Grapheme-to-phoneme conversion (G2P) is an important part of speech synthesis. It can map the character sequence in the text to the corresponding pronunciation symbols, enabling the speech synthesis system to more accurately reproduce the pronunciation of the target language. If there is a standard grapheme-to-phoneme conversion tool for the target language, such as the International Phonetic Alphabet (IPA) converter or a G2P converter based on a statistical model, these tools can be directly used to convert the character symbols into phonetic symbols, and an extended vocabulary can be generated by combining the character symbols and phonetic symbols. This extended vocabulary not only contains the character-level mapping relationship but can also record pronunciation variants to adapt to different contexts and pronunciation habits.
[0210] If there is no available grapheme-to-phoneme conversion tool for the target language, the extended vocabulary can only contain character symbols. In this case, the speech synthesis system needs to rely on data-driven methods to learn the mapping relationship between characters and phonemes through the model. For example, the way of sharing parameters in multiple languages can be used to enable the low-resource language to transfer and learn the correspondence between characters and phonemes from high-resource languages, or a subword unit-based modeling method can be adopted to expand the character-level vocabulary into more generalizable subword units to improve the adaptability of the system.
[0211] Finally, the constructed extended vocabulary will be used for text-to-speech conversion in the subsequent speech synthesis process, enabling the system to improve the accuracy and fluency of speech synthesis based on the complete character set and phonetic symbol set.
[0212] The process of constructing the extended vocabulary in this embodiment can effectively enhance the speech synthesis ability of low-resource languages, enabling the system to more accurately match character symbols with phonetic symbols and improving the accuracy of text-to-speech conversion. By methods such as character extraction, character-to-phonetic symbol mapping, and data-driven statistical learning, the vocabulary coverage rate of the system can be increased, the probability of out-of-vocabulary words can be reduced, and the speech synthesis can be made more fluent and natural. If the target language lacks a standardized character-to-phonetic symbol conversion tool, using data-driven methods for learning can also optimize the mapping relationship between characters and phonetic symbols to a certain extent and improve the adaptability of the system.
[0213] In one embodiment, the above S60 includes:
[0214] S601, preprocess the target input text to remove punctuation marks and non-target language characters, and generate a normalized text;
[0215] S602, segment the normalized text into character units or sub-word units through a sub-word tokenizer;
[0216] S603, traverse each character unit or sub-word unit after segmentation, and detect whether each character unit or sub-word unit exists in the extended vocabulary;
[0217] S604, if the current character unit or sub-word unit exists in the extended vocabulary, map the current character unit or sub-word unit to the corresponding vocabulary token;
[0218] S605, if the current character unit or sub-word unit does not exist in the extended vocabulary, split the current character unit or sub-word unit into in-vocabulary sub-word units based on the sub-word segmentation algorithm, and map each in-vocabulary sub-word unit to the corresponding sub-word token;
[0219] S606, splice all vocabulary tokens and sub-word tokens according to the text order of the target input text to generate an input token sequence.
[0220] In this embodiment, the process of converting the target input text into an input token sequence is an important step in the speech synthesis system, ensuring that the text can be transmitted to the subsequent speech processing module in a standardized form, and improving the accuracy and stability of text-to-speech conversion. Since the data scale of low-resource languages is small, the text data may contain problems such as inconsistent formats, abnormal characters, out-of-vocabulary (OOV) words, etc. Therefore, text preprocessing, word segmentation, and vocabulary mapping need to be performed before tokenization to make it meet the input requirements of the system.
[0221] First, the target input text needs to be preprocessed to remove punctuation marks and non-target language characters to generate normalized text. The main goal of preprocessing is to unify the text format, reduce unnecessary character interference, and ensure that the input received by the speech synthesis model is clean and standard language data. During this process, the system will remove punctuation marks such as full stops, commas, quotation marks, special symbols, etc., and delete non-target language characters, such as other language words, numbers, or special characters mixed in the text. If specific punctuation marks are allowed to affect semantics in the target language, certain symbols can be retained according to the language characteristics. For example, the apostrophe (') in French may need to be retained to maintain the correct spelling of words.
[0222] After the text preprocessing is completed, a subword tokenizer is used to split the normalized text into character units or subword units. Subword tokenization is a key technology in low-resource language speech synthesis, which can reduce the occurrence of out-of-vocabulary words and improve the generalization ability of the system. Common subword tokenization methods include BPE (Byte Pair Encoding) and Unigram language models. The former merges high-frequency character sequences based on statistics, and the latter dynamically selects the optimal subword segmentation method based on probability modeling. The goal of subword tokenization is to decompose the input text into appropriate language units according to the structural characteristics of the language, so that it can be effectively matched with the vocabulary items in the extended vocabulary.
[0223] After the word segmentation is completed, each character unit or subword unit after segmentation needs to be traversed, and it is detected whether it exists in the extended vocabulary. If the current character unit or subword unit already exists in the extended vocabulary, it can be directly mapped to the corresponding vocabulary token and used as part of the input sequence. If the current character unit or subword unit is not in the extended vocabulary, the subword segmentation algorithm needs to be used to further split it into smaller language units to find the closest in-vocabulary subword and map it to the corresponding subword token. This can ensure that even when encountering new words, the system can still reasonably disassemble them to meet the input requirements of the speech synthesis model.
[0224] After all character units and sub-word units are mapped, according to the text order of the target input text, all vocabulary tokens and sub-word tokens are concatenated to finally form an input token sequence. This input token sequence will be used as the input of the speech synthesis system, enabling the model to generate high-quality speech data based on standardized language units.
[0225] In this embodiment, by performing standardization processing on the target input text and combining sub-word tokenization and vocabulary matching, the speech synthesis system can more stably parse low-resource language texts and improve the accuracy of text-to-speech conversion. Text preprocessing can effectively reduce noise in the corpus and improve the robustness of the model; sub-word tokenization can reduce the impact of out-of-vocabulary words, enabling the system to handle new words reasonably; based on the matching of the extended vocabulary and the sub-word splitting algorithm, it is ensured that all input texts can be mapped to a token sequence recognizable by the system, thereby optimizing the naturalness and fluency of speech synthesis.
[0226] In one embodiment, the above S100 includes:
[0227] S1001, concatenating the context semantic features and the pronunciation rule features according to the channel dimension to generate fused features;
[0228] S1002, mapping the fused features to Mel spectrogram frames through the fully connected layer of the decoder;
[0229] S1003, stacking the Mel spectrogram frames along the time axis to generate an initial Mel spectrogram frame sequence;
[0230] S1004, generating target duration parameters according to the semantic information of the target input text through the duration prediction module of the decoder;
[0231] S1005, performing linear interpolation processing on the time axis of the Mel spectrogram frames according to the target duration parameters;
[0232] S1006, adjusting the number of frames of the interpolated Mel spectrogram frame sequence to align the time axis of the Mel spectrogram frame sequence with the target duration parameters, generating a temporally continuous Mel spectrogram acoustic feature sequence.
[0233] In this embodiment, the decoder is responsible for converting the encoded text features into an acoustic feature sequence for speech synthesis in the speech synthesis task, ensuring that the generated speech conforms to the pronunciation rules of the target language and maintaining a coherent and natural listening experience. This process involves fusing context semantic features and pronunciation rule features, enabling the decoder to comprehensively consider the semantic consistency and pronunciation accuracy of speech, thereby generating a high-quality acoustic feature sequence.
[0234] During the decoding process, it is first necessary to concatenate the context semantic features and the pronunciation rule features along the channel dimension to generate fused features. The context semantic features come from the output of the encoder and contain the overall semantic information of the text, syntactic structure, and dependencies between words. The pronunciation rule features, on the other hand, are derived from the phonetic conversion process and contain the phonological features of the target language, phoneme relationships, and possible pronunciation variants. The concatenation method along the channel dimension enables the decoder to simultaneously focus on semantic information and phonological features when processing the input data, improving the accuracy and naturalness of pronunciation. The concatenated fused features are usually processed using Layer Normalization to maintain the stability of the numerical distribution and reduce computational errors caused by feature conflicts.
[0235] After the fused features are generated, it is necessary to map the fused features to Mel spectrogram frames through the fully connected layer of the decoder. Mel spectrogram is a commonly used acoustic feature representation method that can reflect the frequency distribution perceived by the human ear, making the synthesized speech more natural. The role of the fully connected layer is to map the high-dimensional fused features to the acoustic feature space so that the output of the decoder can match the input requirements of speech synthesis. During the mapping process, non-linear activation functions (such as ReLU, Tanh) are usually combined to enhance the expressive ability of the decoder and improve the fitting accuracy of the Mel spectrogram.
[0236] After the Mel spectrogram frames are generated, the Mel spectrogram frames need to be stacked along the time axis to generate an initial sequence of Mel spectrogram frames. Since the output of the decoder is usually generated frame by frame, and natural speech is continuous, directly using the features predicted frame by frame may lead to unstable speech quality. Therefore, by stacking along the time axis, the continuity of temporal information can be increased, making the synthesized speech smoother. During the stacking process, a sliding window method or a time step expansion strategy can be adopted to ensure the rationality of the speech duration and reduce the break phenomenon in the speech.
[0237] In terms of duration modeling, the duration prediction module of the decoder generates target duration parameters based on the semantic information of the target input text. The duration of speech is closely related to the structure, prosody, and pause rules of the text. The duration prediction module needs to model the duration of each syllable, word, or phrase by combining the characteristics of the target language. Common methods include statistical-based duration prediction (modeling by analyzing the syllable duration distribution in the corpus) and deep learning-based duration prediction (using neural networks such as Transformer, BiLSTM to predict syllable duration). The goal of this module is to ensure that the speech duration conforms to the natural prosody of the target language and improve the intelligibility of speech synthesis.
[0238] After the duration prediction is completed, it is necessary to perform linear interpolation on the time axis of the Mel spectrogram frames according to the target duration parameter. Since the original acoustic feature sequence generated by the decoder may not accurately match the actual duration of the target speech, interpolation adjustment is required to make the Mel spectrogram conform to the pronunciation rhythm of the target language in the time dimension. Linear interpolation is an efficient method that can smoothly expand the features without introducing additional noise. In addition, for languages with large variations in pronunciation duration (such as Vietnamese and Thai), a dynamic duration adjustment strategy can be combined, that is, context information is combined during the interpolation process to dynamically adjust the interpolation parameters to optimize the smoothness of syllable transitions.
[0239] Finally, it is necessary to adjust the number of frames in the interpolated Mel spectrogram frame sequence so that its time axis is aligned with the target duration parameter to generate a temporally continuous Mel spectrogram acoustic feature sequence. Through this process, it is ensured that the output of speech synthesis conforms to the speech rhythm of the target language, improving the coherence and naturalness of the synthesized speech.
[0240] In this embodiment, the decoder fuses context semantic features and pronunciation rule features, optimizing the pronunciation accuracy and naturalness of the speech synthesis system. The channel dimension splicing strategy is adopted, enabling the decoder to simultaneously focus on semantic information and phonological features when processing text input, improving the coherence and stability of pronunciation. The Mel spectrogram frames are mapped through a fully connected layer to ensure the high fidelity of speech features, making the synthesized speech closer to real pronunciation. The duration prediction module is used to optimize the syllable duration, enabling speech synthesis to better conform to the natural rhythm of the target language and avoiding rhythm imbalance problems during the speech generation process.
[0241] In one embodiment, the above S110 includes:
[0242] S1101, dividing the acoustic feature sequence into Mel spectrogram frames with a fixed duration according to a preset time window;
[0243] S1102, processing each Mel spectrogram frame frame by frame through a pre-trained vocoder to generate a speech waveform segment corresponding to each frame of the Mel spectrogram;
[0244] S1103, performing window function weighted summation processing on the overlapping regions at the time axis connection points of adjacent speech waveform segments to generate a continuous speech waveform;
[0245] S1104, performing amplitude normalization processing on the continuous speech waveform to generate the target speech data.
[0246] In this embodiment, converting the acoustic feature sequence into target speech data is a key step in the speech synthesis system. This process further processes the acoustic features generated by the decoder and finally outputs a playable high-quality speech waveform. Since the acoustic feature sequence only contains the spectral information of the speech, it needs to be processed by a vocoder to convert it into real-time domain waveform data and perform subsequent optimization to improve the naturalness and clarity of the speech.
[0247] First, the acoustic feature sequence needs to be segmented into Mel spectrogram frames of a fixed duration according to a preset time window. The acoustic feature sequence is usually a continuous time series, and each frame represents the speech spectral information of a specific time period. To adapt to the input format of the vocoder, it needs to be segmented according to a fixed time window to ensure that the spectral information of each time step is complete and meets the input requirements of the neural network. The size of the time window can be adjusted according to the target application. For example, a short time window can improve the clarity of the speech, while a long time window can reduce the computational overhead and improve the efficiency of speech synthesis. After segmentation, the spectral information of each time step will be used as an independent input frame for the vocoder to process frame by frame.
[0248] After the segmentation is completed, it is necessary to process the Mel spectrogram frames frame by frame through a pre-trained vocoder to generate speech waveform segments corresponding to each frame of the Mel spectrogram. The main function of the vocoder is to convert spectral information into a time-domain signal and maintain the naturalness and clarity of the speech. Common vocoders include:
[0249] WaveNet: Based on an autoregressive model, it can generate high-quality speech waveforms, but has a large computational overhead and is suitable for high-precision speech synthesis tasks.
[0250] WaveGlow: Based on a flow-based method, it has high computational efficiency and is suitable for real-time speech synthesis applications.
[0251] HiFi-GAN: Based on a generative adversarial network (GAN), it can generate high-quality and coherent speech waveforms and is suitable for low-latency and high-quality speech synthesis tasks.
[0252] The vocoder adopts a frame-by-frame processing method, and each frame of the Mel spectrogram will be mapped to the corresponding speech waveform segment to ensure the time consistency and sound quality stability of the speech data.
[0253] After the voice waveform segment generated by the vocoder, it is necessary to perform window function weighted summation on the overlapping region at the time axis connection of adjacent voice waveform segments to generate a continuous voice waveform. Since the vocoder usually generates frames one by one, while natural speech is continuous, directly splicing the voice waveforms may produce discontinuous boundary effects, resulting in breaks or unevenness in the synthesized speech. Therefore, in the overlapping region of adjacent waveform segments, it is necessary to use the method of window function weighted summation for smoothing, making the transition at the connection more natural and reducing the problem of waveform discontinuity. Common window functions include Hann Window, Hamming Window, and Gaussian Window, which can effectively smooth the waveform transition and improve the coherence of the speech.
[0254] Finally, it is necessary to perform amplitude normalization on the continuous voice waveform to generate the target voice data. The main purpose of the normalization process is to adjust the amplitude range of the speech, avoid excessive or too low volume, and at the same time improve the clarity and audibility of the speech. Normalization can be performed using peak normalization or root mean square (RMS) normalization, so that the final voice data meets the expected volume standard and can obtain a consistent listening experience on different playback devices.
[0255] In this embodiment, by converting the acoustic feature sequence into voice data, the speech synthesis system can output high-quality voice waveforms, improving the naturalness and clarity of speech synthesis. By adopting the Mel spectrum segmentation method, it is ensured that the vocoder can process the input data frame by frame and ensure the stability of the calculation. Using a pre-trained vocoder can optimize the voice quality of speech synthesis and improve the adaptability of the system. Through window function weighted smoothing, the connection of adjacent voice segments is more natural, reducing the problem of waveform discontinuity. Finally, through amplitude normalization processing, the volume consistency of the speech is optimized, improving the audibility of the speech and making it suitable for different playback environments.
[0256] In one embodiment, after the above S80, it further includes:
[0257] S8001, when the extended vocabulary contains character symbols but does not contain phonetic symbols, generating Mel spectrum acoustic features based on the context semantic features by the decoder;
[0258] S8002, inputting the Mel spectrum acoustic features into a pre-trained vocoder to generate initial voice waveform data;
[0259] S8003, performing amplitude normalization on the initial voice waveform data to generate the target voice data.
[0260] In this embodiment, in a speech synthesis system, the context semantic features extracted by the encoder are used to model the global dependencies of the input text and provide semantic information for the decoder to generate high-quality speech signals. When the extended vocabulary only contains character symbols and does not contain phonetic symbols, the speech synthesis system needs to rely on semantic features to predict pronunciation characteristics to ensure that the generated speech conforms to the natural pronunciation pattern of the target language.
[0261] In this case, the decoder directly generates Mel-spectrum acoustic features based on the context semantic features output by the encoder. The Mel-spectrum is a compact and efficient acoustic feature representation that can reflect the frequency distribution of speech signals and match the characteristics of human auditory perception. Since the input text lacks explicit phonetic information, the decoder needs to combine context information to infer pronunciation rules and dynamically adjust speech rhythm, prosody, stress, and pauses during the generation of the Mel-spectrum, making the synthesized speech more natural. To this end, the decoder can adopt:
[0262] Autoregressive decoding based on Transformer, which learns context associations in the text through the attention mechanism to improve the coherence of speech synthesis.
[0263] Based on non-autoregressive models (such as FastSpeech), directly predict the entire Mel-spectrum sequence to improve the generation speed and ensure the stability of speech rhythm.
[0264] Combine a duration prediction module to optimize the pronunciation time between syllables and avoid uneven speech speed or abnormal rhythm caused by the lack of phonetic information.
[0265] After the Mel-spectrum is generated, it needs to be input into a pre-trained vocoder to generate initial speech waveform data. The role of the vocoder is to map the Mel-spectrum to the time-domain waveform and restore the natural speech signal. Different types of vocoders are suitable for different speech synthesis requirements:
[0266] WaveNet adopts an autoregressive modeling method to generate high-quality speech, but has a high computational complexity and is suitable for high-precision speech synthesis tasks.
[0267] HiFi-GAN is based on a generative adversarial network (GAN) and can generate high-quality speech at a low computational cost, making it suitable for low-latency real-time speech applications.
[0268] WaveGlow adopts a streaming generation method, which is computationally efficient and suitable for scenarios with high requirements for generation speed.
[0269] After generating the initial speech waveform data, it is necessary to perform amplitude normalization on it to generate the target speech data. The purpose of normalization is to ensure that the dynamic range of the speech signal is moderate, improve the volume consistency, and eliminate possible amplitude fluctuation problems. The methods of normalization processing include:
[0270] Peak Normalization, which adjusts the maximum amplitude of the speech signal to the standard range to avoid excessive or too small volume differences.
[0271] Root Mean Square (RMS) Normalization, which normalizes according to the overall energy of the signal to ensure the volume consistency between different synthesized segments and improve the listening experience.
[0272] Finally, the normalized target speech data can be used in application scenarios such as intelligent voice assistants, medical reading, and financial news broadcasts to ensure that the generated speech has a stable sound quality and an appropriate volume level.
[0273] In this embodiment, the Mel-spectrum acoustic features are generated by a decoder and combined with a pre-trained vocoder for waveform conversion, enabling the speech synthesis system to still generate natural and fluent speech in the absence of phonetic information. By adopting pronunciation modeling based on context semantic features, the system can correctly predict the prosody, rhythm, and pauses of speech under unknown words or complex syntactic structures, improving the naturalness of the synthesized speech. Combining different types of vocoders to optimize the speech quality enables the system to generate speech close to human pronunciation in high-precision scenarios and also achieve efficient speech synthesis under low computational resource conditions. Finally, the dynamic range of the speech signal is adjusted by amplitude normalization to improve the audibility of the speech data.
[0274] In one embodiment, a speech generation device is provided, which corresponds one-to-one with the speech generation method in the above embodiment. Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of a preferred embodiment of the speech generation device of the present invention. The model construction module 10, the data acquisition module 20, the vocabulary construction module 30, the unsupervised training module 40, the supervised training module 50, the text processing module 60, the embedding representation module 70, the feature extraction module 80, the phoneme feature extraction module 90, the feature fusion module 100, and the speech generation module 110. The detailed description of each functional module is as follows:
[0275] The model construction module 10 is used to construct a multi-language speech synthesis model including a language perception embedding layer, an encoder, and a decoder;
[0276] The data acquisition module 20 is used to acquire the plain text dataset of the target language and the corresponding paired speech text data;
[0277] A glossary construction module 30 for constructing an extended glossary of the target language;
[0278] An unsupervised training module 40 for updating the parameters of the language-aware embedding layer using the plain text dataset through a masked language model;
[0279] A supervised training module 50 for updating the parameters of the multilingual speech synthesis model using the paired speech-text data;
[0280] A text processing module 60 for converting the target input text into an input token sequence according to the extended glossary;
[0281] An embedding representation module 70 for converting the input token sequence into an initial vector representation through the language-aware embedding layer;
[0282] A feature extraction module 80 for extracting context semantic features of the initial vector representation through the encoder;
[0283] A phoneme feature extraction module 90 for extracting pronunciation rule features from the phonetic symbols when the extended glossary contains phonetic symbols;
[0284] A feature fusion module 100 for fusing the context semantic features and the pronunciation rule features through the decoder to generate an acoustic feature sequence;
[0285] A speech generation module 110 for converting the acoustic feature sequence into target speech data.
[0286] In one embodiment, the model construction module 10 is specifically configured to:
[0287] Construct a language-aware embedding layer and initialize the parameters of the language-aware embedding layer;
[0288] Construct an encoder composed of multiple stacked Transformer encoding layers, each encoding layer including a multi-head self-attention mechanism and a feed-forward network;
[0289] Construct a decoder composed of multiple stacked Transformer decoding layers, each decoding layer including a masked multi-head self-attention mechanism, a cross-attention mechanism, and a feed-forward network;
[0290] Configure the output dimension of the language-aware embedding layer to be consistent with the input dimension of the encoder;
[0291] Configure the output dimension of the encoder to be consistent with the input dimension of the decoder.
[0292] In one embodiment, the glossary construction module 30 is specifically configured to:
[0293] Extract character symbols from the plain text dataset of the target language;
[0294] Detect whether there is an effective character-to-phonetic conversion tool for the target language;
[0295] When it is detected that there is an effective character-to-phonetic conversion tool for the target language, convert the character symbols into corresponding phonetic symbols through the character-to-phonetic conversion tool, and generate an extended vocabulary based on the character symbols and phonetic symbols;
[0296] When it is detected that there is no effective character-to-phonetic conversion tool for the target language, construct an extended vocabulary containing only the character symbols.
[0297] In one embodiment, the text processing module 60 is specifically configured to:
[0298] Preprocess the target input text to remove punctuation marks and non-target language characters, and generate a normalized text;
[0299] Segment the normalized text into character units or sub-word units through a sub-word tokenization tool;
[0300] Traverse each character unit or sub-word unit after segmentation, and detect whether each character unit or sub-word unit exists in the extended vocabulary;
[0301] If the current character unit or sub-word unit exists in the extended vocabulary, map the current character unit or sub-word unit to the corresponding vocabulary token;
[0302] If the current character unit or sub-word unit does not exist in the extended vocabulary, split the current character unit or sub-word unit into registered sub-word units based on the sub-word segmentation algorithm, and map each registered sub-word unit to the corresponding sub-word token;
[0303] Concatenate all vocabulary tokens and sub-word tokens according to the text order of the target input text to generate an input token sequence.
[0304] In one embodiment, the feature fusion module 100 is specifically configured to:
[0305] Concatenate the context semantic feature and the pronunciation rule feature according to the channel dimension to generate a fused feature;
[0306] Map the fused feature to Mel spectrogram frames through the fully connected layer of the decoder;
[0307] Stack the Mel spectrogram frames along the time axis to generate an initial Mel spectrogram frame sequence;
[0308] Generate a target duration parameter according to the semantic information of the target input text through the duration prediction module of the decoder;
[0309] Perform linear interpolation processing on the time axis of the Mel spectrogram frames according to the target duration parameter;
[0310] Adjust the number of frames of the interpolated Mel spectrogram frame sequence to align the time axis of the Mel spectrogram frame sequence with the target duration parameter, and generate a temporally continuous Mel spectrogram acoustic feature sequence.
[0311] In one embodiment, the speech generation module 110 is specifically configured to:
[0312] Divide the acoustic feature sequence into Mel spectrogram frames with a fixed duration according to a preset time window;
[0313] Process the Mel spectrogram frames frame by frame through a pre-trained vocoder to generate speech waveform segments corresponding to each frame of the Mel spectrogram;
[0314] Perform window function weighted summation processing on the overlapping regions at the time axis connection points of adjacent speech waveform segments to generate a continuous speech waveform;
[0315] Perform amplitude normalization processing on the continuous speech waveform to generate the target speech data.
[0316] In one embodiment, the feature extraction module 80 is specifically configured to:
[0317] When the extended vocabulary contains character symbols but does not contain phonetic symbols, generate Mel spectrogram acoustic features through the decoder based on the context semantic features;
[0318] Input the Mel spectrogram acoustic features into a pre-trained vocoder to generate initial speech waveform data;
[0319] Perform amplitude normalization processing on the initial speech waveform data to generate the target speech data.
[0320] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice generation method.
[0321] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a voice generation method
[0322] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0323] Construct a multi-language speech synthesis model including a language perception embedding layer, an encoder, and a decoder;
[0324] Obtain a pure text dataset of the target language and the corresponding paired speech text data;
[0325] Construct an extended vocabulary of the target language;
[0326] Use the pure text dataset to update the parameters of the language perception embedding layer through a masked language model;
[0327] Use the paired speech text data to update the parameters of the multi-language speech synthesis model;
[0328] Convert the target input text into an input token sequence according to the extended vocabulary;
[0329] Convert the input token sequence into an initial vector representation through the language perception embedding layer;
[0330] Extract the context semantic features of the initial vector representation through the encoder;
[0331] When the extended vocabulary contains phonetic symbols, extract the pronunciation rule features from the phonetic symbols;
[0332] Fuse the context semantic features and the pronunciation rule features through the decoder to generate an acoustic feature sequence;
[0333] Convert the acoustic feature sequence into target speech data.
[0334] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0335] Construct a multi-language speech synthesis model including a language perception embedding layer, an encoder, and a decoder;
[0336] Obtain a plain text data set of the target language and the corresponding paired speech text data;
[0337] Construct the extended vocabulary of the target language;
[0338] Use the plain text data set to update the parameters of the language perception embedding layer through a masked language model;
[0339] Use the paired speech text data to update the parameters of the multi-language speech synthesis model;
[0340] Convert the target input text into an input token sequence according to the extended vocabulary;
[0341] Convert the input token sequence into an initial vector representation through the language perception embedding layer;
[0342] Extract the context semantic features of the initial vector representation through the encoder;
[0343] When the extended vocabulary contains phonetic symbols, extract the pronunciation rule features from the phonetic symbols;
[0344] Fuse the context semantic features and the pronunciation rule features through the decoder to generate an acoustic feature sequence;
[0345] Convert the acoustic feature sequence into target speech data.
[0346] It should be noted that for the functions or steps that can be achieved by the above-mentioned computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0347] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0348] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0349] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.< / mask>
Claims
1. A speech generation method, characterized in that: The following steps are involved: Build a multilingual speech synthesis model that includes a language-aware embedding layer, encoder, and decoder; Obtain a plain text dataset in the target language and the corresponding paired speech and text data; constructing an extended vocabulary of the target language; Using the plain text dataset, updating parameters of the language-aware embedding layer through a masked language model; Using the paired speech-text data to update the parameters of the multilingual speech synthesis model; converting a target input text into a sequence of input tokens according to the expanded vocabulary; Converting the input token sequence into an initial vector representation through the language-aware embedding layer; Extracting contextual semantic features represented by the initial vector through the encoder; When the extended vocabulary includes phonetic symbols, extracting pronunciation rule features from the phonetic symbols; fusing the context semantic features with the pronunciation rule features through the decoder to generate an acoustic feature sequence; The acoustic feature sequence is converted into target speech data.
2. The speech generation method according to claim 1, characterized in that: Build a multilingual speech synthesis model with language-aware embedding layers, encoders, and decoders, including: Constructing a language-aware embedding layer and initializing parameters of the language-aware embedding layer; Construct an encoder consisting of multiple stacked Transformer encoding layers, each of which contains a multi-head self-attention mechanism and a feed-forward network; Build a decoder consisting of multiple stacked Transformer decoding layers, each of which contains a masked multi-head self-attention mechanism, a criss-cross attention mechanism, and a feed-forward network; Configuring the output dimension of the language-aware embedding layer to be consistent with the input dimension of the encoder; The output dimension of the encoder is configured to be consistent with the input dimension of the decoder.
3. The speech generation method according to claim 1, wherein: Building an extended vocabulary of the target language, including: Extracting character symbols from a plain text dataset in the target language; Detecting whether there is an effective character-to-phonetic symbol conversion tool in the target language; When it is detected that there is an effective character-to-phonetic symbol conversion tool in the target language, the character symbols are converted into corresponding phonetic symbols by the character-to-phonetic symbol conversion tool, and an extended vocabulary is generated based on the character symbols and the phonetic symbols; When it is detected that there is no valid character-to-phonetic symbol conversion tool for the target language, an extended vocabulary containing only the character symbols is constructed.
4. The speech generation method according to claim 1, wherein: Converting the target input text into an input token sequence according to the expanded vocabulary comprises: Preprocessing the target input text to remove punctuation marks and non-target language characters to generate a standardized text; Segmenting the normalized text into character units or subword units by a subword segmentation tool; Traversing each character unit or subword unit after segmentation, and detecting whether each character unit or subword unit exists in the extended vocabulary; If the current character unit or subword unit exists in the extended vocabulary, mapping the current character unit or subword unit to a corresponding vocabulary token; If the current character unit or subword unit does not exist in the extended vocabulary, split the current character unit or subword unit into registered subword units based on a subword segmentation algorithm, and map each registered subword unit to a corresponding subword tag; All vocabulary tags and subword tags are concatenated according to the text order of the target input text to generate an input tag sequence.
5. The speech generation method according to claim 1, wherein: The decoder fuses the context semantic features with the pronunciation rule features to generate an acoustic feature sequence, including: The context semantic feature and the pronunciation rule feature are spliced according to the channel dimension to generate a fusion feature; Mapping the fused features into a Mel spectrum frame through a fully connected layer of the decoder; Stacking the Mel spectrum frames along the time axis to generate an initial Mel spectrum frame sequence; Generate a target duration parameter according to the semantic information of the target input text through the duration prediction module of the decoder; According to the target duration parameter, a linear interpolation process is performed on the time axis of the Mel spectrum frame; The number of frames of the interpolated Mel spectrum frame sequence is adjusted so that the time axis of the Mel spectrum frame sequence is aligned with the target duration parameter, and a temporally continuous Mel spectrum acoustic feature sequence is generated.
6. The speech generation method according to claim 1, wherein: Converting the acoustic feature sequence into target speech data comprises: Dividing the acoustic feature sequence into Mel-spectrogram frames of fixed length according to a preset time window; Processing the mel spectrum frames frame by frame by a pre-trained vocoder to generate a speech waveform segment corresponding to each mel spectrum frame; Performing window function weighted addition processing on the overlapping areas of adjacent speech waveform segments at the connection of the time axis to generate a continuous speech waveform; Amplitude normalization is performed on the continuous speech waveform to generate the target speech data.
7. The speech generation method according to claim 1, characterized in that: After extracting the contextual semantic features represented by the initial vector through the encoder, the method further includes: When the extended vocabulary includes character symbols but does not include phonetic symbols, generating Mel-spectrogram acoustic features based on the contextual semantic features through the decoder; Inputting the Mel-spectrogram acoustic features into a pre-trained vocoder to generate initial speech waveform data; The initial speech waveform data is subjected to amplitude normalization processing to generate target speech data.
8. A speech generating device, characterized in that: The speech generating device comprises: A model building module for building a multilingual speech synthesis model that includes a language-aware embedding layer, encoder, and decoder; A data acquisition module, used to acquire a plain text data set in a target language and corresponding paired speech and text data; A vocabulary building module, used for building an extended vocabulary of the target language; an unsupervised training module for updating parameters of the language-aware embedding layer through a masked language model using the plain text dataset; A supervised training module, used to update the parameters of the multilingual speech synthesis model using the paired speech-text data; a text processing module, configured to convert a target input text into an input token sequence according to the expanded vocabulary; An embedding representation module, configured to convert the input token sequence into an initial vector representation via the language-aware embedding layer; A feature extraction module, used for extracting contextual semantic features represented by the initial vector through the encoder; A phoneme feature extraction module, used for extracting pronunciation rule features from the phonetic symbols when the extended vocabulary contains phonetic symbols; A feature fusion module, used for fusing the context semantic features with the pronunciation rule features through the decoder to generate an acoustic feature sequence; The speech generation module is used to convert the acoustic feature sequence into target speech data.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a speech generation program stored in the memory and executable on the processor. When the speech generation program is executed by the processor, the steps of the speech generation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a speech generation program, which, when executed by a processor, implements the steps of the speech generation method according to any one of claims 1 to 7.
Citation Information
Cited By
Audio real-time conversion and analysis management system and method based on artificial intelligence
CN120564735A
Audio real-time conversion and analysis management system and method based on artificial intelligence
CN120564735B
Multi-language TTS real-time synthesis method based on deep learning
CN120580987A
A Deep Learning-Based Real-Time Multilingual TTS Synthesis Method
CN120580987B
Data monitoring system and method for voice processing
CN120727038A