Knowledge distillation-based text-to-voice method, apparatus and device, and medium

Through the lightweight text-to-speech method based on knowledge distillation, combined with structured pruning and parameter quantization, the problems of huge models, poor adaptability and high energy consumption in the prior art are solved, and efficient and low-latency speech generation in low-resource environments are achieved.

CN120220641APending Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510417640.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing text-to-speech technology model is huge, has poor adaptability and high energy consumption, making it difficult to achieve efficient and low-latency speech generation in a low-resource environment.

Method used

Using a knowledge distillation-based method, lightweight text encoder and non-autoregressive acoustic feature prediction module are used to combine structured pruning and parameter quantization to reduce model volume and calculation overhead and improve inference speed.

Benefits of technology

While maintaining the quality of speech generation, it significantly reduces model size and computing overhead, improves inference speed, enhances cross-device adaptability, and enables TTS systems to achieve efficient, low-power, and real-time voice synthesis in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220641A_ABST
    Figure CN120220641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes in the fields of medical health, financial science and technology, barrier-free service and the like, and discloses a text-to-voice method based on knowledge distillation, which comprises the following steps: carrying out standardization processing on an input text to generate a standard text sequence; the lightweight text encoder encodes the standard text sequence to generate a text implicit vector; the non-autoregressive acoustic feature prediction module maps the text implicit vector into a student acoustic feature sequence, and calculates alignment loss through knowledge distillation; and performing structured pruning and parameter quantization based on alignment loss, generating an acoustic feature sequence by the optimized model, and converting the acoustic feature sequence into a voice waveform by a vocoder. According to the method, through knowledge distillation, pruning optimization and parameter quantification, the reasoning speed and the cross-equipment adaptability are improved while the model size and the calculation requirement are reduced, so that the TTS system can realize efficient, low-delay and low-power-consumption voice generation in a resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a text-to-speech method, apparatus, device, and storage medium based on knowledge distillation. Background Art

[0002] In the field of medical and health services, TTS technology is gradually being applied to scenarios such as intelligent medical assistants, telemedicine consultations, and electronic medical record reading assistance to improve the accessibility and interaction experience of medical services. However, the current TTS solutions still have various limitations in the medical industry. The speech synthesis requirements in the medical field usually involve complex medical terms, medical record content, and patient consultation records. When existing TTS models process highly specialized medical texts, they often fail to accurately express medical terms, which can easily lead to misunderstandings in information transmission. In addition, applications such as telemedicine and intelligent health assistants require real-time speech generation to ensure smooth communication between doctors and patients. However, due to the slow inference speed of existing TTS models, the speech generation process may experience stuttering or delays, affecting the efficiency of medical services. At the same time, there is a high degree of device diversity in the medical industry, and the TTS system needs to adapt to different platforms such as hospital information systems, mobile health devices, and voice interaction terminals. Existing models still have deficiencies in device adaptability. Since the medical environment has high requirements for speech quality, the speech synthesis quality of existing TTS solutions may decline in a noisy environment, affecting the effective communication between doctors and patients. In addition, the medical industry has strict requirements for data security and privacy protection. Most existing TTS solutions rely on cloud computing, and medical data involves patient privacy. Directly using cloud TTS may pose data security risks, limiting its popularization and application in medical scenarios.

[0003] In the field of fintech business, TTS technology is widely used in interactive scenarios such as intelligent customer service, voice broadcasting, and automatic transaction reminders to provide efficient information transmission and user services. However, existing TTS solutions still have obvious limitations in the application of financial services. First of all, many financial service scenarios require real-time responses, such as intelligent voice customer service systems, risk control warning broadcasts, etc. However, due to the slow inference speed of current TTS solutions, it is difficult to meet the business requirements of high concurrency and low latency. In addition, the voice interaction systems in the financial field often involve highly personalized information, such as users' account data, transaction details, etc. Existing TTS models lack adaptive optimization of business-specific terms during the voice generation process, resulting in insufficient professionalism and accuracy of the voice output. At the same time, the financial system needs to deploy voice synthesis systems on different platforms and devices, but current TTS models still have problems in cross-platform adaptability. For example, existing models can provide high-quality voice synthesis on the server side, but when running on mobile devices, ATM terminals or other embedded devices, due to limited computing resources, it is often difficult to maintain the same quality of voice output. In addition, the financial industry has extremely high requirements for data security and privacy protection. Traditional TTS solutions usually rely on cloud computing, which may increase the risk of user data leakage. Since voice generation involves sensitive information, current cloud-based TTS solutions are difficult to fully meet the strict requirements of the financial business for privacy and compliance.

[0004] In the field of accessibility services, text-to-speech (TTS) technology is widely used to provide voice assistance for visually impaired people, people with reading disabilities, and elderly users. However, existing TTS systems still face many challenges in practical applications. Mainstream TTS solutions, such as Google TTS, Amazon Polly, Microsoft Azure TTS, and open-source systems (such as Tacotron, FastSpeech), although they have reached a relatively high level in terms of speech synthesis quality, still have the following deficiencies when deployed on resource-constrained devices or in real-time interaction scenarios. First, current high-quality TTS models usually rely on large-scale neural networks with a large number of parameters and high computational requirements, making it difficult to run efficiently on mobile devices or embedded terminals. This greatly limits real-time voice generation on the device side and is difficult to meet the application requirements of low power consumption and high response speed. In addition, many end-to-end TTS systems still rely on step-by-step decoding or complex post-processing steps during the inference process, resulting in a slow system response speed and an insufficiently smooth interaction experience. For voice assistance systems that require instant feedback, this delay may affect the user experience and even reduce the usability of the system. At the same time, existing TTS models are usually optimized for large-scale training data, but they have insufficient stability when adapting to low-resource environments. When deployed to different types of terminal devices or facing complex environments (such as background noise, device computing power differences), the quality of the synthesized voice may decline, affecting the voice understanding and information acquisition of accessibility users. In addition, the high computational and storage requirements of high-performance TTS systems not only increase the energy consumption and cost of cloud computing but also limit the feasibility of large-scale promotion, making it difficult to popularize low-cost and low-power accessibility applications. Summary of the Invention

[0005] The main object of the present invention is to provide a text-to-speech method, device, equipment, and storage medium based on knowledge distillation, aiming to solve the technical problems that the existing text-to-speech technology has a large model, poor adaptability, and high energy consumption, and it is difficult to achieve efficient and low-latency voice generation in a low-resource environment.

[0006] To achieve the above object, the present invention provides a text-to-speech method based on knowledge distillation, including:

[0007] Performing normalization processing on the input text to generate a standard text sequence;

[0008] Encoding the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0009] Mapping the text hidden vector into a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module;

[0010] Encoding the standard text sequence and performing acoustic feature prediction processing through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0011] Determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module;

[0012] Performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0013] Performing parameter quantization processing on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module;

[0014] Encoding the standard text sequence through the lightweight text encoder after parameter quantization processing to generate a compressed text hidden vector;

[0015] Mapping the compressed text hidden vector to an optimized acoustic feature sequence through the non-autoregressive acoustic feature prediction module after parameter quantization processing;

[0016] Converting the optimized acoustic feature sequence into a speech waveform through a vocoder.

[0017] Furthermore, to achieve the above object, the present invention provides a text-to-speech device based on knowledge distillation, including:

[0018] A text preprocessing module for performing standardization processing on the input text to generate a standard text sequence;

[0019] A lightweight text encoding module for encoding the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0020] A non-autoregressive acoustic feature prediction module for mapping the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module;

[0021] A teacher model module for encoding the standard text sequence and performing acoustic feature prediction processing through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0022] A knowledge distillation module for determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module;

[0023] A structured pruning module for performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0024] A parameter quantization module for performing parameter quantization processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module after pruning processing;

[0025] The quantized lightweight text encoding module is used to encode the standard text sequence through the lightweight text encoder after parameter quantization processing to generate a compressed text hidden vector;

[0026] The quantized non-autoregressive acoustic feature prediction module is used to map the compressed text hidden vector into an optimized acoustic feature sequence through the non-autoregressive acoustic feature prediction module after parameter quantization processing;

[0027] The vocoder module is used to convert the optimized acoustic feature sequence into a speech waveform through the vocoder.

[0028] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a knowledge distillation-based text-to-speech program stored on the memory and executable on the processor. When the knowledge distillation-based text-to-speech program is executed by the processor, the steps of the knowledge distillation-based text-to-speech method as described above are implemented.

[0029] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a knowledge distillation-based text-to-speech program is stored. When the knowledge distillation-based text-to-speech program is executed by a processor, the steps of the knowledge distillation-based text-to-speech method as described above are implemented.

[0030] Beneficial effects: The present invention relates to the technical field of speech processing and can be applied to business scenarios such as medical health, fintech, and barrier-free services. A text-to-speech method based on knowledge distillation is disclosed, including: performing standardization processing on the input text to generate a standard text sequence; encoding the standard text sequence through a lightweight text encoder to generate a text hidden vector; mapping the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module; encoding and performing acoustic feature prediction processing on the standard text sequence by a pre-trained teacher model to generate a teacher acoustic feature sequence; a knowledge distillation module determining an alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence; performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss; performing parameter quantization processing on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module; the parameter-quantized lightweight text encoder encoding the standard text sequence to generate a compressed text hidden vector; the parameter-quantized non-autoregressive acoustic feature prediction module mapping the compressed text hidden vector to an optimized acoustic feature sequence; and a vocoder converting the optimized acoustic feature sequence into a speech waveform. By means of knowledge distillation, structured pruning, and parameter quantization, the present invention effectively reduces the model volume and computational overhead while maintaining the speech generation quality. The inference speed is improved and the speech generation latency is reduced through non-autoregressive acoustic feature prediction. The optimized lightweight text encoder and lightweight vocoder enhance cross-device adaptability, enabling the TTS system to achieve efficient, low-power, and real-time speech synthesis in resource-constrained environments and meeting the requirements for low-latency and high-quality speech output in fields such as barrier-free services, fintech, and medical health. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0032] Figure 1 is a schematic diagram of an application environment of a text-to-speech method based on knowledge distillation according to an embodiment of the present invention;

[0033] Figure 2 is a schematic flowchart of an embodiment of the text-to-speech method based on knowledge distillation of the present invention;

[0034] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of a text-to-speech device based on knowledge distillation of the present invention;

[0035] Figure 4 is a schematic diagram of a structure of a computer device according to an embodiment of the present invention;

[0036] Figure 5 is another schematic diagram of a structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.

[0038] The text-to-speech method based on knowledge distillation provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 . Among them, the client communicates with the server through the network. The server can perform standardization processing on the input text through the client to generate a standard text sequence; encode the standard text sequence through a lightweight text encoder to generate a text hidden vector; map the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module; the pre-trained teacher model encodes and performs acoustic feature prediction processing on the standard text sequence to generate a teacher acoustic feature sequence; the knowledge distillation module determines the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence; performs structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss; performs parameter quantization processing on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module; the lightweight text encoder after parameter quantization encodes the standard text sequence to generate a compressed text hidden vector; the non-autoregressive acoustic feature prediction module after parameter quantization maps the compressed text hidden vector to an optimized acoustic feature sequence; the vocoder converts the optimized acoustic feature sequence into a speech waveform. The present invention effectively reduces the model volume and computational overhead while maintaining the speech generation quality through knowledge distillation, structured pruning, and parameter quantization. The inference speed is improved and the speech generation latency is reduced through non-autoregressive acoustic feature prediction. The optimized lightweight text encoder and lightweight vocoder improve cross-device adaptability, enabling the TTS system to achieve efficient, low-power, and real-time speech synthesis in resource-constrained environments, meeting the requirements for low-latency and high-quality speech output in fields such as barrier-free services, fintech, and healthcare. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0039] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the text-to-speech method based on knowledge distillation provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.

[0040] As Figure 2 shown, the text-to-speech method based on knowledge distillation proposed by the present invention includes the following steps:

[0041] S10. Standardize the input text to generate a standard text sequence;

[0042] In this embodiment, the purpose of standardizing the input text is to unify text data from different sources, in different formats, and with different encodings into a standard format to ensure stable parsing and encoding in subsequent processing steps. The sources of the text can include user input, database storage, API interface data, web crawler content, etc. Different sources may have issues such as inconsistent encoding formats, mismatched character sets, and chaotic text structures. The core tasks of standardization processing include character encoding conversion, text symbol unification, word segmentation, punctuation normalization, special character filtering, language detection, etc.

[0043] Character encoding conversion requires format recognition for different input sources to ensure that text data is uniformly converted into a processable encoding format, such as UTF-8. Some text data may exist in formats such as GBK, Big5, ISO-8859-1, etc. Without conversion, character scrambling or loss may occur in subsequent processing. The conversion method can be based on an automatic encoding recognition algorithm or can be parsed through text meta-information.

[0044] The task of text symbol unification is mainly to process non-standard characters in different input sources, such as the conversion between full-width and half-width characters and the mapping of equivalent symbols in different languages. For example, in a financial scenario, there may be multiple expressions for currency symbols, such as "$100", "100 US dollars", "RMB100", etc., which need to be standardized into a unified format for subsequent parsing. In the medical and health scenario, the use of units may also be inconsistent. For example, "mg" and "milligram", and standardization processing can ensure data consistency.

[0045] Word segmentation involves splitting a continuous text stream into the smallest parseable language units, and different languages use different word segmentation strategies. For example, word segmentation of English text is mainly based on spaces or punctuation marks, while Chinese text needs to be segmented based on a word library match, statistical model, or deep learning algorithm. For medical texts, word segmentation also needs to combine a medical terminology dictionary to ensure the integrity of terms. For example, in medical literature or electronic medical records, "hypertension" should be regarded as a whole, rather than simply split into two separate words, "high" and "blood pressure".

[0046] The purpose of punctuation normalization is to ensure the clear structure of sentences so that subsequent processing modules can correctly parse sentence components. For example, different input devices may produce different quotation mark styles, such as “” and "", which need to be unified during text standardization. In addition, punctuation marks in Japanese and Korean may also be different from those in English and require special handling.

[0047] Special character filtering is mainly used to remove unnecessary control characters, invisible characters, or interfering symbols, such as zero-width space (ZWSP), non-printable characters, HTML escape characters, etc. In the financial and medical fields, some documents may contain format control characters that can affect text parsing and conversion, so they need to be identified and removed during the standardization process.

[0048] Language detection is particularly important in a multilingual environment, especially when dealing with cross-regional text processing. It is necessary to determine the language used in the text to decide on subsequent processing methods. For example, in a globalized financial service or a multinational medical system, a patient's health record may contain descriptions in multiple languages. Language detection needs to be performed first, and then the corresponding text parsing strategy can be selected. If the input text contains multiple languages, a deep learning-based multilingual recognition model can be used for analysis and subsequent processing can be carried out according to different languages.

[0049] Example illustration: In the field of healthcare, when doctors use a voice assistant to enter an electronic medical record, the text may contain different abbreviations, units, and terms. For example, "BP 120 / 80" means "blood pressure 120 / 80 mmHg", and "Na+135" means "sodium ion concentration 135 mmol / L". Text standardization can parse and convert these terms into structured data so that the TTS system can accurately read the medical record information, ensuring that medical staff can quickly understand the patient's health status. In addition, in remote medical consultations, patients may use dialects or non-standard expressions. For example, "the sugar is a bit high" may correspond to "the blood sugar level is on the high side". After standardization processing, more standardized text can be generated, enabling doctors to more accurately judge the condition.

[0050] In financial data processing, a stock market quotation system may need to parse various financial data sources, such as "Apple's stock price has increased by 2.5%" or "the NASDAQ index has dropped by 150 points". Standardization processing can ensure that the TTS system generates consistent voice broadcast content, avoiding information confusion caused by format differences. In a bank voice customer service, when a user queries the account balance, they may use different expressions, such as "check my money" or "how much is left in the account". The standardization module can convert these inputs so that the TTS system can correctly reply with a standardized voice output, such as "Your account balance is 1000 yuan".

[0051] In accessible service applications, the text sources may include different types of data such as web pages, news, e-books, etc. Standardization processing can remove unnecessary advertising symbols, HTML tags, redundant spaces, etc., making the voice content heard by blind users clearer. For example, the news title on a web page may contain redundant information. For social media comments, the system can remove emojis or meaningless characters to ensure the coherence of the voice broadcast and improve the comprehensibility of the information.

[0052] Through text standardization processing, the data format of the input text can be ensured to be consistent, improving the accuracy and stability of the subsequent TTS processing module. Unifying character encoding can avoid garbled characters and character loss. Unifying text symbols and word segmentation processing can improve the model's understanding ability. Punctuation normalization can ensure the integrity of sentence structure. Special character filtering can reduce interference factors. Language detection can improve the adaptability of multilingual text processing. These optimization measures can reduce the computational complexity of the model, improve the inference efficiency, and ensure the compatibility of different data sources, thus achieving higher-quality speech synthesis in different fields.

[0053] S20, encoding the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0054] In this embodiment, the role of the lightweight text encoder is to convert the standard text sequence into a high-dimensional hidden vector representation suitable for subsequent speech synthesis, so as to more effectively capture the semantics, syntactic structure, and pronunciation features of the text. The key objective of text encoding is to retain the necessary text information while reducing the computational overhead and storage requirements to achieve efficient text-to-speech conversion.

[0055] Compared with traditional Transformer encoders or LSTM structures, the lightweight text encoder mainly adopts technical means such as parameter optimization, model pruning, low-rank decomposition, and sparse attention to reduce the computational complexity and improve the inference speed. The encoding process usually includes the following key steps:

[0056] The text input is first converted into a vector representation of a fixed dimension through an embedding layer. The main role of the embedding layer is to convert discrete characters, words, or phonemes into continuous numerical vectors that can be used for deep learning calculations. The weights of the embedding layer can be pre-trained or jointly trained with the entire TTS system in an end-to-end manner.

[0057] The vector converted by the embedding layer is input to the multi-head self-attention layer, whose role is to capture the long-range dependencies in the text and ensure that the context information is fully expressed in speech generation. Since the multi-head attention calculation of the standard Transformer encoder is computationally intensive, the lightweight text encoder usually reduces the number of attention heads or uses techniques such as factorized attention and sparse attention to reduce the computational complexity.

[0058] Subsequently, the text hidden vector is normalized through a layer normalization layer to stabilize the training process and improve the model convergence speed. The role of the normalization layer is to ensure that the input data of different batches have similar distributions, reducing the risk of gradient vanishing or gradient explosion.

[0059] The normalized hidden vector undergoes information fusion through a residual connection layer, combining the initial embedding information with the context features to enhance the stability of text encoding. Residual connections can effectively alleviate the vanishing gradient problem in deep networks and improve the expressive power of the model.

[0060] Subsequently, the hidden vector is added with position information through a position encoding module to ensure that the model can correctly capture the sequential relationship of the text when processing sequential data. Due to the lightweight design requirement of reducing computational overhead, fixed position encoding (such as sine / cosine encoding) or learnable position encoding can be adopted, and the storage overhead can be reduced through dimensionality reduction optimization.

[0061] Finally, through a feed-forward network for non-linear transformation, the hidden vector can more effectively adapt to the subsequent speech feature prediction task. The feed-forward network can adopt low-rank transformation or grouped convolution to reduce the parameter scale, thereby optimizing the computational efficiency.

[0062] Through the optimized design of the lightweight text encoder, while reducing the computational complexity, the information retention ability of text encoding can be improved, ensuring that the key information in the text can be effectively transmitted to the subsequent speech synthesis stage. Reducing the number of attention heads and optimizing the feed-forward network structure can improve the inference speed of the model and meet the requirements of real-time speech synthesis. Adopting pruning, quantization, or knowledge distillation methods can reduce the model size while maintaining the speech quality, adapting to the operating environment of low-resource devices. In addition, using position encoding and residual connections ensures that the text order information is not lost in lightweight modeling, improving the cross-device and cross-scenario adaptation ability.

[0063] S30, mapping the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module;

[0064] In this embodiment, the role of the non-autoregressive acoustic feature prediction module is to map the text hidden vector to a student acoustic feature sequence, that is, to convert the text information processed by the lightweight text encoder into acoustic features that can be used for speech synthesis. Acoustic features usually adopt Mel-spectrogram or other parametric acoustic feature forms for further processing by the vocoder to generate the final speech waveform.

[0065] Traditional acoustic feature prediction usually adopts an Autoregressive (AR) model, such as Tacotron. This method generates an acoustic feature at each time step and relies on the prediction result of the previous step. Although the autoregressive method can provide high-quality sound, the inference process is slow, there are cumulative errors, and it is difficult to meet the requirements of real-time applications. The Non-Autoregressive (NAR) method effectively improves the inference speed by generating the entire acoustic feature sequence in parallel, while reducing the impact of cumulative errors.

[0066] During the non-autoregressive acoustic feature prediction process, the text hidden vector is first input into the parallel decoder, and local features are extracted through the convolutional layer. The role of the convolutional layer is to capture short-range patterns in the text hidden vector, such as the transition relationship between phonemes, and reduce the computational complexity. Since convolutional operations can be parallelized, it can significantly improve the computational efficiency.

[0067] Next, the global context attention layer is used for temporal dimension modeling to capture long-range dependencies at the sentence level. Compared with the standard Transformer attention mechanism, lightweight non-autoregressive models usually adopt low-rank attention or block attention strategies to reduce the computational resource requirements while ensuring the ability to understand the global information of the text.

[0068] Then, the context information is fused through the residual connection layer to ensure that while optimizing the computational complexity, the original information of the input hidden vector can still be maintained. Residual connections help with gradient stability and reduce information loss in deep networks.

[0069] Subsequently, the non-autoregressive acoustic feature prediction module uses a feed-forward network for non-linear transformation to convert the fused features into a representation more suitable for acoustic modeling. The feed-forward network usually adopts grouped fully connected layers or low-rank matrix factorization methods to reduce the parameter scale and ensure efficient computation.

[0070] Finally, through the linear projection layer, the transformed features are mapped into a Mel spectrogram sequence, which serves as the student acoustic feature sequence for subsequent vocoders. The role of linear projection is to compress the high-dimensional hidden vector into a dimension suitable for speech generation and ensure that the features match the target acoustic space.

[0071] The non-autoregressive acoustic feature prediction module can improve the inference speed and reduce the computational cost while ensuring high-quality speech generation. Compared with the autoregressive method, the non-autoregressive generation method eliminates the dependence on the previous time steps, making parallel computing possible and significantly reducing the synthesis latency. Through convolutional layers and global context modeling, the fluency and stability of speech generation are improved, and the lightweight design ensures feasibility in embedded devices and low-power scenarios. In addition, combined with personalized speech modeling, it can adapt to different speech style requirements and improve the adaptability and user experience of the TTS system.

[0072] S40, encoding and acoustically feature predicting the standard text sequence through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0073] In this embodiment, the pre-trained teacher model serves as a benchmark model in the text-to-speech process. Its main objective is to encode the standard text sequence and predict acoustic features to provide a high-quality teacher acoustic feature sequence for the subsequent knowledge distillation process. The teacher model is usually a large-scale and high-precision TTS model that contains rich language and speech knowledge and has strong generalization ability through large-scale training data. Compared with the lightweight student model, the teacher model can generate more natural, clear, and emotional acoustic feature sequences, but its computational cost is high and the inference speed is slow. Therefore, it is mainly used for offline training and knowledge distillation, rather than directly for real-time speech synthesis.

[0074] When processing the standard text sequence, the teacher model first encodes the input text. This process includes embedding layer mapping, multi-layer self-attention mechanism, positional encoding, and feed-forward network transformation, and finally generates a high-dimensional text hidden vector. This text hidden vector is used to represent the semantic information of the text and retains syntactic structure and pronunciation features to ensure higher accuracy in subsequent acoustic feature prediction.

[0075] In the acoustic feature prediction stage, the teacher model uses an autoregressive acoustic feature prediction network to gradually predict the acoustic features of each frame. Different from the non-autoregressive method adopted by the student model, the autoregressive prediction method of the teacher model depends on the output of the previous time step at each time step, so that it can learn more complex temporal dependencies and make the generated acoustic features smoother and more natural. However, the autoregressive method has a high computational cost and a slow inference speed. Therefore, the teacher model is usually used in the training stage, rather than in actual inference applications.

[0076] The final output of the teacher model is a high-quality acoustic feature sequence, such as a Mel spectrogram sequence. These sequences contain complete speech frequency, temporal information, and pronunciation features and can be used as target data for training the student model to enable it to generate speech features as close as possible to those of the teacher model under limited computational resources.

[0077] The teacher model can provide a high-quality speech feature benchmark, enabling the student model to learn a speech synthesis ability close to that of the teacher model through knowledge distillation, while significantly reducing the computational resource requirements. Since the teacher model usually adopts an autoregressive structure, the generated acoustic feature sequence is more fluent and natural, thus ensuring the accuracy of the learning target of the student model. In addition, through the multi-language and multi-emotion adaptation design, the teacher model can enhance the generalization ability of the TTS system, enabling it to meet the requirements of different application scenarios.

[0078] S50, determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through the knowledge distillation module;

[0079] In this embodiment, the core role of the knowledge distillation module is to compare the acoustic feature outputs of the teacher model and the student model, and calculate the alignment loss between them, which is used as the training target for optimizing the student model. Since the student model adopts a lightweight design, it may not be able to directly achieve the synthesis quality of the teacher model under limited computational resources. Therefore, knowledge distillation constructs a loss function to enable the student model to learn the high-quality acoustic features of the teacher model while reducing the computational complexity.

[0080] The calculation of the alignment loss in knowledge distillation usually involves several key steps. First, obtain the student acoustic feature sequence and the teacher acoustic feature sequence, ensuring that their time steps and feature dimensions are consistent. Since the generation method of the teacher model is autoregressive while the student model is non-autoregressive, during the inference process, the student model may produce an output that is shorter or longer than that of the teacher model. Therefore, a time alignment strategy, such as dynamic time warping (DTW) or interpolation method, is needed to align the feature sequences for subsequent loss calculation.

[0081] Next, calculate the probability distribution difference between the two sequences through the relative entropy loss function (KL divergence) to measure the similarity between the output of the student model and the teacher model. KL divergence is used to measure the information loss between two probability distributions, and can adjust the soft target learning degree of the teacher model's prediction value during the distillation process, enabling the student model to be as close as possible to the feature output of the teacher model while reducing computational resources.

[0082] In addition, the mean squared error (MSE) loss can be used to calculate the Euclidean distance between the teacher acoustic feature sequence and the student acoustic feature sequence, measuring the deviation between the two in the feature space. MSE is suitable for regression tasks and can ensure that the features generated by the student model are as close as possible to the output of the teacher model, improving the stability of speech generation.

[0083] After calculating the initial alignment loss, it is also necessary to weight the loss through a weight adjustment module. The role of the weight adjustment module is to balance the contributions of different loss terms and adapt to the optimization requirements of different training stages. For example, in the early stage of training, the predictions of the student model may deviate significantly, and the weight of the KL divergence loss can be increased to ensure that it can quickly approach the output of the teacher model. In the later stage of training, the influence of the KL divergence can be reduced while increasing the weight of the MSE loss, so that the student model is closer to the teacher model in terms of fine structure.

[0084] Finally, the result of the alignment loss needs to be verified, that is, to check whether the calculated loss is higher than the preset threshold. If the loss is high, it means that the generation quality of the student model is still poor and the weight parameters need to be further optimized. If the loss has converged to an acceptable range, stop optimizing the alignment loss and enter the next pruning and quantization processing stage.

[0085] Calculating the alignment loss through the knowledge distillation module can effectively reduce the computational complexity of the student model while maintaining the speech quality. Compared with directly training a lightweight TTS model, the knowledge distillation method can obtain an output closer to that of the teacher model under the same computing resources, improving the naturalness and clarity of the speech. In addition, through dynamic weight adjustment, the alignment loss can be adaptively optimized at different training stages to ensure the convergence speed and stability of the student model.

[0086] S60, perform structured pruning on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0087] In this embodiment, the purpose of the structured pruning is to reduce the computational amount and storage requirements of the lightweight text encoder and the non-autoregressive acoustic feature prediction module while maintaining the speech generation quality. In a text-to-speech system, the computational complexity of the model is usually affected by the text encoding and acoustic feature prediction stages. Pruning optimization can remove the computationally redundant neural network structure, reduce the computational load, improve the inference speed, and enable the model to operate efficiently on low-power devices or in real-time application environments.

[0088] The pruning process is based on the alignment loss, which means that during the pruning process, it is necessary to ensure that the outputs of the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module still meet the alignment goal to reduce the impact of pruning on the speech synthesis quality. Specifically, the pruning mainly includes the following key steps:

[0089] First, calculate the weight contributions of the multi-head self-attention layer of the lightweight text encoder based on the alignment loss, and determine the importance of each attention head. The pruning strategy can adopt weight importance analysis, that is, calculate the contribution of each attention head to the final acoustic feature prediction, and set the first pruning threshold based on the contribution value to remove the attention heads with low contribution. By reducing the number of attention heads, the computational complexity can be significantly reduced, while the storage overhead is also reduced.

[0090] Then, for the convolutional layer of the non-autoregressive acoustic feature prediction module, analyze the activation value distribution of its different channels, and calculate the second pruning threshold based on the average contribution of the channels. The feature channels with low contribution will be removed to reduce the computational complexity while maintaining the necessary acoustic feature information. In this process, the L1 regularization pruning method can be combined to make the channel pruning process smoother and ensure that the output features still contain sufficient information.

[0091] After pruning, it is necessary to verify the output errors of the pruned lightweight text encoder and the non-autoregressive acoustic feature prediction module to ensure that pruning does not cause excessive information loss. The verification process can set the first preset tolerance range (for the lightweight text encoder) and the second preset tolerance range (for the non-autoregressive acoustic feature prediction module), and check whether the output differences before and after pruning are within the acceptable range respectively.

[0092] If the output error of the pruned model exceeds the tolerance range, it is necessary to dynamically adjust the pruning threshold and iteratively perform pruning. The adjustment strategy can include increasing the number of retained attention heads, reducing the number of pruned channels, or adjusting the distribution of the pruning area to ensure that the pruned model can still meet the requirements of speech synthesis quality.

[0093] Example illustration: In the field of healthcare, TTS is required to perform real-time speech synthesis on medical records and patient interactions. For example, in a hospital self-service system, a patient may query "recent blood glucose level trends", and TTS needs to quickly generate a voice response such as "The last three blood glucose tests were 5.8, 6.2, and 6.0 mmol / L respectively". In such a scenario, the TTS system runs on low-power devices and needs to respond quickly and ensure voice quality. Structured pruning can effectively reduce the computational overhead, enabling the device to read medical data in real-time in a low-computing-power environment and improving the patient's interaction experience.

[0094] In the financial field, bank customer service and market data broadcast systems need to operate in a high-concurrency environment. For example, in stock market broadcasts, the system may need to read aloud in real time "The Nasdaq index has risen 1.2%, and the S&P 500 index has fallen 0.5%". If the model has not been pruned and optimized, the computational cost may limit the concurrency ability, resulting in delayed system responses. Through structured pruning, while ensuring speech quality, the concurrency performance can be improved, enabling voice customer service and market broadcasts to quickly respond to user requests and provide a smooth speech synthesis experience.

[0095] In accessibility services, blind users rely on TTS to read web pages, e-books, or text message notifications. For example, a user may request voice reading of news such as "Today's Headlines, the Global Climate Summit is being held in Switzerland". If the TTS system runs on a portable accessibility device (such as a Braille reader), the computing resources are limited, and an efficient speech generation model is required. Structured pruning can reduce the model size and improve the inference speed, ensuring that users can quickly obtain the reading information without speech delay or quality degradation due to insufficient computing resources.

[0096] Through structured pruning, the computational complexity of the model can be effectively reduced, enabling the lightweight text encoder and non-autoregressive acoustic feature prediction module to still operate efficiently in low-resource environments. Compared with the unpruned model, after pruning optimization, the computational amount can be reduced, the inference speed can be increased, and the storage requirements can be reduced, enabling the TTS system to be efficiently deployed in applications such as mobile devices, smart speakers, and cloud speech synthesis.

[0097] S70, perform parameter quantization processing on the lightweight text encoder and non-autoregressive acoustic feature prediction module after pruning processing;

[0098] In this embodiment, the core goal of parameter quantization processing is to reduce the model storage requirements and computational complexity while maintaining the speech synthesis quality. Since neural network models usually use 32-bit floating-point numbers (FP32) for calculation and storage, which occupy a large amount of storage space and have a high computational overhead, and the TTS task often requires efficient inference, parameter quantization methods are needed to convert the floating-point weights into a low-bit format to reduce the computational amount and optimize the storage efficiency.

[0099] Parameter quantization mainly involves the weight parameter processing of the lightweight text encoder and non-autoregressive acoustic feature prediction module. First, it is necessary to determine the quantization bit width of the floating-point weight parameters, that is, set the first quantization bit width for the weight parameters of the lightweight text encoder after pruning processing, and set the second quantization bit width for the weight parameters of the non-autoregressive acoustic feature prediction module after pruning processing. The choice of quantization bit width directly affects the storage size and computational accuracy of the model. Generally, 8-bit (INT8) or 16-bit (FP16) formats can be selected to balance the accuracy loss and computational efficiency.

[0100] Next, based on the determined quantization bit width, a linear quantization method is used to convert the floating-point weight parameters. The basic principle of linear quantization is to map floating-point numbers to the integer space and reconstruct them through a scale factor and a zero point to minimize the quantization error as much as possible. Specifically, for the floating-point weight parameters of the lightweight text encoder, linear quantization processing with the first quantization bit width is adopted to convert the FP32 weights into INT8 or FP16 format. Similarly, for the floating-point weight parameters of the non-autoregressive acoustic feature prediction module, linear quantization processing with the second quantization bit width is adopted to ensure the efficiency of the calculation during the acoustic feature prediction process.

[0101] After the initial quantization is completed, fine-tuning optimization is required to calibrate the quantized weight parameters. Since the quantization process will introduce certain errors, which may affect the speech synthesis quality, post-training quantization (PTQ) or quantization-aware training (QAT) can be used for optimization. Among them, PTQ is applicable to the already trained models, and the quantization parameters are adjusted through statistical analysis of the weight distribution, while QAT simulates the quantization process during training, enabling the model to gradually adapt to the quantized calculation method to reduce the impact of errors on performance.

[0102] Finally, it is necessary to verify the quantized model and check whether the output errors of the lightweight text encoder and the non-autoregressive acoustic feature prediction module meet the quantization tolerance range. If the quantization error exceeds the acceptable range, it is necessary to readjust the quantization bit width and iteratively execute the quantization and calibration operations. For example, the error can be reduced by increasing the quantization bit width (such as from INT8 to FP16) or optimizing the quantization strategy (such as mixed-precision quantization, that is, different precisions are used for some layers) to ensure that the quality of speech generation is not affected.

[0103] Through the parameter quantization process, the model storage requirements can be significantly reduced, the inference speed can be improved, and the power consumption can be reduced. Compared with the unquantized model, the quantized and optimized model can achieve efficient speech generation in low-resource environments and is applicable to scenarios such as mobile devices, embedded systems, and cloud computing.

[0104] S80, the lightweight text encoder after parameter quantization processes encodes the standard text sequence to generate a compressed text hidden vector;

[0105] In this embodiment, the goal of encoding the standard text sequence by the lightweight text encoder after parameter quantization is to maintain high-quality text representation ability while reducing the consumption of computing resources, ensuring that clear and natural speech can be generated in the subsequent speech synthesis stage. Since the lightweight text encoder has been pruned and parameter quantized, its computational complexity has been greatly reduced. Therefore, it is necessary to ensure that the quantized text encoder can still efficiently extract text semantic information and generate compressed text hidden vectors suitable for acoustic feature prediction.

[0106] The first step of the encoding process is text embedding, that is, converting the standard text sequence into a vector representation of a fixed dimension. Since the text encoder has been parameter quantized, the weights of the embedding layer also use quantized representations (such as INT8 or FP16) to reduce the computational burden. The text embedding can be transformed based on the character, phoneme, or sub-word level, and combined with the pre-trained weights of the language model to ensure that the key features of the text input information are not lost due to quantization.

[0107] The embedded vector is then input into the quantized multi-head self-attention layer. Compared with the unquantized attention layer, the calculation of this layer has been simplified, and usually low-bit matrix operations are used, such as INT8 multiplication or mixed-precision calculation (FP16 acceleration). Since quantization may cause the numerical range of the attention scores to become smaller, a scaling factor can be introduced for adjustment to ensure that the model can still effectively focus on the key dependencies in the text, such as long-range dependency information, pause patterns, and syntactic structures.

[0108] Subsequently, the text hidden vector undergoes a non-linear transformation through a quantized feedforward network (QFFN). The parameters of the feedforward network have been quantized, so the calculation method needs to be adapted to low-bit operations, such as using matrix multiplication optimization (Low-rank Decomposition) or grouped convolution to reduce the amount of calculation.

[0109] In the quantized encoder, positional encoding still plays a key role to ensure that the order information of the input text sequence is not lost. Since the quantization operation may affect the numerical precision, fixed positional encoding (such as sine / cosine encoding) can be used, or discrete positional encoding can be used to adapt to the low-bit computing environment.

[0110] Finally, the quantized text encoder outputs a compressed text hidden vector, which has the same dimension as the original text hidden vector. However, due to the optimized numerical operations during the encoding process, its computational complexity and storage occupancy are significantly reduced. This hidden vector will be passed to the non-autoregressive acoustic feature prediction module to further predict acoustic features and generate speech.

[0111] The text encoder optimized by quantization can significantly reduce storage occupancy, improve inference efficiency, and reduce power consumption. At the same time, it can still ensure the semantic expression ability of the text, enabling the TTS system to operate efficiently in environments such as mobile devices, embedded devices, and cloud computing. Compared with the unquantized encoder, the quantized text encoder can provide text processing capabilities close to those of high-precision models in low-compute-resource environments, while significantly improving the response speed of real-time speech synthesis.

[0112] S90, the non-autoregressive acoustic feature prediction module after parameter quantization maps the compressed text hidden vector to an optimized acoustic feature sequence;

[0113] In this embodiment, the main task of the non-autoregressive acoustic feature prediction module after parameter quantization is to map the compressed text hidden vector to an optimized acoustic feature sequence, that is, to ensure the accuracy of acoustic features and speech quality while reducing computational resource consumption. Since this module has undergone structured pruning and parameter quantization, its computational complexity has been reduced. Therefore, it is necessary to ensure that in a low-bit computing environment, it can still efficiently and accurately predict acoustic features that can be used for speech synthesis.

[0114] First, the compressed text hidden vector is used as input and undergoes feature extraction through the convolutional layer of the quantized parallel decoder. The weights of this convolutional layer have been quantized (such as INT8 or FP16), so low-bit matrix operations or block calculations will be used during the calculation process to reduce the computational burden. Since quantization may lead to a contraction of the numerical range, a scaling factor can be introduced to ensure computational stability and reduce computational complexity while retaining local features.

[0115] Subsequently, the extracted features enter the quantized global context attention layer to model the long-range dependencies of the input text. The self-attention calculation of traditional Transformers usually uses high-bit floating-point numbers, while in a quantized environment, low-rank attention or low-bit matrix multiplication (INT8 MatMul) can be used for calculation to reduce storage requirements. This attention layer can improve the inference efficiency of the model in a parallel computing mode while retaining the global semantic information of the input text to ensure the accuracy of subsequent acoustic feature predictions.

[0116] Next, through the quantized residual connection layer, the output of the context attention layer is fused with the initial convolutional features. Since quantization may lead to a reduction in the range of feature values, a dynamic scaling strategy can be used to adjust the calculation method of the residual connection, ensuring that the model can converge stably and reducing numerical errors without affecting the generation of acoustic features.

[0117] After that, the non-autoregressive acoustic feature prediction module uses a quantized feedforward network (QFFN) for non-linear transformation. The parameters of the QFFN are quantized and optimized using grouped fully connected or low-rank decomposition methods to reduce the computational load while ensuring the accuracy of feature transformation.

[0118] Finally, the optimized features are input into the quantized linear projection layer, which maps them to a Mel-spectrogram or other acoustic feature sequences for subsequent vocoder to synthesize speech. Since quantization may cause loss of numerical precision, this projection layer can incorporate an interpolation strategy or an error compensation mechanism to ensure the smoothness and stability of the output features.

[0119] The non-autoregressive acoustic feature prediction module after parameter quantization can reduce the computational resource requirements while maintaining the speech synthesis quality. Compared with the unquantized model, this optimized version can provide speech synthesis capabilities close to those of high-precision models in low-computational-resource environments and improve the response speed of real-time speech generation.

[0120] S100, convert the optimized acoustic feature sequence into a speech waveform through a vocoder.

[0121] In this embodiment, the core task of the vocoder is to convert the optimized acoustic feature sequence into a high-quality speech waveform. This process involves multiple steps such as denoising, enhancement, and time-domain reconstruction to ensure that the generated speech is natural, fluent, clear, and applicable to different application scenarios, such as intelligent voice assistants, financial customer service, medical voice broadcasts, etc.

[0122] First, optimize the preprocessing layer of the lightweight vocoder for the input of the acoustic feature sequence (such as Mel spectrogram). The preprocessing layer mainly performs normalization in the time dimension and context alignment processing to ensure that the input acoustic features have a smooth transition and eliminate noise interference. The normalization method can be Batch Normalization or Layer Normalization to ensure that the amplitude distributions of the acoustic features at different time steps are consistent, reducing the overfitting of the model to specific frequency components.

[0123] Subsequently, the preprocessed acoustic feature sequence passes through the depthwise separable convolution layer of the lightweight vocoder to extract multi-scale speech features in the frequency domain and time domain. Depthwise separable convolution can reduce the computational complexity while maintaining sufficient feature extraction ability. This convolution layer is used to parse speech features such as phoneme boundaries, pauses, and accents, and combines the dependencies of time steps to ensure that the finally generated speech waveform remains stable in terms of time continuity and prosodic rhythm.

[0124] After the speech feature extraction is completed, the model enters the low-rank convolution kernel compression stage. Since quantization and pruning may lead to the loss of feature information, the role of this stage is to recover some of the lost speech details while reducing the computational complexity. The low-rank convolution kernel uses methods such as Matrix Factorization or Low-rank Approximation to reduce the computational complexity and improve the expression ability of frequency features, making the timbre of the speech more natural and reducing the over-smoothing effect.

[0125] After that, the speech features are input into the residual connection layer to fuse multi-scale features, making the speech features after low-rank compression complementary to the original feature information. The existence of the residual connection layer can ensure the integrity of the speech waveform, improve the sound quality, and reduce the possibility of information loss. In addition, residual learning can also improve the training efficiency of the model and reduce the vanishing gradient problem, enabling the lightweight vocoder after quantization and pruning to still have a high speech restoration ability.

[0126] Finally, upsampling is performed through the transposed convolution layer to convert the speech features into the final time-domain waveform. In this process, the role of the transposed convolution layer is to map the low-dimensional features back to the high-dimensional audio waveform and optimize the speech quality in combination with the multi-scale discriminator. This discriminator can perform fine-grained discrimination on the quality of the synthesized speech based on the adversarial training method, thereby improving the authenticity and clarity of the finally generated speech.

[0127] With a lightweight vocoder, high-quality speech synthesis can be achieved in a low-computing-resource environment. Compared with traditional vocoders, it can reduce the computational complexity, improve the synthesis speed, and at the same time reduce the storage requirements, enabling the speech synthesis system to operate efficiently in various computing environments.

[0128] The present invention relates to the technical field of speech processing and can be applied to business scenarios such as medical health, fintech, and barrier-free service fields. It discloses a text-to-speech method based on knowledge distillation, including: performing standardization processing on the input text to generate a standard text sequence; encoding the standard text sequence through a lightweight text encoder to generate a text hidden vector; mapping the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module; encoding and performing acoustic feature prediction processing on the standard text sequence by a pre-trained teacher model to generate a teacher acoustic feature sequence; a knowledge distillation module determining an alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence; performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss; performing parameter quantization processing on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module; the lightweight text encoder after parameter quantization encoding the standard text sequence to generate a compressed text hidden vector; the non-autoregressive acoustic feature prediction module after parameter quantization mapping the compressed text hidden vector to an optimized acoustic feature sequence; and a vocoder converting the optimized acoustic feature sequence into a speech waveform. The present invention effectively reduces the model size and computational overhead while maintaining the speech generation quality through knowledge distillation, structured pruning, and parameter quantization. The inference speed is improved and the speech generation latency is reduced through non-autoregressive acoustic feature prediction. The optimized lightweight text encoder and lightweight vocoder improve cross-device adaptability, enabling the TTS system to achieve efficient, low-power, and real-time speech synthesis in resource-constrained environments, meeting the requirements for low-latency and high-quality speech output in fields such as barrier-free services, fintech, and medical health.

[0129] In one embodiment, the above step S20 includes:

[0130] S201, inputting the standard text sequence into the embedding layer of the lightweight text encoder to generate a character vector;

[0131] S202, extracting context features of the character vector through the multi-head self-attention layer of the lightweight text encoder;

[0132] S203, performing normalization processing on the context features through the layer normalization layer of the lightweight text encoder;

[0133] S204, adding the context features after normalization processing to the character vector through the residual connection layer of the lightweight text encoder to generate a fused feature;

[0134] S205, add position information to the fusion feature through the position encoding module of the lightweight text encoder;

[0135] S206, perform a non-linear transformation on the fusion feature through the feed-forward network of the lightweight text encoder to generate the text hidden vector.

[0136] In this embodiment, the core objective of the lightweight text encoder is to efficiently extract the semantic information of the text sequence while reducing the computational complexity, and generate a text hidden vector for use by the subsequent acoustic feature prediction module. The text encoder adopts a lightweight design to reduce the computational overhead and ensure that the encoded text representation has complete context information, syntactic structure, and positional relationship.

[0137] The first step of the text encoder is the Embedding Layer, that is, after the standard text sequence is input into the lightweight text encoder, it is converted into a character vector. The character vector is a numerical representation of the text and can be mapped based on different granularities (such as character level, phoneme level, sub-word level). Common methods include Embedding Lookup or CNN-based Embedding. The former is suitable for a fixed vocabulary, and the latter is suitable for open vocabulary scenarios. The weights of the embedding layer can be initialized through a pre-trained language model to enhance the text feature extraction ability.

[0138] Next, the text hidden vector enters the lightweight Multi-Head Self-Attention Layer to extract the context information of the text. The key role of the self-attention mechanism is to model long-range dependencies to ensure that characters or words that are far apart in the text can still influence each other. Due to the lightweight design, this self-attention layer reduces the number of attention heads, reduces the computational complexity, and at the same time uses the Low-Rank Factorization technology to reduce the parameter scale to ensure efficient calculation.

[0139] Then, through the Layer Normalization Layer, the text hidden vector is normalized. The main role of normalization is to stabilize the training process, improve the convergence speed, and reduce the internal covariate shift. The calculation method of layer normalization is to calculate the mean and standard deviation on each feature dimension and perform a scaling transformation to ensure the stability of the data distribution.

[0140] After normalization, the text hidden vector enters the Residual Connection Layer. The role of the residual connection layer is to retain the information of the original character vector, fuse it with the extracted context features, prevent information loss, and enhance the text representation ability. The residual connection method usually adopts Element-wise Addition, that is, adding the normalized context features and the original character vector according to the corresponding dimensions to generate the fused text features.

[0141] After that, the text features are input into the Positional Encoding Module to supplement the position information in the text sequence, ensuring that the text encoder can distinguish characters or words in different orders. Positional encoding can adopt Sinusoidal Positional Encoding or Learnable Positional Embedding. The former is suitable for sequences of fixed length, and the latter is suitable for variable-length inputs.

[0142] Finally, the fused features undergo a non-linear transformation through a Feedforward Network (FFN) to generate the final text hidden vector. The feedforward network usually includes two fully connected layers and non-linear activation functions such as ReLU or GELU to enhance the model's expressive ability and improve the abstraction ability of text features.

[0143] In this embodiment, text encoding is performed through a lightweight text encoder, which can reduce the computational complexity, improve the inference efficiency, and maintain the speech synthesis quality. Compared with the traditional Transformer encoder, this optimization reduces the computational resource occupancy and improves the inference speed of the model, enabling the TTS system to operate efficiently in environments such as mobile devices, embedded devices, and cloud computing. In addition, by combining positional encoding and residual connection, it can ensure that text features still maintain stable encoding capabilities in long text processing, low-computation environments, and multilingual tasks.

[0144] In one embodiment, the above step S30 includes:

[0145] S301, inputting the text hidden vector into the convolutional layer of the parallel decoder of the non-autoregressive acoustic feature prediction module to generate initial acoustic features;

[0146] S302, performing temporal dimension modeling on the initial acoustic features through the global context attention layer of the non-autoregressive acoustic feature prediction module to generate context acoustic features;

[0147] S303. Add the context acoustic features and the initial acoustic features through the residual connection layer of the non-autoregressive acoustic feature prediction module to generate fused acoustic features;

[0148] S304. Perform a non-linear transformation on the fused acoustic features through the feed-forward network of the non-autoregressive acoustic feature prediction module to generate transformed acoustic features;

[0149] S305. Map the transformed acoustic features to a Mel spectrogram sequence through the linear projection layer of the non-autoregressive acoustic feature prediction module as the student acoustic feature sequence.

[0150] In this embodiment, the core task of the non-autoregressive acoustic feature prediction module is to map text hidden vectors to an acoustic feature sequence in a parallel computing manner for the subsequent vocoder to generate the final speech waveform. Compared with the autoregressive method, the non-autoregressive method can significantly improve the inference speed and reduce the generation latency, and is suitable for low-compute resource environments and real-time speech synthesis scenarios.

[0151] First, the text hidden vector is input into the convolutional layer of the parallel decoder of the non-autoregressive acoustic feature prediction module to generate initial acoustic features. The role of the convolutional layer is to extract short-term acoustic patterns and capture spectral features at the phoneme or sub-phoneme level. To improve the computational efficiency and reduce the number of parameters, depthwise separable convolution can be used to reduce the computational complexity while maintaining good feature extraction capabilities.

[0152] Next, the initial acoustic features are input into the global context attention layer for temporal dimension modeling to generate context acoustic features. Since speech has temporal dependence, that is, subsequent pronunciations are affected by previous phonemes or words, the model needs to capture global information rather than just local patterns. The global context attention layer calculates the attention distribution at all time steps in a non-autoregressive manner and combines low-rank factorization or sparse attention mechanism to reduce the computational complexity while maintaining the model performance.

[0153] Then, the context acoustic features enter the residual connection layer and are added to the initial acoustic features to generate fused acoustic features. The role of the residual connection is to retain the original feature information, avoid over-smoothing of features, and enhance gradient propagation to ensure the stability of model training. The residual connection layer uses element-wise addition for feature fusion and combines layer normalization to improve the training convergence speed and prediction stability of the model.

[0154] Subsequently, the fused acoustic features are input into a Feedforward Network (FFN) for non-linear transformation to further enhance the model's feature extraction ability. The FFN usually consists of two fully connected layers and uses activation functions such as ReLU and GELU to enhance the expressive ability of the features. To improve computational efficiency, the FFN can adopt Grouped Fully Connected or Low-rank Matrix Decomposition to reduce the number of parameters and accelerate the calculation.

[0155] Finally, the transformed acoustic features are input into a Linear Projection Layer, which maps them to a Mel-spectrogram sequence as the final student acoustic feature sequence. The main function of the Linear Projection Layer is dimensionality reduction, converting high-dimensional acoustic features into low-dimensional Mel-spectrogram parameters and combining an Interpolation Strategy to ensure smooth spectral transitions and avoid a decline in sound quality caused by quantization errors.

[0156] In this embodiment, the non-autoregressive acoustic feature prediction module can improve the speech synthesis speed, reduce the inference latency, and lower the computational resource consumption. Compared with the autoregressive method, this optimization avoids the computational overhead of step-by-step decoding, enables parallel computing, improves the inference efficiency, and allows the TTS system to operate efficiently in real-time speech generation, low-computation environments, and multi-lingual tasks. Additionally, combining residual connections and linear projections can ensure high-quality generation of acoustic features in long text processing, low-computation environments, and multi-lingual tasks.

[0157] In one embodiment, the above step S50 includes:

[0158] S501, analyzing the probability distribution difference between the student acoustic feature sequence and the teacher acoustic feature sequence through a relative entropy loss function to generate an initial alignment loss value;

[0159] S502, performing weighted processing on the initial alignment loss value through a weight adjustment module to generate a final alignment loss;

[0160] S503, verifying whether the alignment loss is higher than a preset threshold;

[0161] S504, if the alignment loss is higher than the preset threshold, adjusting the weight parameters of the lightweight text encoder and the non-autoregressive acoustic feature prediction module through a gradient descent strategy until the alignment loss is not higher than the preset threshold or the number of iterations in the weight parameter adjustment process reaches a preset maximum number of iterations, and then terminating the adjustment process.

[0162] In this embodiment, the core task of the knowledge distillation module is to guide the student model to learn through the teacher model, so as to reduce the number of model parameters while maintaining the speech synthesis quality. This module calculates the difference between the student acoustic feature sequence and the teacher acoustic feature sequence, and narrows the gap between the two through an optimization strategy to ensure that the student model can still generate high-quality speech features under the condition of light weight.

[0163] First, the relative entropy loss function (Kullback-Leibler Divergence, KLDivergence) is used to analyze the probability distribution difference between the student acoustic feature sequence and the teacher acoustic feature sequence, and an initial alignment loss value is generated. KL divergence is an index to measure the difference between two probability distributions, which is used to compare the deviation between the prediction results of the student model and the output of the teacher model. Since the acoustic features in the TTS task usually have complex temporal correlation and context dependence, dynamic time warping (DTW) can be used when calculating KL divergence to ensure that feature sequences of different lengths can be reasonably aligned and improve the matching accuracy.

[0164] Secondly, the calculated initial alignment loss value enters the weight adjustment module for weighted processing to generate the final alignment loss. The role of the weight adjustment module is to adjust the contribution of the alignment loss according to the importance of different types of acoustic features, so as to avoid some features having too much influence on the overall optimization process. For example, the acoustic features in the high-frequency band have a greater impact on clarity, while the low-frequency band features have a greater impact on the naturalness of speech. Therefore, adaptive weighting can be used to dynamically adjust the loss contribution of different frequency bands.

[0165] Then, it is verified whether the final alignment loss is higher than the preset threshold to judge whether the performance of the student model has approached that of the teacher model. The preset threshold is usually determined based on a large amount of experimental data, and dynamic threshold adjustment can be used, that is, the threshold is adjusted according to the downward trend of the loss during the training process to avoid overfitting or non-convergence of the training.

[0166] If the alignment loss is higher than the preset threshold, the weight parameters of the lightweight text encoder and the non-autoregressive acoustic feature prediction module are adjusted through the Gradient Descent Strategy. The optimization method can adopt an adaptive learning rate optimizer (such as Adam, RMSprop), combined with Gradient Clipping, to prevent the problems of gradient explosion or gradient vanishing and improve the stability of optimization.

[0167] During the optimization process, the new alignment loss is calculated and checked in each iteration. If the alignment loss is lower than the preset threshold, or the number of training iterations reaches the maximum allowed number, the optimization process is terminated to ensure the training efficiency and prevent overfitting at the same time.

[0168] In this embodiment, through the knowledge distillation module, while reducing the number of model parameters, the quality of speech synthesis can be improved. It can ensure that the synthesized speech of the student model is close to the original large model in terms of timbre, prosody, clarity, etc., and at the same time significantly reduce the computational overhead, enabling the TTS system to run in a low-computing resource environment (such as mobile devices, embedded devices).

[0169] In one embodiment, the above step S60 includes:

[0170] S601, determining a first pruning threshold for each attention head in the multi-head self-attention layer of the lightweight text encoder and a second pruning threshold for each feature channel in the convolutional layer of the non-autoregressive acoustic feature prediction module based on the alignment loss;

[0171] S602, removing redundant attention heads in the lightweight text encoder with weight values lower than the first pruning threshold;

[0172] S603, removing redundant feature channels in the non-autoregressive acoustic feature prediction module with the number of output channels lower than the second pruning threshold;

[0173] S604, verifying whether the output error of the lightweight text encoder after pruning processing meets the first preset tolerance range, and verifying whether the output error of the non-autoregressive acoustic feature prediction module after pruning processing meets the second preset tolerance range;

[0174] S605, if the output error of the lightweight text encoder exceeds the first preset tolerance range, adjusting the first pruning threshold and iteratively performing the pruning processing of the lightweight text encoder;

[0175] S606, if the output error of the non-autoregressive acoustic feature prediction module exceeds the second preset tolerance range, adjusting the second pruning threshold and iteratively performing the pruning processing of the non-autoregressive acoustic feature prediction module.

[0176] In this embodiment, the main objective of the structured pruning process is to reduce the computational load of the lightweight text encoder and the non-autoregressive acoustic feature prediction module while maintaining the model performance. By pruning, unnecessary computational units such as redundant attention heads and ineffective feature channels can be removed, thereby reducing the storage and computational costs, enabling the TTS system to operate in resource-constrained environments such as low-power devices, embedded systems, and mobile devices.

[0177] First, it is necessary to determine the specific pruning method based on the alignment loss. The alignment loss measures the difference between the student acoustic feature sequence and the teacher acoustic feature sequence, and it can be used to evaluate the impact of different network components (such as attention heads and feature channels) on the final speech synthesis quality.

[0178] In the multi-head self-attention layer of the lightweight text encoder, calculate the weight contribution of all attention heads and set the first pruning threshold. If the contribution of a certain attention head is lower than the threshold, it indicates that this attention head has a relatively small impact on the overall effect of the text encoder and can be removed to reduce the computational cost.

[0179] In the convolutional layer of the non-autoregressive acoustic feature prediction module, calculate the contribution of all feature channels and set the second pruning threshold. If the number of output channels of a certain feature channel is lower than the threshold, it means that this channel contributes less during the feature transformation process and can be removed to reduce the computational amount.

[0180] Next, the pruning operation starts to execute:

[0181] Remove the redundant attention heads in the lightweight text encoder whose weight values are lower than the first pruning threshold. In the standard Transformer structure, each multi-head self-attention layer contains multiple independent attention heads, and each attention head is responsible for different text features. During pruning, it is possible to determine which attention heads contribute less to the final output based on weight entropy analysis or saliency score and remove them to reduce the computational overhead.

[0182] Remove the redundant feature channels in the non-autoregressive acoustic feature prediction module whose number of output channels is lower than the second pruning threshold. When pruning the convolutional layer, the channel pruning method can be adopted to remove the channels with smaller weights or lower contribution degrees to reduce the computational amount.

[0183] After pruning is completed, it is necessary to perform model verification to ensure that the model can still maintain good performance after pruning.

[0184] Verify whether the output error of the lightweight text encoder meets the first preset tolerance range. If the pruning causes the loss of semantic information of the text hidden vector to exceed the acceptable range, the pruning threshold needs to be adjusted.

[0185] Verify whether the output error of the non-autoregressive acoustic feature prediction module meets the second preset tolerance range. If the quality of the acoustic features drops too much after pruning, the pruning parameters need to be adjusted to ensure that the model can still generate high-quality acoustic features.

[0186] If the output error exceeds the tolerance range, enter the iterative optimization process:

[0187] If the output error of the lightweight text encoder exceeds the first preset tolerance range, the first pruning threshold needs to be readjusted and a new pruning process is performed, that is, reducing the pruning intensity and retaining more attention heads to ensure the stability of text encoding; if the output error of the non-autoregressive acoustic feature prediction module exceeds the second preset tolerance range, the second pruning threshold needs to be adjusted and pruning is performed again, that is, retaining more feature channels to ensure the integrity of the acoustic features.

[0188] Through structured pruning in this embodiment, the computational amount can be reduced, the inference speed can be improved, and the speech synthesis quality can be maintained at the same time. It can reduce the consumption of computing resources and maintain the synthesis quality close to that of the unpruned model, enabling the TTS system to operate efficiently in environments such as mobile devices, cloud platforms, and low-power devices.

[0189] In one embodiment, the above step S70 includes:

[0190] S701, determine the first quantization bit width of the floating-point weight parameters of the lightweight text encoder after pruning, and the second quantization bit width of the floating-point weight parameters of the non-autoregressive acoustic feature prediction module after pruning;

[0191] S702, perform linear quantization processing on the floating-point weight parameters of the lightweight text encoder based on the first quantization bit width to generate quantization weight parameters in the first low-bit integer format;

[0192] S703, perform linear quantization processing on the floating-point weight parameters of the non-autoregressive acoustic feature prediction module based on the second quantization bit width to generate quantization weight parameters in the second low-bit integer format;

[0193] S704, calibrate the quantization weight parameters in the first low-bit integer format and the quantization weight parameters in the second low-bit integer format through a fine-tuning optimization module.

[0194] In this embodiment, the core objective of parameter quantization processing is to maintain the speech synthesis quality while reducing the model storage and computational complexity. Since pruning has already reduced the redundant structures of the lightweight text encoder and the non-autoregressive acoustic feature prediction module, further parameter quantization can reduce the computational precision requirements and storage space occupancy, enabling the model to operate efficiently in low-computing-power environments (such as mobile devices and embedded devices).

[0195] First, determine the quantization bit widths of the floating-point weight parameters of the lightweight text encoder and the non-autoregressive acoustic feature prediction module after pruning. The quantization bit width determines the storage precision of the model parameters. Usually, INT8, INT4, or even INT2 quantization is used to replace the original FP32 floating-point weights, thereby reducing the consumption of computing resources.

[0196] The first quantization bit width: for the floating-point weight parameters of the lightweight text encoder. Usually, for modules that need to accurately capture text semantics, such as the text encoder, an appropriate quantization bit width can ensure that the language understanding ability of the model does not decline.

[0197] The second quantization bit width: for the floating-point weight parameters of the non-autoregressive acoustic feature prediction module. Since this module is mainly responsible for modeling in the time dimension and has a great impact on speech synthesis quality, it is necessary to ensure that the quantization bit width can still maintain sufficient information expression ability while reducing the computational requirements.

[0198] Next, based on the first quantization bit width, perform linear quantization processing on the floating-point weight parameters of the lightweight text encoder to generate quantization weight parameters in the first low-bit integer format. The core idea of linear quantization is to map floating-point numbers to integer representations within a fixed range to reduce storage requirements and computational complexity. Common linear quantization methods include:

[0199] Symmetric Quantization: Map the floating-point weights evenly within a fixed range to integers. For example, [-1, 1] is mapped to [-128, 127] (INT8 quantization).

[0200] Asymmetric Quantization: For data with uneven weight distributions, use a zero point and a scale factor to optimize the quantization precision and reduce quantization errors.

[0201] Similarly, based on the second quantization bitwidth, linear quantization is performed on the floating-point weight parameters of the non-autoregressive acoustic feature prediction module to generate quantization weight parameters in the second low-bit integer format. Since acoustic features involve a large amount of time series modeling, using Perceptual Quantization or Dynamic Quantization can effectively reduce information loss. For example, in vocoders such as WaveRNN or HiFi-GAN, Per-Channel Quantization is often used to ensure the accuracy of the audio spectrum.

[0202] To avoid the cumulative effect of quantization errors on the final speech synthesis quality, it is necessary to calibrate the quantization weight parameters through a Fine-tuning Optimization Module. The ways of fine-tuning optimization include:

[0203] Distillation Fine-tuning: Guide the quantized student model through the teacher model so that the quantized model can still maintain high-quality speech synthesis capabilities.

[0204] Perceptual Loss Optimization: By comparing the speech features before and after quantization, minimize the differences in speech perception and optimize the performance of the quantized model.

[0205] After that, verify whether the output error of the quantized lightweight text encoder and the non-autoregressive acoustic feature prediction module meets the quantization tolerance range corresponding to the first quantization bitwidth and the second quantization bitwidth. The quantization tolerance range determines whether the output error of the model is within an acceptable range. If the quantized model significantly degrades in terms of speech clarity, prosody accuracy, etc., it is necessary to adjust the quantization parameters.

[0206] If the output error exceeds the quantization tolerance range, readjust the first quantization bitwidth and the second quantization bitwidth, and iteratively perform quantization and calibration operations. This process usually adopts Progressive Quantization, that is, first perform high-bit quantization (such as INT8), then gradually reduce the number of bits (such as INT4), and perform fine-tuning at each stage to ensure the robustness of the model.

[0207] Through parameter quantization processing, this embodiment can significantly reduce the storage requirements and computational complexity of the TTS model, improve the inference efficiency, reduce the storage requirements of model parameters, and maintain the speech synthesis quality, enabling the TTS system to operate efficiently in environments such as mobile devices, embedded devices, and cloud computing. In addition, combined with fine-tuning optimization, it can ensure that the quantized model still has high-quality speech synthesis capabilities in different application scenarios.

[0208] In one embodiment, the above step S100 includes:

[0209] S1001, input the optimized acoustic feature sequence into the preprocessing layer of the lightweight vocoder, and perform normalization processing and context alignment processing on the optimized acoustic feature sequence in the time dimension;

[0210] S1002, extract frequency-domain and time-domain features from the optimized acoustic features after normalization processing and context alignment processing through the depthwise separable convolutional layer of the lightweight vocoder to generate multi-scale speech features;

[0211] S1003, compress the parameter scale of the multi-scale speech features through the low-rank convolutional kernel of the lightweight vocoder to generate compressed speech features;

[0212] S1004, fuse the multi-scale speech features and the compressed speech features through the residual connection layer of the lightweight vocoder to generate fused speech features;

[0213] S1005, perform upsampling and time-domain waveform reconstruction on the fused speech features through the transposed convolutional layer of the lightweight vocoder to generate the speech waveform.

[0214] In this embodiment, the core objective of the speech waveform generation process is to convert the optimized acoustic feature sequence into a high-quality speech signal to achieve clear, natural, and low-latency speech synthesis. Since neural network vocoders (such as HiFi-GAN, WaveRNN) have high computational complexity, a lightweight vocoder is adopted, combined with parameter optimization, feature extraction and compression, residual learning, and transposed convolutional techniques, to reduce the computational overhead while maintaining the speech quality.

[0215] First, the optimized acoustic feature sequence is input into the preprocessing layer of the lightweight vocoder for normalization processing and context alignment processing in the time dimension.

[0216] Normalization in the time dimension: Optimizing the numerical range of the acoustic feature sequence may result in large fluctuations. If directly input into the vocoder, it may lead to unstable learning. Therefore, the batch normalization or layer normalization method is adopted to normalize the data to a specific range, making it have better numerical stability.

[0217] Context alignment processing: Since the non-autoregressive generation method may cause alignment offsets in time steps, it is necessary to adjust the time step matching based on the dynamic time warping (DTW) method to make the input features more stable in the time sequence dimension and ensure smooth transitions in the time domain.

[0218] Next, through the depthwise separable convolution layer, frequency domain and time domain features are extracted to generate multi-scale speech features.

[0219] Depthwise separable convolution is an efficient convolution method that can maintain high feature extraction ability while reducing the computational load, and is particularly suitable for TTS vocoders.

[0220] Multi-scale convolution is adopted to extract features in different frequency ranges in parallel, enhancing the detailed control ability of speech synthesis. For example, convolution kernels with different window sizes are used to process short-term and long-term information, enabling the model to more accurately synthesize details such as voiceless sounds, voiced sounds, and pauses in speech.

[0221] Then, the parameter scale of the multi-scale speech features is compressed through low-rank convolution kernels to generate compressed speech features.

[0222] Low-rank convolution kernels use matrix factorization techniques (such as SVD decomposition) to reduce the number of parameters while maintaining efficient information expression ability.

[0223] Low-rank convolution can further reduce the computational complexity of the vocoder, enabling the model to run on mobile devices or low-power devices while ensuring that the speech synthesis quality is not significantly affected.

[0224] After that, the multi-scale speech features and the compressed speech features are fused through the residual connection layer to generate fused speech features.

[0225] Residual connection allows the original feature information to be retained during the optimization process while enhancing gradient propagation, making the training more stable.

[0226] The Adaptive Weight Fusion mechanism is adopted, that is, according to the contribution degrees of different types of speech features, the weights of multi-scale speech features and compressed speech features are dynamically adjusted to improve the adaptability of the model to different voice styles.

[0227] Finally, the fused speech features are input into the transposed convolution layer for upsampling and time-domain waveform reconstruction, and finally the speech waveform is generated.

[0228] Transposed Convolution is a commonly used technique in neural network vocoders, which can map low-resolution features to high-resolution speech waveforms and restore complete time-domain information.

[0229] Multi-stage Upsampling is adopted, that is, the resolution is gradually increased to make the speech waveform smoother and avoid the problem of deteriorated sound quality caused by excessive quantization.

[0230] In this embodiment, the lightweight vocoder can reduce the computational amount, improve the inference speed, and at the same time maintain high-quality speech synthesis. It can reduce the consumption of computing resources and maintain high-naturalness speech output, enabling the TTS system to operate efficiently in environments such as mobile devices, the cloud, and low-power devices. In addition, combined with multi-scale feature extraction and low-rank convolution kernel optimization, it can ensure that the model still maintains high-quality speech output in a low-computing-resource environment.

[0231] In one embodiment, a text-to-speech device based on knowledge distillation is provided, and the text-to-speech device based on knowledge distillation corresponds one-to-one to the text-to-speech method based on knowledge distillation in the above embodiment. Refer to Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the text-to-speech device based on knowledge distillation of the present invention. The text preprocessing module 10, the lightweight text encoding module 20, the non-autoregressive acoustic feature prediction module 30, the teacher model module 40, the knowledge distillation module 50, the structured pruning module 60, the parameter quantization module 70, the quantized lightweight text encoding module 80, the quantized non-autoregressive acoustic feature prediction module 90, and the vocoder module 100. The detailed description of each functional module is as follows:

[0232] The text preprocessing module 10 is used to perform standardization processing on the input text to generate a standard text sequence;

[0233] The lightweight text encoding module 20 is used to encode the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0234] The non-autoregressive acoustic feature prediction module 30 is used to map the text hidden vector into a student acoustic feature sequence through the non-autoregressive acoustic feature prediction module;

[0235] The teacher model module 40 is used to encode the standard text sequence and perform acoustic feature prediction processing through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0236] The knowledge distillation module 50 is used to determine the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through the knowledge distillation module;

[0237] The structured pruning module 60 is used to perform structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0238] The parameter quantization module 70 is used to perform parameter quantization processing on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module;

[0239] The quantized lightweight text encoding module 80 is used to encode the standard text sequence through the quantized lightweight text encoder to generate a compressed text hidden vector;

[0240] The quantized non-autoregressive acoustic feature prediction module 90 is used to map the compressed text hidden vector into an optimized acoustic feature sequence through the quantized non-autoregressive acoustic feature prediction module;

[0241] The vocoder module 100 is used to convert the optimized acoustic feature sequence into a speech waveform through the vocoder.

[0242] In one embodiment, the lightweight text encoding module 20 is specifically used for:

[0243] Input the standard text sequence into the embedding layer of the lightweight text encoder to generate character vectors;

[0244] Extract the context features of the character vectors through the multi-head self-attention layer of the lightweight text encoder;

[0245] Normalize the context features through the layer normalization layer of the lightweight text encoder;

[0246] Add the normalized context features to the character vectors through the residual connection layer of the lightweight text encoder to generate fused features;

[0247] Add position information to the fused features through the position encoding module of the lightweight text encoder;

[0248] Perform a non - linear transformation on the fused feature through the feed - forward network of the lightweight text encoder to generate the text hidden vector.

[0249] In one embodiment, the non - autoregressive acoustic feature prediction module 30 is specifically configured to:

[0250] Input the text hidden vector into the convolutional layer of the parallel decoder of the non - autoregressive acoustic feature prediction module to generate an initial acoustic feature;

[0251] Perform temporal - dimension modeling on the initial acoustic feature through the global context attention layer of the non - autoregressive acoustic feature prediction module to generate a context acoustic feature;

[0252] Add the context acoustic feature and the initial acoustic feature through the residual connection layer of the non - autoregressive acoustic feature prediction module to generate a fused acoustic feature;

[0253] Perform a non - linear transformation on the fused acoustic feature through the feed - forward network of the non - autoregressive acoustic feature prediction module to generate a transformed acoustic feature;

[0254] Map the transformed acoustic feature to a Mel - spectrum sequence through the linear projection layer of the non - autoregressive acoustic feature prediction module as the student acoustic feature sequence.

[0255] In one embodiment, the knowledge distillation module 50 is specifically configured to:

[0256] Analyze the probability distribution difference between the student acoustic feature sequence and the teacher acoustic feature sequence through the relative entropy loss function to generate an initial alignment loss value;

[0257] Perform weighted processing on the initial alignment loss value through the weight adjustment module to generate a final alignment loss;

[0258] Verify whether the alignment loss is higher than a preset threshold;

[0259] If the alignment loss is higher than the preset threshold, adjust the weight parameters of the lightweight text encoder and the non - autoregressive acoustic feature prediction module through the gradient - descent strategy until the alignment loss is not higher than the preset threshold or the number of iterations in the weight - parameter adjustment process reaches the preset maximum number of iterations, then terminate the adjustment process.

[0260] In one embodiment, the structured pruning module 60 is specifically configured to:

[0261] Determine the first pruning threshold for each attention head in the multi - head self - attention layer of the lightweight text encoder and the second pruning threshold for each feature channel in the convolutional layer of the non - autoregressive acoustic feature prediction module based on the alignment loss;

[0262] Remove redundant attention heads in the lightweight text encoder with weight values lower than the first pruning threshold;

[0263] Remove redundant feature channels in the non-autoregressive acoustic feature prediction module with the number of output channels lower than the second pruning threshold;

[0264] Verify whether the output error of the lightweight text encoder after pruning processing meets the first preset tolerance range, and verify whether the output error of the non-autoregressive acoustic feature prediction module after pruning processing meets the second preset tolerance range;

[0265] If the output error of the lightweight text encoder exceeds the first preset tolerance range, adjust the first pruning threshold and iteratively execute the pruning process of the lightweight text encoder;

[0266] If the output error of the non-autoregressive acoustic feature prediction module exceeds the second preset tolerance range, adjust the second pruning threshold and iteratively execute the pruning process of the non-autoregressive acoustic feature prediction module.

[0267] In one embodiment, the parameter quantization module 70 is specifically configured to:

[0268] Determine the first quantization bit width of the floating-point weight parameters of the lightweight text encoder after pruning processing, and the second quantization bit width of the floating-point weight parameters of the non-autoregressive acoustic feature prediction module after pruning processing;

[0269] Perform linear quantization processing on the floating-point weight parameters of the lightweight text encoder based on the first quantization bit width to generate quantization weight parameters in the first low-bit integer format;

[0270] Perform linear quantization processing on the floating-point weight parameters of the non-autoregressive acoustic feature prediction module based on the second quantization bit width to generate quantization weight parameters in the second low-bit integer format;

[0271] Calibrate the quantization weight parameters in the first low-bit integer format and the quantization weight parameters in the second low-bit integer format through the fine-tuning optimization module.

[0272] In one embodiment, the vocoder module 100 is specifically configured to:

[0273] Input the optimized acoustic feature sequence into the preprocessing layer of the lightweight vocoder, and perform normalization processing and context alignment processing on the optimized acoustic feature sequence in the time dimension;

[0274] Performing frequency-domain and time-domain feature extraction on the optimized acoustic features after normalization processing and context alignment processing through the depthwise separable convolutional layer of the lightweight vocoder to generate multi-scale speech features;

[0275] Compressing the parameter scale of the multi-scale speech features through the low-rank convolutional kernel of the lightweight vocoder to generate compressed speech features;

[0276] Fusing the multi-scale speech features and the compressed speech features through the residual connection layer of the lightweight vocoder to generate fused speech features;

[0277] Performing upsampling and time-domain waveform reconstruction on the fused speech features through the transposed convolutional layer of the lightweight vocoder to generate the speech waveform.

[0278] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 4 The figure shows. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a text-to-speech method based on knowledge distillation.

[0279] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as shown in Figure 5 The figure shows. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a text-to-speech method based on knowledge distillation

[0280] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0281] Perform standardization processing on the input text to generate a standard text sequence;

[0282] Encode the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0283] Map the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module;

[0284] Encode and perform acoustic feature prediction processing on the standard text sequence through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0285] Determine the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module;

[0286] Perform structured pruning on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0287] Perform parameter quantization on the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module;

[0288] Encode the standard text sequence through the parameter-quantized lightweight text encoder to generate a compressed text hidden vector;

[0289] Map the compressed text hidden vector to an optimized acoustic feature sequence through the parameter-quantized non-autoregressive acoustic feature prediction module;

[0290] Convert the optimized acoustic feature sequence into a speech waveform through a vocoder.

[0291] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0292] Perform standardization processing on the input text to generate a standard text sequence;

[0293] Encode the standard text sequence through a lightweight text encoder to generate a text hidden vector;

[0294] Map the text hidden vector to a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module;

[0295] Encode and perform acoustic feature prediction processing on the standard text sequence through a pre-trained teacher model to generate a teacher acoustic feature sequence;

[0296] Determine the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module;

[0297] Perform structured pruning on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss;

[0298] Perform parameter quantization on the lightweight text encoder and the non-autoregressive acoustic feature prediction module after pruning;

[0299] Encode the standard text sequence through the lightweight text encoder after parameter quantization to generate a compressed text hidden vector;

[0300] Map the compressed text hidden vector to an optimized acoustic feature sequence through the non-autoregressive acoustic feature prediction module after parameter quantization;

[0301] Convert the optimized acoustic feature sequence into a speech waveform through a vocoder.

[0302] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0303] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0304] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0305] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A text-to-speech method based on knowledge distillation, characterized in that: The following steps are involved: Standardize the input text to generate a standard text sequence; Encoding the standard text sequence through a lightweight text encoder to generate a text latent vector; Mapping the text latent vector into a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module; The standard text sequence is encoded and acoustic feature predicted by a pre-trained teacher model to generate a teacher acoustic feature sequence; Determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module; Performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss; Quantize the parameters of the pruned lightweight text encoder and non-autoregressive acoustic feature prediction module; Encoding the standard text sequence through a lightweight text encoder after parameter quantization processing to generate a compressed text latent vector; Mapping the compressed text latent vector into an optimized acoustic feature sequence through a non-autoregressive acoustic feature prediction module after parameter quantization processing; The optimized acoustic feature sequence is converted into a speech waveform by a vocoder.

2. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: The standard text sequence is encoded by a lightweight text encoder to generate a text latent vector, including: Input the standard text sequence into the embedding layer of the lightweight text encoder to generate a character vector; Extracting contextual features of the character vector through a multi-head self-attention layer of the lightweight text encoder; Normalizing the context features through a layer normalization layer of the lightweight text encoder; Adding the normalized context feature to the character vector through the residual connection layer of the lightweight text encoder to generate a fused feature; Adding position information to the fused feature through a position encoding module of the lightweight text encoder; The fused features are nonlinearly transformed through the feedforward network of the lightweight text encoder to generate the text latent vector.

3. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: The text latent vector is mapped into a student acoustic feature sequence through a non-autoregressive acoustic feature prediction module, including: Inputting the text latent vector into the convolutional layer of the parallel decoder of the non-autoregressive acoustic feature prediction module to generate initial acoustic features; Modeling the initial acoustic features in time dimension through the global context attention layer of the non-autoregressive acoustic feature prediction module to generate contextual acoustic features; Adding the contextual acoustic feature to the initial acoustic feature through the residual connection layer of the non-autoregressive acoustic feature prediction module to generate a fused acoustic feature; Performing a nonlinear transformation on the fused acoustic features through a feedforward network of the non-autoregressive acoustic feature prediction module to generate transformed acoustic features; The transformed acoustic features are mapped into a Mel-spectrogram sequence through a linear projection layer of the non-autoregressive acoustic feature prediction module as the student acoustic feature sequence.

4. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: Determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through a knowledge distillation module includes: Analyzing the probability distribution difference between the student acoustic feature sequence and the teacher acoustic feature sequence through a relative entropy loss function to generate an initial alignment loss value; The initial alignment loss value is weighted by a weight adjustment module to generate a final alignment loss; Verifying whether the alignment loss is above a preset threshold; If the alignment loss is higher than the preset threshold, the weight parameters of the lightweight text encoder and the non-autoregressive acoustic feature prediction module are adjusted by the gradient descent strategy until the alignment loss is no higher than the preset threshold or the number of iterations of the weight parameter adjustment process reaches the preset maximum number of iterations, and the adjustment process is terminated.

5. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: The lightweight text encoder and the non-autoregressive acoustic feature prediction module are subjected to structured pruning processing according to the alignment loss, including: Determine a first pruning threshold of each attention head in the multi-head self-attention layer of the lightweight text encoder and a second pruning threshold of each feature channel in the convolution layer of the non-autoregressive acoustic feature prediction module based on the alignment loss; Removing redundant attention heads whose weight values ​​are lower than the first pruning threshold in the lightweight text encoder; Removing redundant feature channels in the non-autoregressive acoustic feature prediction module whose output channel number is lower than the second pruning threshold; Verify whether the output error of the lightweight text encoder after pruning satisfies a first preset tolerance range, and verify whether the output error of the non-autoregressive acoustic feature prediction module after pruning satisfies a second preset tolerance range; If the output error of the lightweight text encoder exceeds the first preset tolerance range, adjusting the first pruning threshold and iteratively performing pruning processing of the lightweight text encoder; If the output error of the non-autoregressive acoustic feature prediction module exceeds the second preset tolerance range, the second pruning threshold is adjusted and the pruning process of the non-autoregressive acoustic feature prediction module is iteratively performed.

6. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: The pruned lightweight text encoder and non-autoregressive acoustic feature prediction module are parameterized, including: Determine a first quantization bit width of a floating-point weight parameter of a lightweight text encoder after pruning, and a second quantization bit width of a floating-point weight parameter of a non-autoregressive acoustic feature prediction module after pruning; Performing linear quantization processing on the floating-point weight parameter of the lightweight text encoder based on the first quantization bit width to generate a quantization weight parameter in a first low-bit integer format; Performing linear quantization processing on the floating-point weight parameters of the non-autoregressive acoustic feature prediction module based on the second quantization bit width to generate quantization weight parameters in a second low-bit integer format; The quantization weight parameters in the first low-bit integer format and the quantization weight parameters in the second low-bit integer format are calibrated by a fine-tuning optimization module.

7. The text-to-speech method based on knowledge distillation according to claim 1, characterized in that: Converting the optimized acoustic feature sequence into a speech waveform by a vocoder comprises: Inputting the optimized acoustic feature sequence into a preprocessing layer of a lightweight vocoder, and performing normalization processing of the time dimension and context alignment processing on the optimized acoustic feature sequence; Extracting frequency-domain and time-domain features from the optimized acoustic features after normalization and context alignment through the depth-separable convolutional layer of the lightweight vocoder to generate multi-scale speech features; Compressing the parameter scale of the multi-scale speech feature by the low-rank convolution kernel of the lightweight vocoder to generate a compressed speech feature; fusing the multi-scale speech feature and the compressed speech feature through a residual connection layer of the lightweight vocoder to generate a fused speech feature; The fused speech features are upsampled and the time domain waveform is reconstructed through the deconvolution layer of the lightweight vocoder to generate the speech waveform.

8. A text-to-speech device based on knowledge distillation, characterized in that: The text-to-speech device based on knowledge distillation includes: The text preprocessing module is used to standardize the input text and generate a standard text sequence; A lightweight text encoding module, used for encoding the standard text sequence through a lightweight text encoder to generate a text latent vector; A non-autoregressive acoustic feature prediction module, used for mapping the text latent vector into a student acoustic feature sequence through the non-autoregressive acoustic feature prediction module; A teacher model module, used for encoding and acoustic feature prediction processing of the standard text sequence through a pre-trained teacher model to generate a teacher acoustic feature sequence; A knowledge distillation module, used for determining the alignment loss between the student acoustic feature sequence and the teacher acoustic feature sequence through the knowledge distillation module; A structured pruning module, used for performing structured pruning processing on the lightweight text encoder and the non-autoregressive acoustic feature prediction module according to the alignment loss; A parameter quantization module is used to perform parameter quantization on the lightweight text encoder and the non-autoregressive acoustic feature prediction module after the pruning process; A quantized lightweight text encoding module, used for encoding the standard text sequence through a lightweight text encoder after parameter quantization processing to generate a compressed text latent vector; A quantized non-autoregressive acoustic feature prediction module, used to map the compressed text latent vector into an optimized acoustic feature sequence through the non-autoregressive acoustic feature prediction module after parameter quantization processing; The vocoder module is used to convert the optimized acoustic feature sequence into a speech waveform through a vocoder.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a text-to-speech program based on knowledge distillation stored in the memory and executable on the processor. When the text-to-speech program based on knowledge distillation is executed by the processor, the steps of the text-to-speech method based on knowledge distillation as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores a text-to-speech program based on knowledge distillation, and when the text-to-speech program based on knowledge distillation is executed by a processor, the steps of the text-to-speech method based on knowledge distillation as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Non-autoregressive optimized data sequence processing method and device, equipment and medium

    CN120808746A

  • Medical care voice structured input method and system based on self-supervised learning

    CN121034289A

  • Compression method, system and equipment of speech synthesis model

    CN121075309A

  • Model weight quantification method, electronic device and program product

    CN121351913A

  • Quantization method of model weights, electronic device, and program product

    CN121351913B