Method and device for converting text into voice, equipment and medium
Through the artificial intelligence-driven text conversion speech method, the dual autoregression architecture and quantitative reconstruction processing technology are used to solve the problems of insufficient accuracy and inefficiency in the existing technology, and more efficient and accurate speech synthesis is achieved to adapt to the voice style needs in complex fields.
Patent Information
- Application Number
- CN202510441902.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-27
AI Technical Summary
Existing text-to-speech technologies have problems inadequate accuracy, inefficiency and inability to flexibly adjust voice styles in the fields of healthcare and financial technology, especially when dealing with complex medical and financial terms, it is difficult to ensure accurate pronunciation and natural pronunciation.
Using the text conversion speech method driven by artificial intelligence, the target text word segmentation processing and part-of-speech annotation is performed on the target text, the output encoding is generated using the dual autoregression architecture, the decoder generates the Mel spectrum, and the encoding architecture is quantized and reconstructed, and the encoder architecture is optimized to generate high-quality audio.
It improves the accuracy and efficiency of speech synthesis, can better handle complex professional terms, adapt to the voice style needs of different fields and scenarios, and improves the codebook processing efficiency and the quality of speech conversion.
Smart Images

Figure CN120220645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of language signal processing, fintech, and healthcare, and particularly to a method, apparatus, device, and medium for text-to-speech conversion. Background Art
[0002] In the fields of healthcare and finance, the application of text-to-speech technology is becoming increasingly widespread. However, the existing technologies still expose many problems when meeting the special needs of various fields and dealing with complex situations. In the field of healthcare, medical text information is rich and diverse, covering medical records, doctor's orders, medical research reports, etc. Due to the professionalism, complexity, and diversity of medical terms, the existing text-to-speech methods are difficult to accurately grasp their pronunciation rules and semantic focuses. For example, for the names of some rare diseases and complex pharmaceutical chemical names, the converted speech often has pronunciation errors or unnatural intonations, which not only affects the accurate acquisition of information by medical staff but may also cause misunderstandings when patients receive health guidance. At the same time, the demand for emotional expression and personalization of speech in medical scenarios is relatively high. For example, when communicating the condition to patients, a gentle and concerned tone is required. However, the existing technologies have limited capabilities in emotional simulation and cannot flexibly adjust the speech style to adapt to different communication scenarios.
[0003] In the field of fintech, text-to-speech technology is commonly used in scenarios such as financial news broadcasting, customer service guidance, and transaction information prompts. Financial text content contains a large number of precise numbers, complex financial terms, and real-time changing market dynamic information. When the existing text-to-speech methods process this content, problems such as incorrect digital concatenation and non-standard pronunciation of terms are likely to occur, affecting the accuracy of information transmission. For example, when broadcasting key data such as stock price trends and interest rate adjustments, a voice error may lead investors to make wrong decisions. In addition, in the face of different customer groups and financial business scenarios, such as high-end wealth management services requiring a professional and steady voice style, while inclusive financial services may be more inclined to an amiable and easy-to-understand voice style, the existing technologies are difficult to quickly and flexibly achieve the customization and switching of diverse voice styles. The existing text-to-speech technologies are mainly divided into rule-based methods and deep learning-based methods. Rule-based methods rely on pre-set speech rules and pronunciation dictionaries, and have poor adaptability to newly emerging words or complex language structures, and are difficult to cope with the constantly updated professional terms in the medical and financial fields. While deep learning-based methods can learn a large number of language patterns, when training on small-sample, specific-domain datasets, overfitting is likely to occur, resulting in insufficient generalization ability for unseen professional texts, unstable speech synthesis quality, and low efficiency. Summary of the Invention
[0004] The present invention provides a method, apparatus, device and medium for converting text to speech in artificial intelligence to solve the problems of low efficiency and insufficient accuracy in existing text-to-speech methods.
[0005] In a first aspect, a method for converting text to speech is provided, including:
[0006] Performing text tokenization processing and part-of-speech tagging processing on the target text to obtain a preprocessed text;
[0007] Generating an output code according to the preprocessed text by using a preset double autoregressive architecture;
[0008] Generating a Mel spectrogram according to the output code by using a preset decoder;
[0009] Performing quantization reconstruction processing on the Mel spectrogram by using a preset encoding architecture to obtain a quantization tensor;
[0010] Determining the tensor loss value between the quantization tensor and the Mel spectrogram, and optimizing the parameters of the encoder architecture according to the tensor loss value based on the backpropagation algorithm to obtain an optimized encoding architecture;
[0011] Generating a prompt code according to the text to be processed obtained in advance based on the optimized encoding architecture;
[0012] Generating a Mel spectrogram to be processed by combining the prompt code and the text to be processed;
[0013] Generating a target audio according to the Mel spectrogram to be processed by using a preset vocoder.
[0014] In a second aspect, a device for converting text to speech is provided, including:
[0015] A text processing module, configured to perform text tokenization processing and part-of-speech tagging processing on the target text to obtain a preprocessed text;
[0016] A spectrogram generation module, configured to generate an output code according to the preprocessed text by using a preset double autoregressive architecture, and generate a Mel spectrogram according to the output code by using a preset decoder;
[0017] A quantization reconstruction module, configured to perform quantization reconstruction processing on the Mel spectrogram by using a preset encoding architecture to obtain a quantization tensor;
[0018] A parameter optimization module, configured to determine the tensor loss value between the quantization tensor and the Mel spectrogram, and optimize the parameters of the encoder architecture according to the tensor loss value based on the backpropagation algorithm to obtain an optimized encoding architecture;
[0019] An audio generation module is configured to generate a prompt code based on the optimized encoding architecture according to the pre-acquired text to be processed, generate a to-be-processed Mel spectrogram by combining the prompt code and the text to be processed, and generate a target audio according to the to-be-processed Mel spectrogram by using a preset vocoder.
[0020] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above text-to-speech method are implemented.
[0021] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above text-to-speech method are implemented.
[0022] In the solution implemented by the above text-to-speech method, apparatus, computer device, and storage medium, text tokenization processing and part-of-speech tagging processing can be performed on the target text to obtain a preprocessed text, an output code is generated according to the preprocessed text by using a preset double autoregressive architecture, a Mel spectrogram is generated according to the output code by using a preset decoder, quantization reconstruction processing is performed on the Mel spectrogram by using a preset encoding architecture to obtain a quantization tensor, a tensor loss value between the quantization tensor and the Mel spectrogram is determined, the parameters of the encoder architecture are optimized based on the backpropagation algorithm according to the tensor loss value to obtain an optimized encoding architecture, a prompt code is generated based on the optimized encoding architecture according to the pre-acquired text to be processed, a to-be-processed Mel spectrogram is generated by combining the prompt code and the text to be processed, and a target audio is generated according to the to-be-processed Mel spectrogram by using a preset vocoder. The accuracy of speech synthesis is improved, the codebook processing efficiency and utilization rate are improved while maintaining high-quality output, and the efficiency of speech conversion is improved. Description of the Drawings
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0024] Figure 1 It is a schematic diagram of an application environment of the text-to-speech method in an embodiment of the present invention;
[0025] Figure 2 It is a schematic flowchart of the text-to-speech method in an embodiment of the present invention;
[0026] Figure 3It is a schematic structural diagram of a text-to-speech device in an embodiment of the present invention;
[0027] Figure 4 It is a schematic structural diagram of a computer device in an embodiment of the present invention;
[0028] Figure 5 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] The text-to-speech method provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server through a network. The server can perform text tokenization processing and part-of-speech tagging processing on the target text through the client to obtain a preprocessed text, generate an output code according to the preprocessed text using a preset dual autoregressive architecture, generate a Mel spectrogram according to the output code using a preset decoder, perform quantization reconstruction processing on the Mel spectrogram using a preset encoding architecture to obtain a quantization tensor, determine the tensor loss value between the quantization tensor and the Mel spectrogram, optimize the parameters of the encoder architecture based on the tensor loss value according to the backpropagation algorithm to obtain an optimized encoding architecture, generate a prompt code according to the pre-obtained text to be processed based on the optimized encoding architecture, generate a Mel spectrogram to be processed in combination with the prompt code and the text to be processed, and generate a target audio according to the Mel spectrogram to be processed using a preset vocoder, improving the efficiency and accuracy of text-to-speech. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.
[0031] Please refer to Figure 2 as shown in Figure 2 which is a flowchart of a text-to-speech method provided by an embodiment of the present invention, including the following steps:
[0032] S1. Perform text tokenization processing and part-of-speech tagging processing on the target text to obtain a preprocessed text.
[0033] Example illustration: In the scenario of telemedicine diagnosis in the field of healthcare, doctors need to promptly understand the patient's condition and symptoms. However, sometimes patients may not be able to clearly describe their situation in writing. Through text-to-speech technology, the system can convert the text description of the patient's condition input into speech, facilitating doctors to quickly listen to and understand it during their busy work. For example, an elderly patient with heart disease may input text such as "I feel my heart beating very fast today and also have some chest tightness" when using a medical software. After the system converts it into speech, doctors can more conveniently obtain information during ward rounds, consultations, etc., without having to stop their work to read the text, improving the diagnosis efficiency.
[0034] Similarly, in the scenario of intelligent investment advisor services in the fintech field, investors need to promptly understand market dynamics and investment advice. However, sometimes it may be inconvenient to read long analysis reports. Through text-to-speech technology, the system can convert text content such as investment analysis reports and financial news into speech, facilitating investors to listen to it in scenarios such as driving and exercising. For example, an investment analysis report on a popular stock, which includes the company's financial situation, industry prospects, risk assessment, etc., after the system converts it into speech, investors can listen to it through headphones during their commute to work and promptly obtain important information to make more informed investment decisions. At the same time, in financial customer service, the questions consulted by customers and the answers from customer service are usually recorded in text form in the chat window. By using text-to-speech technology, important information can be converted into speech, facilitating customer service staff and customers to review and confirm at any time. For example, when a customer consults questions related to credit card repayment and the customer service staff replies, the system can convert the answer of the customer service staff into speech, and the customer can listen to it again to confirm, avoiding disputes caused by misunderstanding the text content. This method improves the quality and efficiency of customer service and enhances customer satisfaction.
[0035] Specifically, in the field of healthcare, the target text can be the text description of the patient's condition input by the patient.
[0036] Specifically, in the fintech field, the target text can be an analysis report on the investment market.
[0037] In the embodiments of the present invention, the text tokenization process is an important step in natural language processing, aiming to decompose a continuous text string into meaningful units such as words, phrases, or characters, which can improve the readability of the text.
[0038] In the embodiments of the present invention, the part-of-speech tagging process refers to assigning a part-of-speech tag (such as noun, verb, adjective, etc.) to each word obtained through text tokenization, which can help the computer understand the semantics of the text.
[0039] In the embodiments of the present invention, by performing text word segmentation processing and part-of-speech tagging processing on the pre-acquired target text, the accuracy of subsequent speech synthesis can be improved.
[0040] S2. Generate an output code according to the preprocessed text by using a preset dual autoregressive architecture.
[0041] In the embodiments of the present invention, the dual autoregressive architecture is a TTS (Text-to-Speech) system designed specifically for processing complex language features, multi-syllable words, and natural and fluent multi-language synthesis. The dual autoregressive architecture consists of two sequential autoregressive Transformer modules: a slow Transformer and a fast Transformer. This design can efficiently process the global and detailed features of speech synthesis.
[0042] Specifically, the core of the dual autoregressive architecture consists of two sequentially arranged autoregressive Transformer modules, namely a slow Transformer and a fast Transformer. The slow Transformer module is mainly responsible for capturing the global features in speech synthesis, such as intonation, rhythm, and semantic structure, etc. Through deep learning technology, it conducts high-level semantic understanding and analysis on the input text, thus providing a solid foundation for the entire speech synthesis process. The slow Transformer can identify complex language features in the text, such as grammatical structure, semantic association, and emotional color, etc., and convert this information into the global guiding signals required for speech synthesis.
[0043] Specifically, the fast Transformer module focuses on processing the detailed features of speech synthesis, such as the pronunciation details of phonemes, the subtle changes in pitch, and the coherence of speech, etc. Under the global guidance provided by the slow Transformer, the fast Transformer conducts fine-grained modeling and adjustment on each phoneme to ensure that the generated speech achieves a natural and fluent effect in terms of sound quality, timbre, and rhythm. It can accurately process the pronunciation of multi-syllable words, avoiding inaccurate or rigid pronunciation during the synthesis process, thereby improving the overall quality of speech synthesis. The design of this dual autoregressive architecture cleverly combines the advantages of global and detailed processing, enabling the TTS system to exhibit excellent performance when dealing with complex language and multi-language synthesis tasks.
[0044] Furthermore, the dual autoregressive architecture enhances the stability and computational efficiency of codebook processing during the sequence generation process, especially when using grouped finite scalar vector quantization.
[0045] In the embodiments of the present invention, by introducing the dual autoregressive architecture, the processing of multiple language texts and the speech conversion of multiple language texts can be realized.
[0046] Specifically, the Transformer is a deep learning model architecture based on the attention mechanism. The Transformer completely abandons the traditional recurrent neural network structure and can capture the relationships between different positions in the sequence. The self-attention mechanism allows the model to consider the information of the entire sequence while processing each element, thus better capturing long-range dependencies. And it can perform parallel processing on the entire sequence, which greatly speeds up the training speed and inference speed of the model.
[0047] In the embodiment of the present invention, the generating the output encoding according to the preprocessed text by using the preset double autoregressive architecture includes:
[0048] Generating a text embedding and a codebook embedding according to the preprocessed text;
[0049] Generating a hidden state vector according to the text embedding by using a preset slow Transformer layer;
[0050] Concatenating the hidden state vector and the codebook embedding to obtain a concatenated vector;
[0051] Generating an output encoding according to the concatenated vector by using a preset fast Transformer layer.
[0052] In the embodiment of the present invention, the generating the text embedding and the codebook embedding according to the preprocessed text is to generate a text embedding according to the preprocessed text by using a word embedding model; and extracting the features of the preprocessed text, discretizing the preprocessed text into continuous codewords, and then generating a codebook in the feature space according to the continuous codewords by using a clustering algorithm, mapping the extracted features into the codebook to obtain a sequence of codewords, and converting the sequence of codewords into an embedding vector to obtain a codebook embedding.
[0053] In detail, in a TTS system, generating high-quality speech output requires in-depth analysis and processing of the input text. The process includes generating text embeddings and codebook embeddings based on the preprocessed text, which involves multiple key steps to ensure the naturalness and accuracy of the final speech. First, a word embedding model is used to generate text embeddings based on the preprocessed text. The word embedding model captures the semantic relationship and contextual information between words by mapping each word or phrase in the text to a vector in a high-dimensional space. This embedding can provide rich semantic features for subsequent speech synthesis, so that the generated speech is not only accurate in pronunciation, but also closer to natural language in intonation and emotion. Next, extract the features of the preprocessed text. This step involves in-depth analysis of the text to identify speech features such as phonemes, tones, and rhythms. These features are the basis of speech synthesis and determine the naturalness and comprehensibility of speech. Then, the preprocessed text is discretized into continuous codewords. Codewords are discrete representations of text in feature space, which can better capture the nuances and complex structures of text. By discretizing the text, more refined input can be provided for subsequent clustering and embedding generation. Next, a codebook is generated in the feature space based on continuous codewords through a clustering algorithm. The clustering algorithm clusters similar codewords together to form a codebook. Each codeword in the codebook represents a feature set in the text, and they have similar properties in the feature space. Finally, the extracted features are mapped to the codebook to obtain a codeword sequence. This codeword sequence is a compact representation of the original text, which retains the key feature information of the text. The codeword sequence is converted into an embedding vector to obtain a codebook embedding. Codebook embedding can provide more precise guidance for speech synthesis, making the generated speech closer to natural language in details.
[0054] In detail, the generating a hidden state vector according to the text embedding using a preset slow Transformer layer includes:
[0055] Converting the text embedding into an embedding vector using an embedding layer in the slow Transformer layer;
[0056] Calculating the attention weight of the embedding vector based on a multi-head self-attention mechanism;
[0057] Generate a context vector based on the attention weight and the embedding vector;
[0058] Performing a nonlinear transformation on the context vector to obtain a hidden state;
[0059] The hidden state is converted into a labeled logarithm based on preset weights and bias terms using a normalization function to obtain a hidden state vector.
[0060] Specifically, the multi-head self-attention mechanism is a key component in the Transformer model, which allows the model to capture information in different representation subspaces.
[0061] Specifically, calculating the attention weights of the embedding vectors based on the multi-head self-attention mechanism uses the softmax function to calculate the attention weights according to the query vector and the key vector.
[0062] Specifically, generating the context vector according to the attention weights and the embedding vectors is to multiply the attention weights by the embedding vectors.
[0063] Specifically, performing a non-linear transformation on the context vector to obtain the hidden state is a non-linear transformation based on a feed-forward neural network.
[0064] Specifically, the normalization function can be the softmax function.
[0065] Specifically, concatenating the hidden state vector and the codebook embedding to obtain the concatenated vector is based on dimensions.
[0066] Specifically, using the preset fast Transformer layer to generate the output encoding according to the concatenated vector includes:
[0067] Performing a linear transformation on the concatenated vector to obtain a linear transformation matrix;
[0068] Performing residual connection and layer normalization processing on the linear transformation matrix to obtain a normalized matrix;
[0069] Based on a feed-forward neural network, concatenating the normalized matrix and the linear transformation matrix to obtain a feed-forward neural network concatenated matrix;
[0070] Performing secondary residual connection and layer normalization processing on the feed-forward neural network concatenated matrix to obtain the output encoding.
[0071] Specifically, performing a linear transformation on the concatenated vector to obtain a linear transformation matrix is to use the multi-head self-attention mechanism to convert the input concatenated vector into a linear transformation matrix of query, key, and value.
[0072] In the embodiment of the present invention, performing residual connection and layer normalization processing on the linear transformation matrix means performing residual connection processing on the linear transformation matrix first and then performing layer normalization processing, specifically including adding skip connections between layers of the linear transformation matrix so that signals can directly propagate across layers, and then normalizing the activation values of all neurons of a single training sample to improve training stability and accelerate convergence.
[0073] Specifically, in a deep learning model, performing residual connection and layer normalization on the linear transformation matrix is a common optimization strategy aimed at enhancing the training stability of the model and accelerating the convergence speed. Specifically in the said process, first perform residual connection processing on the linear transformation matrix, and then perform layer normalization processing. These two steps play important roles respectively.
[0074] Specifically, the residual connection, also known as the skip connection or shortcut connection, is a special connection method added between layers of the model. When performing residual connection processing, the output of the linear transformation matrix is directly added to the input to form a residual block. This design allows signals to directly propagate across layers in the deep structure of the model, thus alleviating the problems of vanishing gradients and exploding gradients in the training of deep networks. Through the residual connection, the model can more easily learn the identity mapping, making the training of deep networks more feasible and efficient. In addition, the residual connection can also promote the flow of information and the reuse of features, helping the model capture more complex feature representations.
[0075] Specifically, the layer normalization is a normalization method for a single training sample, which normalizes all neuron activation values of a single sample. Specifically, when performing layer normalization, calculate the mean and variance of the activation values of each sample, and then standardize the activation values according to these statistics so that they have zero mean and unit variance. The advantage of layer normalization is that it does not depend on the batch size, so it is still effective when processing small batches of data or a single sample. In addition, layer normalization helps to stabilize the training process, reduce the instability of gradients, and thus accelerate the convergence speed of the model. Through layer normalization, the model can better adapt to different data distributions and improve the generalization ability of the model.
[0076] Specifically, the feedforward neural network is a basic and widely used neural network structure, whose characteristic is that information passes from the input layer through one or more hidden layers and finally reaches the output layer, and there is no cycle or feedback in the whole process.
[0077] In the embodiment of the present invention, by using a preset double autoregressive architecture to generate an output code according to the preprocessed text, the efficiency of generating the output code is improved.
[0078] S3. Use a preset decoder to generate a Mel spectrogram according to the output code.
[0079] In the embodiments of the present invention, the Mel spectrogram is an advanced audio signal processing technology that adopts a special spectral representation method to more accurately capture and simulate the way the human auditory system perceives sound. The core of the Mel spectrogram lies in its basis on the Mel Frequency Scale, which is designed according to the perceptual sensitivity of the human ear to sounds of different frequencies and can be closer to the auditory characteristics of humans.
[0080] Specifically, compared with the traditional linear frequency scale, the Mel frequency scale is denser in the low-frequency region and relatively sparse in the high-frequency region. This is because the human ear has a stronger ability to distinguish low-frequency sounds and a relatively weaker ability to distinguish high-frequency sounds. Through this non-linear frequency distribution, the Mel spectrogram can better reflect the perceptual differences of the human auditory system to sound frequencies, thus providing more accurate and effective spectral information in audio signal processing.
[0081] Furthermore, in the process of generating the Mel spectrogram, the audio signal is first converted into a frequency-domain representation through the short-time Fourier transform, and then the spectrum is filtered through a Mel filter bank. The Mel filter bank consists of a series of filters with specific frequency responses, which are evenly distributed on the Mel frequency scale and can map the spectral energy of the audio signal onto the Mel frequency scale. After being processed by the Mel filter bank, the obtained Mel spectrogram can highlight the key frequency components in the audio signal while suppressing noise and other unnecessary frequency components, providing clearer and more reliable spectral features for subsequent audio analysis and processing.
[0082] In the embodiments of the present invention, the decoder is an autoregressive recurrent neural network that can predict one frame of the Mel spectrogram from the encoded input sequence at a time.
[0083] S4. Use a preset coding architecture to perform quantization and reconstruction processing on the Mel spectrogram to obtain a quantization tensor.
[0084] In the embodiments of the present invention, the coding architecture includes an encoder and a quantizer.
[0085] In the embodiments of the present invention, using a preset coding architecture to perform quantization and reconstruction processing on the Mel spectrogram to obtain a quantization tensor means grouping the Mel spectrogram based on the channel dimension, performing scalar quantization on each grouping result, and then re-splicing them.
[0086] In the embodiments of the present invention, using a preset coding architecture to perform quantization and reconstruction processing on the Mel spectrogram to obtain a quantization tensor includes:
[0087] Group the Mel spectrogram according to the channel dimension to obtain a set of grouping results;
[0088] Perform scalar quantization on each grouped tensor in the grouped result set to obtain a quantization result;
[0089] Perform finite quantization on the quantization result to obtain a finite quantization result;
[0090] Concatenate the finite quantization results of all groups along the channel dimension to obtain a quantization tensor.
[0091] In the embodiment of the present invention, by grouping the Mel spectrogram according to the channel dimension to obtain a grouped result set, similar features can be clustered together so that subsequent quantization operations can more effectively capture the local structure of the data.
[0092] Specifically, the performing scalar quantization on each grouped tensor in the grouped result set means mapping each element in each group to a vector in a predefined codebook, and each element in each group will be replaced by the vector closest to it in the codebook. The codebook is obtained by performing clustering analysis on the training data.
[0093] Specifically, the performing finite quantization on the quantization result to obtain a finite quantization result retains the indices of the top K maximum values in the quantization result of each group, where K is a preset value, which can reduce the computational complexity and storage requirements while retaining the key features of the data.
[0094] In the embodiment of the present invention, by using a preset encoding architecture to perform quantization reconstruction processing on the Mel spectrogram to obtain a quantization tensor, the feature expression ability of the quantization tensor is improved.
[0095] S5. Determine the tensor loss value between the quantization tensor and the Mel spectrogram, and optimize the parameters of the encoder architecture based on the backpropagation algorithm according to the tensor loss value to obtain an optimized encoding architecture.
[0096] In the embodiment of the present invention, the determining the tensor loss value between the quantization tensor and the Mel spectrogram is to calculate the tensor loss value by using a preset objective function.
[0097] In the embodiment of the present invention, the determining the tensor loss value between the quantization tensor and the Mel spectrogram includes:
[0098] Calculate the tensor loss value by using the following formula:
[0099]
[0100] where N is the loss value, B represents the preset batch size, C represents the number of channels, L represents the sequence length, represents the value at the l-th position of the c-th channel in the b-th batch of the quantization tensor, z(b,c,l) represents the value at position l in the b-th batch and the c-th channel of the said Mel spectrogram.
[0101] Specifically, the determination of the tensor loss value between the quantization tensor and the Mel spectrogram includes:
[0102] Obtain the batch size, number of channels, and sequence length of the quantization tensor;
[0103] Using the batch size, number of channels, and sequence length as indices, determine the squared difference between the data in each sequence of each channel in all batches of the quantization tensor and the corresponding data in the Mel spectrogram, obtaining a set of squared difference calculation results;
[0104] Sum up the set of squared difference calculation results to obtain a summation result;
[0105] Calculate the reciprocal of the product of the batch size, number of channels, and sequence length, and multiply the reciprocal by the summation result to obtain the tensor loss value.
[0106] In an embodiment of the present invention, the optimization of the parameters of the encoder architecture based on the backpropagation algorithm according to the tensor loss value includes:
[0107] The backpropagation algorithm is represented by the following formula:
[0108]
[0109] where θ new is the updated parameter of the encoder architecture, θ old represents the original parameter of the encoder architecture, η is a preset learning rate, is the gradient value of the loss function with respect to the parameters of the encoder architecture.
[0110] S6. Generate a prompt code based on the optimized encoding architecture according to the pre-obtained text to be processed.
[0111] In an embodiment of the present invention, in the medical field, the text to be processed can be the text that is real-time feedback for the user's question in a medical service software, or the text that the user inputs in real-time according to their own condition.
[0112] In an embodiment of the present invention, in the fintech field, the text to be processed can be an investment analysis report generated in real-time based on the market investment situation.
[0113] In an embodiment of the present invention, the generation of a prompt code based on the optimized encoding architecture according to the pre-obtained text to be processed includes:
[0114] Perform a double autoregressive process on the text to be processed to obtain an output encoding to be processed.
[0115] Use a decoder to convert the output encoding to be processed into a Mel spectrogram to be processed.
[0116] Use the encoding architecture to generate a prompt encoding based on the Mel spectrogram to be processed.
[0117] In an embodiment of the present invention, the prompt encoding is an encoding that helps identify text features and can improve the accuracy of subsequent speech synthesis.
[0118] In an embodiment of the present invention, the step of using the encoding architecture to generate a prompt encoding based on the Mel spectrogram to be processed is to group the Mel spectrogram to be processed, quantize the data of each group, and finally splice the quantization results of each group.
[0119] S7. Combine the prompt encoding and the text to be processed to generate a Mel spectrogram to be processed.
[0120] In an embodiment of the present invention, the step of combining the prompt encoding and the text to be processed to generate a Mel spectrogram to be processed is to preprocess the text to be processed (perform word segmentation and part-of-speech tagging), combine the preprocessing results with the prompt encoding, and perform a double autoregressive process on the combined data to finally obtain a Mel spectrogram to be processed.
[0121] Specifically, the combination of the preprocessing results and the prompt encoding can be achieved by splicing the preprocessing results and the prompt encoding.
[0122] S8. Use a preset vocoder to generate a target audio based on the Mel spectrogram to be processed.
[0123] In an embodiment of the present invention, the vocoder is a model for analyzing and synthesizing speech signals and can generate audio signals based on the input spectral signal.
[0124] Specifically, the vocoder adopts an enhanced convolutional structure, including depthwise separable convolution and dilated convolution, which improves the model's ability to capture and synthesize complex audio features.
[0125] Specifically, the vocoder introduces depthwise separable convolution and dilated convolution. Depthwise separable convolution separates spatial convolution and channel convolution, reducing the computational complexity while maintaining the model's expressive power. Dilated convolution, on the other hand, increases the receptive field, helping the model better capture long-range dependencies and enabling it to more effectively capture multi-scale features in audio data.
[0126] In an embodiment of the present invention, the vocoder includes a ConvNext encoder and a generator.
[0127] In an embodiment of the present invention, generating the target audio according to the to-be-processed Mel spectrogram by using a preset vocoder includes:
[0128] Performing filtering processing and normalization processing on the Mel spectrogram to obtain a preprocessed spectrogram;
[0129] Extracting the spectral features of the preprocessed spectrogram through short-time Fourier transform;
[0130] Using the ConvNext encoder in the vocoder to generate a quantized Mel spectrogram according to the spectral features;
[0131] Using the generator in the vocoder to generate the target audio according to the quantized Mel spectrogram.
[0132] Specifically, the filtering processing is a method for removing noise in the Mel spectrogram, which can improve the data quality.
[0133] Specifically, the normalization processing is a method of mapping the data size in the Mel spectrogram to between 0 and 1 to improve the efficiency of subsequent processing.
[0134] Specifically, the ConvNext encoder is an encoder architecture based on a convolutional neural network, aiming to combine some optimization techniques in the Transformer model to improve the performance of the convolutional network in image recognition tasks.
[0135] In an embodiment of the present invention, the ParallelBlock is used to replace the traditional Multi-ReceptiveField module in the architecture, which can improve the processing efficiency of the codebook input.
[0136] Specifically, the ParallelBlock is a structure used for parallel processing of multiple sub-modules in a deep learning model. It allows the input data to be simultaneously passed to multiple sub-modules and combines the outputs of these sub-modules.
[0137] Furthermore, the ParallelBlock realizes parallel processing of different convolutional kernel sizes and dilation rates, which enables the model to more flexibly adapt to different audio features. Different from the traditional direct addition operation, the ParallelBlock uses a stacking and averaging mechanism to process the outputs from three ResBlocks. This method not only improves the feature extraction ability but also enhances the stability and generalization ability of the model. The ParallelBlock provides a wider receptive field coverage and stronger feature extraction ability, and at the same time has higher configurability. These characteristics together improve the quality of audio synthesis, making the generated audio more natural and realistic.
[0138] It can be seen that in the above solution, for the problem of converting multi-language text to speech, the target text is preprocessed, and then the output encoding is generated according to the preprocessed text by using a dual autoregressive architecture. Then, the Mel spectrogram is generated by using a decoder according to the output encoding. Based on the Mel spectrogram, the parameters of the encoding architecture are optimized. The optimized encoding architecture is used to generate a prompt encoding according to the text to be processed obtained, and then the text to be processed is combined with the prompt encoding. According to the combined result, a Mel spectrogram to be processed is generated. Finally, a preset vocoder is used to synthesize audio according to the Mel spectrogram to be processed. This method can support multiple languages, improve the accuracy of speech synthesis, improve the codebook processing efficiency and utilization rate while maintaining high-quality output, and improve the efficiency of speech conversion.
[0139] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0140] In one embodiment, a text-to-speech device is provided, and the text-to-speech device corresponds one-to-one with the text-to-speech method in the above embodiment. As Figure 3 shown, the text-to-speech device includes a text processing module 101, a spectrogram generation module 102, a quantization reconstruction module 103, a parameter optimization module 104, and an audio generation module 105. The detailed description of each functional module is as follows:
[0141] The text processing module 101 is configured to perform text word segmentation processing and part-of-speech tagging processing on the target text to obtain a preprocessed text;
[0142] The spectrogram generation module 102 is configured to generate an output encoding according to the preprocessed text by using a preset dual autoregressive architecture, and generate a Mel spectrogram according to the output encoding by using a preset decoder;
[0143] The quantization reconstruction module 103 is configured to perform quantization reconstruction processing on the Mel spectrogram by using a preset encoding architecture to obtain a quantization tensor;
[0144] The parameter optimization module 104 is configured to determine the tensor loss value between the quantization tensor and the Mel spectrogram, and optimize the parameters of the encoder architecture according to the tensor loss value based on the backpropagation algorithm to obtain an optimized encoding architecture;
[0145] The audio generation module 105 is configured to generate a prompt encoding according to the text to be processed obtained in advance based on the optimized encoding architecture, combine the prompt encoding and the text to be processed to generate a Mel spectrogram to be processed, and generate a target audio according to the Mel spectrogram to be processed by using a preset vocoder.
[0146] In one embodiment, the spectrum generation module 102 is specifically configured to:
[0147] Generate a text embedding and a codebook embedding according to the preprocessed text;
[0148] Generate a hidden state vector according to the text embedding by using a preset slow Transformer layer;
[0149] Concatenate the hidden state vector and the codebook embedding to obtain a concatenated vector;
[0150] Generate an output encoding according to the concatenated vector by using a preset fast Transformer layer.
[0151] In one embodiment, the spectrum generation module 102 is specifically configured to:
[0152] Convert the text embedding into an embedding vector by using the embedding layer in the slow Transformer layer;
[0153] Calculate the attention weights of the embedding vector based on the multi-head self-attention mechanism;
[0154] Generate a context vector according to the attention weights and the embedding vector;
[0155] Perform a non-linear transformation on the context vector to obtain a hidden state;
[0156] Convert the hidden state into a token logarithm based on a preset weight and bias term by using a normalization function to obtain a hidden state vector.
[0157] In one embodiment, the spectrum generation module 102 is specifically configured to:
[0158] Perform a linear transformation on the concatenated vector to obtain a linear transformation matrix;
[0159] Perform a residual connection and a layer normalization process on the linear transformation matrix to obtain a normalized matrix;
[0160] Concatenate the normalized matrix and the linear transformation matrix based on a feed-forward neural network to obtain a feed-forward neural network concatenated matrix;
[0161] Perform a secondary residual connection and a layer normalization process on the feed-forward neural network concatenated matrix to obtain an output encoding.
[0162] In one embodiment, the quantization reconstruction module 103 is specifically configured to:
[0163] Group the Mel spectrogram according to the channel dimension to obtain a set of grouping results;
[0164] Perform scalar quantization on each grouped tensor in the grouped result set to obtain a quantization result;
[0165] Perform finite quantization on the quantization result to obtain a finite quantization result;
[0166] Concatenate the finite quantization results of all groups along the channel dimension to obtain a quantization tensor.
[0167] In one embodiment, the parameter optimization module 104 is specifically configured to:
[0168] Obtain the batch size, the number of channels, and the sequence length of the quantization tensor;
[0169] Using the batch size, the number of channels, and the sequence length as indices, determine the squared difference between the data in each sequence in each channel in all batches in the quantization tensor and the corresponding data in the Mel spectrogram, to obtain a set of squared difference calculation results;
[0170] Sum the set of squared difference calculation results to obtain a summation result;
[0171] Calculate the reciprocal of the product of the batch size, the number of channels, and the sequence length, and multiply the reciprocal by the summation result to obtain the tensor loss value.
[0172] In one embodiment, the audio conversion module 104 is specifically configured to:
[0173] Perform filtering processing and normalization processing on the Mel spectrogram to obtain a preprocessed spectrogram;
[0174] Extract the spectral features of the preprocessed spectrogram through short-time Fourier transform;
[0175] Use the ConvNext encoder in the vocoder to generate a quantized Mel spectrogram according to the spectral features;
[0176] Use the generator in the vocoder to generate a target audio according to the quantized Mel spectrogram.
[0177] The present invention provides a text-to-speech device. For the problem of multi-language text-to-speech, by preprocessing the target text, generating an output code according to the preprocessed text using a dual autoregressive architecture, generating a Mel spectrogram according to the output code using a decoder, and then optimizing the parameters of the coding architecture based on the Mel spectrogram. Using the optimized coding architecture to generate a prompt code according to the obtained text to be processed, then combining the text to be processed with the prompt code, generating a Mel spectrogram to be processed according to the combined result, and finally synthesizing an audio according to the Mel spectrogram to be processed using a preset vocoder. This method can support multiple languages, improve the accuracy of speech synthesis, and while maintaining high-quality output, improve the codebook processing efficiency and utilization rate, and improve the efficiency of speech conversion.
[0178] For the specific limitations of the text-to-speech device, reference can be made to the limitations of the text-to-speech method in the above text, which will not be elaborated here. Each module in the above text-to-speech device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0179] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a text-to-speech method.
[0180] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a text-to-speech method
[0181] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0182] Perform text tokenization processing and part-of-speech tagging processing on the target text to obtain a preprocessed text;
[0183] Generate an output code according to the preprocessed text by using a preset dual autoregressive architecture;
[0184] Generate a Mel spectrogram according to the output code by using a preset decoder;
[0185] Perform quantization reconstruction processing on the Mel spectrogram by using a preset encoding architecture to obtain a quantization tensor;
[0186] Determine the tensor loss value between the quantization tensor and the Mel spectrogram, and optimize the parameters of the encoder architecture according to the tensor loss value based on the backpropagation algorithm to obtain an optimized encoding architecture;
[0187] Generate a prompt code according to the pre-obtained text to be processed based on the optimized encoding architecture;
[0188] Combine the prompt code and the text to be processed to generate a Mel spectrogram to be processed;
[0189] Generate a target audio according to the Mel spectrogram to be processed by using a preset vocoder.
[0190] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:
[0191] Perform text tokenization processing and part-of-speech tagging processing on the target text to obtain a preprocessed text;
[0192] Generate an output code according to the preprocessed text by using a preset dual autoregressive architecture;
[0193] Generate a Mel spectrogram according to the output encoding using a preset decoder;
[0194] Perform quantization reconstruction processing on the Mel spectrogram using a preset encoding architecture to obtain a quantization tensor;
[0195] Determine the tensor loss value between the quantization tensor and the Mel spectrogram, and optimize the parameters of the encoder architecture based on the tensor loss value according to the backpropagation algorithm to obtain an optimized encoding architecture;
[0196] Generate a prompt encoding based on the optimized encoding architecture according to the text to be processed obtained in advance;
[0197] Combine the prompt encoding and the text to be processed to generate a Mel spectrogram to be processed;
[0198] Generate a target audio according to the Mel spectrogram to be processed using a preset vocoder.
[0199] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0200] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0201] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0202] Finally, it should be noted that if there are software tools or components of other companies in the application embodiments, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A text-to-speech method, characterized in that: include: Perform text segmentation and part-of-speech tagging on the target text to obtain a preprocessed text; Generate an output encoding based on the preprocessed text using a preset dual autoregressive architecture; Using a preset decoder to generate a Mel spectrum according to the output code; Quantizing and reconstructing the Mel spectrum using a preset coding architecture to obtain a quantized tensor; Determine tensor loss values of the quantized tensor and the Mel spectrum, and optimize parameters of the encoder architecture according to the tensor loss values based on a back propagation algorithm to obtain an optimized encoding architecture; Generate prompt codes based on the optimized coding architecture and the pre-acquired text to be processed; Combining the prompt code and the text to be processed to generate a Mel spectrum to be processed; A preset vocoder is used to generate target audio according to the to-be-processed Mel spectrum.
2. The text-to-speech method according to claim 1, wherein: The method of generating an output code according to the preprocessed text using a preset dual autoregressive architecture includes: Generating text embedding and codebook embedding according to the preprocessed text; Generate a hidden state vector based on the text embedding using a preset slow Transformer layer; Embedding and concatenating the hidden state vector and the codebook to obtain a concatenated vector; The output encoding is generated based on the concatenated vector using a preset fast Transformer layer.
3. The text-to-speech method according to claim 2, wherein: The using a preset slow Transformer layer to generate a hidden state vector according to the text embedding includes: Converting the text embedding into an embedding vector using an embedding layer in the slow Transformer layer; Calculating the attention weight of the embedding vector based on a multi-head self-attention mechanism; Generate a context vector based on the attention weight and the embedding vector; Performing a nonlinear transformation on the context vector to obtain a hidden state; The hidden state is converted into a labeled logarithm based on preset weights and bias terms using a normalization function to obtain a hidden state vector.
4. The text-to-speech method according to claim 2, wherein: The step of using a preset fast Transformer layer to generate an output code according to the concatenated vector includes: Performing a linear transformation on the splicing vector to obtain a linear transformation matrix; Performing residual connection and layer normalization processing on the linear transformation matrix to obtain a normalized matrix; Based on a feedforward neural network, the normalized matrix and the linear transformation matrix are concatenated to obtain a feedforward neural network concatenated matrix; The feedforward neural network concatenation matrix is subjected to secondary residual connection and layer normalization processing to obtain output coding.
5. The text-to-speech method according to claim 1, wherein: The quantization and reconstruction processing of the Mel spectrum using a preset coding architecture to obtain a quantized tensor includes: Grouping the Mel spectrum according to the channel dimension to obtain a grouping result set; Performing scalar quantization on each grouping tensor in the grouping result set to obtain a quantization result; Performing finite quantization on the quantization result to obtain a finite quantization result; The finite quantization results of all groups are concatenated along the channel dimension to obtain a quantized tensor.
6. The text-to-speech method according to claim 1, wherein: The determining of the tensor loss value of the quantized tensor and the Mel spectrum comprises: Get the batch size, number of channels and sequence length of the quantized tensor; Taking the batch size, the number of channels and the sequence length as indexes, determining the square difference between the data in each sequence in each channel in all batches in the quantized tensor and the corresponding data in the Mel spectrum, and obtaining a set of square difference calculation results; Summing the set of square difference calculation results to obtain a summation result; The reciprocal of the product of the batch size, the number of channels, and the sequence length is calculated, and the reciprocal is multiplied by the summation result to obtain the tensor loss value.
7. The text-to-speech method according to claim 1, wherein: The method of using a preset vocoder to generate a target audio according to the to-be-processed Mel spectrum comprises: Performing filtering and normalization processing on the Mel spectrum to obtain a preprocessed spectrum; Extracting the spectrum features of the preprocessed spectrum by short-time Fourier transform; Using a ConvNext encoder in the vocoder to generate a quantized Mel spectrum according to the frequency spectrum features; A generator in the vocoder is used to generate target audio according to the quantized Mel spectrum.
8. A text-to-speech device, characterized in that: include: A text processing module is used to perform text segmentation and part-of-speech tagging on the target text to obtain a preprocessed text; A spectrum generation module, used to generate an output code according to the preprocessed text using a preset dual autoregressive architecture, and to generate a Mel spectrum according to the output code using a preset decoder; A quantization and reconstruction module, used for performing quantization and reconstruction processing on the Mel spectrum using a preset coding architecture to obtain a quantized tensor; A parameter optimization module, used to determine the tensor loss value of the quantized tensor and the Mel spectrum, and optimize the parameters of the encoder architecture according to the tensor loss value based on a back propagation algorithm to obtain an optimized encoding architecture; The audio generation module is used to generate a prompt code according to the pre-acquired text to be processed based on the optimized coding architecture, generate a Mel spectrum to be processed by combining the prompt code and the text to be processed, and generate a target audio according to the Mel spectrum to be processed using a preset vocoder.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the text-to-speech method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text-to-speech method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Security defense method and device for audio language model
CN121331159A