Speech synthesis method and apparatus, and device, storage medium and program product
By extracting semantic tokens and acoustic tokens that prompt audio and input text, the problem of excessive feature span from text to acoustic tokens in the prior art is solved, and the accuracy and efficiency of speech synthesis are improved.
Patent Information
- Application Number
- PCT/CN2024/113350
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-05
AI Technical Summary
The existing speech synthesis system has too large feature span from text to acoustic tokens, resulting in high requirements for labeled data for model training, which affects the accuracy of speech synthesis.
By extracting the features of the prompt audio, a prompt semantic token and prompt acoustic token are obtained, and a feature extraction is performed on the input text to obtain the input semantic token. Based on these tokens, the feature span in the prediction process from text to acoustic token is reduced.
It improves the accuracy of speech synthesis, reduces the dependence on labeled data, reduces the complexity of model training, and realizes a more efficient speech synthesis process.
Smart Images

Figure CN2024113350_05062025_PF_FP_ABST
Abstract
Description
Speech synthesis method, device, equipment, storage medium and program product
[0001] This application claims priority to Chinese patent application number 202311403590.8, filed on October 25, 2023, entitled “Speech synthesis method, device, equipment, storage medium and program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, apparatus, device, storage medium, and program product. Background Art
[0003] Speech synthesis refers to the process of converting text into audio. In this process, speech synthesis is usually performed using a speech synthesis system based on an AI (Artificial Intelligence) model.
[0004] In related technologies, a speech synthesis system can input the text of the speech content and a prompt audio into an acoustic token extraction model. The acoustic tokens are extracted and used as acoustic features of the audio to be generated. The acoustic tokens are then input into a sound decoder to generate the final audio. The speech content of the generated audio is derived from the text, and the timbre, emotion, and other characteristics of the audio are derived from the prompt audio.
[0005] The above solution directly predicts acoustic tokens from text and prompt audio. The feature span from text to acoustic tokens is too large, resulting in high requirements for labeled data in the training process of the acoustic token extraction model, which limits the accuracy of the acoustic token extraction model and further affects the accuracy of speech synthesis.
[0006] Summary of the Invention
[0007] The present application provides a speech synthesis method, apparatus, device, storage medium and program product, which can improve the accuracy of speech synthesis; the technical solution is as follows.
[0008] According to one aspect of the present application, a speech synthesis method is provided, the method being executed by a computer device, the method comprising:
[0009] Get input text and prompt audio;
[0010] Extracting features of the prompt audio to obtain prompt semantic tokens and prompt acoustic tokens, wherein the prompt semantic tokens are used to indicate semantic features of the prompt audio at various time points, and the prompt acoustic tokens are used to indicate acoustic features of the prompt audio at various time points;
[0011] Extracting features of the input text to obtain input semantic tokens, where the input semantic tokens are used to indicate semantic features of speech corresponding to the input text at various time points;
[0012] Based on the prompt semantic token, the prompt acoustic token and the input semantic token, obtaining an input acoustic token, wherein the input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point;
[0013] Based on the input acoustic tokens, an output audio of the input text is obtained.
[0014] According to one aspect of the present application, a speech synthesis device is provided, the device comprising:
[0015] Acquisition module, used to obtain input text and prompt audio;
[0016] a first extraction module, configured to extract features of the prompt audio to obtain prompt semantic tokens and prompt acoustic tokens, wherein the prompt semantic tokens are used to indicate semantic features of the prompt audio at various time points, and the prompt acoustic tokens are used to indicate acoustic features of the prompt audio at various time points;
[0017] A second extraction module is used to extract features of the input text and obtain input semantic tokens, where the input semantic tokens are used to indicate semantic features of the speech corresponding to the input text at various time points;
[0018] An input acoustic token acquisition module, configured to acquire an input acoustic token based on the prompt semantic token, the prompt acoustic token, and the input semantic token, wherein the input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point;
[0019] An output audio acquisition module is used to acquire the output audio of the input text based on the input acoustic token.
[0020] In some embodiments, the first extraction module is configured to input the prompt audio into a semantic token extractor to obtain a prompt semantic token obtained by the semantic token extractor processing the prompt audio; input the prompt audio into an acoustic token extractor to obtain a prompt acoustic token obtained by the acoustic token extractor processing the prompt audio;
[0021] A second extraction module is used to input the input text into the text-to-semantic token model to obtain input semantic tokens obtained by the text-to-semantic token model processing the input text;
[0022] An input acoustic token acquisition module is used to input the prompt semantic token, the prompt acoustic token, and the input semantic token into the semantic token-to-acoustic token model to obtain the input acoustic token output by the semantic token-to-acoustic token model;
[0023] The output audio acquisition module is used to input the input acoustic token into the sound decoder and obtain the output audio output by the sound decoder.
[0024] In some embodiments, the semantic token extractor includes a convolution branch and a first converter; the first extraction module is used to input the prompt audio into the convolution branch to obtain the hidden layer features of the prompt audio at each time point output by the convolution branch; the hidden layer features of the prompt audio at each time point are processed by the first converter to obtain the intermediate layer features of the prompt audio at each time point output by the intermediate layer of the first converter; the intermediate layer features of the prompt audio at each time point are clustered separately to obtain the prompt semantic token.
[0025] In some embodiments, the device also includes: a semantic token extractor training module, which is used to obtain a first audio sample and a semantic token label of the first audio sample; input the first audio sample into the convolution branch to obtain hidden feature samples of the first audio sample at each time point output by the convolution branch; partially mask the hidden feature samples of the first audio sample at each time point to obtain partially masked hidden feature samples; process the partially masked hidden feature samples through the first converter to obtain intermediate layer features of the first audio sample at each time point output by the intermediate layer of the first converter; cluster the intermediate layer features of the first audio sample at each time point to obtain semantic token samples of the first audio sample; and update the parameters of the semantic token extractor based on the semantic token samples of the first audio sample and the semantic token label of the first audio sample.
[0026] In some embodiments, the text-to-semantic token model includes a text encoder, a duration predictor, an upsampling branch, and a decoder; a second extraction module is used to input the input text into the text encoder to obtain a hidden text encoding representation of the input text; input the hidden text encoding representation into the duration predictor to obtain the playback duration of the speech corresponding to the input text predicted by the duration predictor; upsample the hidden text encoding representation to the number of frames corresponding to the playback duration through the upsampling branch to obtain the upsampled hidden text encoding representation; and decode the upsampled hidden text encoding representation through the decoder to obtain the input semantic token.
[0027] In some embodiments, the device also includes: a text-to-semantic token model training module, which is used to obtain a second audio sample and a speech text of the second audio sample when the semantic token extractor training is completed; input the second audio sample into the semantic token extractor to obtain a semantic token label of the second audio sample output by the semantic token extractor; input the speech text of the second audio sample into the text-to-semantic token model to obtain a semantic token sample of the second audio sample output by the text-to-semantic token model; and update the parameters of the text-to-semantic token model based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample.
[0028] In some embodiments, the text-to-semantic token model training module is further used to input the speech text of the second audio sample into a text encoder to obtain a hidden text encoding representation sample of the speech text of the second audio sample; input the hidden text encoding representation sample into a duration predictor to obtain a first playback duration sample of the speech corresponding to the speech text of the second audio sample predicted by the duration predictor; input the hidden text encoding representation sample into an attention branch to obtain a second playback duration sample of the speech corresponding to the speech text of the second audio sample output by the attention branch; upsample the hidden text encoding representation sample to the number of frames corresponding to the second playback duration sample through an upsampling branch to obtain an upsampled hidden text encoding representation sample; decode the upsampled hidden text encoding representation sample through a decoder to obtain a semantic token sample of the second audio sample; obtain a loss function value of the text-to-semantic token model based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample; and update the parameters of the text-to-semantic token model based on the loss function value of the text-to-semantic token model.
[0029] In some embodiments, the text-to-semantic token model training module is used to obtain a first loss function value of the text-to-semantic token model based on the difference between the first playback duration sample and the second playback duration sample; obtain a second loss function value of the text-to-semantic token model based on the difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample; and determine the loss function value of the text-to-semantic token model based on the first loss function value of the text-to-semantic token model and the second loss function value of the text-to-semantic token model.
[0030] In some embodiments, the semantic token to acoustic token model includes a second converter; an input acoustic token acquisition module, which is used to obtain a prefix by combining the prompt semantic token, the input semantic token, and the prompt acoustic token in order; through the second converter, starting from the prefix, the acoustic features of the speech corresponding to the input text at each time point are predicted in a self-recursive manner to obtain the input acoustic token.
[0031] In some embodiments, the order of the prompt acoustic token and the input acoustic token is 2.
[0032] In some embodiments, the device also includes: a semantic token-to-acoustic token model training module, which is used to obtain a third audio sample and a fourth audio sample when the semantic token extractor and the acoustic token extractor are trained; the third audio sample and the fourth audio sample are two non-overlapping audio segments in the same audio; the semantic token label of the third audio sample and the semantic token label of the fourth audio sample are respectively extracted by the semantic token extractor; the acoustic token label of the third audio sample and the acoustic token label of the fourth audio sample are respectively extracted by the acoustic token extractor; the prefix sample is obtained by sequentially combining the semantic token label of the third audio sample, the semantic token label of the fourth audio sample, and the acoustic token label of the third audio sample; the acoustic token sample of the fourth audio sample is predicted in a self-recursive manner starting from the prefix sample through the second converter; the parameters of the semantic token-to-acoustic token model are updated based on the acoustic token sample of the fourth audio sample and the acoustic token label of the fourth audio sample.
[0033] According to another aspect of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the speech synthesis method described above.
[0034] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the speech synthesis method described above.
[0035] According to another aspect of the present application, a computer program product is provided, which includes computer instructions stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium to implement the speech synthesis method described above.
[0036] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0037] First, the input text and prompt audio are obtained; second, feature extraction is performed on the prompt audio to obtain prompt semantic tokens and prompt acoustic tokens, and feature extraction is performed on the input text to obtain input semantic tokens; then, based on the prompt semantic tokens, prompt acoustic tokens, and input semantic tokens, input acoustic tokens are obtained; finally, based on the input acoustic tokens, the output audio of the input text is obtained, thereby achieving rapid conversion from acoustic tokens to audio; through the above scheme, the processing of the input text and prompt audio is divided into two stages: first, the semantic tokens of the input text, the semantic tokens of the prompt audio, and the acoustic tokens of the prompt audio are obtained through the input text and prompt audio, and then the final decoded acoustic tokens are predicted through the above semantic tokens of the input text, the semantic tokens of the prompt audio, and the acoustic tokens of the prompt audio. The extraction process of the semantic tokens is introduced as a transition. This is conducive to reducing the feature span of each process in the prediction process from the input text and prompt audio to the final acoustic token, thereby improving the accuracy of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] FIG1 is a schematic diagram of a computer system for a speech synthesis method provided by an exemplary embodiment of the present application;
[0039] FIG2 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0040] FIG3 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0041] FIG4 is a flowchart of an implementation of a speech synthesis method provided by an exemplary embodiment of the present application;
[0042] FIG5 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0043] FIG6 is a schematic diagram of a semantic token extractor provided by an exemplary embodiment of the present application;
[0044] FIG7 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0045] FIG8 is a schematic diagram of a text-to-semantic token model provided by an exemplary embodiment of the present application;
[0046] FIG9 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0047] FIG10 is a schematic diagram of a semantic token-to-acoustic token model provided by an exemplary embodiment of the present application;
[0048] FIG11 is an exemplary training and reasoning flow chart of the speech synthesis system involved in this application;
[0049] FIG12 is a schematic diagram of an exemplary application scenario of the speech synthesis system involved in this application;
[0050] FIG13 is a block diagram of a speech synthesis apparatus according to an exemplary embodiment of the present application;
[0051] FIG14 is a structural block diagram of a computer device provided by an exemplary embodiment of the present application.
[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. DETAILED DESCRIPTION
[0053] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0054] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0055] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0056] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the attack operations and other object behaviors involved in this application were obtained with full authorization.
[0057] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter without departing from the scope of this disclosure. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0058] The following are some definitions of terms involved in this application:
[0059] Spectrograms: A spectrum represents a time-domain signal in the frequency domain. This spectrum is obtained by performing a Fourier transform on the signal. The resulting spectrum is a graph with amplitude and phase on the vertical axis and frequency on the horizontal axis. In speech synthesis applications, phase information is often omitted, retaining only the amplitude information corresponding to different frequencies.
[0060] Fundamental frequency: In sound, fundamental frequency refers to the frequency of the fundamental tone in a complex tone, represented by the symbol FO. Among the multiple tones that make up a complex tone, the fundamental tone has the lowest frequency and the greatest intensity. The high or low fundamental frequency determines the pitch of a sound. When we refer to the frequency of speech, we generally refer to the frequency of the fundamental tone.
[0061] Vocoder: Derived from the abbreviation of voice encoder, also known as speech signal analysis and synthesis system, its function is to convert acoustic features into sound.
[0062] Hidden Markov Model (HMM): A statistical analysis model used to describe a Markov process with hidden unknown parameters. In an HMM, the state is not directly visible, but some variables (observations) affected by the state are visible.
[0063] Deep Neural Network (DNN): A discriminative model, a multilayer perceptron (MLP) with more than two hidden layers. Except for the input node, each node is a neuron with a nonlinear activation function. Like MLP, DNN can be trained using the backpropagation algorithm.
[0064] Convolutional Neural Network (CNN): A feedforward neural network whose neurons respond to units within their receptive field. CNNs typically consist of multiple convolutional layers topped by a fully connected layer. By sharing parameters, they reduce the number of model parameters, making them widely used in image and speech recognition.
[0065] Recurrent Neural Network (RNN): A type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain-like manner.
[0066] Long Short-Term Memory (LSTM) is a recurrent neural network that incorporates a cell in its algorithm to determine whether information is useful. Each cell contains an input gate, a forget gate, and an output gate. After information enters the LSTM, it is judged based on rules to determine whether it is useful. Only information that meets the algorithm's criteria is retained; information that does not meet these criteria is forgotten through the forget gate. This network is suitable for processing and predicting important events in time series with relatively long intervals and delays.
[0067] The Gate Recurrent Unit (GRU) is a type of recurrent neural network. Like the LSTM, it was designed to address issues such as long-term memory and gradients in backpropagation. Compared to the LSTM, the GRU lacks an internal gate and has fewer parameters. In most cases, it can achieve comparable performance to the LSTM while significantly reducing computational time.
[0068] Loss function: Also known as the cost function, this function is used to evaluate the difference between the predicted value and the true value of a neural network model. The smaller the loss function value, the better the performance of the neural network model. The model training process is the process of minimizing the loss function value by adjusting the model parameters. Different neural network models use different loss functions. Common loss functions include 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, perceptual loss function, cross entropy loss function, KL divergence loss function, triplet loss function, and so on.
[0069] Text to Speech (TTS): also known as text-to-speech, its function is to convert computer-generated or externally input text information into understandable and fluent speech and read it aloud.
[0070] With the rapid development of smart devices (such as smartphones and smart speakers), voice interaction technology is increasingly being used as a natural form of interaction. As a key component of voice interaction technology, speech synthesis technology has also made significant progress. In recent years, large language models based on semi-supervised learning have achieved significant success in natural language processing tasks.
[0071] Semi-supervised learning utilizes a large amount of unlabeled data for pre-training, and then uses a small amount of labeled data for fine-tuning or training specific modules. Semi-supervised learning lies between unsupervised learning (where all training data is unlabeled) and supervised learning (where all training data is labeled), effectively alleviating the problem of limited labeled training data.
[0072] Please refer to FIG1 , which shows a schematic diagram of a computer system for a speech synthesis method provided by an exemplary embodiment of the present application. The computer system may include: a terminal device 110 and a server 120 .
[0073] The terminal device 110 is an electronic device provided with a speech synthesis function.
[0074] The terminal device 110 includes but is not limited to a smart phone, a tablet computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal device, a laptop computer or a desktop computer, etc.
[0075] The terminal device 110 may run a client that provides a speech synthesis function. The client may be an instant messaging application, a music player application, a reading application, etc. The embodiment of the present application does not limit the specific type of the client.
[0076] The server 120 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. In the embodiment of the present application, the server is a backend server that provides a speech synthesis client in the terminal device 110 and can convert text into speech.
[0077] Among them, data communication is carried out between the terminal device 110 and the server 120 through a communication network. In some embodiments, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network and a wide area network.
[0078] In the method provided in the embodiments of the present application, the execution entity of each step may be a computer device. The computer device may be any electronic device capable of storing and processing data. For example, the computer device may be the terminal device 110 in Figure 1 or the server 120.
[0079] Please refer to Figure 2, which shows a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. The method is performed by a computer device. Optionally, the computer device can be the server 120 or terminal device 110 in the system shown in Figure 1, or the computer device can be another electronic device with computing capabilities. As shown in Figure 2, the method can include at least one of the following steps 210, 220, 230, 240, and 250.
[0080] Step 210: Obtain input text and prompt audio.
[0081] In some embodiments, a computer device obtains input text and prompt audio input by a terminal device, where the input text contains text content of the output audio (or output voice) that the terminal device wants to synthesize, and the prompt audio is an audio (or voice) that contains sound information such as the timbre, rhythm, and emotion of the terminal device user.
[0082] For example, the output audio is the dubbing of a video, the input text is the complete dubbing text, and the prompt audio can be a 10-second dubbing.
[0083] For example, the output audio is an 800-word poem recitation, and the input text is a complete 800-word poem text. The prompt audio can be a 5-second poem recitation with 15 words.
[0084] Step 220: Extract the features of the prompt audio and obtain a prompt semantic token and a prompt acoustic token. The prompt semantic token is used to indicate the semantic features of the prompt audio at each time point, and the prompt acoustic token is used to indicate the acoustic features of the prompt audio at each time point.
[0085] In some embodiments, the computer device performs feature extraction on the prompt audio obtained in step 210 using a pre-trained extraction model to obtain a prompt semantic token corresponding to the prompt audio.
[0086] The prompt semantic token is used to indicate the semantic features of the prompt audio at each time point.
[0087] In some embodiments, the prompt semantic token may be a serial number for encoding a semantic unit corresponding to a text contained in the prompt audio, where the semantic unit is the smallest semantic object in a semantic codebook.
[0088] For example, a 1-second prompt audio is converted into 50 prompt semantic tokens after inference by the above extraction model.
[0089] In some embodiments, the computer device performs feature extraction on the prompt audio obtained in step 210 using a pre-trained extraction model to obtain a prompt acoustic token.
[0090] Among them, the prompt acoustic token is used to indicate the acoustic characteristics of the prompt audio at each time point. In some embodiments, each time point is a time stamp in the prompt audio. Exemplarily, based on the length of the prompt audio, the duration interval of the prompt audio is determined. Exemplarily, a time stamp is set every threshold length in the duration interval, and each time stamp set in the duration interval is considered to be a time point here. Exemplarily, the threshold length is 1s. For the time points in other positions described below, please refer to the explanation here and will not be repeated here.
[0091] In some embodiments, the prompt acoustic token may be a serial number for encoding a sound unit corresponding to a sound contained in the prompt audio, where a sound unit is the smallest sound object in a sound codebook.
[0092] For example, a 1-second 24kHz prompt audio is converted into 2×75 prompt acoustic tokens after inference by the above extraction model.
[0093] Step 230: extracting features of the input text and obtaining input semantic tokens, where the input semantic tokens are used to indicate semantic features of the speech corresponding to the input text at various time points.
[0094] In some embodiments, the computer device obtains input semantic tokens from the input text obtained in step 210 using a pre-trained extraction model.
[0095] The input semantic token is used to indicate the semantic features of the speech corresponding to the input text at each time point.
[0096] In some embodiments, the input semantic token may be a serial number for encoding a semantic unit corresponding to the input text, where a semantic unit is the smallest semantic object in a semantic codebook.
[0097] For example, an input text of thousands of words is converted into tens of thousands of input semantic tokens after being inferred by the above extraction model.
[0098] Step 240: Based on the prompt semantic token, the prompt acoustic token and the input semantic token, an input acoustic token is obtained; the input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point.
[0099] In some embodiments, the computer device processes and infers the prompt semantic token obtained in step 220, the prompt acoustic token obtained in step 220, and the input semantic token obtained in step 230 through a pre-trained conversion model to predict the input acoustic token.
[0100] The input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point.
[0101] In some embodiments, the input acoustic token may be a sequence number for encoding a sound unit corresponding to the input text, where a sound unit is the smallest sound object in a sound codebook.
[0102] Step 250: Based on the input acoustic tokens, obtain output audio of the input text.
[0103] In some embodiments, when obtaining the output audio of the input text based on the input acoustic token, the computer device can decode the input acoustic token obtained in step 240 through a pre-trained decoder to convert the input acoustic token into the output audio corresponding to the input text.
[0104] The sound information such as timbre, rhythm, emotion, etc. in the above-mentioned output audio comes from the prompt audio, and the voice content in the above-mentioned output audio comes from the input text.
[0105] In summary, in the embodiment of the present application, the computer device first obtains the input text and the prompt audio; secondly, it performs feature extraction on the prompt audio to obtain the prompt semantic token and the prompt acoustic token, and performs feature extraction on the input text to obtain the input semantic token; then, based on the prompt semantic token, the prompt acoustic token and the input semantic token, the input acoustic token is obtained; finally, based on the input acoustic token, the output audio of the input text is obtained to achieve rapid conversion from acoustic token to audio; through the above scheme, the processing process of the input text and the prompt audio is divided into two stages, first, the semantic token of the input text, the semantic token of the prompt audio and the acoustic token of the prompt audio are obtained through the input text and the prompt audio, and then the final decoded acoustic token is predicted through the above semantic token of the input text, the semantic token of the prompt audio and the acoustic token of the prompt audio, and the extraction process of the semantic token is introduced as a transition. This is conducive to reducing the feature span of each process in the prediction process from the input text and the prompt audio to the final acoustic token. Reduce the model's requirements for the amount and quality of labeled data, so that it can be trained with the help of a large amount of unlabeled data and a small amount of labeled data, thereby ensuring the accuracy of the model and improving the accuracy of speech synthesis.
[0106] By adopting the solution provided in the embodiment of the present application, it is possible to mine the semantics, timbre, rhythm, emotion and other information in the prompt audio; at the same time, based on the transition of obtaining semantic tokens from text, it is possible to alleviate the one-to-many problem faced when directly obtaining acoustic tokens from text, thereby achieving the purpose of zero-time synthesis of speech through prompt audio.
[0107] Based on the embodiment shown in FIG2 , please refer to FIG3 , which shows a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. As shown in FIG3 , step 220 in the embodiment shown in FIG2 can be implemented as at least one of step 220a1 and step 220a2, step 230 can be implemented as step 230a, step 240 can be implemented as step 240a, and step 250 can be implemented as at least one of step 250a.
[0108] Step 220a1: input the prompt audio into the semantic token extractor to obtain the prompt semantic token obtained by the semantic token extractor processing the prompt audio.
[0109] Among them, the above-mentioned voice token extraction module can be a machine learning model that is pre-trained through audio samples in an unsupervised learning manner. Its function is to extract the semantic features of the voice content in the audio from the input audio and obtain the corresponding semantic tokens.
[0110] In some embodiments, the semantic token extractor is a machine learning model for extracting semantic features from audio. Exemplarily, the semantic token extractor is a trained machine learning model for extracting semantic features from audio. In some embodiments, the semantic token extractor takes as input the prompt audio and outputs the prompt semantic tokens.
[0111] Step 220a2: Input the prompt audio into the acoustic token extractor to obtain the prompt acoustic token obtained by the acoustic token extractor processing the prompt audio.
[0112] The acoustic token extractor can be a machine learning model pre-trained using audio samples in an unsupervised learning manner. Its function is to extract acoustic features from the input audio and generate corresponding acoustic tokens. The acoustic features can include semantics, timbre, emotion, rhythm, and other features.
[0113] In some embodiments, the acoustic token extractor is a machine learning model for extracting acoustic features from audio. Exemplarily, the acoustic token extractor is a trained machine learning model for extracting acoustic features from audio. In some embodiments, the acoustic token extractor takes as input the prompt audio and outputs the prompt acoustic tokens.
[0114] Step 230a: Input the input text to the text-to-semantic token model, and obtain input semantic tokens obtained by the text-to-semantic token model processing the input text.
[0115] Among them, the above-mentioned text-to-semantic token model can be a machine learning model trained in a supervised learning manner using a trained semantic token extractor and labeled audio samples. The function of the text-to-semantic token model is to predict the semantic features of the input text at various time points after the text is converted into speech, and obtain the corresponding semantic tokens.
[0116] In some embodiments, the text-to-semantic tokenization model is a machine learning model for extracting semantic features from text. Exemplarily, the text-to-semantic tokenization model is a machine learning model for extracting semantic features from text. In some embodiments, the input of the text-to-semantic tokenization model is input text, and the output is input semantic tokens.
[0117] Step 240a: Input the prompt semantic token, the prompt acoustic token, and the input semantic token into the semantic token-to-acoustic token model to obtain the input acoustic token output by the semantic token-to-acoustic token model.
[0118] Among them, the above-mentioned semantic token to acoustic token model is a machine learning model trained in an unsupervised learning manner using the trained semantic token extractor, acoustic token extractor, and audio samples. Its function is to predict the acoustic token corresponding to another semantic token through the semantic token and acoustic token of the same audio segment and another semantic token.
[0119] In some embodiments, the semantic token-to-acoustic token model is a machine learning model for converting semantic features into acoustic features. Exemplarily, the semantic token-to-acoustic token model is a trained machine learning model for converting semantic features into acoustic features. In some embodiments, the input of the semantic token-to-acoustic token model is a prompt semantic token, a prompt acoustic token, and an input semantic token, and the output is an input acoustic token.
[0120] Step 250a: Input the input acoustic token into the sound decoder to obtain the output audio output by the sound decoder.
[0121] Among them, the above-mentioned sound decoder can be a trained acoustic token extractor and an unlabeled audio sample. The machine learning model trained in an unsupervised learning manner decodes the input acoustic token to generate the audio corresponding to the acoustic token.
[0122] In an embodiment of the present application, a computer device obtains a prompt semantic token of the prompt audio through a semantic token extractor, obtains a prompt acoustic token of the prompt audio through an acoustic token extractor, and obtains an input semantic token of the input text through a text-to-semantic token model; then, through the semantic token-to-acoustic token model, based on the prompt semantic token, the prompt acoustic token and the input semantic token, the input acoustic token of the input text is obtained; finally, through the sound decoder, the input acoustic token is converted into sound, and the output audio corresponding to the input text is obtained, thereby realizing a rapid conversion from acoustic token to audio, thereby providing a two-stage conversion scheme from input text and prompt audio to semantic tokens and then to acoustic tokens through a machine learning model. Through the above scheme, the semantic token extractor, acoustic token extractor and text-to-semantic token model can be used to mine the semantic, timbre, rhythm and emotion information in the prompt audio; at the same time, using text to predict semantic tokens can alleviate the one-to-many problem faced when directly predicting acoustic tokens from text, thereby achieving the purpose of zero-time synthesis of speech through prompt audio.
[0123] Please refer to Figure 4, which shows a flowchart of an implementation of a speech synthesis method provided by an exemplary embodiment of the present application. As shown in Figure 4, the specific process is as follows:
[0124] After the computer device obtains the prompt audio 301, it inputs the prompt audio 301 into the semantic token extractor 310. After the semantic token extractor 310 infers the prompt audio 301, it outputs the prompt semantic token 303 corresponding to the prompt audio 301.
[0125] After the computer device obtains the prompt audio 301, it inputs the prompt audio 301 into the acoustic token extractor 320. After the acoustic token extractor 320 infers the prompt audio 301, it outputs the prompt acoustic token 304 corresponding to the prompt audio 301.
[0126] After the computer device obtains the input text 302, the input text 302 is input into the text-to-semantic token model 330. After the text-to-semantic token model 330 infers the input text 302, the input semantic token 305 corresponding to the input text 302 is output;
[0127] The computer device inputs the obtained prompt semantic token 303, prompt acoustic token 304 and input semantic token 305 into the semantic token to acoustic token model 340. After the semantic token to acoustic token model 340 infers the prompt semantic token 303, prompt acoustic token 304 and input semantic token 305, it outputs the input acoustic token 306 corresponding to the input text 302.
[0128] The computer device inputs the input acoustic token 306 obtained above into the sound decoder 350 . After the sound decoder 350 infers the input acoustic token 306 , it outputs the output audio 307 corresponding to the input text 302 .
[0129] Based on the embodiment shown in FIG3 , please refer to FIG5 , which illustrates a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. As shown in FIG5 , the semantic token extractor includes a convolution branch and a first transformer. Step 220a1 in the embodiment shown in FIG3 can be implemented as at least one of step 220a1-1, step 220a1-2, and step 220a1-3.
[0130] Step 220a1-1: Input the prompt audio into the convolution branch to obtain the hidden layer features of the prompt audio at each time point output by the convolution branch.
[0131] In some embodiments, the convolution branch is a neural network layer for implementing a convolution operation. Exemplarily, the convolution branch includes at least one convolution layer. Exemplarily, different convolution layers correspond to different convolution kernels. Exemplarily, the at least one convolution layer implements a convolution process that first upsamples and then downsamples input features.
[0132] In an embodiment of the present application, for the prompt audio, the semantic token extractor can extract features thereof through a convolutional layer to obtain hidden layer features at each time point.
[0133] Step 220a1-2: Process the hidden layer features of the prompt audio at each time point through the first converter to obtain the intermediate layer features of the prompt audio at each time point output by the intermediate layer of the first converter.
[0134] In the embodiment of the present application, the intermediate layer features of the prompt audio at each time point output by the intermediate layer of the above-mentioned first converter may refer to the features output by a specified layer in the first converter.
[0135] Alternatively, the above-mentioned intermediate layer features can also be replaced by the features finally output by the first converter.
[0136] In some embodiments, the first converter is a neural network model for implementing a conversion function, and the neural network model includes multiple neural network layers. In some embodiments, the above-mentioned first converter can be a Transformer network. In some embodiments, the intermediate layer of the first converter refers to the output of any one of the neural network layers included in the Transformer network. Of course, the first converter can also be other neural networks other than the Transformer network, including but not limited to at least one of the BERT network and the U-net network. In some embodiments, the intermediate layer of the first converter can be specified in advance. In some embodiments, the target neural network layer among the multiple neural network layers in the Transformer network can be specified in advance as the intermediate layer of the Transformer network. In some embodiments, the first converter and the second converter described below are the same or different converters.
[0137] Step 220a1-3: Cluster the intermediate layer features of the prompt audio at each time point to obtain prompt semantic tokens.
[0138] In an embodiment of the present application, for the intermediate layer features output by the first converter, for the intermediate layer features at each time point, the semantic category to which the intermediate layer features corresponding to the time point belong is determined by feature clustering, thereby determining the semantic token corresponding to the time point, and then obtaining a prompt semantic token.
[0139] The above solution provides a solution for extracting semantic tokens by clustering audio after feature extraction, ensuring the feasibility of semantic token extraction through the model.
[0140] In some embodiments, the above method further comprises:
[0141] Obtain a first audio sample and a semantic token tag for the first audio sample. In some embodiments, the semantic token tag for the first audio sample is extracted using a pre-trained semantic token extractor. Exemplarily, the semantic token tag for the first audio sample is obtained by clustering Mel-frequency cepstral features of the first audio sample.
[0142] Exemplarily, the first audio sample is an acquired audio segment, and the audio segment is used as a sample to obtain the first audio sample.
[0143] Inputting the first audio sample into the convolution branch to obtain hidden feature samples of the first audio sample at each time point output by the convolution branch;
[0144] partially masking the hidden feature samples of the first audio sample at each time point to obtain partially masked hidden feature samples;
[0145] Exemplarily, partial masking can also be considered as partial masking. Exemplarily, by partially masking the hidden feature samples of the first audio sample at each time point, the diversity of the hidden feature samples is increased, thereby improving the training effect of the model.
[0146] Processing the partially masked hidden layer feature samples by the first converter to obtain intermediate layer features of the first audio sample at each time point output by the intermediate layer of the first converter;
[0147] Clustering the intermediate layer features of the first audio sample at each time point to obtain a semantic token sample of the first audio sample;
[0148] Parameters of a semantic token extractor are updated based on the semantic token sample of the first audio sample and the semantic token label of the first audio sample.
[0149] In some embodiments, a loss function value of a semantic token extractor is obtained based on a semantic token sample of the first audio sample and a semantic token label of the first audio sample; the semantic token label of the first audio sample is obtained by clustering the Mel-cepstral features of the first audio sample;
[0150] In some embodiments, a loss function value for the semantic token extractor is determined based on a difference between the semantic token sample of the first audio sample and the semantic token label of the first audio sample.
[0151] Based on the loss function value of the semantic token extractor, the parameters of the semantic token extractor are updated.
[0152] Exemplarily, the parameters of the semantic token extractor are updated with the goal of minimizing the loss function value. The present application does not limit the specific category of the loss function, such as the loss function is cross entropy loss, 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, perceptual loss function, etc. Exemplarily, the parameters of each module in the semantic token extractor are updated with the goal of minimizing the loss function value. Exemplarily, the parameters of the target module in each module in the semantic token extractor are updated with the goal of minimizing the loss function value, such as the target module is the convolution branch or the first converter. Exemplarily, the parameters of the first converter are kept unchanged, and only the parameters of the convolution branch are updated. In this way, the training cost can be reduced and the training efficiency can be improved.
[0153] In an embodiment of the present application, when the computer device trains the semantic token extractor, it can extract the mel-cepstral features of the first audio sample, and then determine the semantic token label of the first audio sample by clustering the mel-cepstral features of the first audio sample. The hidden feature samples output by the convolution branch are partially masked, and the partially masked hidden feature samples are predicted by the first converter, and the loss is calculated with the semantic token label of the first audio sample. In this way, the convolution branch and the first converter are trained in their ability to extract semantic features, thereby providing a solution for unsupervised learning of the semantic token extractor through unlabeled audio, which does not rely on labeled data, reduces the requirements for training data, and ensures the accuracy of the model.
[0154] Please refer to Figure 6, which shows a schematic diagram of a semantic token extractor provided by an exemplary embodiment of the present application. As shown in Figure 6, the semantic token extractor is composed of a CNN-based convolution module 610 and a Transformer module 620.
[0155] The convolution module 610 downsamples the input audio 601 and outputs X n Hidden layer representation; Transformer module 620 pairs of X n Hidden layer representation is predicted, and Z is obtained n Predicted labels.
[0156] For example, the convolution module 610 converts one second of audio into 50 frames of hidden layer representation with a dimension of D; the Transformer module 620 predicts the input 50 frames of hidden layer representation to obtain 50 predicted labels.
[0157] When training the semantic token extractor, a large amount of unlabeled data can be used for training. Among them, the original audio 601 is used as the input of the convolution module 610. After the convolution module 610 processes the input audio 601, the output of the convolution module 610 is randomly masked (masked) and then input into the Transformer module 620. The Transformer module 620 is required to be able to predict the label of the missing part according to the context when the input is missing, so as to enhance the context capture ability of the model. Among them, the Mel-scale Frequency Cepstral Coefficients (MFCC) can be extracted from the original audio 601 and then unsupervised K-mean clustering 630 can be performed to obtain the corresponding label and the predicted label to construct a loss function, and the parameters of the semantic token extractor are updated.
[0158] When extracting semantic tokens through semantic token extractor inference, audio 601 is input, downsampled by convolution module 610, and directly input into Transformer module 620. The intermediate layer features of Transformer module 620 are obtained for clustering, and the category obtained by clustering each frame is used as the semantic token of the frame.
[0159] For example, one second of audio is converted into 50 frames of hidden layer representation after passing through the convolution module 610. This is then input into the Transformer module 620, where the Lth layer output (also 50 frames) is taken for K-class clustering 630. If the clustering result for the first frame belongs to the third class, the semantic token for that frame is 3. In summary, one second of audio is converted into 50 semantic tokens.
[0160] Based on the embodiments shown in Figures 3 or 5, please refer to Figure 7, which shows a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. As shown in Figure 7, the text-to-semantic tokenization model includes a text encoder, a duration predictor, an upsampling branch, and a decoder. Step 230a in the embodiment shown in Figure 3 can be implemented as steps 230a1, 230a2, 230a3, and 230a4.
[0161] Step 230a1: Input the input text to the text encoder to obtain a hidden text encoding representation of the input text.
[0162] In an embodiment of the present application, the text-to-semantic token model first encodes the input text through a text encoder to obtain a hidden text encoding representation, which can be a feature vector or feature matrix of the input text.
[0163] In some embodiments, the text encoder is a neural network model (or neural network unit) for encoding text. In some embodiments, the text encoder is a trained neural network model (or neural network unit) for encoding text.
[0164] Step 230a2: Input the hidden text encoding representation into the duration predictor to obtain the playback duration of the speech corresponding to the input text predicted by the duration predictor.
[0165] In an embodiment of the present application, the text-to-semantic token model encodes and represents the hidden text through a duration predictor to predict the playback duration of the speech converted from the input text, so as to subsequently determine the length / number of semantic tokens to be predicted based on the predicted playback duration.
[0166] In some embodiments, the duration predictor is a neural network model (or neural network unit) for predicting duration. In some embodiments, the duration predictor is a trained neural network model (or neural network unit) for predicting duration.
[0167] Step 230a3: up-sample the hidden text encoding representation to the number of frames corresponding to the playback duration through the up-sampling branch to obtain the up-sampled hidden text encoding representation.
[0168] In some embodiments, the upsampling branch is a neural network model (or neural network unit) for encoding. In some embodiments, the upsampling branch is a trained neural network model (or neural network unit) for encoding.
[0169] In an embodiment of the present application, after predicting the playback duration of the speech corresponding to the input text, the text-to-semantic token model upsamples the hidden text encoding representation through an upsampling branch, so that the number of frames corresponding to the hidden text encoding representation is aligned with the playback duration of the speech corresponding to the input text, so that the upsampled hidden text encoding representation can be used to predict the number of semantic tokens that match the playback duration of the speech corresponding to the input text.
[0170] Step 230a4: Decode the upsampled hidden text encoding representation through the decoder to obtain the input semantic token.
[0171] In an embodiment of the present application, after obtaining the upsampled hidden text encoding representation, the text-to-semantic token model decodes the upsampled hidden text encoding representation through a decoder to obtain a number of input semantic tokens that match the playback duration of the speech corresponding to the input text.
[0172] In some embodiments, the decoder is a neural network model (or neural network unit) for decoding. In some embodiments, the decoder is a trained neural network model (or neural network unit) for decoding.
[0173] Through the scheme shown in the above embodiment of the present application, the text representation is converted into a series of semantic tokens through the sequential processing of the text encoder, duration predictor, upsampling branch and decoder. The number of semantic tokens is aligned with the playback duration of the speech converted from the input text, thereby ensuring that the input semantic tokens can subsequently match the length of the audio to be generated, thereby ensuring the accuracy of the semantic tokens extracted from the text.
[0174] In some embodiments, the above method further comprises:
[0175] When the semantic token extractor is trained, a second audio sample and a speech text of the second audio sample are obtained; the second audio sample is input into the semantic token extractor to obtain a semantic token label of the second audio sample output by the semantic token extractor; wherein the semantic token label refers to a semantic token extracted from the second audio sample;
[0176] Inputting the speech text of the second audio sample into the text-to-semantic token model to obtain a semantic token sample of the second audio sample output by the text-to-semantic token model;
[0177] Based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample, parameters of the text-to-semantic token model are updated.
[0178] Exemplarily, the loss function value is determined based on the difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample. With the goal of minimizing the loss function value, the parameters of the text-to-semantic token model are updated. The present application does not limit the specific category of the loss function, such as the loss function is a cross entropy loss, a 0-1 loss function, an absolute value loss function, a logarithmic loss function, an exponential loss function, a perceptual loss function, and the like. Exemplarily, with the goal of minimizing the loss function value, the parameters of each module in the text-to-semantic token model are updated. Exemplarily, with the goal of minimizing the loss function value, the parameters of the target module in each module in the text-to-semantic token model are updated, such as the target module is at least one of the text encoder, the duration predictor, the upsampling branch, and the decoder. Exemplarily, the parameters of the text encoder and the decoder are kept unchanged, and only the parameters of the duration predictor and the upsampling branch are updated. In this way, the training cost can be reduced and the training efficiency can be improved.
[0179] In an embodiment of the present application, the above-mentioned text-to-semantic token model is trained by supervised learning with the help of a trained semantic token extractor and audio annotated with text (that is, the second audio sample corresponding to the voice text, wherein the voice text is the annotated text, and the voice text can be manually pre-annotated), so as to ensure the accuracy of the text-to-semantic token model. In the above-mentioned supervised learning, the semantic tokens used as labels are extracted by the semantic token extractor from the audio annotated with text.
[0180] In some embodiments, the process of inputting the speech text of the second audio sample into the text-to-semantic token model to obtain the semantic token sample of the second audio sample output by the text-to-semantic token model can be the same as the above steps 230a1 to 230a4 and will not be repeated here.
[0181] In some embodiments, inputting the speech text of the second audio sample into a text-to-semantic token model to obtain a semantic token sample of the second audio sample output by the text-to-semantic token model includes:
[0182] Inputting the speech text of the second audio sample into a text encoder to obtain a hidden text encoding representation sample of the speech text of the second audio sample;
[0183] Inputting the hidden text encoding representation sample into a duration predictor to obtain a first playback duration sample of the speech corresponding to the speech text of the second audio sample predicted by the duration predictor;
[0184] Inputting the hidden text encoding representation sample into the attention branch, obtaining a second playback duration sample of the speech corresponding to the speech text of the second audio sample output by the attention branch;
[0185] Upsampling the hidden text encoding representation sample to the number of frames corresponding to the second playback duration sample through the upsampling branch to obtain an upsampled hidden text encoding representation sample;
[0186] Decoding the upsampled hidden text encoding representation sample through a decoder to obtain a semantic token sample of the second audio sample;
[0187] Obtaining a loss function value of a text-to-semantic token model based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample, including:
[0188] A loss function value of a text-to-semantic token model is obtained based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample.
[0189] In some embodiments, during the training process, an auxiliary learning network module, that is, the above-mentioned attention branch, can be introduced into the text-to-semantic token model to assist in the prediction of the playback time through the attention branch. Specifically, during the training process, the speech text of the second audio sample is input into the text encoder, and after obtaining the hidden text encoding representation sample of the speech text of the second audio sample, the hidden text encoding representation sample is input into the duration predictor to obtain the first playback time sample predicted by the duration predictor. At the same time, the hidden text encoding representation sample is also input into the attention branch, and the second playback time sample is predicted by the attention prediction branch. Subsequently, after the second playback time sample and the hidden text encoding representation sample are input into the upsampling branch for upsampling, the semantic token sample of the second audio sample is predicted by the decoder. When calculating the loss function, the first playback time sample, the second playback time sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample are used for calculation at the same time, which expands the available loss and improves the accuracy of model training.
[0190] In some embodiments, obtaining a loss function value of a text-to-semantic token model based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample includes:
[0191] Obtaining a first loss function value of a text-to-semantic token model based on a difference between the first playback duration sample and the second playback duration sample;
[0192] A second loss function value of the text-to-semantic token model is obtained based on a difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample.
[0193] Based on the first loss function value of the text-to-semantic token model and the second loss function value of the text-to-semantic token model, a loss function value of the text-to-semantic token model is determined.
[0194] Exemplarily, the sum of the first loss function value and the second loss function value is directly used as the loss function value of the text-to-semantic token model. Exemplarily, the first loss function value and the second loss function value of the text-to-semantic token model are weighted and summed to obtain the loss function value of the text-to-semantic token model. Exemplarily, the weights of the first loss function value and the second loss function value can be set in advance.
[0195] For example, when calculating the loss function value of the text-to-semantic token model, the computer device can calculate the difference between the first playback time sample and the second playback time sample through a preset loss function to obtain the above-mentioned first loss function value.
[0196] Similarly, the computer device may calculate the difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample using a preset loss function to obtain the second loss function value.
[0197] Among them, the above-mentioned first loss function value can be used to update the parameters of the duration predictor, or can be used to update the parameters of the duration predictor and the text encoder; the above-mentioned second loss function value can be used to update the parameters of the text encoder, attention branch, upsampling branch and decoder.
[0198] During model training, the computer device can use the difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample to update the text encoder, attention branch, upsampling branch and decoder, so that the accuracy of the attention branch can gradually increase with the training process. At the same time, the second playback duration sample output by the attention branch is used as a label for duration predictor training, and the difference between the second playback duration sample and the second playback duration sample output by the duration predictor itself is calculated to update the parameters of the duration predictor, or the duration predictor and the text encoder. The prediction ability of the duration predictor is close to that of the attention branch, thereby achieving the simultaneous use of the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample for calculation, expanding the available loss, and thus improving the accuracy of model training.
[0199] In addition, the network complexity of the above-mentioned duration predictor can be lower than the network complexity of the attention branch. That is to say, during the model training process, the duration is predicted through an attention branch with higher complexity to ensure the accuracy of the duration prediction. At the same time, through the first loss function, the duration predictor can learn the prediction ability of the attention branch with higher complexity to ensure the accuracy of the duration predictor. At the same time, since the network complexity of the duration predictor is lower, the efficiency of duration prediction can be improved in the subsequent reasoning process.
[0200] Please refer to FIG8 , which shows a schematic diagram of a text-to-semantic token model provided by an exemplary embodiment of the present application.
[0201] After the semantic token extractor is trained, it can extract semantic tokens from any audio with text annotations (in this application's technical solution, only this module undergoes supervised training) and train the text-to-semantic token prediction module. As shown in Figure 8, considering both training convenience and inference efficiency, the text-to-semantic token model primarily consists of five components: a text encoder 810, a duration predictor 820, an upsampling module 830, a parallel decoder 840, and an attention module 850.
[0202] Text encoder 810: Encodes the input text 801 to obtain a hidden text encoding representation 802. The required synthetic text (such as "I am customer service Amy, employee number 1001, happy to serve you.") is preprocessed to obtain a regular text representation (such as pinyin), and the regular text representation is input into the text encoder 810. The specific structure of the text encoder 810 can be an RNN-based CBHG encoder (Tacotron) or a Transformer block-based encoder (Fastspeech). The text encoder 810 abstracts the regularized text representation layer by layer into a hidden text encoding representation 802 for use by subsequent modules.
[0203] Duration predictor 820: Inputs hidden text encoding representation 802 and predicts the predicted pronunciation duration 803 of each hidden text encoding representation 802. Due to the length difference between the text to be synthesized and the final acoustic feature (which can be understood as the different pronunciation durations of each word, and the corresponding number of acoustic feature frames), duration predictor 820 is required to predict the number of acoustic feature frames (or pronunciation duration) corresponding to each hidden text representation, so that the hidden text representation can be upsampled to the corresponding number of frames. The specific structure of duration predictor 820 can be a pure CNN network or a CNN+RNN network.
[0204] Upsampling module 830: according to the predicted duration 803 of the duration predictor 820, the hidden text encoding representation 802 is expanded to the corresponding number of frames (for example, if the predicted duration of a hidden text representation is 5, it will be copied 5 times).
[0205] Parallel decoder 840: The input of parallel decoder 840 is the upsampled hidden text representation. Through multiple nonlinear transformations, it finally obtains the input semantic token 804 corresponding to the text to be synthesized. Among them, parallel decoder 840 can be a Transformer structure or a pure CNN structure.
[0206] In this embodiment of the present application, the same text 801 can be input into a trained semantic token extractor to obtain semantic token labels corresponding to the text 801. Based on the semantic token labels and the input semantic tokens 804, a semantic token loss is determined. Based on the semantic token loss, the parallel decoder 840, upsampling module 830, duration predictor 820, and text encoder 810 are trained.
[0207] Attention module 850: Contains two parts: an attention mechanism 8501 and an auxiliary decoder 8502. The attention mechanism 8501 can be any common attention mechanism, such as the location-sensitive attention mechanism used in Tacotron or the Gaussian Mixture Model (GMM)-based attention mechanism, which determines which hidden text representations will be used in each decoding step. The auxiliary decoder 8502 can be a two-layer RNN structure. The alignment matrix between the hidden text encoding representation 802 and the acoustic features is obtained through the attention module 850 and converted into the corresponding duration information 805 (number of acoustic feature frames) for each input text.
[0208] The attention module 850 is used only during training. Its primary function is to obtain duration information 805 of the hidden text encoded representation 802. This duration information 805 is used as a label for training the duration predictor 820 (a process known as distillation, transferring the duration prediction capability learned by the attention module 850 to the duration predictor 820). Furthermore, this duration information 805 is input into the upsampling module 830 to upsample the hidden text encoded representation 802. During testing, the duration predictor 820 is used to directly predict the duration information 805, and the output of the text encoder 810 is then upsampled.
[0209] In the embodiment of the present application, the duration prediction loss can be determined based on the predicted duration 803 and the duration information 805; and the duration predictor 820 and the text encoder 810 can be trained based on the duration prediction loss.
[0210] In summary, as shown in Figure 8, the training process of the text-to-semantic token prediction module is as follows:
[0211] After the computer device obtains the text 801, it outputs the hidden text encoding representation 802 corresponding to the text 801 through the text encoder 810. The hidden text encoding representation 802 will be sent to the duration predictor 820, the upsampling module 830 and the attention module 850 respectively, so that the attention mechanism 8501 in the attention module 850 determines the alignment matrix and attention weight between the hidden text encoding representation 802 and the semantic token label based on the hidden text encoding representation 802 and the semantic token label corresponding to the text 801.
[0212] The attention module 850 then determines the duration information 805 corresponding to the hidden text encoding representation 802 based on the alignment matrix, and the auxiliary decoder 8502 obtains the semantic token 806 based on the attention weight, the hidden text encoding representation 802 and the semantic token label.
[0213] The duration information 805 determined by the attention mechanism 8501 is sent to the duration predictor 820 and the upsampling module 830. The duration predictor 820 generates a predicted duration 803 based on the hidden text encoded representation 802. The upsampling module 830 upsamples the hidden text encoded representation 802 based on the duration information 805 to obtain an extended hidden text representation. The parallel decoder 840 then decodes the extended hidden text representation to obtain an input semantic token 804.
[0214] Finally, the computer device determines the duration prediction loss based on the duration information 805 and the predicted duration 803, and determines the semantic token prediction loss based on the semantic token label and the semantic token 806; determines the second semantic token prediction loss based on the semantic token label and the input semantic token 804, and then based on these three losses, trains the text encoder 810, duration predictor 820, attention module 850 and parallel decoder 840 in an end-to-end manner, and constructs a text-to-semantic token prediction module based on the trained text encoder 810, duration predictor 820 and parallel decoder 840.
[0215] Based on the embodiments shown in Figures 3, 5, or 7, please refer to Figure 9, which shows a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. As shown in Figure 9, the semantic token-to-acoustic token model includes a second converter, and step 240a in the embodiment shown in Figure 3 can be implemented as steps 240a1 and 240a2.
[0216] Step 240a1: Combine the prompt semantic token, input semantic token, and prompt acoustic token in order to obtain a prefix.
[0217] In an embodiment of the present application, the computer device may concatenate the prompt semantic token, the input semantic token, and the prompt acoustic token in sequence to obtain the prefix. For example, the concatenation order may be randomly determined or preset in advance.
[0218] Step 240a2: Using the second converter, starting from the prefix, predict the acoustic features of the speech corresponding to the input text at each time point in a self-recursive manner to obtain the input acoustic token.
[0219] The second converter predicts the acoustic features of the speech corresponding to the input text at each time point in a self-recursive manner starting from the prefix to obtain the input acoustic token.
[0220] In an embodiment of the present application, a computer device processes the prefix through a second converter (Transformer network), predicts the acoustic token of the first time point of the speech corresponding to the input text, and then splices the acoustic token at the first time point to the prefix, re-inputs the second converter, and obtains the acoustic token of the second time point of the speech corresponding to the input text, and then splices the acoustic token at the first time point to the acoustic token at the first time point, re-inputs the second converter, and obtains the acoustic token of the third time point of the speech corresponding to the input text, and so on, until the acoustic features of the speech corresponding to the input text at all time points are predicted and the above-mentioned input acoustic token is obtained.
[0221] In some embodiments, the second converter is a neural network model for implementing the conversion function, and the neural network model includes multiple neural network layers. In some embodiments, the second converter can be a Transformer network. Of course, the second converter can also be other neural networks other than the Transformer network, including but not limited to at least one of a BERT network and a U-Net network.
[0222] In an embodiment of the present application, a feasible solution is proposed for predicting input acoustic tokens by prompting semantic tokens, inputting semantic tokens, and prompting acoustic tokens, thereby ensuring the feasibility of converting semantic tokens into acoustic tokens.
[0223] In some embodiments, the order of the prompt acoustic token and the input acoustic token is 2.
[0224] In an embodiment of the present application, the order of the above-mentioned acoustic tokens only needs to be set to 2 to meet the accuracy requirements of speech synthesis. Compared with the related technology that requires acoustic tokens of about order 8, the scheme shown in the embodiment of the present application can greatly reduce the complexity of the model and improve the processing efficiency of the model.
[0225] In some embodiments, the above method further comprises:
[0226] When the semantic token extractor and the acoustic token extractor are trained, a third audio sample and a fourth audio sample are obtained; the third audio sample and the fourth audio sample are two non-overlapping audio segments in the same audio;
[0227] extracting a semantic token label of the third audio sample and a semantic token label of the fourth audio sample respectively through a semantic token extractor;
[0228] extracting, by an acoustic token extractor, an acoustic token label of the third audio sample and an acoustic token label of the fourth audio sample;
[0229] Combining the semantic token label of the third audio sample, the semantic token label of the fourth audio sample, and the acoustic token label of the third audio sample in that order to obtain a prefix sample;
[0230] predicting, by the second transformer, an acoustic token sample of a fourth audio sample in a self-recursive manner starting from the prefix sample;
[0231] Parameters of the semantic token-to-acoustic token model are updated based on the acoustic token sample of the fourth audio sample and the acoustic token label of the fourth audio sample.
[0232] In some embodiments, obtaining a loss function value of a semantic token-to-acoustic token model based on the acoustic token sample of the fourth audio sample and the acoustic token label of the fourth audio sample;
[0233] The parameters of the semantic token-to-acoustic token model are updated based on the loss function value of the semantic token-to-acoustic token model.
[0234] Exemplarily, the parameters of the semantic token-to-acoustic token model are updated with the goal of minimizing the loss function value. This application does not limit the specific category of the loss function, such as the loss function is cross entropy loss, 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, perceptual loss function, etc. Exemplarily, the parameters of each module in the semantic token-to-acoustic token model are updated with the goal of minimizing the loss function value. Exemplarily, the parameters of the target module in each module in the semantic token-to-acoustic token model are updated with the goal of minimizing the loss function value. In this way, the training cost can be reduced and the training efficiency can be improved.
[0235] The solution shown in the embodiment of the present application, with the help of a semantic token extractor and an acoustic token extractor, can use non-overlapping segments in the same audio as samples of prompt audio and text respectively, so as to calculate the loss in the process of predicting acoustic tokens by the semantic token-to-acoustic token model, and then realize unsupervised training of the semantic token-to-acoustic token model, without relying on labeled data, reducing the requirements for training data and ensuring the accuracy of the model.
[0236] Please refer to Figure 10, which shows a schematic diagram of a semantic token-to-acoustic token model provided by an exemplary embodiment of the present application. As shown in Figure 10, after the semantic token and acoustic token extractors are trained, semantic tokens and acoustic tokens can be extracted simultaneously for a single audio file to train the semantic token-to-acoustic token prediction module. This process is also unsupervised training, requiring only a large amount of unlabeled audio data.
[0237] The semantic token-to-acoustic token model is a 12-layer Transformer architecture 1010 with 12 heads and a dimension of 768. It uses the same training method as a language model, inputting tokens 1 to t-1 and predicting the tth token. Mutual entropy loss is used as the loss function.
[0238] During training, two non-overlapping segments of the same audio are taken (one as the prompt segment and the other as the actual segment), and semantic tokens and acoustic tokens are extracted respectively. The semantic token of the prompt segment, the semantic token of the actual segment, and the acoustic token of the prompt segment are used as prefixes 1001 (prefix), and the acoustic token of the actual segment is self-recursively predicted 1002.
[0239] For example, when prefix1001 and the first substantial segment acoustic token X1 are known, the second substantial segment acoustic token X2 is predicted; when prefix1001, the first substantial segment acoustic token X1 and the second substantial segment acoustic token X2 are known, the third substantial segment acoustic token X3 is predicted; and so on.
[0240] During inference, semantic and acoustic tokens are extracted from an output audio segment. The semantic tokens corresponding to the text to be synthesized are arranged in the same order to form a prefix, and the acoustic tokens to be synthesized are recursively predicted. Since the target segment is not in the training set, it is considered a zero-shot synthesis.
[0241] In addition, the acoustic token extractor is a convolution-based encoder-decoder structure, in which the encoder consists of a one-dimensional convolutional layer with C channels and a kernel size of 7 - four convolutional blocks - two LSTM layers - a one-dimensional convolutional layer with D channels and a kernel size of 7. Each of the above convolutional blocks contains two convolutional layers with a kernel size of 3 and a convolutional layer with a stride of S. The strides of the four convolutional blocks are set to (2, 4, 5, 8) respectively. After passing through the convolutional layer with a stride of S, the length will become 1 / S of the original, and the number of channels is set to double. After passing through the encoder, the length is downsampled by 320 times, that is, for one second of 24kHz audio (24,000 samples), the encoder outputs the corresponding 75 frames, with a hidden layer representation of dimension D. The decoder is a mirror image of the encoder, except that the convolutional layers with a stride of S in the convolutional blocks are replaced with deconvolutional layers to achieve the corresponding upsampling multiple, that is, the quantized D-dimensional hidden layer representation of 75 frames is upsampled back to 24,000 sampling points.
[0242] The codec is connected to a residual vector quantizer (RVQ), which quantizes the encoder output before feeding it into the decoder. The quantization process primarily maps the encoder's output latent representation to the object with the smallest distance in the codebook. RVQ uses multiple codebooks and performs multiple cyclic quantization cycles, each quantizing the residual from the previous pass.
[0243] The technical solution of the embodiment of the present application adopts 8 codebooks of size K and dimension D. The result obtained by the first quantization is subjected to a residual operation with the original hidden layer representation, which is used as the input of the second quantization. The result obtained by the second quantization is subjected to a residual operation with the input of the second quantization, which is used as the input of the third quantization. This process is repeated eight times, and the quantization output of each time is added together as the final quantized hidden layer representation, which is input into the decoder. During training, a large amount of unlabeled audio is used for training, and the reconstruction error between the input audio and the output audio is used as the loss function. During inference, only the encoder and the residual vector quantizer are used to extract acoustic tokens. For one second of 24kHz audio, the encoder outputs 75 frames of hidden layer representation with a dimension of D, and only the first two quantizations are performed, and the quantization subscript is used as the value of the acoustic token. For example, if the hidden layer representation of the first frame is closest to the third vector in the first codebook, it is recorded as 3. If the hidden layer representation of the first frame is closest to the seventh vector in the second codebook after the residual of the third vector in the first codebook, it is recorded as 7. Therefore, the acoustic token corresponding to the hidden layer representation of the first frame is recorded as (3, 7). In summary, one second of 24kHz audio will be converted into 2×75 acoustic tokens.
[0244] After the acoustic token extractor is trained, the corresponding acoustic token can be extracted for any audio, and the sound decoder can be trained unsupervised to achieve fast conversion from acoustic tokens to audio.
[0245] The above-mentioned sound decoder is a parallel vocoder based on acoustic tokens. The structure of the parallel vocoder based on acoustic tokens is similar to the high-speed neural vocoder (HiFiGAN) based on generative adversarial networks (GAN), except that the input is acoustic tokens instead of Mel acoustic features. It is necessary to embed acoustic tokens of different orders (the technical solution of this application is 2nd order) separately to obtain a matrix of the number of frames × 2nd order × Ed and input it into the generator. The rest of the structure remains consistent with HiFiGAN.
[0246] The generator mainly consists of two parts: one is the upsampling structure, which is specifically composed of one-dimensional transposed convolution (the technical solution of this application requires upsampling the acoustic token by 320 times); the other is the Multi-Receptive Field Fusion (MRF) module, which is mainly responsible for optimizing the sampling points obtained by upsampling, and is specifically composed of a residual network.
[0247] There are two discriminators, namely multi-scale and multi-period discriminators, which identify speech from two different perspectives:
[0248] The multi-scale discriminator continuously averages and pools the speech sequence, gradually halving the length of the speech sequence, then applies several layers of convolution at different scales of the speech, and finally flattens it as the output of the multi-scale discriminator;
[0249] The multi-cycle discriminator folds the one-dimensional audio sequence into a two-dimensional plane with different sequence lengths and applies two-dimensional convolution on the two-dimensional plane.
[0250] Speech synthesis technology converts text into corresponding audio content using specific rules or model algorithms. Traditional speech synthesis techniques are primarily based on concatenation or statistical parameter methods. With the continuous breakthroughs achieved in speech recognition using deep learning, leading internet companies both domestically and internationally have begun to apply deep learning to speech synthesis, achieving significant progress.
[0251] For example, the related technology uses massive audio data to train an audio codec (codec) in an unsupervised manner, and uses the intermediate quantization values of the codec as acoustic tokens; then extracts acoustic tokens from audio data with text annotations, and trains the text-to-acoustic token module. In actual use, the acoustic tokens are predicted from the text, and then the acoustic tokens are input into the decoding part of the audio codec to generate the final audio. As mentioned above, the related technical solutions have the following problems that need to be solved:
[0252] First, predicting acoustic tokens directly from text has a large span, so a large amount of labeled data is required for training;
[0253] Second, using the decoding portion of an audio codec to convert acoustic tokens into audio requires predicting a higher order of acoustic tokens from the text (e.g., eighth-order residual vector quantization) to achieve good synthesis quality. Consequently, the text-to-acoustic tokenization module is complex, requiring two prediction phases: an autoregressive phase and a non-autoregressive phase, resulting in lower overall computational efficiency.
[0254] To address the above issues, the technical solution of this application introduces semantic tokens as a transition, which can alleviate the one-to-many problem faced when predicting acoustic tokens directly from text and reduce dependence on labeled data.
[0255] In addition, the technical solution of this application also introduces a parallel vocoder based on two-order acoustic tokens. On the one hand, it can reduce the order of acoustic tokens required for prediction, so that the semantic token-to-acoustic token model only requires a single autoregressive stage; on the other hand, the parallel vocoder can significantly reduce the conversion time required from acoustic tokens to audio.
[0256] Based on the above embodiments of the present application, a semi-supervised speech synthesis system can be constructed. The system comprises five components: a text-to-semantic token model, a semantic token extractor, an acoustic token extractor, a semantic token-to-acoustic token model, and an acoustic token vocoder. Except for the text-to-semantic token model, which requires a small amount of audio data with text annotations for training, the other four components only require a large amount of unlabeled audio for training.
[0257] The semi-supervised speech synthesis system effectively utilizes massive amounts of unlabeled audio data. Unsupervised training of semantic and acoustic token extractors extracts information such as semantics, timbre, prosody, and emotion from the audio data, enabling zero-shot speech synthesis using target prompts. Furthermore, using text to predict semantic tokens alleviates the one-to-many problem faced when predicting acoustic features directly from text, significantly reducing the amount of labeled data required for training. Finally, a parallel vocoder based on acoustic tokens enables rapid conversion from acoustic tokens to audio. This innovative semi-supervised speech synthesis system, on the one hand, fully utilizes readily available unlabeled audio data, significantly reducing its reliance on labeled audio data. On the other hand, while maintaining operational efficiency, it achieves the ability to control generated content using prompts, similar to large language models. This speech synthesis system can also control the synthesized audio using target prompts, achieving zero-shot synthesis.
[0258] For example, a prompt segment containing a target timbre (such as a cartoon character A) and a target emotion (happy) is used to control the system to synthesize the corresponding audio (the timbre of the happy cartoon character A has never appeared in the training set, so it is a zero-time synthesis).
[0259] Please refer to FIG11 , which shows an exemplary training and reasoning flowchart of the speech synthesis system involved in this application.
[0260] As shown in FIG11 , an exemplary semi-supervised training process of the speech synthesis system involved in this application is as follows:
[0261] Step A1: Using massive unlabeled audio data, perform unsupervised training on the semantic token extractor 1110;
[0262] Step A2: Using massive unlabeled audio data, perform unsupervised training on the acoustic token extractor 1120;
[0263] Step A3: Based on the semantic token extractor 1110 trained in step A1, a small amount of audio data with text annotations is used to perform supervised training on the text-to-semantic token model 1130;
[0264] Step A4: Based on the acoustic token extractor 1120 trained in step A2, the sound decoder 1140 is unsupervisedly trained using massive unlabeled audio data;
[0265] Step A5: Based on the semantic token extractor 1110 trained in step A1 and the acoustic token extractor 1120 trained in step A2, the semantic token to acoustic token model 1150 is unsupervisedly trained using massive unlabeled audio data.
[0266] As shown in FIG11 , an exemplary reasoning process of the speech synthesis system involved in this application is as follows:
[0267] Step B1: Input the prompt audio 1101 into the semantic token extractor 1110. After the semantic token extractor 1110 performs inference on the prompt audio 1101, it can obtain the prompt semantic token corresponding to the prompt audio 1101.
[0268] Step B2: Input the prompt audio 1101 into the acoustic token extractor 1120. After the acoustic token extractor 1120 performs inference on the prompt audio 1101, it can obtain the prompt acoustic token corresponding to the prompt audio 1101.
[0269] Step B3: Input the input text 1102 into the text-to-semantic token model 1130. After the text-to-semantic token model 1130 performs inference on the input text 1102, it can obtain the input semantic token corresponding to the input text 1102.
[0270] Step B4: Input the prompt semantic token obtained in step B1, the prompt acoustic token obtained in step B2, and the input semantic token obtained in step B3 into the semantic token-to-acoustic token model 1150. After the semantic token-to-acoustic token model 1150 performs inference, the input acoustic token corresponding to the input text 1102 can be obtained;
[0271] Step B5: Input the input acoustic token obtained in the above step B4 into the sound decoder 1140. After the sound decoder 1140 infers the input acoustic token, it can obtain the output audio 1103 corresponding to the input text 1102.
[0272] The application scenarios of this application are wide-ranging. The semi-supervised trained speech synthesis system can be placed on the cloud service as a basic technology to empower users of the cloud service.
[0273] Please refer to Figure 12, which shows an exemplary application scenario diagram of the speech synthesis system involved in this application. As shown in Figure 12, the speech synthesis system is deployed to a cloud service to provide customers with controllable speech synthesis services.
[0274] The specific calling process is as follows:
[0275] 1. The customer uploads the required synthesized text and prompt audio via the device 1210 connected to the cloud service;
[0276] 2. After the server 1220 performs rapid synthesis based on the speech synthesis system, it sends the corresponding synthesized audio to the device 1210 in the form of streaming or whole sentence return.
[0277] FIG13 is a block diagram of a speech synthesis device according to an exemplary embodiment of the present application. The device may be used to perform all or part of the steps performed by a computer device in the method shown in FIG2 , FIG3 , or FIG4 . The device includes:
[0278] Acquisition module 1301, used to acquire input text and prompt audio;
[0279] A first extraction module 1302 is configured to extract features of the prompt audio and obtain prompt semantic tokens and prompt acoustic tokens, wherein the prompt semantic tokens are used to indicate the semantic features of the prompt audio at various time points, and the prompt acoustic tokens are used to indicate the acoustic features of the prompt audio at various time points;
[0280] The second extraction module 1303 is used to extract features of the input text and obtain input semantic tokens, where the input semantic tokens are used to indicate semantic features of the speech corresponding to the input text at various time points;
[0281] An input acoustic token acquisition module 1304 is configured to acquire an input acoustic token based on the prompt semantic token, the prompt acoustic token, and the input semantic token; the input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point;
[0282] The output audio acquisition module 1305 is configured to acquire the output audio of the input text based on the input acoustic token.
[0283] In some embodiments, the first extraction module 1302 is configured to input the prompt audio into a semantic token extractor to obtain prompt semantic tokens obtained by the semantic token extractor processing the prompt audio, where the semantic token extractor is a machine learning model for extracting semantic features from audio;
[0284] Input the prompt audio into the acoustic token extractor to obtain prompt acoustic tokens obtained by the acoustic token extractor processing the prompt audio. The acoustic token extractor is a machine learning model for extracting acoustic features.
[0285] A second extraction module 1303 is configured to input the input text into a text-to-semantic tokenization model to obtain input semantic tokens obtained by processing the input text by the text-to-semantic tokenization model, where the text-to-semantic tokenization model is a machine learning model for extracting semantic features from text;
[0286] An input acoustic token acquisition module 1304 is configured to input the prompt semantic token, the prompt acoustic token, and the input semantic token into a semantic token-to-acoustic token model to obtain an input acoustic token output by the semantic token-to-acoustic token model. The semantic token-to-acoustic token model is a machine learning model for converting semantic features into acoustic features.
[0287] The output audio acquisition module 1305 is used to input the input acoustic token into the sound decoder to obtain the output audio output by the sound decoder.
[0288] In some embodiments, the semantic token extractor comprises a convolution branch and a first transformer; a first extraction module 1302 for,
[0289] Input the prompt audio into the convolution branch to obtain the hidden layer features of the prompt audio at each time point output by the convolution branch;
[0290] Processing the hidden layer features of the prompt audio at each time point by the first converter to obtain the intermediate layer features of the prompt audio at each time point output by the intermediate layer of the first converter;
[0291] The intermediate layer features of the prompt audio at each time point are clustered separately to obtain the prompt semantic tokens.
[0292] In some embodiments, the apparatus further comprises: a semantic token extractor training module for,
[0293] Inputting the first audio sample into the convolution branch to obtain hidden feature samples of the first audio sample at each time point output by the convolution branch;
[0294] partially masking the hidden feature samples of the first audio sample at each time point to obtain partially masked hidden feature samples;
[0295] Processing the partially masked hidden layer feature samples by the first converter to obtain intermediate layer features of the first audio sample at each time point output by the intermediate layer of the first converter;
[0296] Clustering the intermediate layer features of the first audio sample at each time point to obtain a semantic token sample of the first audio sample;
[0297] Parameters of a semantic token extractor are updated based on the semantic token sample of the first audio sample and the semantic token label of the first audio sample.
[0298] In some embodiments, the text-to-semantic tokenization model includes a text encoder, a duration predictor, an upsampling branch, and a decoder;
[0299] The second extraction module 1303 is used to:
[0300] Input the input text to the text encoder to obtain the hidden text encoding representation of the input text;
[0301] The hidden text encoding representation is input into the duration predictor to obtain the playback duration of the speech corresponding to the input text predicted by the duration predictor;
[0302] Through the upsampling branch, the hidden text encoding representation is upsampled to the number of frames corresponding to the playback duration to obtain the upsampled hidden text encoding representation;
[0303] The upsampled hidden text encoding representation is decoded by the decoder to obtain the input semantic token.
[0304] In some embodiments, the apparatus further comprises: a text-to-semantic token model training module, configured to:
[0305] When the semantic token extractor is trained, obtaining a second audio sample and a speech text of the second audio sample;
[0306] Inputting the second audio sample into the semantic token extractor to obtain a semantic token label of the second audio sample output by the semantic token extractor;
[0307] Inputting the speech text of the second audio sample into the text-to-semantic token model to obtain a semantic token sample of the second audio sample output by the text-to-semantic token model;
[0308] Based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample, parameters of the text-to-semantic token model are updated.
[0309] In some embodiments, the text-to-semantic token model training module is further configured to:
[0310] Inputting the speech text of the second audio sample into a text encoder to obtain a hidden text encoding representation sample of the speech text of the second audio sample;
[0311] Inputting the hidden text encoding representation sample into a duration predictor to obtain a first playback duration sample of the speech corresponding to the speech text of the second audio sample predicted by the duration predictor;
[0312] Inputting the hidden text encoding representation sample into the attention branch, obtaining a second playback duration sample of the speech corresponding to the speech text of the second audio sample output by the attention branch;
[0313] Upsampling the hidden text encoding representation sample to the number of frames corresponding to the second playback duration sample through the upsampling branch to obtain an upsampled hidden text encoding representation sample;
[0314] Decoding the upsampled hidden text encoding representation sample through a decoder to obtain a semantic token sample of the second audio sample;
[0315] Obtaining a loss function value of a text-to-semantic token model based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample;
[0316] Based on the loss function value of the text-to-semantic tokenization model, the parameters of the text-to-semantic tokenization model are updated.
[0317] In some embodiments, the text-to-semantic token model training module is used to:
[0318] Obtaining a first loss function value of a text-to-semantic token model based on a difference between the first playback duration sample and the second playback duration sample;
[0319] obtaining a second loss function value of the text-to-semantic token model based on a difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample;
[0320] Based on the first loss function value of the text-to-semantic token model and the second loss function value of the text-to-semantic token model, a loss function value of the text-to-semantic token model is determined.
[0321] In some embodiments, the semantic token to acoustic token model includes a second converter; an input acoustic token acquisition module 1304, which is used to:
[0322] Combine the prompt semantic token, input semantic token, and prompt acoustic token in this order to get the prefix;
[0323] The second converter predicts the acoustic features of the speech corresponding to the input text at each time point in a self-recursive manner starting from the prefix to obtain the input acoustic token.
[0324] In some embodiments, the order of the prompt acoustic token and the input acoustic token is 2.
[0325] In some embodiments, the apparatus further comprises: a semantic token to acoustic token model training module, configured to:
[0326] When the semantic token extractor and the acoustic token extractor are trained, a third audio sample and a fourth audio sample are obtained; the third audio sample and the fourth audio sample are two non-overlapping audio segments in the same audio;
[0327] extracting a semantic token label of the third audio sample and a semantic token label of the fourth audio sample respectively through a semantic token extractor;
[0328] extracting, by an acoustic token extractor, an acoustic token label of the third audio sample and an acoustic token label of the fourth audio sample;
[0329] Combining the semantic token label of the third audio sample, the semantic token label of the fourth audio sample, and the acoustic token label of the third audio sample in that order to obtain a prefix sample;
[0330] predicting, by the second transformer, an acoustic token sample of a fourth audio sample in a self-recursive manner starting from the prefix sample;
[0331] Parameters of the semantic token-to-acoustic token model are updated based on the acoustic token sample of the fourth audio sample and the acoustic token label of the fourth audio sample.
[0332] FIG14 shows a block diagram of a computer device 1400 according to an exemplary embodiment of the present application. The computer device can be implemented as the server in the above-mentioned solution of the present application. The computer device 1400 includes a central processing unit (CPU) 1401, a system memory 1404 including a random access memory (RAM) 1402 and a read-only memory (ROM) 1403, and a system bus 1405 connecting the system memory 1404 and the central processing unit 1401. The computer device 1400 also includes a mass storage device 1406 for storing an operating system 1409, application programs 1410, and other program modules 1411.
[0333] The mass storage device 1406 is connected to the central processing unit 1401 via a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1406 and its associated computer-readable media provide non-volatile storage for the computer device 1400. In other words, the mass storage device 1406 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0334] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, Erasable Programmable Read Only Memory (EPROM), Electronically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 1404 and mass storage device 1406 can be collectively referred to as memory.
[0335] According to various embodiments of the present disclosure, the computer device 1400 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1400 may be connected to a network 1408 via a network interface unit 1407 connected to the system bus 1405. Alternatively, the network interface unit 1407 may be used to connect to other types of networks or remote computer systems (not shown).
[0336] The memory further includes at least one computer program, which is stored in the memory. The central processing unit 1401 implements all or part of the steps in the methods shown in the above embodiments by executing the at least one computer program.
[0337] In an exemplary embodiment, a chip is also provided. The chip includes a programmable logic circuit and / or program instructions. When the chip runs on a computer device, it is used to implement the speech synthesis method in the above aspect.
[0338] In an exemplary embodiment, a computer program product is also provided. The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the speech synthesis method provided in each of the above method embodiments.
[0339] In an exemplary embodiment, a computer-readable storage medium is further provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the speech synthesis method provided by the above-mentioned method embodiments.
[0340] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0341] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0342] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A speech synthesis method, the method being executed by a computer device, the method comprising: Get input text and prompt audio; Extracting features of the prompt audio to obtain prompt semantic tokens and prompt acoustic tokens, wherein the prompt semantic tokens are used to indicate semantic features of the prompt audio at various time points, and the prompt acoustic tokens are used to indicate acoustic features of the prompt audio at various time points; Extracting features of the input text to obtain input semantic tokens, where the input semantic tokens are used to indicate semantic features of speech corresponding to the input text at various time points; Based on the prompt semantic token, the prompt acoustic token and the input semantic token, obtaining an input acoustic token; The input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point; Based on the input acoustic tokens, output audio of the input text is obtained.
2. The method according to claim 1, wherein: The step of extracting the features of the prompt audio and obtaining the prompt semantic token and the prompt acoustic token comprises: Inputting the prompt audio into a semantic token extractor to obtain the prompt semantic token obtained by the semantic token extractor processing the prompt audio, wherein the semantic token extractor is a machine learning model for extracting semantic features from audio; Inputting the prompt audio into an acoustic token extractor to obtain the prompt acoustic token obtained by the acoustic token extractor processing the prompt audio, wherein the acoustic token extractor is a machine learning model for extracting acoustic features from audio; The step of extracting the features of the input text and obtaining input semantic tokens includes: Inputting the input text into a text-to-semantic token model to obtain the input semantic token obtained by the text-to-semantic token model processing the input text, wherein the text-to-semantic token model is a machine learning model for extracting semantic features from text; The acquiring the input acoustic token based on the prompt semantic token, the prompt acoustic token and the input semantic token comprises: Inputting the prompt semantic token, the prompt acoustic token and the input semantic token into a semantic token-to-acoustic token model to obtain the input acoustic token output by the semantic token-to-acoustic token model, wherein the semantic token-to-acoustic token model is a machine learning model for converting semantic features into acoustic features; The step of obtaining the output audio of the input text based on the input acoustic token comprises: The input acoustic token is input into a sound decoder to obtain the output audio output by the sound decoder.
3. The method according to claim 2, wherein: The semantic token extractor comprises a convolutional branch and a first transformer; The step of inputting the prompt audio into a semantic token extractor to obtain the prompt semantic token obtained by the semantic token extractor through processing the prompt audio comprises: Inputting the prompt audio into the convolution branch to obtain hidden features of the prompt audio at each time point output by the convolution branch; Processing the hidden layer features of the prompt audio at each time point by the first converter to obtain the intermediate layer features of the prompt audio at each time point output by the intermediate layer of the first converter; The intermediate layer features of the prompt audio at each time point are clustered respectively to obtain the prompt semantic token.
4. The method according to claim 3, wherein: The method further comprises: Obtaining a first audio sample and a semantic token label of the first audio sample; Inputting the first audio sample into the convolution branch to obtain hidden layer feature samples of the first audio sample at each time point output by the convolution branch; Partially masking the hidden layer feature samples of the first audio sample at each time point to obtain the partially masked hidden layer feature samples; The first converter processes the partially masked hidden layer feature samples to obtain the intermediate The intermediate layer features of the first audio sample at each time point output by the intermediate layer; Clustering the intermediate layer features of the first audio sample at each time point to obtain semantic token samples of the first audio sample; Parameters of the semantic token extractor are updated based on the semantic token sample of the first audio sample and the semantic token label of the first audio sample.
5. The method according to any one of claims 2 to 4, wherein: The text-to-semantic token model includes a text encoder, a duration predictor, an upsampling branch, and a decoder; The step of inputting the input text into a text-to-semantic token model to obtain the input semantic token obtained by processing the input text by the text-to-semantic token model includes: Inputting the input text into the text encoder to obtain a hidden text encoding representation of the input text; Inputting the hidden text encoding representation into the duration predictor to obtain the playback duration of the speech corresponding to the input text predicted by the duration predictor; The hidden text encoding representation is upsampled to the number of frames corresponding to the playback duration by the upsampling branch to obtain the upsampled hidden text encoding representation; The upsampled hidden text encoding representation is decoded by the decoder to obtain the input semantic token.
6. The method according to claim 5, wherein: The method further comprises: When the semantic token extractor is trained, obtaining a second audio sample and a speech text of the second audio sample; Inputting a second audio sample into the semantic token extractor to obtain a semantic token label of the second audio sample output by the semantic token extractor; Inputting the speech text of the second audio sample into the text-to-semantic token model to obtain a semantic token sample of the second audio sample output by the text-to-semantic token model; Based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample, the parameters of the text-to-semantic token model are updated.
7. The method according to claim 6, wherein: The step of inputting the speech text of the second audio sample into the text-to-semantic token model to obtain the semantic token sample of the second audio sample output by the text-to-semantic token model comprises: Inputting the speech text of the second audio sample into the text encoder to obtain a hidden text encoding representation sample of the speech text of the second audio sample; Inputting the hidden text encoding representation sample into the duration predictor to obtain a first playback duration sample of the speech corresponding to the speech text of the second audio sample predicted by the duration predictor; Input the hidden text encoding representation sample into the attention branch, and obtain a second playback duration sample of the speech corresponding to the speech text of the second audio sample output by the attention branch; Upsampling the hidden text encoding representation sample to the number of frames corresponding to the second playback duration sample through the upsampling branch to obtain the upsampled hidden text encoding representation sample; Decoding the upsampled hidden text encoding representation sample by the decoder to obtain the semantic token sample of the second audio sample; The updating of the parameters of the text-to-semantic token model based on the semantic token sample of the second audio sample and the semantic token label of the second audio sample includes: Based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample, obtaining a loss function value of the text-to-semantic token model; Based on the loss function value of the text-to-semantic token model, the parameters of the text-to-semantic token model are updated.
8. The method according to claim 7, wherein: The acquiring the loss function value of the text-to-semantic token model based on the first playback duration sample, the second playback duration sample, the semantic token sample of the second audio sample, and the semantic token label of the second audio sample includes: Based on the difference between the first playback duration sample and the second playback duration sample, obtaining a first loss function value of the text-to-semantic token model; acquiring a second loss function value of the text-to-semantic token model based on a difference between the semantic token sample of the second audio sample and the semantic token label of the second audio sample; Based on the first loss function value of the text-to-semantic token model and the second loss function value of the text-to-semantic token model, a loss function value of the text-to-semantic token model is determined.
9. The method according to any one of claims 2 to 8, wherein: The semantic token-to-acoustic token model includes a second converter; The step of inputting the prompt semantic token, the prompt acoustic token, and the input semantic token into a semantic token-to-acoustic token model to obtain the input acoustic token output by the semantic token-to-acoustic token model comprises: Combining the prompt semantic token, the input semantic token, and the prompt acoustic token in order to obtain a prefix; The second converter predicts the acoustic features of the speech corresponding to the input text at each time point in a self-recursive manner starting from the prefix to obtain the input acoustic token.
10. The method according to claim 9, wherein: The order of the prompt acoustic token and the input acoustic token is 2.
11. The method according to claim 9 or 10, wherein: The method further comprises: When the semantic token extractor and the acoustic token extractor are trained, obtaining a third audio sample and a fourth audio sample; the third audio sample and the fourth audio sample are two non-overlapping audio segments in the same audio; extracting the semantic token label of the third audio sample and the semantic token label of the fourth audio sample respectively by the semantic token extractor; extracting the acoustic token tag of the third audio sample and the acoustic token tag of the fourth audio sample respectively by the acoustic token extractor; Combining the semantic token label of the third audio sample, the semantic token label of the fourth audio sample, and the acoustic token label of the third audio sample in this order to obtain a prefix sample; predicting, by the second converter, acoustic token samples of the fourth audio sample in a self-recursive manner starting from the prefix sample; Based on the acoustic token sample of the fourth audio sample and the acoustic token label of the fourth audio sample, parameters of the semantic token-to-acoustic token model are updated.
12. A speech synthesis device, comprising: The acquisition module is used to obtain input text and prompt audio; A first extraction module, used to extract features of the prompt audio, and obtain prompt semantic tokens and prompt acoustic tokens, wherein the prompt semantic tokens are used to indicate semantic features of the prompt audio at various time points, and the prompt acoustic tokens are used to indicate acoustic features of the prompt audio at various time points; A second extraction module is used to extract the features of the input text and obtain input semantic tokens, where the input semantic tokens are used to indicate the semantic features of the speech corresponding to the input text at various time points; An input acoustic token acquisition module, used to acquire an input acoustic token based on the prompt semantic token, the prompt acoustic token and the input semantic token, wherein the input acoustic token is used to indicate the acoustic features of the speech corresponding to the input text at each time point; The output audio acquisition module is used to acquire the output audio of the input text based on the input acoustic token.
13. A computer device, comprising a processor and a memory, wherein the memory stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to implement the speech synthesis method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, wherein at least one computer instruction is stored in the computer-readable storage medium, and the computer instruction is loaded and executed by a processor to implement the speech synthesis method according to any one of claims 1 to 11.
15. A computer program product, comprising computer instructions stored in a computer-readable storage medium; the computer instructions are read and executed by a processor of a computer device to implement the method for speech synthesis as described in any one of claims 1 to 11.