Speech synthesis method and device, equipment and storage medium

By immediately performing speech synthesis during the receiving of streaming text, the problem of insufficient timeliness of speech synthesis in the prior art is solved, and faster voice response and smoother human-computer interaction are achieved.

CN120375798APending Publication Date: 2025-07-25IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510292668.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing voice synthesis technology is poor in human-computer interaction scenarios, and users need to wait for a long time to hear the reply voice, which affects the user experience.

Method used

During the process of receiving streaming text, the encoding module of the speech synthesis model is immediately input to the first target number of text characters to encode, generate encoding features, and determine the second voice based on the new text characters and the previous encoding features each time a new text character is received, and the first target number is less than the total number of target text.

Benefits of technology

It improves the timeliness of voice synthesis, reduces user waiting time, and improves the intelligence and fluency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375798A_ABST
    Figure CN120375798A_ABST
Patent Text Reader

Abstract

The invention provides a speech synthesis method, device and equipment and a storage medium, which are applied to the technical field of speech processing, and the method comprises the following steps: in the process of receiving a streaming text, under the condition that a first target number of text characters in a target text are received, obtaining a first target number of text characters in the target text; inputting the first target number of text characters into a coding module of a speech synthesis model to obtain coding features, the first target number being smaller than the total number of text characters in the target text; inputting the coding feature into a decoding module of a speech synthesis model to obtain a first speech; and under the condition that a new text character is received each time, determining a second voice based on the new text character, all the previously received text characters and the previously obtained coding characteristics. According to the invention, when a part of text characters are received, the part of text characters are input into the coding module for coding, so that the phenomenon that in the prior art, speech synthesis is started after a complete text is received is avoided, and the timeliness of speech synthesis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a speech synthesis method, device, equipment, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence and speech technology, text-to-speech (TTS) technology has been widely used in many fields, such as intelligent voice assistants, online education, navigation systems, etc. The speech synthesis system based on large language models can synthesize synthetic speech with high naturalness and high anthropomorphism, improving the user interaction experience.

[0003] Among them, in the human-computer interaction scenario based on intelligent voice, usually the received complete text to be synthesized is input into the speech synthesis system for speech synthesis. However, the above method results in poor timeliness of speech synthesis. Summary of the Invention

[0004] The present invention provides a speech synthesis method, device, equipment, and storage medium to solve the defect of poor timeliness in speech synthesis in the prior art and achieve improved timeliness of speech synthesis.

[0005] The present invention provides a speech synthesis method, including: During the process of receiving streaming text, when the first target number of text characters in the target text is received, input the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; Input the encoding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; When a new text character is received each time, determine a second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously, and the combination of the first speech and each second speech is used as the target speech corresponding to the target text.

[0006] According to the speech synthesis method provided by the present invention, the first target number is determined based on the number of words pronounced by the user per second on average, the preset encoding feature length, and the frame rate for determining encoding features by the encoding module, and the average text length covered by the preset encoding feature length is less than or equal to the length of the text characters input into the encoding module each time.

[0007] A voice synthesis method provided by the present invention, when receiving a new text character each time, determining a second voice based on the new text character, all previously received text characters, and previously obtained encoding features, includes: When the length of the encoding feature output by the encoding module last time reaches the preset encoding feature length and a new text character is received, determining the second voice based on the new text character, all previously received text characters, and previously obtained encoding features.

[0008] A voice synthesis method provided by the present invention, determining the second voice based on the new text character, all previously received text characters, and previously obtained encoding features, includes: Concatenating the new text character and all previously received text characters to obtain a concatenated character; Inputting the concatenated character and the previously obtained encoding features into the encoding module to obtain new encoding features output by the encoding module, where the new encoding features are encoding features corresponding to other characters after the text character to which the encoding feature obtained last time belongs; Inputting the new encoding features into the decoding module to obtain the second voice output by the decoding module.

[0009] A voice synthesis method provided by the present invention, the encoding module is trained based on the following method: Determining a second target quantity based on the frame number corresponding to the current predicted encoding feature that the initial encoding module needs to output, and obtaining the second target quantity of sample text characters in the sample text; Inputting the sample text characters and the previously obtained predicted encoding features into the initial encoding module to obtain the current predicted encoding feature output by the initial encoding module; Constructing loss information based on the current predicted encoding feature and the voice encoding feature in the sample voice corresponding to the sample text character, where the sample voice is the voice corresponding to the sample text; Iteratively optimizing the parameters of the initial encoding module based on the loss information to obtain the encoding module.

[0010] A voice synthesis method provided by the present invention, determining the second target quantity based on the frame number corresponding to the current predicted encoding feature output by the initial encoding module, includes: Determining the second target quantity based on the frame number corresponding to the current predicted encoding feature, the length of the sample text characters input into the initial encoding module per second, and the number of forward operation steps that the initial encoding module can execute per second.

[0011] A speech synthesis method provided by the present invention, determining the second target quantity based on the frame number corresponding to the current prediction coding feature, the length of the sample text characters input into the initial coding module per second, and the number of forward operations that the initial coding module can perform per second, includes: Determining the target value obtained by rounding down the ratio between the product of the frame number corresponding to the current prediction coding feature and the length of the sample text characters input into the initial coding module per second, and the number of forward operations that the initial coding module can perform per second; When the target value is equal to 0, determining the second target quantity as 1; When the target value is greater than 0, determining the second target quantity as the target value.

[0012] The present invention also provides a speech synthesis device, including: An input module, configured to, during the process of receiving streaming text, when receiving the first target quantity of text characters in the target text, input the first target quantity of text characters into the coding module of the speech synthesis model to obtain the coding features output by the coding module, where the first target quantity is less than the total number of text characters in the target text; The input module is further configured to input the coding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; A determination module, configured to, when receiving new text characters each time, determine a second speech based on the new text characters, all the text characters received previously, and the coding features obtained previously, and the combination of the first speech and each second speech is used as the target speech corresponding to the target text.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the speech synthesis method described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the speech synthesis method described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the speech synthesis method described in any one of the above is implemented.

[0016] The speech synthesis method, device, equipment, and storage medium provided by the present invention, during the process of receiving streaming text, when the first target number of text characters in the target text is received, by inputting the first target number of text characters into the encoding module of the speech synthesis model, the encoding features output by the encoding module are obtained, where the first target number is less than the total number of text characters in the target text, and the encoding features are input into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module. In each case of receiving new text characters, based on the new text characters, all the text characters received previously, and the encoding features obtained previously, a second speech is determined. The combination of the first speech and each second speech is used as the target speech corresponding to the target text. Since when a part of the text characters in the target text is received, the received part of the text characters can be input into the encoder of the speech synthesis model for encoding processing, and thus speech is generated based on the obtained encoding features, it avoids the phenomenon in the prior art where it is necessary to wait for the complete text to be received before starting speech synthesis, and improves the timeliness of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of one of the speech synthesis methods provided by the embodiments of the present invention.

[0019] Figure 2 It is a schematic flowchart of another speech synthesis method provided by the embodiments of the present invention.

[0020] Figure 3 It is a schematic structural diagram of the speech synthesis device provided by the embodiments of the present invention.

[0021] Figure 4 It is a schematic physical structure diagram of an electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0023] In recent years, with the development of large language model technology, the development of speech synthesis technology has been further promoted. The speech synthesis system based on large language models can synthesize synthetic speech with high naturalness and high anthropomorphism, significantly improving the user's interaction experience. In addition to the synthesis effect, in the human-computer interaction scenario based on intelligent speech, the response time of the speech synthesis system is also one of the key factors affecting the user experience. When the traditional speech synthesis system performs speech synthesis, it needs to receive the complete text to be synthesized before starting the speech synthesis. During the entire speech synthesis process, the timeliness of speech synthesis is poor, resulting in a long waiting time for the user to hear the reply speech in the human-computer interaction scenario, thus reducing the intelligence of human-computer interaction.

[0024] Considering the above problems, the embodiment of the present invention proposes a speech synthesis method. In this method, it is not necessary to wait for the complete text to be synthesized before performing speech synthesis. Instead, during the process of receiving the streaming text, once the first target number of text characters in the target text is received, speech synthesis will be performed, where the first target number is less than the total number of text characters in the complete target text. In this way, since it is not necessary to wait for the complete target text to perform speech synthesis, the timeliness of speech synthesis can be improved. In the human-computer interaction scenario, the waiting time of the user can be reduced, and the intelligence of human-computer interaction is improved.

[0025] The following combines Figure 1 and Figure 2 to describe the speech synthesis method provided by the embodiment of the present invention. The embodiment of the present invention can be applied to any scenario that requires speech synthesis, such as intelligent voice assistants, online education, or navigation systems, and can be particularly applied to the human-computer interaction scenario. The execution subject of this method can be an electronic device such as a terminal device, a computer, a server, a server cluster, or a specially designed speech synthesis device, or it can also be a speech synthesis device set in the electronic device, and the speech synthesis device can be implemented by software, hardware, or a combination of both.

[0026] Figure 1 One of the flow diagrams of the speech synthesis method provided by the embodiment of the present invention is shown in Figure 1 as follows. This method includes: Step 101: During the process of receiving the streaming text, when the first target number of text characters in the target text is received, input the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text.

[0027] In this step, streaming text refers to the text content that is gradually transmitted and processed in the form of a data stream, rather than static text received all at once. In streaming transmission, the text data arrives at the speech synthesis model in segments and in chronological order.

[0028] When performing speech synthesis, streaming text is received in real time. When performing speech synthesis for the first time, once the first target number of text characters is received, these first target number of text characters will be input into the encoding module of the speech synthesis model for processing, so as to obtain the encoded features output by the encoding module. Among them, the text characters can include Chinese characters or phonemes in the text, etc.

[0029] The encoding module is used to encode text characters into discrete encoded features. This encoding module can adopt the method of extracting semantic representations based on Hubert and then further performing kmeans clustering, or can adopt the residual vector quantization (RVQ) layer of codec models such as Soundstream, Encodec, or SpeechTokenizer for discrete encoding, or can also adopt other quantization encoding methods, etc. The embodiments of the present invention do not make specific limitations.

[0030] In addition, the above-mentioned first target number is less than the total number of text characters included in the complete target text. In this way, when performing speech synthesis, there is no need to wait for the complete text to be received, but speech synthesis will start when part of the text is received, making the speech synthesis more timely and the response speed faster.

[0031] Step 102: Input the encoded features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module.

[0032] In this step, the encoded features are the intermediate representations obtained by processing the input text characters through the encoding module, which usually include phoneme sequences, prosodic features, and semantic embeddings, etc. Further, in order to recover speech based on the discrete encoded features, a speech prediction model is also needed, that is, the decoding module. The encoded features can be input into the decoding module of the speech synthesis model, and the decoding module decodes the encoded features, so as to generate the corresponding speech waveform. Specifically, the decoder can be used to convert the encoded features into a mel spectrogram, and a vocoder can be used to convert the mel spectrogram into a speech waveform, so as to obtain the first speech. In a human-computer interaction scenario, the generated first speech will be played in real time through a speaker.

[0033] Among them, the decoding module can be a decoding model based on the HiFiGAN vocoder, which can directly input the discretized encoded features into the decoding module to obtain the predicted speech waveform. It can also first pass the discretized encoded features through a neural network to predict acoustic features, such as FBK, and then input the predicted acoustic features into the vocoder to restore the speech waveform. To achieve low latency, the structures of these decoding models need to be designed as models with only historical receptive fields.

[0034] Step 103: In the case of receiving a new text character each time, based on the new text character, all the text characters received previously, and the encoded features obtained previously, determine a second speech, and combine the first speech and each second speech as the target speech corresponding to the target text.

[0035] In this step, since during the streaming text transmission, the text will be gradually input into the speech synthesis model in the form of a character stream. When the encoded features obtained last time meet the preset length requirement and a new text character is received, the received new text character and all the text characters received previously will be concatenated to update the text sequence to be synthesized, so as to input the updated text sequence to be synthesized and all the encoded features obtained previously into the encoding module to obtain new discretized encoded features, and input the new discretized encoded features into the decoding module to obtain the second speech output by the decoding module.

[0036] Among them, the first speech and each obtained second speech jointly constitute the target speech corresponding to the streaming target text. In a human-computer interaction scenario, once the speech is generated, it will be transmitted to the speaker for playback. Since the text is input into the speech synthesis model in a streaming manner, the speech synthesized by the speech synthesis model is continuous. Additionally, after outputting the first speech and each second speech, they can also be mixed and stored to ensure smoothness and coherence when generating new speech segments.

[0037] The speech synthesis method provided by the embodiment of the present invention, during the process of receiving streaming text, when the first target number of text characters in the target text is received, by inputting the first target number of text characters into the encoding module of the speech synthesis model, the encoded features output by the encoding module are obtained. The first target number is less than the total number of text characters in the target text, and the encoded features are input into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module. When a new text character is received each time, based on the new text character, all the text characters received previously, and the encoded features obtained previously, a second speech is determined. The combination of the first speech and each second speech is used as the target speech corresponding to the target text. Since when a part of the text characters in the target text is received, the received part of the text characters can be input into the encoder of the speech synthesis model for encoding processing, and thus speech is generated based on the obtained encoded features, avoiding the phenomenon in the prior art that it is necessary to wait until the complete text is received before starting speech synthesis, and improving the timeliness of speech synthesis.

[0038] Exemplarily, on the basis of the above embodiment, the first target number is determined based on the number of words pronounced by the user per second on average, the preset length of the encoded features, and the frame rate at which the encoding module determines the encoded features. The preset length of the encoded features corresponding to the average text length covered by the speech is less than or equal to the length of the text characters input into the encoding module each time.

[0039] Specifically, when performing speech synthesis for the first time, it is necessary to wait until at least the first target number of text characters are received before starting speech synthesis. The size of this first target number is related to the duration of the first response. Therefore, how to select an appropriate first target number, which can not only achieve accurate speech synthesis based on the received text characters, but also improve the response timeliness and reduce the waiting duration of the user, is very important. In this embodiment, it can be based on the number of words pronounced by the user per second on average , the preset length of the encoded features and the frame rate at which the encoding module determines the encoded features to jointly determine the first target number.

[0040] Among them, the number of words pronounced by the user per second on average can also be understood as the speech rate. By collecting the speech rates of multiple different users and averaging the speech rates of different users, , for example, this can be 5 words / second.

[0041] The preset length of the encoded features represents the length of the discretized encoded features set during speech synthesis. In practical applications, when the encoding module generates the preset length of the encoded features After the encoding features, these encoding features will be input into the decoding module for decoding, and the newly received text characters will be concatenated with the previous characters. Then, the concatenated text and the previously obtained encoding features will be input into the encoding module for processing again. Therefore, the preset length of the encoding features determines the duration of the first response. Thus, the preset length of the encoding features should not be set too large. Moreover, since the prediction duration of the encoding features cannot be displayed in the speech synthesis model, in order to avoid generating empty encoding features, it is necessary to set the average text length covered by the preset encoding features to be less than or equal to the length of the text characters input into the encoding module each time. It can be understood that for the first speech synthesis process, the text characters input into the encoding module are the text characters received this time, while for subsequent speech synthesis processes, the text characters input into the encoding module are the text characters received this time concatenated with all the text characters received previously.

[0042] The encoding module determines the frame rate of the encoding features which is the frequency of generating encoding features by the encoding module. For example, if the frame rate of the encoding module is 25 Hz, then 25 frames of features need to be processed per second.

[0043] Exemplarily, in this embodiment, the first target quantity can be determined according to Suppose the average number of words pronounced by the user per second is 5, the preset length of the encoding features is 10 characters, and the frame rate of the encoding features is 25 Hz. Then the first target quantity can be determined to be 2. That is, in the first speech synthesis, when two text characters are received, these two text characters can be input into the encoding module for encoding without waiting to receive all the text, which can greatly improve the timeliness and efficiency of speech synthesis.

[0044] In this embodiment, since the first target quantity can be determined based on the average number of words pronounced by the user per second, the preset length of the encoding features, and the frame rate of the encoding features determined by the encoding module, the first target quantity can be dynamically adjusted according to the user's speech rate, so it can be applied to scenarios with different speech rates and improve the adaptability. In addition, by dynamically adjusting the first target quantity, speech synthesis processing can start immediately after receiving a small number of text characters without waiting for the complete text input. In the scenario of human-computer interaction, the waiting time of the user for the speech output can be reduced, and the fluency of real-time interaction is improved.

[0045] Exemplarily, based on the above solution, when a new text character is received each time, to determine the second voice based on the new text character, all the text characters received previously, and the encoding features obtained previously, the following method can be used: When the length of the encoded feature output by the encoding module last time reaches the preset encoded feature length and a new text character is received, a second voice is determined based on the new text character, all the text characters received previously, and the encoded feature obtained previously.

[0046] Specifically, by setting the preset encoded feature length once the length of the encoded feature output by the encoding module reaches the preset encoded feature length if a new text character is received at this time, the newly received text character and all the text characters received previously will be concatenated to update the text sequence to be synthesized. The updated text sequence to be synthesized and the previously generated encoded feature are input into the encoding module together, and the encoding module processes the text characters that have not been encoded previously to obtain a new encoded feature. The second voice can be generated by decoding the new encoded feature through a decoder.

[0047] For example, assume that the first target quantity is 2, when it is 20, when the two text characters "today" are received, "today" will be input into the encoding module for encoding. After the number of encoded features output by the encoding module reaches 20 tokens, if the new text character "is Sunday" is received at this time, then "is Sunday" will be concatenated after "today" to obtain the updated text sequence to be synthesized "today is Sunday", and "today is Sunday" and the previously obtained 20-token encoded feature are input into the encoding module for processing together.

[0048] It should be noted that if the preset encoded feature length is set to the total length of the text sequence to be synthesized, at this time, the voice synthesis method in the present invention can be compatible with the voice synthesis scheme of the prior art, improving the compatibility of voice synthesis. In practical applications, it can be selected according to specific scenario requirements, improving the flexibility of voice synthesis.

[0049] In this embodiment, by setting the preset encoded feature length, once the length of the encoded feature output by the encoding module reaches the preset encoded feature length, voice restoration will be performed, improving the real-time performance of voice generation. In addition, by setting the preset encoded feature length, the encoding module processes encoded features of a fixed length each time, which can not only avoid the computational bottleneck caused by too large input data volume, but also prevent resource waste caused by too small input data volume.

[0050] Figure 2This is the second flowchart of the speech synthesis method provided by the embodiments of the present invention. On the basis of the foregoing embodiments, the process of determining the second voice based on the new text characters, all the text characters received previously, and the encoding features obtained previously is described in detail. As Figure 2 shown, the method includes: Step 201: Concatenate the new text characters and all the text characters received previously to obtain concatenated characters.

[0051] In this step, after receiving the new text characters, the new text characters are concatenated behind the text characters received previously to obtain concatenated characters. The concatenated characters are a new text sequence, that is, the text sequence of the voice to be synthesized. By concatenating the new text characters and the text characters received previously, more complete and contextually informative concatenated characters can be obtained, and the accuracy of speech synthesis can be improved when performing speech synthesis based on the concatenated characters subsequently.

[0052] In addition, the concatenated characters need to be saved. After receiving new text characters next time, the new text characters received next time are concatenated behind the concatenated characters again until a terminator indicating the end of the text is received.

[0053] Step 202: Input the concatenated characters and the encoding features obtained previously into an encoding module to obtain new encoding features output by the encoding module. The new encoding features are the encoding features corresponding to the other characters after the text characters to which the encoding features obtained previously belong.

[0054] In this step, the encoding features obtained previously can be all the encoding features obtained previously or the encoding features obtained last time. After inputting the concatenated characters and the encoding features obtained previously into the encoding module, the encoding module will start from the position where encoding was performed last time and continue to encode the text characters in the concatenated characters by using the context information of the encoding features obtained previously, so as to obtain new encoding features. By this means, the phenomenon of repeatedly encoding the characters that have been encoded previously can be avoided, and the encoding efficiency is improved.

[0055] Continuing with the above example, if the previous text characters are "Today", and the new text characters "is Sunday" are received subsequently, then "is Sunday" will be concatenated behind "Today" to obtain the updated text sequence to be synthesized "Today is Sunday", and "Today is Sunday" and the 20 token encoding features obtained previously will be input into the encoding module for processing, and based on the previous encoding features, the text characters in the concatenated characters will be encoded continuously.

[0056] Step 203: Input the new encoding features into a decoding module to obtain the second voice output by the decoding module.

[0057] In this step, after obtaining the new encoded feature, the new encoded feature can be input into the decoding module for decoding, so as to predict the speech waveform and obtain the second speech. Alternatively, the acoustic feature can be predicted from the new encoded feature through a neural network, and the acoustic feature can be input into the decoder to restore the speech waveform and obtain the second speech.

[0058] In this embodiment, by splicing the new text characters and the previously received text characters, the spliced characters are made more complete. When performing speech synthesis based on the spliced text, the accuracy of speech synthesis can be improved. In addition, inputting the spliced characters and the previously obtained encoded features into the encoding module for processing, referring to the context information in the previously obtained encoded features, can further improve the accuracy of speech synthesis.

[0059] Exemplarily, based on the above embodiments, the foregoing encoding module is trained in the following manner: Determine the second target quantity based on the frame number corresponding to the currently predicted encoded feature that the initial encoding module needs to output, obtain the second target quantity of sample text characters in the sample text, and input the sample text characters and the previously obtained predicted encoded features into the initial encoding module to obtain the currently predicted encoded feature output by the initial encoding module. Based on the currently predicted encoded feature and the speech encoded feature corresponding to the sample text characters in the sample speech, construct loss information, where the sample speech is the speech corresponding to the sample text. Iteratively optimize the parameters of the initial encoding module based on the loss information to obtain the encoding module.

[0060] Specifically, a large number of training sample pairs will be collected in advance. Each training sample pair includes a sample text and the sample speech corresponding to the sample text. For each sample speech, it can be encoded into a discretized representation sequence, that is, the speech encoded feature, by using a speech encoding model. Among them, the speech encoding model can extract semantic representations based on Hubert and then further perform kmeans clustering, or can perform discretized encoding using the RVQ layer of codec models such as Soundstream, Encodec, or SpeechTokenizer, or can use other quantization encoding methods. The embodiments of the present invention do not limit this.

[0061] Let represent the codeword number sequence of the speech encoded feature obtained after the above discretized encoding of the sample speech, where T represents the frame length after encoding the sample speech. Let represent the sample text corresponding to the sample speech, where N represents the text length.

[0062] In the prior art, when training an encoding module, a loss function is usually constructed according to formula (1): (1) wherein, represents the encoding feature to be output at the current moment, t represents the frame number of the encoding feature to be output at the current moment. Taking all the text information as conditional features and given the historical encoding features before t, the encoding feature is predicted accordingly and the training of the encoding module is realized by minimizing the loss function L.

[0063] In the above method, since it is necessary to wait until the entire sentence of text is received for speech synthesis, the latency is usually relatively high. To reduce the latency of speech synthesis, in this embodiment, a modeling method based on text streaming synthesis is adopted, that is, only part of the text information is given as conditional features for each frame of speech for modeling.

[0064] Specifically, based on the frame number t corresponding to the current predicted encoding feature that the initial encoding module is about to output, the second target quantity can be determined first. The second target quantity can be used to represent the number of characters that the initial encoding module should see when correctly synthesizing text characters. After determining the second target quantity, obtain the second target quantity of sample text characters from the sample text, and input these sample text characters and all the previously obtained predicted encoding features into the initial encoding module to obtain the current predicted encoding feature. Further, based on the frame number of the current predicted encoding feature, among the speech encoding features obtained by discretizing the sample speech corresponding to the sample text, search for the speech encoding feature with the same frame number as the current predicted encoding feature, so as to construct loss information based on the current predicted encoding feature and the speech encoding feature.

[0065] Exemplarily, the loss function L can be as shown in formula (2): (2) wherein, T represents the frame length after encoding the sample speech, t represents the frame number corresponding to the current predicted encoding feature to be output, i represents the second target quantity, and the second target quantity i is less than the number of text characters included in the complete text.

[0066] Based on the above loss information, the parameters of the initial encoding module are optimized. By repeating the above process until the model converges or the loss is minimized, the finally obtained model is determined as the trained encoding module.

[0067] In this embodiment, when training the encoding module, it is not necessary to train based on the complete text. Instead, after giving a part of the text, the part of the text can be used as a condition for modeling, and then the encoding module can be trained. The encoding module trained in this way can also perform speech synthesis when receiving a part of the text during subsequent speech synthesis, greatly reducing the latency of speech interaction and improving the user experience.

[0068] Exemplarily, based on the frame number corresponding to the current predicted encoding feature output by the initial encoding module, when determining the second target quantity, the second target quantity can be determined based on the frame number corresponding to the current predicted encoding feature, the length of the sample text characters input into the initial encoding module per second, and the number of forward operations that the initial encoding module can perform per second.

[0069] Specifically, the frame number corresponding to the current predicted encoding feature is the frame number of the current predicted encoding feature that is about to be output. The number of forward operations that the initial encoding module can perform per second can represent the processing ability of the encoding module. In this embodiment, the second target quantity can be dynamically determined according to the frame number t corresponding to the current predicted encoding feature that is about to be output, which can ensure that during the process of encoding text characters by the encoding module, it can see enough text characters to correctly encode the text characters, and can also avoid the phenomenon of waiting to receive too many text characters before starting to encode. This can not only improve the encoding efficiency but also improve the encoding accuracy.

[0070] In a possible implementation manner, when determining the second target quantity based on the frame number corresponding to the current predicted encoding feature, the length of the sample text characters input into the initial encoding module per second, and the number of forward operations that the initial encoding module can perform per second, it can be performed in the following manner: Determine the ratio between the product of the frame number corresponding to the current predicted encoding feature and the length of the sample text characters input into the initial encoding module per second, and the number of forward operations that the initial encoding module can perform per second, and take the floor value of the ratio to obtain the target value. When the target value is equal to 0, determine the second target quantity as 1. When the target value is greater than 0, determine the second target quantity as the target value.

[0071] Specifically, the second target quantity can be determined by the following formula (3): (3) where i represents the second target quantity, represents the length of the sample text characters input into the initial encoding module per second, represents the ratio between the number of forward operations that the initial encoding module can perform per second, and t represents the frame number corresponding to the current predicted encoding feature, Denotes the floor function.

[0072] When occurs, it indicates that t is 0, that is, the text characters are encoded for the first time. To ensure that the acoustic features of the first few frames can see their corresponding text characters, the second target quantity i can be set to 1 at this time. In addition, to ensure that each frame of speech can see the text characters corresponding to the acoustic information of the current frame and the historical text characters, in this embodiment, it is necessary to ensure that the generation speed of the synthesized text is at least greater than or equal to the synthesis speed of the speech, that is, it is necessary to , where represents the average number of words pronounced by the user per second, represents the frame rate at which the encoding module determines the encoding features. Usually, due to the high frame rate of acoustics, the above conditions can be met. For example, the frame rate of the discretized encoding features obtained based on Hubert is 50 Hz, that is, it is necessary to perform 50 forward operations of the speech synthesis model to synthesize one second of speech.

[0073] In addition, since in the subsequent stage of speech synthesis, and values will have certain dynamic changes. To be compatible with this situation and improve the robustness of speech synthesis, when training the encoding module, the value of the second target quantity i can be perturbed to a certain extent. For example, i can be perturbed by 5% left and right, that is, randomly increase or decrease the value of i by 5% and then take the floor. Specifically, it can be designed according to the and value range.

[0074] By training the encoding module in the above manner, since the current predicted encoding features can see the current text characters and historical text characters, and since usually the speed of obtaining the text to be synthesized is faster, more future text characters can be seen as the generation progresses. Therefore, based on the speech synthesis in the embodiments of the present invention, compared with the prior art, the speech quality can be significantly improved in terms of timeliness without changing the speech quality.

[0075] The speech synthesis device provided by the present invention will be described below. The speech synthesis device described below can be mutually referred to the speech synthesis method described above.

[0076] Figure 3 is a schematic structural diagram of the speech synthesis device provided by the embodiments of the present invention. As Figure 3 shown, the speech synthesis device 300 includes: An input module 11, configured to, during the process of receiving streaming text, when the first target number of text characters in the target text are received, input the first target number of text characters into an encoding module of a speech synthesis model to obtain encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; The input module 11 is further configured to input the encoding features into a decoding module of the speech synthesis model to obtain a first speech output by the decoding module; A determination module 12, configured to, when a new text character is received each time, determine a second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously, and the combination of the first speech and each of the second speeches is used as a target speech corresponding to the target text.

[0077] In an exemplary embodiment, the first target number is determined based on the number of words pronounced by a user per second on average, a preset encoding feature length, and a frame rate for determining encoding features by the encoding module, and the average text length covered by the preset encoding feature length is less than or equal to the length of the text characters input into the encoding module each time.

[0078] In an exemplary embodiment, the determination module 12 is specifically configured to: When the length of the encoding features output by the encoding module last time reaches the preset encoding feature length and a new text character is received, determine the second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously.

[0079] In an exemplary embodiment, the determination module 12 is specifically configured to: Concatenate the new text character and all the text characters received previously to obtain concatenated characters; Input the concatenated characters and the encoding features obtained previously into the encoding module to obtain new encoding features output by the encoding module, where the new encoding features are encoding features corresponding to other characters after the text characters to which the encoding features obtained last time belong; Input the new encoding features into the decoding module to obtain the second speech output by the decoding module.

[0080] In an exemplary embodiment, the encoding module is trained based on the following method: Determine a second target number based on the frame number corresponding to the current predicted encoding features that need to be output by the initial encoding module, and obtain the second target number of sample text characters in the sample text; Input the sample text characters and the previously obtained predicted coding features into the initial coding module to obtain the current predicted coding features output by the initial coding module; Construct loss information based on the current predicted coding features and the speech coding features corresponding to the sample text characters in the sample speech, where the sample speech is the speech corresponding to the sample text; Iteratively optimize the parameters of the initial coding module based on the loss information to obtain the coding module.

[0081] In an exemplary embodiment, determining the second target quantity based on the frame number corresponding to the current predicted coding features output by the initial coding module includes: Determine the second target quantity based on the frame number corresponding to the current predicted coding features, the length of the sample text characters input into the initial coding module per second, and the number of forward operations that the initial coding module can execute per second.

[0082] In an exemplary embodiment, determining the second target quantity based on the frame number corresponding to the current predicted coding features, the length of the sample text characters input into the initial coding module per second, and the number of forward operations that the initial coding module can execute per second includes: Determine the target value obtained by rounding down the ratio between the product of the frame number corresponding to the current predicted coding features and the length of the sample text characters input into the initial coding module per second, and the number of forward operations that the initial coding module can execute per second; When the target value is equal to 0, determine that the second target quantity is 1; When the target value is greater than 0, determine that the second target quantity is the target value.

[0083] The device in this embodiment can be used to execute the method of any one of the method embodiments on the speech synthesis side. Its specific implementation process and technical effects are similar to those in the method embodiments on the speech synthesis side. For specific details, reference can be made to the detailed introduction in the method embodiments on the speech synthesis side, which will not be elaborated here.

[0084] Figure 4 This is a schematic physical structure diagram of an electronic device provided by an embodiment of the present invention, as Figure 4As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a speech synthesis method, which includes: during the process of receiving streaming text, when the first target number of text characters in the target text is received, inputting the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; inputting the encoding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; and when each new text character is received, determining a second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously, and combining the first speech and each of the second speeches as the target speech corresponding to the target text.

[0085] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0086] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech synthesis method provided by each of the above methods. The method includes: during the process of receiving streaming text, when the first target number of text characters in the target text are received, input the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; input the encoding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; and when each new text character is received, determine a second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously, and use the combination of the first speech and each of the second speeches as the target speech corresponding to the target text.

[0087] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the speech synthesis method provided by each of the above methods. The method includes: during the process of receiving streaming text, when the first target number of text characters in the target text are received, input the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; input the encoding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; and when each new text character is received, determine a second speech based on the new text character, all the text characters received previously, and the encoding features obtained previously, and use the combination of the first speech and each of the second speeches as the target speech corresponding to the target text.

[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice synthesis method, characterized in that, Including: During the process of receiving streaming text, when the first target number of text characters in the target text are received, input the first target number of text characters into the encoding module of the speech synthesis model to obtain the encoding features output by the encoding module, where the first target number is less than the total number of text characters in the target text; Input the encoding features into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; When a new text character is received each time, based on the new text character, all the text characters received previously, and the encoding features obtained previously, determine a second speech, and the combination of the first speech and each of the second speeches is used as the target speech corresponding to the target text.

2. The speech synthesis method according to claim 1, wherein The first target number is determined based on the average number of words pronounced by the user per second, the preset encoding feature length, and the frame rate at which the encoding module determines the encoding features, and the average text length covered by the preset encoding feature length is less than or equal to the length of the text characters input into the encoding module each time.

3. The speech synthesis method according to claim 2, characterized in that When a new text character is received each time, based on the new text character, all the text characters received previously, and the encoding features obtained previously, determining a second speech includes: When the length of the encoding features output by the encoding module last time reaches the preset encoding feature length and a new text character is received, based on the new text character, all the text characters received previously, and the encoding features obtained previously, determine the second speech.

4. The speech synthesis method according to claim 1, wherein Based on the new text character, all the text characters received previously, and the encoding features obtained previously, determining a second speech includes: Concatenate the new text character and all the text characters received previously to obtain a concatenated character; Input the concatenated character and the encoding features obtained previously into the encoding module to obtain the new encoding features output by the encoding module, where the new encoding features are the encoding features corresponding to the other characters after the text characters to which the encoding features obtained last time belong; Input the new encoding features into the decoding module to obtain the second speech output by the decoding module.

5. The speech synthesis method according to any one of claims 1-4, characterized in that, The encoding module is trained based on the following method: Determine a second target number based on the frame number corresponding to the currently predicted encoding features that the initial encoding module needs to output, and obtain the second target number of sample text characters in the sample text; Input the sample text characters and the predicted encoding features obtained previously into the initial encoding module to obtain the currently predicted encoding features output by the initial encoding module; Based on the currently predicted encoding features and the speech encoding features in the sample speech corresponding to the sample text characters, construct loss information, where the sample speech is the speech corresponding to the sample text; Iteratively optimize the parameters of the initial encoding module based on the loss information to obtain the encoding module.

6. The speech synthesis method according to claim 5, wherein Based on the frame number corresponding to the currently predicted encoding features output by the initial encoding module, determining a second target number includes: Determine the second target quantity based on the frame number corresponding to the current prediction coding feature, the length of the sample text characters input to the initial coding module per second, and the number of forward operations that the initial coding module can perform per second.

7. The speech synthesis method according to claim 6, wherein The determining the second target quantity based on the frame number corresponding to the current prediction coding feature, the length of the sample text characters input to the initial coding module per second, and the number of forward operations that the initial coding module can perform per second includes: Determine the target value obtained by rounding down the ratio between the product of the frame number corresponding to the current prediction coding feature and the length of the sample text characters input to the initial coding module per second, and the number of forward operations that the initial coding module can perform per second; When the target value is equal to 0, determine that the second target quantity is 1; When the target value is greater than 0, determine that the second target quantity is the target value.

8. A voice synthesis device, characterized in that, including: An input module, configured to, during the process of receiving streaming text, when receiving the first target quantity of text characters in the target text, input the first target quantity of text characters into the coding module of the speech synthesis model to obtain the coding feature output by the coding module, where the first target quantity is less than the total number of text characters in the target text; The input module is further configured to input the coding feature into the decoding module of the speech synthesis model to obtain the first speech output by the decoding module; A determination module, configured to, each time a new text character is received, determine a second speech based on the new text character, all the text characters received previously, and the coding features obtained previously, and use the combination of the first speech and each of the second speeches as the target speech corresponding to the target text.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the computer program, it implements the speech synthesis method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method according to any one of claims 1 to 7.