Speech synthesis processing method, apparatus, and related device
By generating discrete semantic tokens through self-supervised methods and establishing text-speech alignment relationships, the problem of inaccurate speech tagging in traditional technologies is solved, thereby improving the realism and accuracy of synthesized speech in self-service voice services in the fintech field.
Patent Information
- Application Number
- CN202411465306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-10-18
AI Technical Summary
In traditional text-to-speech synthesis processing based on large language models, the generated speech tags are inaccurate, resulting in low similarity between synthesized speech and real human speech, which affects the quality and efficiency of self-service voice services in fintech fields such as insurance.
A self-supervised approach is used to generate discrete semantic tokens corresponding to speech signals. An alignment relationship between text and discrete semantic tokens is established through a pre-set text-speech large language model. Speech signals are generated using a pre-set conditional flow matching model and a speech synthesis decoder to achieve few-sample speech synthesis.
It improves the similarity between synthesized speech and human speech, making the synthesized speech more realistic and accurate, and enhancing the quality and efficiency of self-service voice services in the fintech field.
Smart Images

Figure CN119479608B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of financial technology, and in particular to a speech synthesis processing method and device, computer equipment and a computer readable storage medium. BACKGROUND
[0002] A large language model (English for Large Language Models, abbreviated as LLM) is an artificial intelligence model trained with a large amount of data, aiming to understand and generate natural language text. The large language model is usually based on deep learning technology and can capture the complexity and diversity of language.
[0003] Among them, the text-to-speech synthesis model based on the large language model has become a common processing method for speech synthesis due to its high naturalness and zero-shot capability. In the traditional technology, the text-to-speech synthesis processing based on the large language model generally discretizes the speech signal to generate speech tags, and performs speech synthesis according to the speech tags. The speech tag is metadata that describes the synthesized speech, such as the start and end positions of a sentence or a word in an audio stream. The speech tag plays a key role in the large language model-based speech synthesis model.
[0004] In the traditional technology, the speech tags generated in the text-to-speech synthesis process based on the large language model are learned in an unsupervised manner. Unsupervised learning is a way of automatic learning by automatically discovering the structure and pattern of speech data from unlabeled speech data. Since unsupervised learning cannot be trained with labeled speech data, the speech tags generated in the text-to-speech synthesis process based on the large language model are inaccurate, which in turn leads to a low similarity between the synthesized speech and the human speech. For example, in the field of financial technology, such as in an insurance company, the self-service voice service based on the text-to-speech synthesis of the large language model for insurance business, in the case of generating labeled speech based on an unsupervised manner, will result in a low similarity between the synthesized speech and the human speech for the self-service voice service of the insurance business, reducing the quality and level of the self-service voice service of the insurance business, affecting the self-service level of the insurance business, and failing to effectively reduce the human, material and financial costs required for the development of the insurance business through the self-service voice service of the insurance business.
[0005] Therefore, it is an urgent technical problem to improve the similarity between the speech synthesis based on the large language model and the human speech. SUMMARY
[0006] The present application provides a speech synthesis processing method, device, computer equipment and computer readable storage medium, which can solve the technical problem of low similarity between the synthesized speech and the human speech in the traditional technology.
[0007] In a first aspect, the present application provides a speech synthesis processing method, comprising: obtaining text to be converted into speech, and obtaining a prompt text; determining discrete semantic tokens; establishing an alignment relationship between the text and the discrete semantic tokens according to the prompt text based on a preset text-to-speech large language model, to obtain target discrete semantic tokens corresponding to the text; determining speech spectrum features corresponding to the target discrete semantic tokens based on a preset conditional flow matching model; and generating speech signals from the speech spectrum features based on a preset speech synthesis decoder, to obtain speech corresponding to the text.
[0008] In a second aspect, the present application provides a speech synthesis processing apparatus, comprising: a first obtaining unit configured to obtain text to be converted into speech, and obtain a prompt text; a first determining unit configured to determine discrete semantic tokens; a first establishing unit configured to establish an alignment relationship between the text and the discrete semantic tokens according to the prompt text based on a preset text-to-speech large language model, to obtain target discrete semantic tokens corresponding to the text; a second determining unit configured to determine speech spectrum features corresponding to the target discrete semantic tokens based on a preset conditional flow matching model; and a first generating unit configured to generate speech signals from the speech spectrum features based on a preset speech synthesis decoder, to obtain speech corresponding to the text.
[0009] In a third aspect, the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the speech synthesis processing method when executing the computer program.
[0010] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the processor executes the steps of the speech synthesis processing method when executing the computer program.
[0011] The application provides a speech synthesis processing method and device, computer equipment and a computer readable storage medium. The method comprises the following steps: obtaining text to be converted into speech, obtaining prompt text, determining discrete semantic tokens, establishing an alignment relationship between the text and the discrete semantic tokens based on a preset text-to-speech large language model according to the prompt text, obtaining target discrete semantic tokens corresponding to the text, determining speech spectrum features corresponding to the target discrete semantic tokens based on a preset conditional flow matching model, and generating speech signals from the speech spectrum features based on a preset speech synthesis decoder to obtain speech corresponding to the text. The method realizes few-shot speech synthesis based on self-supervised semantic tokens, can realize semantic self-supervision by using a small amount of input speech signals, can generate text-speech pairs by using a large amount of unlabeled data, makes the synthesized speech more vivid and realistic, and makes the synthesized speech more similar to human speech, thereby improving the similarity between the synthesized speech and human speech and realizing more accurate synthesized speech. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1 The flowchart of the speech synthesis processing method provided by the embodiments of the present application is shown in the figure.
[0014] Figure 2 The first sub-flowchart of the speech synthesis processing method provided by the embodiments of the present application is shown in the figure.
[0015] Figure 3 The second sub-flowchart of the speech synthesis processing method provided by the embodiments of the present application is shown in the figure.
[0016] Figure 4 The schematic block diagram of the speech synthesis processing device provided by the embodiments of the present application is shown in the figure.
[0017] Figure 5 The schematic block diagram of the computer equipment provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0020] This application provides a speech synthesis processing method, which can be applied to, but is not limited to, computer devices such as smartphones, laptops, desktop computers, servers, and intelligent robots, and is used in the field of financial technology, including but not limited to, when providing self-service voice services for financial transactions, for speech synthesis processing.
[0021] To address the technical problem of low similarity between synthesized speech and real human speech in traditional technologies, the inventors propose a speech synthesis processing method according to embodiments of this application. The core idea of this application is to generate semantic tokens representing semantic labels corresponding to speech signals using a self-supervised approach. These semantic tokens are derived from the recognition of a small number of speech signals based on a speech recognition model. Based on the semantic tokens, when it is necessary to generate speech corresponding to text, an alignment relationship between the text and the semantic tokens is established. Then, the semantic tokens are synthesized into speech, realizing a few-sample speech synthesis based on self-supervised semantic tokens, from speech signals to semantic tokens, and then from semantic tokens to the speech corresponding to the text. This method can achieve self-supervision using a small number of input speech signals, thereby using a large amount of unlabeled data to generate text-speech pairs, making the synthesized speech more rich and realistic, and the similarity between the synthesized speech and real human speech is more consistent, thus improving the similarity between synthesized speech and real human speech, and making the synthesized speech more accurate than real human speech.
[0022] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0023] Please see Figure 1 , Figure 1 This is a flowchart illustrating the speech synthesis processing method provided in an embodiment of this application. Figure 1 As shown, the method includes, but is not limited to, the following steps S11-S15:
[0024] S11. Obtain the text to be converted into speech and obtain the prompt text.
[0025] Explanatorily, in application scenarios including but not limited to the financial technology field, such as insurance, banking, securities and other businesses, there is usually a demand for synthesizing text into corresponding speech, in which case, the text to be converted into speech, i.e., the text needs to be converted into speech, is obtained to obtain the speech signal corresponding to the text, which can be further played or combined with images to form a video for real person simulation.
[0026] Exemplarily, in the financial technology field, in an example, users can perform self-service voice services through the telephone platform or APP platform of the insurance business, and there is a scenario of converting text into corresponding speech, and playing the synthesized speech to the user to provide self-service handling of the insurance business. In another example, for the deposit and loan business of a bank, users can consult through a telephone platform, can perform self-service voice handling of the bank business through an APP, and can also perform self-service voice handling of the financial business through a self-service ATM machine, in which case, the corresponding text is converted into synthesized speech, and the synthesized speech is played to the user as much as possible to provide self-service handling of the bank business.
[0027] Not only the text to be converted into speech is obtained, but also the prompt text is obtained. The prompt text is a piece of text input provided for guiding the speech synthesis model to generate corresponding output. The prompt text can be a question, an instruction, a lead or any other form of text, which sets the context and lets the speech synthesis model understand the user's intent and generate a relevant answer or complete a task. The prompt text can be an instant input text or a pre-stored text. Exemplarily, in the financial technology field, for the insurance business, in the business scenario of users performing self-service voice services through the telephone platform or APP platform of the insurance business, the prompt text can be pre-stored for guiding the speech synthesis model to generate corresponding speech output in the case of converting text into corresponding speech by the embodiments of the present application.
[0028] S12, determining discrete semantic tokens.
[0029] Explanatorily, Tokens: In natural language processing, text is first segmented into smaller units called tokens, which can be words, phrases or characters, and models understand and generate text through these tokens.
[0030] Semantic tokens, English for Semantic Tokens, represent the smallest indivisible unit of semantics in speech. Semantic tokens are used to capture local dependencies and global long-term structures in speech based on semantics. The process of obtaining semantic tokens includes identifying predicates and their arguments in speech and labeling these arguments with the semantic roles they play.
[0031] The discrete semantic token represents a voice token discretized from a continuous voice signal, wherein the "discrete" in the discrete semantic token is only used to distinguish different semantic tokens, and does not limit the semantic token, and other similar descriptions are used in the same way.
[0032] The determination of the discrete semantic token includes but is not limited to obtaining a voice signal, discretizing the voice signal to obtain an initial voice token, identifying the recognized text corresponding to the voice signal, then aligning the recognized text with the initial voice token to obtain the discrete semantic token corresponding to the recognized text, thereby discretizing the continuous voice signal to obtain the corresponding discrete token, realizing the determination of the discrete semantic token based on the voice signal, and being able to obtain the discrete semantic token based on the self-supervised manner according to a small amount of voice signal without labeling. The obtained voice signal includes but is not limited to a training sample voice signal and a user input voice signal, and the user input voice signal includes but is not limited to a real voice corresponding to the prompt text, that is, the discrete semantic token includes but is not limited to the following two sources: 1) the training sample voice signal when training the voice synthesis model; and 2) the user input voice signal, and the discrete semantic token can be obtained from the two voice signal sources.
[0033] Exemplarily, in the field of financial technology, for an insurance business, when performing voice synthesis through a self-service voice service of a telephone platform or an APP platform of the insurance business, in order to generate voice meeting the user's inclination, the prompt text and the prompt voice corresponding to the prompt text can be required to be input, and the prompt voice is used as the voice signal to determine the corresponding discrete semantic token.
[0034] Compared with the voice token generated based on unsupervised learning in the prior art, the discrete semantic token used in the embodiment of the present application can improve the accuracy of the synthesized voice, and compared with the voice token generated based on supervised learning in the prior art, a large number of labeled training samples are not required, and the corresponding discrete semantic token can be automatically generated in time according to the voice signal obtained in time, such as the prompt voice, so as to improve the efficiency, flexibility and automation of generating the discrete semantic token.
[0035] S13, based on the preset text voice large language model, establishing an alignment relationship between the text and the discrete semantic token according to the prompt text, to obtain a target discrete semantic token corresponding to the text.
[0036] Explanatorily, the preset text-to-speech large language model, i.e., the preset text-to-speech large language model, represents a large language model for converting text into speech. The large language model (LLM) is an artificial intelligence model trained on a large amount of data, aiming to understand and generate natural language text. The preset text-to-speech large language model is an autoregressive model that establishes an alignment relationship between text and discrete semantic tokens using an autoregressive inference mechanism. The autoregressive inference mechanism predicts the probability distribution of subsequent words based on the previous text, enabling the LLM to fully utilize the information from the previous text to generate coherent and reasonable text, thereby improving the accuracy of the target discrete semantic tokens corresponding to the text. In this embodiment, the alignment relationship is a precise match between the text and the discrete semantic tokens in time, ensuring that the text and the corresponding audio signal are accurately matched when generating the speech corresponding to the text. Each character, word, sentence, etc. in the text corresponds to the correct time period in the audio signal. A text corresponds to several discrete speech tokens, i.e., several discrete speech tokens represent a text, and the text is a single word, a word, a phrase, or a sentence.
[0037] Based on the preset text-to-speech large language model, the autoregressive inference is used to establish the alignment relationship between the text and the discrete semantic tokens according to the guidance of the prompt text, and the target discrete semantic tokens corresponding to the text are obtained, thereby accurately matching the text and the discrete semantic tokens in time to ensure the accuracy of the synthesized speech.
[0038] S14, based on the preset condition flow matching model, determine the speech spectrum feature corresponding to the target discrete semantic token.
[0039] Explanatorily, the preset condition flow matching model, i.e., the preset condition flow matching model, is used to establish the relationship between the target discrete semantic token and the speech spectrum feature based on the preset condition flow matching model to determine the speech spectrum feature corresponding to the target discrete semantic token. Since the target discrete semantic token is the result of aligning the text and the discrete semantic token, the speech spectrum feature corresponding to the target discrete semantic token determined based on the preset condition flow matching model is the relationship between the target discrete semantic token and the continuous acoustic feature. The condition flow refers to determining different speech spectrum feature operations or processes based on given conditions. The condition flow matching model, in English, is a model that matches the target discrete semantic token and the speech spectrum feature based on the condition flow. The condition flow matching model can generate detailed waveforms. The speech spectrum feature represents different components and patterns in speech, and the speech spectrum feature is derived from the training sample speech signal when training the speech synthesis model.
[0040] S15, generate a speech signal based on the preset speech synthesis decoder, and obtain the speech corresponding to the text.
[0041] Illustratively, a decoder including but not limited to HIFI-GAN or HIFTNet, i.e., a preset speech synthesis decoder, is used to restore the speech spectrum feature into a speech signal to obtain the speech corresponding to the text, wherein HIFI-GAN is a neural vocoder based on a generative adversarial network, and HiFTNet is a neural vocoder that combines inverse short-time Fourier transform (iSTFT) and harmonic noise filter.
[0042] Illustratively, in the field of financial technology, for insurance business, when a user performs self-service voice service through a telephone platform or an APP platform of the insurance business, a decoder based on HIFI-GAN or HIFTNet is used to restore the speech spectrum feature into a speech signal to obtain the speech corresponding to the text, and the speech is played to the user to provide self-service voice service.
[0043] In the embodiments of the present application, the text to be converted into speech is obtained, the prompt text is obtained, and the discrete semantic token is determined. Then, based on the preset text-to-speech large language model, the alignment relationship between the text and the discrete semantic token is established according to the prompt text to obtain the target discrete semantic token corresponding to the text. Then, based on the preset conditional flow matching model, the speech spectrum feature corresponding to the target discrete semantic token is determined. Based on the preset speech synthesis decoder, the speech spectrum feature is generated into a speech signal to obtain the speech corresponding to the text. The few-shot speech synthesis based on self-supervised semantic token is realized. The semantic self-supervision can be realized by using a small amount of input speech signal. Therefore, a large amount of unlabeled data can be used to generate text-speech pairs. The synthesized speech is more realistic and rich, and the similarity between the synthesized speech and the real human speech is higher. Therefore, the similarity between the synthesized speech and the real human speech is improved, and the synthesized speech is more accurate relative to the real human speech.
[0044] In an embodiment, please refer to Figure 2 , Figure 2 The first sub-process schematic diagram of the speech synthesis processing method provided by the embodiments of the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the discrete semantic token is determined, including:
[0045] S21, obtain an initial speech signal, and perform down-sampling on the initial speech signal based on a preset speech encoder to obtain a speech latent vector;
[0046] S22, discretize the speech latent vector based on a preset speech quantizer to obtain a discretized speech;
[0047] S23, restore the discretized speech into a speech signal based on a preset speech decoder to obtain a generated speech;
[0048] S24, identify the generated speech based on a preset speech recognition model to obtain recognized text;
[0049] S25, establish an alignment relationship between the discretized speech and the recognized text by using a preset text-to-speech large language model based on an autoregressive manner to obtain discrete semantic tokens.
[0050] Explanatorily, an initial speech signal is acquired. In the case of training a speech synthesis model, the initial speech signal can be a training sample speech signal. In the inference case of using the speech synthesis model for speech synthesis, the initial speech signal can be a speech signal corresponding to a speech provided by a user at will, including but not limited to a prompt speech corresponding to a prompt text. Exemplarily, in the field of financial technology, for an insurance business, a prompt speech can be provided when performing speech synthesis through a self-service voice service of a telephone platform or an APP platform of the insurance business.
[0051] Based on a preset speech encoder, the initial speech signal is down-sampled to obtain a high-dimensional speech latent vector. Exemplarily, a 1024-dimensional speech signal is down-sampled to a 128-dimensional or 256-dimensional continuous speech latent vector. The working principle of the speech encoder is to convert the speech signal into a digital signal through a compression algorithm. Down-sampling (English for Downsampling) means reducing the sampling rate of the speech signal, that is, reducing the number of speech data points.
[0052] Based on a preset speech quantizer, the speech latent vector is discretized to obtain discretized speech. Based on the preset speech quantizer, the speech latent vector is discretized, that is, the continuous speech latent vector is quantized in the amplitude dimension to quantize the continuous amplitude value into discrete. Common quantization methods include linear and nonlinear.
[0053] Then, based on a preset speech decoder, the discretized speech is restored into a speech signal to obtain generated speech. Then, based on a preset speech recognition model, the generated speech is identified to obtain recognized text. Then, based on a preset text-to-speech large language model based on an autoregressive manner, an alignment relationship between the discretized speech and the recognized text is established to obtain discrete semantic tokens. Thus, only based on the speech signal, the discrete semantic tokens are acquired based on a self-supervised manner, so that the discrete semantic tokens are obtained without labeling according to a small amount of speech signal.
[0054] The embodiment of the application can realize semantic self-supervision by using a small amount of input voice signals, and then generate a large amount of text-voice pairs using unlabelled data, so that the synthesized voice is more realistic and accurate, and the similarity between the synthesized voice and the real voice is more matched, thereby improving the similarity between the synthesized voice and the real voice, and realizing more accurate synthesized voice relative to the real voice.
[0055] In an embodiment, the initial voice signal includes a first language and a second language; based on a preset voice recognition model, the generated voice is recognized to obtain a recognized text, including:
[0056] Based on a preset multilingual voice recognition model, the first language included in the generated voice is recognized to obtain a first sub-recognized text, and the second language included in the generated voice is recognized to obtain a second sub-recognized text;
[0057] According to the voice positions of the first language and the second language in the initial voice signal, the first sub-recognized text and the second sub-recognized text are combined into an overall text to obtain a recognized text.
[0058] Explanatorily, a multilingual voice recognition model is preset, i.e., a preset multilingual voice recognition model, which represents a model capable of recognizing voices of multiple languages.
[0059] In the case where the initial voice signal includes but is not limited to the first language and the second language, based on the preset multilingual voice recognition model, the first language included in the generated voice is recognized to obtain a first sub-recognized text, and the second language included in the generated voice is recognized to obtain a second sub-recognized text, wherein the "first" related to the first language and the "second" related to the second language are only used to distinguish different languages, and are not used to limit different languages. Exemplarily, in the field of financial technology, for insurance business, when voice synthesis is performed through self-service voice service of a telephone platform or an APP platform of the insurance business, the initial voice signal includes but is not limited to Chinese, English, or Chinese, French.
[0060] In the case where the initial voice signal includes but is not limited to the first language and the second language, when the generated voice is recognized based on the preset multilingual voice recognition model, the voice positions of the first language and the second language in the initial voice signal are recorded, and according to the voice positions of the first language and the second language in the initial voice signal, the first sub-recognized text and the second sub-recognized text are combined into an overall text to obtain a recognized text.
[0061] The embodiment of the application can generate discrete semantic tokens corresponding to multiple languages, support voice synthesis of multiple languages, and even support text-to-speech of mixed languages, and can further improve the similarity between synthesized speech and real human speech, and realize more accurate synthesized speech relative to real human speech.
[0062] In an embodiment, please refer to Figure 3 , Figure 3 The second sub-flow diagram of the voice synthesis processing method provided by the embodiment of the application is shown in FIG. 6. As shown in the figure, in this embodiment, based on a preset text-to-speech large language model, an alignment relationship between the text and the discrete semantic tokens is established according to the prompt text, to obtain target discrete semantic tokens corresponding to the text, including: Figure 3
[0063] S31, determining a speaker feature corresponding to the text;
[0064] S32, splicing the speaker feature and the text to obtain a spliced text;
[0065] S33, based on a preset text-to-speech large language model, an alignment relationship between the spliced text and the discrete semantic tokens is established according to the prompt text, to obtain target discrete semantic tokens corresponding to the text.
[0066] Explanatorily, the speaker feature corresponding to the text is determined, and the speaker feature includes but is not limited to speaker information and emotional representation. The speaker information includes but is not limited to features such as tone, pitch, and gender of the speaker, which describe the speaker information. The emotional representation includes but is not limited to features such as happiness, sadness, depression, and high pitch, which describe the speaker's emotion.
[0067] Then the speaker feature and the text are spliced to obtain a spliced text. Exemplarily, the speaker information and emotional representation included in the speaker feature are obtained, and then a sequence {discrete semantic token, speaker information, emotional representation, text} is formed to splice the speaker information, emotional representation, and text, so as to embed the speaker feature into the text to obtain the spliced text. The spliced text corresponds to as much rich content as the sound should have as possible, and is not limited to text only. Then, based on the preset text-to-speech large language model, an alignment relationship between the text and the discrete semantic tokens is established according to the prompt text, to obtain target discrete semantic tokens corresponding to the text.
[0068] Further, the speaker feature is spliced with the text to obtain spliced text, including:
[0069] The speaker feature is spliced with the text to obtain spliced text.
[0070] Explanatorily, the discrete semantic token TOKEN, the speaker information, the emotion representation, and the text composition sequence form, that is, {discrete semantic token, speaker information, emotion representation, and text}, are obtained to obtain spliced text, where the form of the spliced text includes but is not limited to a vector and a long vector. Then, based on a preset text voice large language model, an alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text, a target discrete semantic token corresponding to the text is obtained, and then synthesized voice is generated.
[0071] In the embodiments of the present application, the speaker feature is embedded into the text to obtain a corresponding target discrete semantic token, and then synthesized voice is generated, so that the synthesized voice can be more realistic and rich, the similarity between the synthesized voice and the real voice is more matched, the similarity between the synthesized voice and the real voice is improved, and the synthesized voice is more accurate relative to the real voice.
[0072] In an embodiment, the determining of the speaker feature corresponding to the text includes:
[0073] It is determined whether the prompt voice is input.
[0074] In a case where it is determined that the prompt voice is input, the prompt voice is obtained, and the speaker feature corresponding to the speaker included in the prompt voice is extracted to obtain the speaker feature corresponding to the text.
[0075] In a case where it is determined that the prompt voice is not input, the speaker feature corresponding to the text is randomly determined.
[0076] Explanatorily, it is determined whether the prompt voice is input. In a case where it is determined that the prompt voice is input, the prompt voice is obtained. The prompt voice can be voice corresponding to the prompt text, that is, the voice expression of the prompt text. The prompt text is the text expression of the prompt voice. The speaker feature corresponding to the speaker included in the prompt voice can be extracted by using a method including but not limited to voice feature extraction, speaker modeling, speaker recognition, and the like to obtain the speaker feature corresponding to the text. Further, in a case where it is determined that the prompt voice is not input, the speaker feature corresponding to the text is randomly determined.
[0077] The embodiments of the present application can flexibly determine the speaker characteristics corresponding to the text by judging whether the prompt voice is input, can meet the user's voice trend demand for the synthesized voice as much as possible, thereby further making the synthesized voice more vivid and rich, and the synthesized voice and the real voice more similar, thereby improving the similarity of the synthesized voice and the real voice, and realizing that the synthesized voice is more accurate relative to the real voice.
[0078] In an embodiment, based on a preset text-to-speech large language model, an alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text, and a target discrete semantic token corresponding to the text is obtained, including:
[0079] Based on a preset text-to-speech large language model, an alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text, and an initial discrete semantic token corresponding to the text is obtained;
[0080] Determine the length of the voice corresponding to the spliced text;
[0081] According to the length of the voice, the information of the redundant length contained in the initial discrete semantic token is cut off, and the target discrete semantic token corresponding to the text is obtained.
[0082] Explanatorily, based on a preset text-to-speech large language model, an alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text, an initial discrete semantic token corresponding to the text is obtained first, the initial discrete semantic token is the token corresponding to the entire spliced text, and then the length of the voice corresponding to the spliced text is determined according to the prompt text., and then according to the length of the voice, the information of the redundant length contained in the initial discrete semantic token is cut off, and the target discrete semantic token corresponding to the text is obtained. The target discrete semantic token is a subset of the initial discrete semantic token, and then the synthesized voice is generated according to the target discrete semantic token.
[0083] The embodiments of the present application can further make the synthesized voice more accurate by controlling the length of the target discrete semantic token corresponding to the text, thereby making the synthesized voice and the real voice more similar, thereby improving the similarity of the synthesized voice and the real voice, and realizing that the synthesized voice is more accurate relative to the real voice.
[0084] It should be noted that the voice synthesis processing method described in each of the above embodiments can combine the technical features contained in different embodiments as needed to obtain a combined embodiment, but all within the scope of protection claimed by the present application.
[0085] Please refer to Figure 4 , Figure 4A schematic block diagram of a speech synthesis processing apparatus is provided for the embodiments of the present application. Corresponding to the speech synthesis processing method described above, the embodiments of the present application also provide a speech synthesis processing apparatus. As shown in Figure 4 , the speech synthesis processing apparatus includes units for performing the speech synthesis processing method described above, and the speech synthesis processing apparatus can be configured in a computer device. Specifically, please refer to Figure 4 , the speech synthesis processing apparatus 40 includes a first acquisition unit 41, a first determination unit 42, a first establishment unit 43, a second determination unit 44, and a first generation unit 45.
[0086] The first acquisition unit 41 is configured to acquire text to be converted into speech and acquire prompt text.
[0087] The first determination unit 42 is configured to determine discrete semantic tokens.
[0088] The first establishment unit 43 is configured to establish an alignment relationship between the text and the discrete semantic tokens based on a preset text-to-speech large language model according to the prompt text, to obtain target discrete semantic tokens corresponding to the text.
[0089] The second determination unit 44 is configured to determine speech spectrum features corresponding to the target discrete semantic tokens based on a preset condition flow matching model.
[0090] The first generation unit 45 is configured to generate speech signals based on a preset speech synthesis decoder, to obtain speech corresponding to the text.
[0091] In an embodiment, the first determination unit 42 includes:
[0092] A subsampling subunit is configured to acquire an initial speech signal and perform subsampling on the initial speech signal based on a preset speech encoder, to obtain a speech latent vector.
[0093] A discretization subunit is configured to perform discretization on the speech latent vector based on a preset speech quantizer, to obtain a discretized speech.
[0094] A recovery subunit is configured to recover the discretized speech into a speech signal based on a preset speech decoder, to obtain generated speech.
[0095] A first recognition subunit is configured to recognize the generated speech based on a preset speech recognition model, to obtain recognized text.
[0096] A first establishment subunit is configured to establish an alignment relationship between the discretized speech and the recognized text using a preset text-to-speech large language model based on an autoregressive manner, to obtain discrete semantic tokens.
[0097] In an embodiment, the initial speech signal comprises a first language and a second language; the first recognition subunit comprises:
[0098] a second recognition subunit configured to recognize the first language contained in the generated speech based on a preset multi-language speech recognition model to obtain a first sub-recognition text, and recognize the second language contained in the generated speech to obtain a second sub-recognition text;
[0099] a combination subunit configured to combine the first sub-recognition text and the second sub-recognition text into an overall text according to the speech positions of the first language and the second language in the initial speech signal to obtain a recognition text.
[0100] In an embodiment, the first establishment unit 43 comprises:
[0101] a first determination subunit configured to determine a speaker feature corresponding to the text;
[0102] a splicing subunit configured to splice the speaker feature and the text to obtain a spliced text;
[0103] a second establishment subunit configured to establish an alignment relationship between the spliced text and the discrete semantic token based on a preset text-speech large language model according to the prompt text to obtain a target discrete semantic token corresponding to the text.
[0104] In an embodiment, the first determination subunit comprises:
[0105] a judgment subunit configured to judge whether a prompt speech is inputted;
[0106] a first judgment subunit configured to, in a case where it is judged that the prompt speech is inputted, acquire the prompt speech, and extract a speaker feature corresponding to a speaker contained in the prompt speech to obtain the speaker feature corresponding to the text.
[0107] In an embodiment, the first determination subunit further comprises:
[0108] a second judgment subunit configured to, in a case where it is judged that the prompt speech is not inputted, randomly determine the speaker feature corresponding to the text.
[0109] In an embodiment, the second establishment subunit comprises:
[0110] a third establishment subunit configured to establish an alignment relationship between the spliced text and the discrete semantic token based on a preset text-speech large language model according to the prompt text to obtain an initial discrete semantic token corresponding to the text.
[0111] The second determining sub-unit is configured to determine a speech length corresponding to the spliced text.
[0112] The cutting sub-unit is configured to cut information of a redundant length contained in the initial discrete semantic token according to the speech length, to obtain a target discrete semantic token corresponding to the text.
[0113] It should be noted that the specific implementation process of the speech synthesis processing apparatus and each unit can be clearly understood by those skilled in the art, and can refer to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, the details are not described herein.
[0114] Meanwhile, the division and connection mode of each unit in the speech synthesis processing apparatus are only used for illustration, and in other embodiments, the speech synthesis processing apparatus can be divided into different units as needed, or each unit in the speech synthesis processing apparatus can adopt different connection order and mode, to complete all or part of the functions of the speech synthesis processing apparatus.
[0115] The speech synthesis processing apparatus can be implemented in the form of a computer program, which can run on a computer device as shown in the accompanying drawings. Figure 5
[0116] Please refer to Figure 5 , Figure 5 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a desktop computer or a server computer, or can be a component or part of other devices.
[0117] Referring to Figure 5 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501, wherein the memory can include a non-volatile storage medium 503 and an internal memory 504, and the memory can also be a volatile storage medium.
[0118] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can execute an above-described speech synthesis processing method.
[0119] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.
[0120] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an above-described speech synthesis processing method.
[0121] The network interface 505 is configured to perform network communication with other devices. Those skilled in the art can understand that the network interface 505 can be implemented by various network interfaces, such as an Ethernet, a modem, a Bluetooth, a Wi-Fi, a radio frequency, and the like. Figure 5 The structure shown in FIG. 5 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. Specifically, the computer device 500 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. For example, in some embodiments, the computer device can only include a memory and a processor. In such embodiments, the structure and functions of the memory and the processor are consistent with those of the memory 501 and the processor 502 shown in the embodiments, and will not be described here again. Figure 5 The structure shown in FIG. 5 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. Specifically, the computer device 500 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. For example, in some embodiments, the computer device can only include a memory and a processor. In such embodiments, the structure and functions of the memory and the processor are consistent with those of the memory 501 and the processor 502 shown in the embodiments, and will not be described here again.
[0122] The processor 502 is configured to run the computer program 5032 stored in the memory to implement the following steps: obtaining text to be converted into speech, and obtaining prompt text; determining discrete semantic tokens; based on a preset text-to-speech large language model, establishing an alignment relationship between the text and the discrete semantic tokens according to the prompt text, to obtain target discrete semantic tokens corresponding to the text; based on a preset conditional flow matching model, determining speech spectrum features corresponding to the target discrete semantic tokens; and based on a preset speech synthesis decoder, generating speech signals from the speech spectrum features to obtain speech corresponding to the text.
[0123] In an embodiment, when implementing the determination of the discrete semantic tokens, the processor 502 specifically implements the following steps:
[0124] obtaining an initial speech signal, and performing down-sampling on the initial speech signal based on a preset speech encoder to obtain a speech latent vector;
[0125] performing discretization on the speech latent vector based on a preset speech quantizer to obtain discretized speech;
[0126] restoring the discretized speech into a speech signal based on a preset speech decoder to obtain generated speech;
[0127] performing recognition on the generated speech based on a preset speech recognition model to obtain recognized text;
[0128] adopting a preset text-to-speech large language model based on an autoregressive manner to establish an alignment relationship between the discretized speech and the recognized text, to obtain discrete semantic tokens.
[0129] In an embodiment, when implementing that the initial speech signal contains a first language and a second language, and performing recognition on the generated speech based on a preset speech recognition model to obtain recognized text, the processor 502 specifically implements the following steps:
[0130] The first language contained in the generated voice is recognized based on a preset multi-language voice recognition model to obtain a first sub-recognized text, and the second language contained in the generated voice is recognized to obtain a second sub-recognized text.
[0131] The first sub-recognized text and the second sub-recognized text are combined into an overall text according to the voice positions of the first language and the second language in the initial voice signal to obtain a recognized text.
[0132] In an embodiment, when the processor 502 implements the alignment relationship between the text and the discrete semantic token according to the prompt text based on the preset text-voice large language model, the following steps are specifically implemented:
[0133] The speaker feature corresponding to the text is determined.
[0134] The speaker feature and the text are spliced to obtain a spliced text.
[0135] The alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text based on the preset text-voice large language model to obtain the target discrete semantic token corresponding to the text.
[0136] In an embodiment, when the processor 502 implements the determination of the speaker feature corresponding to the text, the following steps are specifically implemented:
[0137] It is judged whether a prompt voice is input.
[0138] In a case where it is judged that the prompt voice is input, the prompt voice is obtained, and the speaker feature corresponding to the speaker contained in the prompt voice is extracted to obtain the speaker feature corresponding to the text.
[0139] In an embodiment, when the processor 502 implements the method, the following steps are further implemented:
[0140] In a case where it is judged that the prompt voice is not input, the speaker feature corresponding to the text is randomly determined.
[0141] In an embodiment, when the processor 502 implements the alignment relationship between the spliced text and the discrete semantic token according to the prompt text based on the preset text-voice large language model to obtain the target discrete semantic token corresponding to the text, the following steps are specifically implemented:
[0142] The alignment relationship between the spliced text and the discrete semantic token is established according to the prompt text based on the preset text-voice large language model to obtain the initial discrete semantic token corresponding to the text.
[0143] determine a speech length corresponding to the spliced text;
[0144] According to the speech length, cut off redundant length information contained in the initial discrete semantic token to obtain a target discrete semantic token corresponding to the text.
[0145] It should be understood that, in the embodiments of the present application, the processor 502 can be a central processing unit (CPU), and the processor 502 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0146] It can be understood by those skilled in the art that all or part of the processes in the method of the above-mentioned embodiments can be completed by a computer program, which can be stored in a computer readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments of the method.
[0147] Therefore, the present application also provides a computer readable storage medium. The computer readable storage medium can be a non-volatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores a computer program, which is executed by a processor to make the processor execute the following steps:
[0148] A computer program product, when running on a computer, causes the computer to execute the steps of the speech synthesis processing method described in the above embodiments.
[0149] The computer readable storage medium can be an internal storage unit of the device, such as a hard disk or a memory of the device. The computer readable storage medium can also be an external storage device of the device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the device.
[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0151] The storage medium is an entity, non-transient storage medium, for example, can be a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a magnetic disk or an optical disk, and various entity storage media that can store computer programs.
[0152] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the foregoing description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0153] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0154] The steps in the method embodiments of the present application can be adjusted, combined and deleted according to actual needs. The units in the device embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0155] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the parts that make contributions to the prior art, or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for making an electronic device (which can be a personal computer, a terminal or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application.
[0156] The non-company software tools or components appearing in the embodiments of the present application are only illustrative and do not represent actual use.
[0157] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A speech synthesis processing method, characterized in that, include: Get the text to be converted into speech, and get the prompt text; Determine a discrete semantic token, wherein the discrete semantic token represents a semantic token obtained by discretizing a continuous speech signal, and the semantic token represents the smallest indivisible unit of semantics in speech; Based on a preset text-speech large language model, an alignment relationship is established between the text and the discrete semantic token according to the prompt text, so as to obtain the target discrete semantic token corresponding to the text; Based on a preset conditional flow matching model, the speech spectrum features corresponding to the target discrete semantic token are determined; Based on a preset speech synthesis decoder, the speech spectrum features are used to generate a speech signal, thereby obtaining the speech corresponding to the text; Among them, determining discrete semantic tokens includes: An initial speech signal is acquired, and based on a preset speech encoder, the initial speech signal is downsampled to obtain a speech latent vector; Based on a preset speech quantizer, the speech latent vector is discretized to obtain discretized speech; Based on a preset speech decoder, the discrete speech is restored into a speech signal to obtain the generated speech; Based on a preset speech recognition model, the generated speech is recognized to obtain the recognized text; A pre-defined text-speech large language model based on autoregression is used to establish the alignment relationship between the discrete speech and the recognized text, thereby obtaining discrete semantic tokens.
2. The speech synthesis processing method according to claim 1, characterized in that, The initial speech signal contains a first language and a second language; based on a preset speech recognition model, the generated speech is recognized to obtain recognized text, including: Based on a preset multilingual speech recognition model, the first language contained in the generated speech is recognized to obtain a first sub-recognized text, and the second language contained in the generated speech is recognized to obtain a second sub-recognized text; Based on the speech positions of the first language and the second language in the initial speech signal, the first sub-recognized text and the second sub-recognized text are combined into a whole text to obtain the recognized text.
3. The speech synthesis processing method according to claim 1 or 2, characterized in that, Based on a pre-defined text-speech large language model, an alignment relationship is established between the text and the discrete semantic token according to the prompt text, to obtain the target discrete semantic token corresponding to the text, including: Determine the speaker characteristics corresponding to the text; The speaker features are concatenated with the text to obtain the concatenated text; Based on a preset text-speech large language model, an alignment relationship is established between the concatenated text and the discrete semantic token according to the prompt text, so as to obtain the target discrete semantic token corresponding to the text.
4. The speech synthesis processing method according to claim 3, characterized in that, Determining the speaker features corresponding to the text includes: Determine if a prompt voice has been entered; If a prompt voice is detected, the prompt voice is acquired, and the speaker features corresponding to the speaker contained in the prompt voice are extracted to obtain the speaker features corresponding to the text.
5. The speech synthesis processing method according to claim 4, characterized in that, The method further includes: If no prompt voice is input, the speaker characteristics corresponding to the text are randomly determined.
6. The speech synthesis processing method according to claim 3, characterized in that, Based on a preset text-speech large language model, an alignment relationship is established between the concatenated text and the discrete semantic token according to the prompt text, to obtain the target discrete semantic token corresponding to the text, including: Based on a preset text-speech large language model, an alignment relationship is established between the concatenated text and the discrete semantic token according to the prompt text, so as to obtain the initial discrete semantic token corresponding to the text; Determine the length of the speech corresponding to the concatenated text; Based on the speech length, the redundant information contained in the initial discrete semantic token is removed to obtain the target discrete semantic token corresponding to the text.
7. A speech synthesis processing device, characterized in that, include: The first acquisition unit is used to acquire the text to be converted into speech and to acquire the prompt text; The first determining unit is used to determine a discrete semantic token, wherein the discrete semantic token represents a semantic token obtained by discretizing a continuous speech signal, and the semantic token represents the smallest indivisible unit of semantics in speech. The first establishment unit is used to establish an alignment relationship between the text and the discrete semantic token based on a preset text-speech large language model and the prompt text, so as to obtain the target discrete semantic token corresponding to the text. The second determining unit is used to determine the speech spectrum features corresponding to the target discrete semantic token based on a preset conditional flow matching model. The first generation unit is used to generate a speech signal from the speech spectrum features based on a preset speech synthesis decoder, thereby obtaining the speech corresponding to the text; The first determining unit includes: The sampling subunit is used to acquire the initial speech signal and, based on a preset speech encoder, downsample the initial speech signal to obtain the speech latent vector. The discretization subunit is used to discretize the speech latent vector based on a preset speech quantizer to obtain discretized speech; The recovery subunit is used to recover the discrete speech into a speech signal based on a preset speech decoder to obtain the generated speech; The first recognition subunit is used to recognize the generated speech based on a preset speech recognition model to obtain recognized text; The first subunit is used to establish the alignment relationship between the discrete speech and the recognized text using a pre-defined text-speech large language model based on autoregression, thereby obtaining a discrete semantic token.
8. A computer device, characterized in that, The computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program to perform the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Artificial-intelligence-based text-to-speech method and apparatus, and computer device and medium
WO2022141870A1
Systems and methods for using neural codec language model for zero-shot cross-lingual text-to-speech synthesis
WO2024178710A1