Speech synthesis method and device
By using polyphonic characters models and target modeling tools in Chinese pronunciation synthesis, Chinese characters are converted into pinyin and combined with audio prompt data, the pronunciation inaccuracy caused by Chinese polyphonic characters is solved, and a more accurate and natural Chinese pronunciation synthesis is achieved.
Patent Information
- Application Number
- CN202510278388.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
AI Technical Summary
There is a problem of polyphonic characters in Chinese, which often leads to inaccurate pronunciation content when synthesizing Chinese pronunciation.
By obtaining the to-process Chinese character text, using the polyphonic character model to convert Chinese characters into corresponding pinyin text, and splicing the pinyin text with the pinyin text in the audio prompt data to generate extended characters. Then, the target voice is output based on the extended characters and audio prompt data using the target modeling tool and vocoder.
It effectively avoids the inaccurate pronunciation content caused by Chinese polyphonic characters, and improves the accuracy and naturalness of Chinese pronunciation synthesis.
Smart Images

Figure CN120089127A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a speech synthesis method and apparatus. Background Art
[0002] Currently, recent text-to-speech (TTS) research has made great progress. With a few seconds of audio prompt data, a TTS model can synthesize speech for any given text, and the synthesized speech can imitate the speaker of the audio prompt data. The synthesized speech can achieve high fidelity and naturalness, and is almost indistinguishable from human speech.
[0003] However, in the usage scenarios of Chinese, due to the existence of polyphonic characters in Chinese, when using a TTS model to synthesize Chinese speech, the problem of inaccurate speech content often occurs. Summary of the Invention
[0004] To solve the above technical problems, the present disclosure provides a speech synthesis method and apparatus.
[0005] In a first aspect, a speech synthesis method is provided. The method includes: obtaining a first text to be processed; using a polyphonic character model to obtain a first phonetic text corresponding to the first text; the first phonetic text is a phonetic text indicating the pronunciation of the first text; splicing the first phonetic text with a second phonetic text to obtain an extended character; the second phonetic text is a phonetic text corresponding to audio prompt data, and the audio prompt data includes audio data of a speaker speaking the second phonetic text; outputting a target speech according to the extended character and the audio prompt data; the target speech is the speech of the speaker speaking the first text.
[0006] In some implementations, the outputting a target speech according to the extended character and the audio prompt data includes: using a target modeling tool to model the extended character to obtain text hidden layer information; the target modeling tool is a neural network constructed by using a convolutional neural network and a transformer; outputting a target speech according to the text hidden layer information and the audio prompt data.
[0007] In some implementations, the outputting a target speech according to the text hidden layer information and the audio prompt data includes: obtaining target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data; using a vocoder to output a target speech corresponding to the target Mel feature data.
[0008] In some implementations, the method further includes: processing the first text to obtain the language information corresponding to the first text; obtaining the target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data, including: at each time step, obtaining the target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information, the audio prompt data, and the language information.
[0009] In a second aspect, a model training method is provided. The method includes: obtaining a training sample set, where the training sample set includes a plurality of training samples, and each training sample includes a corresponding audio-text pair; training a target model using the training samples in the training sample set; the target model includes: a polyphone model, a target modeling tool, and a synthesis module; where the polyphone model is a machine learning model that outputs a phonetic transcription text after inputting a text, and the phonetic transcription text is used to indicate the pronunciation of the text; the target modeling tool is used to model the phonetic transcription text output by the polyphone model to obtain text hidden layer information; the synthesis module is used to obtain second audio data according to the text hidden layer information obtained by the modeling tool and first audio data; where the first audio data and the second audio data are respectively a part of the audio corresponding to the text input to the polyphone model.
[0010] In some implementations, the target modeling tool is a neural network constructed using a convolutional neural network and a transformer.
[0011] In some implementations, in the synthesis module, different languages correspond to different timestep values.
[0012] In a third aspect, a voice synthesis device is provided, including:
[0013] An acquisition unit, configured to acquire a first text to be processed;
[0014] A processing unit, configured to use a polyphone model to obtain a first phonetic transcription text corresponding to the first text; the first phonetic transcription text is a phonetic transcription text indicating the pronunciation of the first text;
[0015] The processing unit is further configured to splice the first phonetic transcription text and a second phonetic transcription text to obtain an extended character; the second phonetic transcription text is the phonetic transcription text corresponding to the audio prompt data, and the audio prompt data includes audio data of a speaker speaking the second phonetic transcription text;
[0016] The processing unit is further configured to output a target voice according to the extended character and the audio prompt data; the target voice is the voice of the speaker speaking the first text.
[0017] In some implementation manners, the processing unit is further configured to output a target voice according to the extended character and the audio prompt data, including:
[0018] The processing unit is further configured to use a target modeling tool to model the extended character to obtain text hidden layer information; the target modeling tool is a neural network constructed by using a convolutional neural network and a transformer;
[0019] The processing unit is further configured to output a target voice according to the text hidden layer information and the audio prompt data.
[0020] In some implementation manners, the processing unit is further configured to output a target voice according to the text hidden layer information and the audio prompt data, including:
[0021] The processing unit is further configured to obtain target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data;
[0022] The processing unit is further configured to use a vocoder to output a target voice corresponding to the target Mel feature data.
[0023] In some implementation manners, the processing unit is further configured to process the first text to obtain language information corresponding to the first text;
[0024] The processing unit is further configured to obtain target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data, including:
[0025] The processing unit is further configured to obtain target Mel feature data corresponding to the text hidden layer information at each time step according to the text hidden layer information, the audio prompt data, and the language information.
[0026] In a fourth aspect, a model training device is provided, including: an acquisition unit, configured to acquire a training sample set, where the training sample set includes a plurality of training samples, and each training sample includes a corresponding audio-text pair;
[0027] A training unit for training a target model using training samples in the training sample set; the target model includes: a polyphonic character model, a target modeling tool, and a synthesis module; wherein, the polyphonic character model is a machine learning model that outputs a phonetic transcription text after inputting a text, and the phonetic transcription text is used to indicate the pronunciation of the text; the target modeling tool is used to model the phonetic transcription text output by the polyphonic character model to obtain text hidden layer information; the synthesis module is used to obtain second audio data according to the text hidden layer information obtained by the modeling tool and first audio data; wherein, the first audio data and the second audio data are respectively a part of the audio corresponding to the text input to the polyphonic character model.
[0028] In some implementation manners, the target modeling tool is a neural network constructed by using a convolutional neural network and a transformer.
[0029] In some implementation manners, in the synthesis module, different languages correspond to different timestep values.
[0030] In a fifth aspect, an electronic device is provided, including a processor and a communication interface, the processor receives or sends data through the communication interface, and the processor is used to implement the method according to the first aspect or any implementation manner in the first aspect or the second aspect or any implementation manner in the second aspect.
[0031] In a sixth aspect, a computer-readable storage medium is provided, and instructions are stored in the computer-readable storage medium. When the instructions run on a processor, the method according to the first aspect or any implementation manner in the first aspect or the second aspect or any implementation manner in the second aspect is implemented.
[0032] In a seventh aspect, a computer program product is provided. When the computer program product runs on a computer, the computer is enabled to implement the method described in the first aspect or any implementation manner in the first aspect or the second aspect or any implementation manner in the second aspect.
[0033] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art:
[0034] In the embodiments of the present application, it is considered that: when converting text to speech, the problem of polyphonic characters in Chinese will greatly reduce the synthesis effect and have a great impact on the user experience. Therefore, in order to avoid the problem of inaccurate speech content caused by the existence of polyphonic characters in Chinese, in the embodiments of the present application, after obtaining the first Chinese character text to be processed, the first pinyin text corresponding to the first Chinese character text is first obtained by using a polyphonic character model, and then the first pinyin text is spliced with the second pinyin text to obtain an extended character, and the target speech is output according to the extended character and the audio prompt data. In this way, the problem of inaccurate speech content caused by the existence of polyphonic characters in Chinese can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a schematic flow chart of the training process and the inference process of an F5-TTS provided by the embodiments of the present disclosure;
[0038] Figure 2 It is one of the schematic flow charts of a speech synthesis method provided by the embodiments of the present disclosure;
[0039] Figure 3 It is another schematic flow chart of a speech synthesis method provided by the embodiments of the present disclosure;
[0040] Figure 4 It is still another schematic flow chart of a speech synthesis method provided by the embodiments of the present disclosure;
[0041] Figure 5 It is one of the schematic flow charts of a model training method provided by the embodiments of the present disclosure;
[0042] Figure 6 It is a schematic structural diagram of a speech synthesis device provided by the embodiments of the present disclosure;
[0043] Figure 7 It is a schematic structural diagram of a model training device provided by the embodiments of the present disclosure;
[0044] Figure 8 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To more clearly understand the above objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0046] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0047] First, the related technologies involved in the embodiments of the present application are introduced:
[0048] Text To Speech (TTS) refers to a speech synthesis solution that synthesizes corresponding speech according to the given text. According to whether all are predicted at one time, the speech synthesis solution can be divided into two technical solutions: auto regressive (AR) and non auto regressive (NAR).
[0049] In the AR-based TTS model, an intuitive way of continuously predicting the next one or more tokens is shown, and promising zero-shot TTS capabilities have been achieved. However, the inherent limitations of AR modeling require additional efforts to solve problems such as inference latency and exposure bias. In addition, the quality of the tokenizer for speech representation is crucial for the AR model to achieve high-fidelity synthesis.
[0050] In the NAR-based TTS model, benefits can be obtained from parallel processing, enabling fast inference and effectively balancing synthesis quality and latency. It is worth noting that the diffusion model has made the greatest contribution to the success of the NAR-based TTS model. In particular, Flow Matching with Optimal Transport path (FM-OT) has been widely used in recent research fields, not only in the field of text-to-speech, but also in image generation and music generation fields.
[0051] Different from AR-based TTS models, aligning the input text with the synthesized speech is crucial and challenging for NAR-based TTS models. Although NaturalSpeech 3 and Voicebox use frame-by-frame phoneme alignment; Matcha-TTS adopts monotonic alignment search and relies on a phoneme-level duration model. However, recent research has pointed out that introducing such rigid and inflexible alignment between text and speech will hinder the model from generating more natural results. Easy end to end diffusion-based text to speech (E3 TTS) abandons the phoneme-level duration and applies cross-attention on the input sequence, but the generated audio quality is limited. Efficient and Scalable Zero-Shot Text-to-Speech with Diffusion Transformer (DiTTo-TTS) uses a diffusion transform (DiT) conditioned on text encoded by a pre-trained language model. To further enhance alignment, it uses a pre-trained language model to fine-tune a neural audio codec and injects semantic information into the generated representation. In contrast, Enhanced End-to-End Text-to-Speech (E2 TTS) based on Voicebox adopts a simpler approach. It removes the phoneme and duration predictors and directly uses the length from character padding tokens to the mel spectrogram as the input. This simple scheme also achieves very natural and realistic synthesis results. However, in the embodiments of this application, it is considered that E2 TTS has robustness issues in text and speech alignment. Seed-TTS adopts a similar strategy and achieves excellent results, but does not elaborate on the model details.
[0052] F5-TTS (Fairytaler Fakes Fluent and Faithful speech with Flowmatching), without using phoneme alignment, duration prediction, text encoder, and semantic decoder models, maintains the simplicity of the process and uses diffusion transform and ConvNeXt V2 (where "ConvNeXt V2" is a convolutional network co-designed and extended using a masked autoencoder) to better solve the text-speech alignment problem in context learning.
[0053] Exemplarily, Figure 1Among them, (a) is a schematic flowchart of the training process of F5-TTS. Among them, let x represent the audio sample in a training sample, and y represent the text sample corresponding to the audio sample x in a training sample, that is, the data pair (x, y) constitutes a training sample. Specifically, the training process of F5-TTS may include: on the one hand, extracting the mel spectrogram feature from the audio sample x, denoted as x 1 , where x 1 satisfies x 1 ∈ R F×N . Among them, F is the dimension of the mel spectrum, and N is the length of the sequence. In addition, according to the mel spectrogram feature x 1 , the noisy speech (1 - t)x 0 + tx 1 and the masked speech (1 - m) ⊙ x 1 are obtained, where x 0 represents the sampled Gaussian noise, t is the sampled time step, and m is the binary time mask in {0, 1} F×N . On the other hand, the text sample y is decomposed into a character sequence and filled with a padding token <f>Pad it to the same frame length as the Mel spectrogram to form an extended sequence z. Additionally, use the ConvNeXt V2 module to model the extended sequence z to obtain the text hidden layer information corresponding to the extended sequence z. Next, use the text hidden layer information corresponding to the extended sequence z, the noisy speech (1 - t)x 0 +tx 1 and the masked speech (1 - m)⊙x 1 to train the synthesis module. Among them, the synthesis module is used to obtain another part of the audio data corresponding to the Chinese character text according to the text hidden layer information and a part of the audio data corresponding to the Chinese character text. For example Figure 1 as shown in (a) of, the synthesis module specifically includes: a Diffusion Transformer (DiT Block), a Final Modulation module, a Linear module, and a Predicted Flow module. Among them, the synthesis module is used to reconstruct m⊙x 1 according to the masked speech (1 - m)⊙x 1 and the text hidden layer information corresponding to the extended sequence z, which is equivalent to learning the target distribution p 1 in the form of P(m⊙x 1 |(1 - m)⊙x 1 ,z) to approximate the true data distribution q.
[0054] Figure 1 In (b) of, it is a schematic flow diagram of the inference process of F5 - TTS. The inference process of F5 - TTS generally includes: obtaining the Mel spectrogram features x ref of the audio prompt data, the text y ref corresponding to the audio prompt data, and the text prompt data y gen . Among them, the audio prompt data includes the audio data of the speaker speaking the text y ref , which can be understood as: the audio prompt data is used to provide the voice characteristics of the speaker. The text prompt data y gen is used to guide the content of the generated speech. After that, referring to the training process, splice y ref and y gen and fill them with padding tokens to obtain the extended character z ref.gen . Then, according to x ref and z ref.gen , and using the synthesis module and an Ordinary Differential Equation (ODE) solver, obtain the Mel feature data. Then, use the vocoder to convert the Mel feature data into speech and output it.
[0055] It can be seen that in the above related technologies, the F5-TTS can be used to achieve the effect of synthesizing corresponding voices according to the given text. However, in the usage scenarios of languages with polyphonic characters such as Chinese, due to the problem of polyphonic characters, when using the F5-TTS to convert text into voice at present, the problem of inaccurate voice content often occurs.
[0056] In view of the above technical problems, an embodiment of the present application provides a voice synthesis method. The implementation process of the voice synthesis method provided by the embodiment of the present application will be introduced below with reference to examples.
[0057] Specifically, the execution subject of the method provided by the embodiment of the present application may be a voice synthesis device. When the voice synthesis device runs, it can be used to execute all or part of the steps in the method provided by the embodiment of the present application. Among them, in the actual application process, the functions of the voice synthesis device can be implemented by electronic devices such as personal computers (including desktop computers, laptop computers, handheld computers, and notebook computers), ultra-mobile personal computers (UMPCs), or smart phones, servers, etc.; alternatively, the functions of the above voice synthesis device can also be implemented by some hardware / software devices in the above electronic devices. The embodiment of the present application does not impose special restrictions on the specific form of the voice synthesis device.
[0058] As Figure 2 shown, the voice synthesis method provided by the embodiment of the present application may include the following steps:
[0059] S101. The voice synthesis device obtains a first text in Chinese characters to be processed.
[0060] Specifically, the first text in Chinese characters may include one or more Chinese characters.
[0061] S102. The voice synthesis device uses a polyphonic character model to obtain a first pinyin text y corresponding to the first text in Chinese characters gen .
[0062] Wherein, the first pinyin text y gen is: a phonetic text indicating the pronunciation of the first text in Chinese characters.
[0063] In practical applications, the polyphone model can be specifically a g2pW model, wherein g2pW is an improved method for grapheme-to-phoneme (g2p) mentioned in the paper "A Conditional WeightedSoftmax BERT for Polyphone Disambiguation in Mandarin".
[0064] Specifically, in the embodiment of the present application, the g2pw model can be trained using text data of polyphones to obtain a model that can capture the pronunciation rules of the text. Thus, based on the g2pW module, the complex pronunciation rules in the text can be learned to deal with the problem of polyphonetic pronunciation errors.
[0065] For example, taking the first Chinese character text "你好" as an example, the first pinyin text y gen For "nǐhǎo".
[0066] S103, the speech synthesis device converts the first pinyin text y gen With the second pinyin text y ref Concatenate and get the extended character z ref.gen .
[0067] Among them, the second pinyin text is the audio prompt data x ref Corresponding pinyin text, audio prompt data x ref The audio data of the speaker speaking the second pinyin text is included.
[0068] For example, the second pinyin text is "jīn tiān tiān qìbùcuò", the audio prompt data x ref The audio data is the speaker saying "The weather is nice today". Then the character z is expanded ref.gen It can include the combination of "nǐhǎo" and "jīntiān tiān qìbùcuò".
[0069] In addition, the extended character z ref.gen A fill marker may also be included for filling.
[0070] S104, the speech synthesis device according to the extended character z ref.gen and audio prompt data x ref , output the target speech.
[0071] The target voice is the voice of the speaker speaking the first Chinese character text.
[0072] For example, the speech synthesis device can refer to the structure of F5-TTS to obtain the first pinyin text y gen corresponding Mel feature data, and then use a vocoder to output the target speech corresponding to the target Mel feature data.
[0073] In some implementation manners, in the embodiments of the present application, it is considered that: in a model structure based on flow matching, during the process of synthesizing audio, it depends on multiple time step iteration processes. If the model performance is more robust, the required timestep will be less. In F5-TTS, generally 32 steps are required to synthesize audio with better performance. In order to shorten the value of the timestep required for synthesizing audio, in the embodiments of the present application, a neural network constructed by using a convolutional neural network and a transformer can be used to construct a target modeling tool, and then the target modeling tool is used to encode the text. With the convolutional and attention mechanisms of the target modeling tool, the ability to model the text can be effectively enhanced, which helps the model better learn the alignment between the text and the audio. Compared with using the ConvNeXtV2 module to encode the text in F5-TTS, in the embodiments of the present application, by using the target modeling tool to encode the text, normal audio can be synthesized in 16 steps or even 8 steps.
[0074] Therefore, as Figure 3 shown, the above S104 may specifically include:
[0075] S1041. The speech synthesis device uses the target modeling tool to model the extended character z ref.gen to obtain text hidden layer information.
[0076] Among them, the target modeling tool is a neural network constructed by using a convolutional neural network and a transformer.
[0077] In the actual application process, the target modeling tool can be a conformer structure. Among them, conformer is a neural network structure that combines the advantages of a convolutional neural network (CNN) and a transformer. On the one hand, the transformer is good at capturing global information interaction based on content. On the other hand, the cnn can effectively utilize local features. Therefore, the combination of the two can better learn the local and global dependencies of the audio sequence. Compared with the ConvNeXTV2 module (using depthwise separable convolution and residual modules) used in F5-TTS, conformer can better model the text context relationship. Therefore, the robustness of the text hidden layer information obtained by using the modeling tool of the conformer structure to model the extended character z ref.gen is stronger.
[0078] S1042. The speech synthesis device outputs a target speech according to the text hidden layer information and the audio prompt data x ref , and outputs a target speech.
[0079] In some designs, as Figure 4 shown, S1042 may include:
[0080] S10421. The speech synthesis device obtains target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data x ref , and obtains target Mel feature data corresponding to the text hidden layer information.
[0081] For example, in the speech synthesis device, the synthesis module and the processing process of the ODE solver used in F5-TTS can be referred to. According to the text hidden layer information and the audio prompt data x ref , the target Mel feature data corresponding to the text hidden layer information is obtained by using the synthesis module and the ODE solver. Among them, the specific structure of the synthesis module can refer to F5-TTS. For example, as Figure 1 shown, the synthesis module may include: modules such as DiT Block, Final Modulation, Linear, and PredictedFlow. Among them, the functions of each module can refer to the content in F5-TTS and will not be elaborated here.
[0082] Furthermore, in some possible designs, in the embodiments of the present application, considering that: there are different pronunciation styles in speech data of different languages, so in order to synthesize more natural and language-style-adapted audio, before executing S10421, the method may further include:
[0083] S105. Process the first text to obtain language information corresponding to the first text.
[0084] For example, the first text can be input into a pre-deployed language detection module, and then the language detection module is used to identify the language information corresponding to the first text.
[0085] Specifically, S10421 may include: at each time step, obtaining target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information, the audio prompt data, and the language information.
[0086] Specifically, in the embodiments of the present application, different languages may correspond to different timestep values. Therefore, at each time step, target Mel feature data corresponding to the text hidden layer information can be obtained according to the text hidden layer information, the audio prompt data, and the language information.
[0087] S10422. The speech synthesis device uses a vocoder to output a target speech corresponding to the target Mel feature data.
[0088] In addition, in the technical solution provided by the embodiments of the present application, a model training method is also provided. The execution subject of the model training method in the embodiments of the present application can be a model training device. When the model training device runs, it can be used to execute all or part of the steps in the model training process provided by the embodiments of the present application. Among them, in the actual application process, the functions of the model training device can be implemented by electronic devices such as personal computers (including desktop computers, laptop computers, handheld computers, and notebook computers), ultra-mobile personal computers (UMPCs), or smart phones, servers, etc.; alternatively, the functions of the above model training device can also be implemented by some hardware / software devices in the above electronic devices. The embodiments of the present application do not impose special restrictions on the specific form of the model training device.
[0089] As Figure 5 shown, the model training method may include:
[0090] S201. The model training device obtains a training sample set.
[0091] Among them, the training sample set includes multiple training samples, and each training sample includes a corresponding audio-text pair.
[0092] S202. The model training device uses the training samples in the training sample set to train a target model.
[0093] Among them, the target model includes: a polyphonic character model, a target modeling tool, and a synthesis module.
[0094] Among them, the polyphonic character model is a machine learning model that outputs phonetic annotation text after inputting the text. Among them, the phonetic annotation text is used to indicate the pronunciation of the text.
[0095] Among them, the target modeling tool is used to model the phonetic annotation text output by the polyphonic character model to obtain text hidden layer information.
[0096] Among them, the synthesis module is used to obtain second audio data according to the text hidden layer information obtained by the modeling tool and the first audio data. Among them, the first audio data and the second audio data are respectively a part of the audio corresponding to the text input into the polyphonic character model.
[0097] For example, the synthesis module may include modules such as DiT Block, Final Modulation, Linear, and PredictedFlow. The functions of each module can refer to the content in F5-TTS and will not be elaborated here.
[0098] Specifically, when the target model runs in the speech synthesis device, the speech synthesis device can use the polyphonic character model, the target modeling tool, and the synthesis module in the target model to execute the above S101-S104. Therefore, the functions of the polyphonic character model, the target modeling tool, and the synthesis module in the target model can refer to the corresponding descriptions of S101-S104 above, and the repeated content will not be elaborated here.
[0099] In some implementation manners, the target modeling tool is a neural network constructed by using a convolutional neural network and a transformer.
[0100] In some implementation manners, in the synthesis module, different languages correspond to different values of the timestep.
[0101] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application further provides a speech synthesis device. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, details in the foregoing method embodiment will not be elaborated one by one here, but it should be clear that the speech synthesis device in this embodiment can correspondingly implement all contents in the foregoing method embodiment.
[0102] As Figure 6 shown, it is a schematic structural diagram of a speech synthesis device provided by an embodiment of the present application. Among them, it includes:
[0103] An obtaining unit 301, configured to obtain a first text to be processed;
[0104] A processing unit 302, configured to use a polyphonic character model to obtain a first phonetic text corresponding to the first text; the first phonetic text is a phonetic text indicating the pronunciation of the first text;
[0105] The processing unit 302 is further configured to splice the first phonetic text and a second phonetic text to obtain an extended character; the second phonetic text is a phonetic text corresponding to audio prompt data, and the audio prompt data includes audio data of a speaker speaking the second phonetic text;
[0106] The processing unit 302 is further configured to output a target voice according to the extended character and the audio prompt data; the target voice is the voice of the speaker speaking the first text.
[0107] In some implementation manners, the processing unit 302 outputs a target voice according to the extended character and the audio prompt data, including:
[0108] The processing unit 302 uses a target modeling tool to model the extended characters, obtaining text hidden layer information; the target modeling tool is a neural network constructed using a convolutional neural network and a transformer.
[0109] The processing unit 302 is further configured to output a target voice according to the text hidden layer information and the audio prompt data.
[0110] In some implementation manners, the processing unit 302 is further configured to output a target voice according to the text hidden layer information and the audio prompt data, including:
[0111] The processing unit 302 is further configured to obtain target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data.
[0112] The processing unit 302 is further configured to use a vocoder to output a target voice corresponding to the target Mel feature data.
[0113] In some implementation manners, the processing unit 302 is further configured to process the first text to obtain language information corresponding to the first text.
[0114] The processing unit 302 is further configured to obtain target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data, including:
[0115] The processing unit 302 is further configured to obtain target Mel feature data corresponding to the text hidden layer information at each time step according to the text hidden layer information, the audio prompt data, and the language information.
[0116] The voice synthesis device 30 provided in the embodiments of the present application can execute the methods of S101 - S104, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0117] Based on the same inventive concept, as an implementation of the above method, the embodiments of the present application further provide a model training device. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be elaborated one by one in this embodiment, but it should be clear that the model training device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.
[0118] As Figure 7 shown, it is a schematic structural diagram of a model training device provided in the embodiments of the present application. Wherein, it includes:
[0119] An acquisition unit 401 acquires a training sample set, where the training sample set includes multiple training samples, and each training sample includes a corresponding audio-text pair.
[0120] A training unit 402 is configured to train a target model using the training samples in the training sample set; the target model includes: a polyphonic character model, a target modeling tool, and a synthesis module; wherein, the polyphonic character model is a machine learning model that outputs a phonetic transcription text after inputting a text, and the phonetic transcription text is used to indicate the pronunciation of the text; the target modeling tool is configured to model the phonetic transcription text output by the polyphonic character model to obtain text hidden layer information; the synthesis module is configured to obtain second audio data according to the text hidden layer information obtained by the modeling tool and first audio data; wherein, the first audio data and the second audio data are respectively a part of the audio corresponding to the text input to the polyphonic character model.
[0121] In some implementation manners, the target modeling tool is a neural network constructed using a convolutional neural network and a transformer.
[0122] In some implementation manners, in the synthesis module, different languages correspond to different timestep values.
[0123] The model training device 40 provided in the embodiments of the present application can execute the methods of S201-S202, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0124] Based on the same inventive concept, the embodiments of the present application further provide an electronic device. Figure 8 As shown in the structure diagram of the electronic device provided in the embodiments of the present application, Figure 8 the electronic device provided in this embodiment includes: a memory 501 and a processor 502, the memory 501 is used to store a computer program, and the processor 502 is used to make the data processing device execute any method provided in the above embodiments when executing the computer program.
[0125] Based on the same inventive concept, the embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computing device is enabled to implement the methods provided in the above embodiments.
[0126] Based on the same inventive concept, the embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, the computing device is enabled to implement the methods provided in the above embodiments.
[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code.
[0128] The processor can be a central processing unit (CPU), or it can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0129] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0130] Computer-readable media include permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.< / f>
Claims
1. A speech synthesis method, characterized in that: The method comprises: Get the first text to be processed; Using the polyphonetic character model, obtaining a first phonetic text corresponding to the first text; the first phonetic text is: a phonetic text indicating the pronunciation of the first text; The first phonetic text is concatenated with the second phonetic text to obtain an extended character; the second phonetic text is a phonetic text corresponding to the audio prompt data, and the audio prompt data includes audio data of a speaker speaking the second phonetic text; A target voice is output according to the extended characters and the audio prompt data; the target voice is the voice of the speaker speaking the first text.
2. The method according to claim 1, characterized in that: The step of outputting a target voice according to the extended characters and the audio prompt data comprises: Using a target modeling tool, the extended characters are modeled to obtain text hidden layer information; the target modeling tool is a neural network constructed using a convolutional neural network and a transformer; Output the target speech according to the text hidden layer information and the audio prompt data.
3. The method according to claim 2, characterized in that Outputting a target voice according to the text hidden layer information and the audio prompt data includes: According to the text hidden layer information and the audio prompt data, obtaining target Mel feature data corresponding to the text hidden layer information; The target speech corresponding to the target Mel feature data is outputted using a vocoder.
4. The method according to claim 3, characterized in that: The method further comprises: Processing the first text to obtain language information corresponding to the first text; The step of obtaining target Mel feature data corresponding to the text hidden layer information according to the text hidden layer information and the audio prompt data includes: At each time step, target Mel feature data corresponding to the text hidden layer information is obtained according to the text hidden layer information, the audio prompt data and the language information.
5. A model training method, characterized in that: The method comprises: Acquire a training sample set, wherein the training sample set includes a plurality of training samples, each training sample includes a corresponding audio-text pair; The target model is trained using the training samples in the training sample set; the target model includes: a polyphone model, a target modeling tool, and a synthesis module; wherein the polyphone model is: a machine learning model that outputs a phonetic text after inputting a text, and the phonetic text is used to indicate the pronunciation of the text; the target modeling tool is used to model the phonetic text output by the polyphone model to obtain text hidden layer information; the synthesis module is used to obtain second audio data based on the text hidden layer information obtained by the modeling tool and the first audio data; wherein the first audio data and the second audio data are respectively a part of the audio corresponding to the text input to the polyphone model.
6. The method according to claim 5, characterized in that The target modeling tool is a neural network constructed using a convolutional neural network and a transformer.
7. The method according to claim 5 or 6, characterized in that: In the synthesis module, different languages correspond to different timestep values.
8. A speech synthesis device, characterized in that: The speech synthesis device comprises: An acquisition unit, used for acquiring a first text to be processed; A processing unit, configured to obtain a first phonetic text corresponding to the first text by using a polyphonetic character model; the first phonetic text is a phonetic text indicating the pronunciation of the first text; The processing unit is further used to concatenate the first phonetic text with the second phonetic text to obtain an extended character; the second phonetic text is a phonetic text corresponding to the audio prompt data, and the audio prompt data includes audio data of a speaker speaking the second phonetic text; The processing unit is further used to output a target voice according to the extended characters and the audio prompt data; the target voice is the voice of the speaker speaking the first text.
9. An electronic device, characterized in that: The method comprises a processor and a communication interface, wherein the processor receives or sends data through the communication interface, and the processor is used to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a processor, the method according to any one of claims 1 to 7 is implemented.