Speech synthesis method and device, electronic equipment and storage medium

By introducing a speech synthesis model trained with target attribute text and multi-attribute speech samples, the problem of cumbersome processes in generating speech of different languages, styles, and dialects in existing technologies is solved, and efficient and controllable diversified speech synthesis effects are achieved.

CN120895024AActive Publication Date: 2025-11-04IFLYTEK CO LTD

Patent Information

Application Number
CN202511439834.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-04
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing large-scale speech synthesis models require training multiple independent models to generate speech in different languages, styles, and dialects, making the speech synthesis process cumbersome and complex, increasing costs and reducing efficiency.

Method used

By introducing target attribute text and multi-attribute speech samples to train the speech synthesis model, and using the target encoder, target encoding module and target decoding module, combined with an autoregressive approach, the target synthesized speech is generated, supporting the control of audio attributes such as language switching, environmental sound effects and dialect generation.

Benefits of technology

It enables control over the expressiveness and prosody of synthesized speech according to user needs, improving the controllability and efficiency of speech synthesis, and supporting the generation of diverse audio attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895024A_ABST
    Figure CN120895024A_ABST
Patent Text Reader

Abstract

The invention provides a speech synthesis method and device, electronic equipment and a storage medium, and relates to the technical field of speech synthesis, and the method introduces a target attribute text in a speech synthesis process, can support speech synthesis with an audio attribute corresponding to the target attribute text, and improves the speech synthesis efficiency. Therefore, the expressive force and rhythm of the target synthetic speech can be controlled according to the user demand, so that the target synthetic speech better meets the user demand, and the user experience is improved. A speech synthesis model is obtained through training of a text sample with an attribute tag, so that the speech synthesis model has an audio attribute control capability during speech synthesis, and audio attributes such as tone, speech style, emotional expression, human settings, mood, rhythm and the like of target synthesis speech can be controlled at the same time; and audio attributes such as language switching, environment sound effect, dialect generation and the like can be supported, and controllability is improved while the generation quality of the speech synthesis model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, in particular to a speech synthesis method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Speech synthesis, also known as text-to-speech (TTS), aims to convert input text into fluent and natural output speech, and is a key technology for realizing intelligent human-computer voice interaction.

[0003] A mature TTS small model aims to synthesize clear speech with high sound quality and low errors, and pursues quality, with weaker prosodic performance and emotional performance of synthesized audio. A speech synthesis large model pursues to reduce the gap between synthesized speech and real communication speech, such as hoping to synthesize speech with large prosodic fluctuations and changes in an audiobook scenario, and hoping to synthesize speech with higher anthropomorphism and empathy in an interactive scenario, and so on, and the tolerance of pronunciation errors such as mouth movement is improved.

[0004] Currently, mainstream speech synthesis large models mainly include an encoder and a decoder. In the training stage, phonemes and discrete acoustic features are spliced as the input of the encoder, and the discrete acoustic features are taken as the training target, and an autoregressive and non-autoregressive combination is used for prediction; the predicted discrete acoustic features are converted into speech waveforms by the decoder. In the test stage, one scheme is to splice the phoneme sequence of the prompt, the phoneme sequence of the text to be synthesized, and the acoustic features of the prompt as the input of the encoder, and to synthesize the audio corresponding to the text to be synthesized with the habit and tone of the prompt through the decoder. Another scheme is to only input the phoneme sequence of the text to be synthesized, and to synthesize the audio with a random tone through the decoder.

[0005] Although the synthesized audio obtained by using the above speech synthesis large model has greater improvement in richness and naturalness compared to the TTS small model, it only supports basic speech generation. If different languages, different styles, different dialects, etc. are needed to be synthesized, multiple independent models generally need to be trained, the process of synthesizing speech is complicated, and the cost of synthesizing speech is increased, and the efficiency of synthesizing speech is reduced. SUMMARY

[0006] The present application provides a speech synthesis method, device, electronic equipment and storage medium to solve the defects in the related art.

[0007] The present application provides a speech synthesis method, comprising: obtaining a text to be synthesized and a target attribute text; inputting the to-be-synthesized text and the target attribute text into a speech synthesis model to obtain target synthesized speech corresponding to the to-be-synthesized text output by the speech synthesis model; The speech synthesis model is trained based on multi-attribute speech samples, text samples with attribute labels corresponding to the multi-attribute speech samples, and discrete acoustic feature label sequences of the multi-attribute speech samples.

[0008] According to the speech synthesis method provided in the application, the speech synthesis model comprises a target encoder, a target encoding module, and a target decoding module. The inputting the to-be-synthesized text and the target attribute text into a speech synthesis model to obtain target synthesized speech corresponding to the to-be-synthesized text output by the speech synthesis model comprises: fusing the to-be-synthesized text and the target attribute text to obtain fused text; inputting the fused text into the target encoder to obtain a first discrete acoustic feature corresponding to the to-be-synthesized text output by the target encoder; inputting the fused text and the first discrete acoustic feature into the target encoding module to obtain a final discrete acoustic feature sequence determined by the target encoding module in an autoregressive manner; inputting the to-be-synthesized text, the target attribute text, and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module.

[0009] According to the speech synthesis method provided in the application, the target decoding module comprises a first encoder, a second encoder, and a decoder; and the inputting the to-be-synthesized text, the target attribute text, and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module comprises: inputting the to-be-synthesized text into the first encoder to obtain a first semantic feature output by the first encoder; inputting the target attribute text into the second encoder to obtain a second semantic feature output by the second encoder; inputting the first semantic feature, the second semantic feature, and the final discrete acoustic feature sequence into the decoder to obtain the target synthesized speech output by the decoder.

[0010] According to the speech synthesis method provided in the application, the target attribute text comprises sentence-level attribute text and local-level attribute text; and the fusing the to-be-synthesized text and the target attribute text to obtain fused text comprises: Splice the sentence-level attribute text with the to-be-synthesized text, and insert the local-level attribute text into the to-be-synthesized text to obtain the fusion text.

[0011] According to the speech synthesis method provided by the application, the training step of the target coding module comprises: A pre-training coding module is obtained, which is trained based on discrete acoustic feature data of multi-attribute speech data and multi-attribute text corresponding to attribute-free labels of the multi-attribute speech data; The text sample and the first discrete acoustic feature sample corresponding to the text sample are input into the pre-training coding module to obtain a final discrete acoustic feature sample sequence determined by the pre-training coding module in an autoregressive manner. Based on the final discrete acoustic feature sample sequence and the discrete acoustic feature label sequence, a first loss is calculated, and the pre-training coding module is iteratively trained based on the first loss to obtain the target coding module.

[0012] According to the speech synthesis method provided by the application, the training step of the target coding module comprises: The non-attribute label and the attribute label in the text sample are input into a first initial encoder and a second initial encoder of an initial decoding module respectively to obtain a first initial semantic feature output by the first initial encoder and a second initial semantic feature output by the second initial encoder. The first initial semantic feature, the second initial semantic feature, and the discrete acoustic feature label sequence are input into an initial decoder of the initial decoding module to obtain an initial synthesized speech output by the initial decoder. Based on the initial synthesized speech and the multi-attribute speech sample, a second loss is calculated, and the initial decoding module is iteratively trained based on the second loss to obtain the target decoding module.

[0013] According to the speech synthesis method provided by the application, the target attribute text comprises text of at least one type of audio attribute in language, dialect, style, emotion, speaker gender, age, voice characteristics, body state, background audio, and inserted audio.

[0014] The application further provides a speech synthesis device, comprising: A text acquisition module is configured to acquire to-be-synthesized text and target attribute text. A speech synthesis module is configured to input the to-be-synthesized text and the target attribute text into a speech synthesis model to obtain target synthesized speech corresponding to the to-be-synthesized text output by the speech synthesis model. The speech synthesis model is trained based on multi-attribute speech samples, text samples with attribute labels corresponding to the multi-attribute speech samples, and discrete acoustic feature label sequences of the multi-attribute speech samples.

[0015] The application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the speech synthesis method according to any one of the above when executing the computer program.

[0016] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the speech synthesis method according to any one of the above.

[0017] The application further provides a computer program product, which includes a computer program, and the computer program is executable on a processor to implement the speech synthesis method according to any one of the above.

[0018] The speech synthesis method, device, electronic device, and storage medium provided by the application first acquire a text to be synthesized and a target attribute text, and then perform speech synthesis on the text to be synthesized by taking the target attribute text as a control signal through a speech synthesis model to obtain target synthesized speech. The method introduces the target attribute text in the speech synthesis process, can support speech synthesis with audio attributes corresponding to the target attribute text, and can further control the expressiveness and rhythm of the target synthesized speech according to user needs, so that the target synthesized speech is more in line with user needs and user experience is improved. The speech synthesis model is trained through text samples with attribute labels, so that the speech synthesis model has the control ability of audio attributes during speech synthesis, can control the timbre, speech style, emotional expression, persona, tone, rhythm, and other audio attributes of the target synthesized speech, and can support audio attributes such as language switching, environmental sound effects, and dialect generation, to ensure the generation quality of the speech synthesis model and improve the controllability. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0020] Figure 1 is a structural schematic diagram of an existing speech synthesis large model.

[0021] Figure 2 is a flowchart of the speech synthesis method provided by the application.

[0022] Figure 3 is a structural schematic diagram of the target decoding module provided by the present application.

[0023] Figure 4 is an input and output schematic diagram of the target encoding module provided by the present application.

[0024] Figure 5 is an input and output schematic diagram of two stages in the training process of the target encoding module provided by the present application.

[0025] Figure 6 is a structural schematic diagram of the speech synthesis device provided by the present application.

[0026] Figure 7 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0027] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0028] As shown in Figure 1 , the existing speech synthesis large model usually includes an encoder and a decoder. The input of the encoder includes the phoneme sequence of the text to be synthesized and the current acoustic feature sequence. The output of the encoder is the acoustic feature of the next frame, which is used to update the current acoustic feature sequence. The updated acoustic feature sequence can be marked by a start identifier (start_id) and an end identifier (end_id). Finally, the final acoustic feature sequence output by the encoder is input to the decoder, which is converted into an audio waveform by the decoder.

[0029] Compared with the TTS small model, the existing speech synthesis large model has greater improvement in the richness and naturalness of synthesized audio, but it only supports basic speech generation. In order to simultaneously realize speech synthesis with different languages, different styles, different dialects, etc., multiple independent models need to be trained, resulting in a complex process of synthesizing speech, increasing the cost of synthesizing speech, and reducing the efficiency of synthesizing speech. Based on this, a speech synthesis method is provided in the embodiments of the present application.

[0030] Figure 2 is a flowchart of the speech synthesis method provided in the embodiments of the present application, as shown in Figure 2 , the method comprises: S1, obtaining a text to be synthesized and a target attribute text; S2, input the to-be-synthesized text and the target attribute text into a speech synthesis model to obtain target synthesized speech corresponding to the to-be-synthesized text output by the speech synthesis model; The speech synthesis model is trained based on multi-attribute speech samples, text samples with attribute labels corresponding to the multi-attribute speech samples, and discrete acoustic feature label sequences of the multi-attribute speech samples.

[0031] Specifically, the speech synthesis method provided in the embodiments of the present application has a speech synthesis device as an execution subject. The device can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not limited here.

[0032] First, step S1 is performed to obtain to-be-synthesized text and target attribute text provided by a user. The to-be-synthesized text refers to text that needs to be used to generate target synthesized speech. The to-be-synthesized text is used as the text content of the generated target synthesized speech.

[0033] The target attribute text is used to describe text of at least one type of audio attribute of the target synthesized speech, such as language, dialect, speech style, emotional expression, speaker gender, age, voice characteristics, physical state, background sound, and inserted sound.

[0034] Then, step S2 is performed to input the to-be-synthesized text and the target attribute text into a speech synthesis model, which can be a large language model (LLM). The speech synthesis model performs speech synthesis on the to-be-synthesized text by taking the target attribute text as control text to obtain and output target synthesized speech corresponding to the to-be-synthesized text.

[0035] Here, the speech synthesis model can be obtained by training an initial synthesis model based on multi-attribute speech samples, text samples with attribute labels corresponding to the multi-attribute speech samples, and discrete acoustic feature label sequences of the multi-attribute speech samples. The multi-attribute speech samples can be speech data of different languages, different dialects, different speech styles, and different emotional expressions obtained by recording or audio and video data of different sources collected.

[0036] The text sample is a text obtained by transcribing the multi-attribute speech sample and annotated with an attribute label. The text sample is used as the text content of the multi-attribute speech sample. The multi-attribute speech sample corresponds to a known discrete acoustic feature label sequence. The discrete acoustic feature label sequence is a sequence obtained by arranging the discrete acoustic features of each frame of speech in the multi-attribute speech sample in time sequence. Here, the discrete acoustic features can be discrete audio encoding.

[0037] A feasible training scheme for the initial synthesis model can include: inputting the text sample and the discrete acoustic feature label sequence into the initial synthesis model to obtain a synthesis result output by the initial synthesis model, then calculating a loss function value according to the synthesis result and the multi-attribute speech sample, and finally updating model parameters of the initial synthesis model according to the loss function value; the input process and the calculation process are iteratively performed until the loss function converges or a preset iteration number is reached, to obtain the speech synthesis model. The preset iteration number can be set as needed, and is not specifically limited here.

[0038] The speech synthesis method provided in the embodiments of the present application first acquires a text to be synthesized and a target attribute text; then performs speech synthesis on the text to be synthesized by taking the target attribute text as a control signal through a speech synthesis model, to obtain target synthesized speech. This method introduces the target attribute text in the speech synthesis process, can support speech synthesis with audio attributes corresponding to the target attribute text, and can further control the expressiveness and prosody of the target synthesized speech according to user needs, so that the target synthesized speech is more in line with user needs and user experience is improved. The speech synthesis model is trained by using text samples with attribute labels, so that the speech synthesis model has the control ability of audio attributes during speech synthesis, can control the timbre, speech style, emotional expression, persona, tone, prosody and other audio attributes of the target synthesized speech, and can support audio attributes such as language switching, environmental sound effects and dialect generation, to ensure the generation quality of the speech synthesis model and improve the controllability.

[0039] and further ensure speech synthesis with audio attributes meeting user needs.

[0040] On the basis of the above embodiments, the speech synthesis model includes a target encoder, a target encoding module and a target decoding module. The inputting of the text to be synthesized and the target attribute text into the speech synthesis model to obtain the target synthesized speech corresponding to the text to be synthesized output by the speech synthesis model includes: fusing the text to be synthesized and the target attribute text to obtain a fused text; inputting the fused text into the target encoder to obtain a first discrete acoustic feature corresponding to the text to be synthesized output by the target encoder; inputting the fused text and the first discrete acoustic feature into the target encoding module to obtain a final discrete acoustic feature sequence determined by the target encoding module in an autoregressive manner; inputting the text to be synthesized, the target attribute text and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module.

[0041] Specifically, before the text to be synthesized and the target attribute text are input into the speech synthesis model, the text to be synthesized and the target attribute text can be fused to obtain a fused text. Here, the fusion manner can be splicing, and the splicing position can be determined according to actual conditions.

[0042] In the embodiment of the present application, the speech synthesis model includes a target encoder, a target encoding module and a target decoding module. The fused text is input into the target encoder, and the target encoder can generate the first discrete acoustic feature corresponding to the text to be synthesized, that is, the first discrete acoustic feature corresponding to the target synthesized speech corresponding to the text to be synthesized.

[0043] Thereafter, the fused text and the first discrete acoustic feature are input into the target encoding module, and the target encoding module can iteratively determine the final discrete acoustic feature sequence corresponding to the text to be synthesized in an autoregressive manner. For example, the target encoding module can predict the second discrete acoustic feature according to the fused text and the first discrete acoustic feature. Thereafter, the target encoding module can continue to predict the third discrete acoustic feature according to the fused text and the existing discrete acoustic feature sequence, that is, the first discrete acoustic feature and the second discrete acoustic feature. The target encoding module can continue to predict the fourth discrete acoustic feature according to the fused text and the existing discrete acoustic feature sequence, that is, the first discrete acoustic feature, the second discrete acoustic feature and the third discrete acoustic feature. The prediction is continued until the last discrete acoustic feature is predicted. Finally, all the obtained discrete acoustic features can be arranged in time sequence to form the final discrete acoustic feature sequence.

[0044] It can be understood that the input of the target encoding module includes the fused text and the discrete acoustic feature sequence, and in particular, the first input discrete acoustic feature sequence only contains the first discrete acoustic feature.

[0045] Thereafter, the text to be synthesized, the target attribute text and the final discrete acoustic feature sequence are input into the target decoding module, and the target decoding module can convert the final discrete acoustic feature sequence into an audio waveform in combination with the text to be synthesized and the target attribute text to obtain the target synthesized speech.

[0046] In the embodiment of the present application, the target encoding module determines the final discrete acoustic feature sequence in an autoregressive manner. Each time, only the fused text and the historical discrete acoustic feature determined need to be input to predict the subsequent discrete acoustic feature, without the need for additional variables or complex data collection processes. Moreover, the processing process of the target decoding module introduces the text to be synthesized and the target attribute text, which can provide a reference for the processing process and reduce the difficulty of the processing process. The specific structure and implementation logic of the speech synthesis model can reduce the difficulty of speech synthesis and improve the efficiency of speech synthesis.

[0047] On the basis of the above-mentioned embodiments, the target decoding module comprises a first encoder, a second encoder and a decoder; the inputting of the text to be synthesized, the target attribute text and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module comprises: inputting the text to be synthesized into the first encoder to obtain the first semantic feature output by the first encoder; inputting the target attribute text into the second encoder to obtain the second semantic feature output by the second encoder; inputting the first semantic feature, the second semantic feature and the final discrete acoustic feature sequence into the decoder to obtain the target synthesized speech output by the decoder.

[0048] Specifically, in the embodiments of the present application, as shown in the figure, Figure 3 the target decoding module can comprise a first encoder (incoder), a second encoder and a decoder (decoder).

[0049] In the process of determining the target synthesized speech, the target decoding module can first input the text to be synthesized and the target attribute text into the first encoder and the second encoder respectively, extract the first semantic feature of the text to be synthesized through the first encoder, and extract the second semantic feature of the target attribute text through the second encoder.

[0050] Thereafter, the first semantic feature, the second semantic feature and the final discrete acoustic feature sequence are inputted into the decoder, and the decoder utilizes the first semantic feature and the second semantic feature to convert the final discrete acoustic feature sequence into the target synthesized speech.

[0051] In the embodiments of the present application, by introducing the first encoder and the second encoder into the target decoding module, the text to be synthesized and the target attribute text are converted into semantic features, which facilitates the participation in the conversion process of the final discrete acoustic feature sequence, so that the obtained target synthesized speech is more in line with the semantics of the text to be synthesized and the target attribute text.

[0052] On the basis of the above-mentioned embodiments, the target attribute text comprises sentence-level attribute text and local-level attribute text; the fusion of the text to be synthesized and the target attribute text to obtain a fusion text comprises: splicing the sentence-level attribute text and the text to be synthesized, and inserting the local-level attribute text into the text to be synthesized to obtain the fusion text.

[0053] Specifically, the target attribute text used in the embodiments of the present invention may include sentence-level attribute text and local-level attribute text. Sentence-level attribute text is used to describe audio attributes applicable to the entire target synthesized speech, while local-level attribute text is used to describe audio attributes corresponding to text near the location of the local-level attribute text in the text to be synthesized.

[0054] When fusing the text to be synthesized with the target attribute text, sentence-level attribute text can be directly concatenated before the text to be synthesized, and local-level attribute text can be inserted into the corresponding positions within the text to be synthesized to obtain the fused text. The text to be synthesized and each target attribute text can be separated by identifiers. Identifiers can include start and end identifiers. Similarly, when inputting the fused text and each discrete acoustic feature sequence into the target encoding module, the fused text and the discrete acoustic feature sequence can also be concatenated and separated by identifiers.

[0055] For example, such as Figure 4 As shown, the input to the target encoding module can be represented as: <attr><Chinese>One wise old man, with low and steady voice, not in a hurry to tell the past< / attr> <Bos_t> I remember back in my prime, when I traveled and encountered injustices, I would always step in to help those in need.<Eos_t><Bos_a> Discrete acoustic feature sequences<Eos_a> ,in" <attr>"and"< / attr> "" represents the start and end identifiers of the sentence-level attribute text, "<coughing sound>" represents the identifier of the local-level attribute text, indicating that a coughing sound appears at the corresponding position in the target synthesized audio, and "<Chinese>" represents the language of the text to be synthesized.<Bos_t> "and"<Eos_t> "These are the start and end identifiers of the text to be synthesized, respectively."<Bos_a> "and"<Eos_a> "" represents the start identifier and end identifier of the discrete acoustic feature sequence, respectively.

[0056] Figure 4 In this context, the discrete acoustic feature sequence can also include a speaker identifier (spkid) to distinguish discrete acoustic feature sequences from different speakers. It should be noted that... Figure 4 The discrete acoustic feature sequence in the middle contains the discrete acoustic feature sequence currently input to the target coding module and the next discrete acoustic feature output by the target coding module.

[0057] Based on the above embodiments, the training steps of the target encoding module include: A pre-trained encoding module is obtained, which is trained based on the discrete acoustic feature data of multi-attribute speech data and the multi-attribute text without attribute labels corresponding to the multi-attribute speech data. input the text sample and a first discrete acoustic feature sample corresponding to the text sample to a pre-training encoding module to obtain a final discrete acoustic feature sample sequence determined by the pre-training encoding module in an autoregressive manner; based on the final discrete acoustic feature sample sequence and the discrete acoustic feature label sequence, calculate a first loss, and based on the first loss, iteratively train the pre-training encoding module to obtain the target encoding module.

[0058] Specifically, in the embodiments of the present application, the target encoding module and the target decoding module in the speech synthesis model can be independently trained or jointly trained. When the target encoding module is trained independently, a pre-training encoding module can be obtained first, which can be obtained by training an initial encoding module using the discrete acoustic feature data of multi-attribute speech data and the multi-attribute text corresponding to the attribute-free label of the multi-attribute speech data. The multi-attribute speech data can be speech data covering enough audio attributes such as languages, dialects, styles, etc., and can include millions of hours of speech data to ensure that the pre-training encoding module has strong general speech synthesis capability.

[0059] The discrete acoustic feature data of the multi-attribute speech data can be obtained by inputting the multi-attribute speech data to the encoder of Encodec.

[0060] During the training process of the pre-training encoding module, the multi-attribute text can be first input to the initial encoding module to obtain the output result of the initial encoding module. According to the output result of the initial encoding module and the discrete acoustic feature data, the loss function value is calculated, and the model parameters of the initial encoding model are updated according to the loss function value. The input process and the calculation process are iteratively executed until the loss function converges or reaches a specified number of iterations to obtain the pre-training encoding module. The specified number of iterations can be set as needed, which is not limited here.

[0061] Subsequently, the text sample and its corresponding first discrete acoustic feature sample are input into the pre-training encoding module, resulting in the final discrete acoustic feature sample sequence determined by the pre-training encoding module using an autoregressive approach. For example, the pre-training encoding module can predict the second discrete acoustic feature sample based on the text sample and the first discrete acoustic feature sample. It can then continue to predict the third discrete acoustic feature sample based on the text sample and the existing sequence (i.e., the first and second discrete acoustic feature samples), and so on, until the last discrete acoustic feature sample is predicted. Finally, all the obtained discrete acoustic feature samples can be arranged in chronological order to form the final discrete acoustic feature sample sequence.

[0062] Subsequently, using the final discrete acoustic feature sample sequence and discrete acoustic feature label sequence, the first loss can be calculated using the following formula: ; Where P is the first loss. This represents a discrete acoustic feature sample sequence containing t+1 discrete acoustic feature samples. Let represent a discrete acoustic feature sample sequence containing t discrete acoustic feature samples, where T represents the difference between the number of discrete acoustic feature samples in the final discrete acoustic feature sample sequence and 1, and x represents a text sample containing attribute labels. These are the model parameters for the pre-trained encoding module. Indicates the use of And the deviation calculated from the corresponding subsequence in the discrete acoustic feature label sequence.

[0063] Finally, the model parameters of the pre-trained encoding module are iteratively trained using the first loss. When the first target number of iterations is reached or the first loss converges, the target encoding module is obtained.

[0064] like Figure 5 As shown, the training process of the target encoding module can include two stages: the first stage is to train the initial encoding module to obtain the pre-trained encoding module, and the second stage is to train the pre-trained encoding module to obtain the target encoding module.

[0065] In this embodiment of the invention, a pre-trained encoding module is first obtained by training discrete acoustic feature data and multi-attribute text, and then the pre-trained encoding module is trained, which can reduce the requirement for the number of text samples and thus reduce the model training cost.

[0066] Based on the above embodiments, the training steps of the target decoding module include: The non-attribute tags and attribute tags in the text sample are respectively input to the first initial encoder and the second initial encoder of the initial decoding module to obtain the first initial semantic feature output by the first initial encoder and the second initial semantic feature output by the second initial encoder. The first initial semantic feature, the second initial semantic feature, and the discrete acoustic feature label sequence are input into the initial decoder of the initial decoding module to obtain the initial synthesized speech output by the initial decoder; Based on the initial synthesized speech and the multi-attribute speech samples, a second loss is calculated, and based on the second loss, the initial decoding module is iteratively trained to obtain the target decoding module.

[0067] Specifically, the text samples contain both non-attribute labels and attribute labels. Non-attribute labels refer to the text in the text sample excluding the attribute labels. When training the target decoding module separately, the non-attribute labels and attribute labels in the text samples can be input into the first initial encoder and the second initial encoder of the initial decoding module, respectively, to obtain the first initial semantic features output by the first initial encoder and the second initial semantic features output by the second initial encoder.

[0068] Then, the first initial semantic features, the second initial semantic features, and the discrete acoustic feature label sequence are input into the initial decoder of the initial decoding module to obtain the initial synthesized speech output by the initial decoder.

[0069] Finally, using the initial synthesized speech and multi-attribute speech samples, the second loss is calculated, and the model parameters of the initial decoding module are iteratively trained using the second loss. When the second target iteration number is reached or the second loss converges, the target decoding module is obtained.

[0070] In this embodiment of the invention, by jointly training the first initial encoder, the second initial encoder, and the initial decoder in the initial decoding module, the training efficiency of the target decoding module can be improved and training time can be saved.

[0071] like Figure 6 As shown, based on the above embodiments, this embodiment of the invention provides a speech synthesis device, including: The text acquisition module 61 is used to acquire the text to be synthesized and the target attribute text; The speech synthesis module 62 is used to input the text to be synthesized and the target attribute text into the speech synthesis model to obtain the target synthesized speech corresponding to the text to be synthesized output by the speech synthesis model; The speech synthesis model is trained based on multi-attribute speech samples, text samples with attribute labels corresponding to the multi-attribute speech samples, and discrete acoustic feature label sequences of the multi-attribute speech samples.

[0072] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application comprises a target encoder, a target encoding module and a target decoding module. The speech synthesis module is specifically configured to: fuse the text to be synthesized and the target attribute text to obtain a fused text; input the fused text into the target encoder to obtain a first discrete acoustic feature corresponding to the text to be synthesized output by the target encoder; input the fused text and the first discrete acoustic feature into the target encoding module to obtain a final discrete acoustic feature sequence determined by the target encoding module in an autoregressive manner; input the text to be synthesized, the target attribute text and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module.

[0073] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application comprises a first encoder, a second encoder and a decoder. The speech synthesis module is specifically configured to: input the text to be synthesized into the first encoder to obtain a first semantic feature output by the first encoder; input the target attribute text into the second encoder to obtain a second semantic feature output by the second encoder; input the first semantic feature, the second semantic feature and the final discrete acoustic feature sequence into the decoder to obtain the target synthesized speech output by the decoder.

[0074] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application comprises a first encoder, a second encoder and a decoder. The speech synthesis module is specifically configured to: splice the sentence-level attribute text and the text to be synthesized, and insert the local-level attribute text into the text to be synthesized to obtain the fused text.

[0075] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application further comprises a first training module configured to: obtain a pre-training encoding module, wherein the pre-training encoding module is trained based on discrete acoustic feature data of multi-attribute speech data and multi-attribute text corresponding to attribute-free labels of the multi-attribute speech data; input the text sample and the first discrete acoustic feature sample corresponding to the text sample into the pre-training encoding module to obtain a final discrete acoustic feature sample sequence determined by the pre-training encoding module in an autoregressive manner; based on the final discrete acoustic feature sample sequence and the discrete acoustic feature label sequence, calculate a first loss, and based on the first loss, iteratively train the pre-training encoding module to obtain the target encoding module.

[0076] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application further comprises a second training module configured to: input the attribute-free labels in the text sample and the attribute labels into a first initial encoder and a second initial encoder of an initial decoding module respectively to obtain first initial semantic features output by the first initial encoder and second initial semantic features output by the second initial encoder; input the first initial semantic features, the second initial semantic features and the discrete acoustic feature label sequence into an initial decoder of the initial decoding module to obtain initial synthesized speech output by the initial decoder; based on the initial synthesized speech and the multi-attribute speech sample, calculate a second loss, and based on the second loss, iteratively train the initial decoding module to obtain the target decoding module.

[0077] On the basis of the above-mentioned embodiments, the speech synthesis device provided in the embodiments of the present application, the target attribute text comprises text of at least one kind of audio attribute in language, dialect, style, emotion, speaker gender, age, voice characteristics, physical state, background audio and inserted audio.

[0078] Specifically, the roles of each module in the speech synthesis device provided in the embodiments of the present application are one-to-one corresponding to the operation processes of each step in the method class embodiments, and the effects achieved are consistent. For details, refer to the above-mentioned embodiments, which will not be described here again in the embodiments of the present application.

[0079] Figure 7 An example of an entity structure schematic diagram of an electronic device is shown as Figure 7As shown, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 complete communications with each other through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute the speech synthesis method provided in each of the embodiments described above.

[0080] In addition, the logical instruction in the memory 830 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the related art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0081] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the speech synthesis method provided in each of the embodiments described above.

[0082] In yet another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the speech synthesis method provided in each of the embodiments described above. The computer readable storage medium can be a non-transitory computer readable storage medium or a transitory computer readable storage medium, which is not specifically limited here.

[0083] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0084] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the related art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the various embodiments or some parts of the embodiments.

[0085] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A speech synthesis method, characterized in that, include: Obtain the text to be synthesized and the target attribute text; The text to be synthesized and the target attribute text are input into the speech synthesis model to obtain the target synthesized speech corresponding to the text to be synthesized output by the speech synthesis model. The speech synthesis model is trained based on multi-attribute speech samples, corresponding text samples with attribute labels, and discrete acoustic feature label sequences of the multi-attribute speech samples.

2. The speech synthesis method according to claim 1, characterized in that, The speech synthesis model includes a target encoder, a target encoding module, and a target decoding module; The step of inputting the text to be synthesized and the target attribute text into the speech synthesis model to obtain the target synthesized speech corresponding to the text to be synthesized output by the speech synthesis model includes: The text to be synthesized is merged with the target attribute text to obtain the merged text; The fused text is input into the target encoder to obtain the first discrete acoustic feature corresponding to the text to be synthesized, which is output by the target encoder. The fused text and the first discrete acoustic feature are input into the target encoding module to obtain the final discrete acoustic feature sequence determined by the target encoding module using an autoregressive approach; The text to be synthesized, the target attribute text, and the final discrete acoustic feature sequence are input into the target decoding module to obtain the target synthesized speech output by the target decoding module.

3. The speech synthesis method according to claim 2, characterized in that, The target decoding module includes a first encoder, a second encoder, and a decoder; the step of inputting the text to be synthesized, the target attribute text, and the final discrete acoustic feature sequence into the target decoding module to obtain the target synthesized speech output by the target decoding module includes: The text to be synthesized is input into the first encoder to obtain the first semantic feature output by the first encoder; The target attribute text is input into the second encoder to obtain the second semantic feature output by the second encoder. The first semantic feature, the second semantic feature, and the final discrete acoustic feature sequence are input into the decoder to obtain the target synthesized speech output by the decoder.

4. The speech synthesis method according to claim 2, characterized in that, The target attribute text includes sentence-level attribute text and local-level attribute text; the process of fusing the text to be synthesized with the target attribute text to obtain fused text includes: The sentence-level attribute text is concatenated with the text to be synthesized, and the local-level attribute text is inserted into the text to be synthesized to obtain the fused text.

5. The speech synthesis method according to claim 2, characterized in that, The training steps for the target encoding module include: A pre-trained encoding module is obtained, which is trained based on the discrete acoustic feature data of multi-attribute speech data and the multi-attribute text without attribute labels corresponding to the multi-attribute speech data. The text sample and the first discrete acoustic feature sample corresponding to the text sample are input into the pre-trained encoding module to obtain the final discrete acoustic feature sample sequence determined by the pre-trained encoding module using an autoregressive method; Based on the final discrete acoustic feature sample sequence and the discrete acoustic feature label sequence, a first loss is calculated, and based on the first loss, the pre-trained coding module is iteratively trained to obtain the target coding module.

6. The speech synthesis method according to claim 3, characterized in that, The training steps for the target decoding module include: The non-attribute tags and attribute tags in the text sample are respectively input to the first initial encoder and the second initial encoder of the initial decoding module to obtain the first initial semantic feature output by the first initial encoder and the second initial semantic feature output by the second initial encoder. The first initial semantic feature, the second initial semantic feature, and the discrete acoustic feature label sequence are input into the initial decoder of the initial decoding module to obtain the initial synthesized speech output by the initial decoder; Based on the initial synthesized speech and the multi-attribute speech samples, a second loss is calculated, and based on the second loss, the initial decoding module is iteratively trained to obtain the target decoding module.

7. The speech synthesis method according to any one of claims 1-6, characterized in that, The target attribute text includes text with at least one type of audio attribute, such as language, dialect, style, emotion, speaker gender, age, voice characteristics, physical condition, background audio, and inserted audio.

8. A speech synthesis device, characterized in that, include: The text acquisition module is used to acquire the text to be synthesized and the target attribute text; The speech synthesis module is used to input the text to be synthesized and the target attribute text into the speech synthesis model to obtain the target synthesized speech corresponding to the text to be synthesized output by the speech synthesis model. The speech synthesis model is trained based on multi-attribute speech samples, corresponding text samples with attribute labels, and discrete acoustic feature label sequences of the multi-attribute speech samples.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the speech synthesis method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-language speech synthesis method and system based on hierarchical rhythm prediction

    CN115547293A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN119763545A

  • Emotional voice generation method and device, equipment and medium

    CN120199282A

  • Speech synthesis method and apparatus, and device and storage medium

    WO2022178941A1

  • Speech synthesis model training method, speech synthesis method, electronic device, and storage medium

    WO2025140054A1

Cited By

  • Voice data generation method based on large model and method for training large model

    CN121393416A