Speech synthesis style transfer method and device, computer device and storage medium
By using multi-level style coding and regularization, speech that better matches the style to be transferred is generated, solving the speech quality problem when the sample style speech is not included in the training samples, and achieving higher quality speech synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2023-05-31
- Publication Date
- 2026-05-19
AI Technical Summary
Existing speech synthesis techniques produce poor-quality synthesized speech when the sample-style speech is not included in the training samples.
A multi-level style encoder is used to obtain multi-level speech style representations of the speech to be transferred, a style predictor is used to predict speech style, and the target speech is generated by multi-style layer regularization and pitch predictor. The style information of the language content is eliminated, and the target speech is generated by combining prosodic variation data.
It improves the fluency and emotional expression of synthesized speech, thus enhancing speech quality.
Smart Images

Figure CN116645951B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis, and more particularly to a speech synthesis style transfer method, apparatus, computer device, and storage medium. Background Technology
[0002] With the development of intelligent speech technology, speech synthesis has made great strides. However, in many cases, the provided sample speech styles are not included in the training samples, resulting in a significant decrease in the quality of synthesized speech. Therefore, we consider introducing a speech style transfer method to handle out-of-domain style synthesis, that is, to use a speech style transfer method to process speech synthesis based on sample speech styles that are not in the training samples. Summary of the Invention
[0003] Therefore, it is necessary to provide a speech synthesis style transfer method, apparatus, computer equipment, and storage medium to address the above-mentioned technical problems, so as to solve the problem of poor synthesized speech quality when the sample style speech is not included in the training samples in existing speech synthesis.
[0004] A speech synthesis style transfer method includes:
[0005] The text data to be converted is input into the speech synthesis encoder to obtain the text encoding;
[0006] A multi-level speech style representation of the speech to be transferred is obtained through a multi-level style encoder.
[0007] The text data to be converted and the multi-level speech style representation are input into the style predictor to predict the speech style, thereby obtaining the predicted speech style.
[0008] The text encoding and the predicted speech style are subjected to multi-style layer regularization to obtain style-regularized text encoding.
[0009] The style-regularized text code is input into the pitch predictor to obtain the prosodic variation data of the text data to be converted;
[0010] The prosodic variation data and the style-regularized text encoding are input into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and the target speech is generated based on the Mel spectrogram.
[0011] A speech synthesis style transfer device, comprising:
[0012] The text encoding module is used to input the text data to be converted into the speech synthesis encoder to obtain the text encoding;
[0013] A multi-level speech style representation module is used to obtain multi-level speech style representations of the speech to be transferred through a multi-level style encoder.
[0014] The speech style prediction module is used to input the text data to be converted and the multi-level speech style representation into the style predictor to predict the speech style and obtain the predicted speech style.
[0015] The style-regularized text encoding module is used to perform multi-style-layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding.
[0016] The prosody variation data module is used to input the style-regularized text encoding into the pitch predictor to obtain the prosody variation data of the text data to be converted.
[0017] The target speech module is used to input the prosodic variation data and the style regularized text encoding into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and to generate the target speech based on the Mel spectrogram.
[0018] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the above-described speech synthesis style transfer method when executing the computer-readable instructions.
[0019] One or more readable storage media storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the speech synthesis style transfer method described above.
[0020] The aforementioned speech synthesis style transfer method, apparatus, computer equipment, and storage medium involve: inputting the text data to be converted into a speech synthesis encoder to obtain text encoding; obtaining multi-level speech style representations of the speech to be transferred through a multi-level style encoder; inputting the text data to be converted and the multi-level speech style representations into a style predictor for speech style prediction to obtain a predicted speech style; performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding; inputting the style-regularized text encoding into a pitch predictor to obtain prosodic variation data of the text data to be converted; inputting the prosodic variation data and the style-regularized text encoding into a speech synthesis decoder to obtain a Mel spectrogram corresponding to the text data to be converted, and generating target speech based on the Mel spectrogram. This invention, through the multi-level speech style representations obtained from the speech to be transferred, fully considers the style representations at each level of the speech to be transferred, that is, it focuses on the style of words and phonemes in the sentence as well as the overall prosodic style of the speaker. Furthermore, multi-style regularization was applied to the multi-level speech style representation and text encoding, eliminating style information in the language content and making the obtained style-regularized text encoding more standardized. Further, prosodic variation prediction was performed based on the standardized style-regularized text encoding, ensuring that the synthesized target speech, based on style-regularized text encoding and prosodic variation data, considers not only the prosodic information of the text data to be converted but also the style of words and phonemes in the transferred style speech, as well as the overall prosodic style of the speaker. This results in a more fluent and expressive synthesized target speech, improving its overall speech quality. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an application environment for the speech synthesis style transfer method in one embodiment of the present invention;
[0023] Figure 2 This is a flowchart illustrating a speech synthesis style transfer method according to an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of a speech synthesis style transfer device according to an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] The speech synthesis style transfer method provided in this embodiment can be applied to, for example, Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0028] In one embodiment, such as Figure 2 As shown, a speech synthesis style transfer method is provided, which is then applied to... Figure 1 Taking the server-side as an example, the explanation includes the following steps:
[0029] S10. Input the text data to be converted into the speech synthesis encoder to obtain the text encoding.
[0030] Understandably, the text data to be converted refers to text data that needs to be converted into speech. This text data contains some textual information.
[0031] S20. Obtain multi-level speech style representations of the speech to be transferred through a multi-level style encoder.
[0032] Understandably, a multi-level style encoder comprises encoders at multiple different levels of style. For example, this multi-level style encoder may include a word-level style encoder, a phoneme-level style encoder, a discourse-level style encoder, and a global style encoder. The speech to be transferred refers to speech containing the style to be transferred. This speech can be obtained from a speech database or determined by the user's selection. The speech database pre-stores several speech samples with different styles. The styles of the speech can be categorized based on the speaker's identity, emotion, prosody, etc. Multi-level speech style representation refers to the representation of the style of the speech to be transferred at multiple different levels.
[0033] S30. The text data to be converted and the multi-level speech style representation are input into the style predictor to perform speech style prediction, thereby obtaining the predicted speech style.
[0034] Understandably, a style predictor is used to extract the contextual semantic information of the text data to be converted, and to predict the speech style based on the contextual semantic information and multi-level speech style representations. Specifically, speech style prediction refers to the process by which the style predictor predicts the speech style of the text data to be converted based on the contextual semantic information and the representation information of multiple different levels of style in the speech to be transferred. Predicting the speech style refers to the speech style predicted based on the contextual semantic information and the representation information of multiple different levels of style in the speech to be transferred.
[0035] S40. Perform multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding.
[0036] Understandably, regularization aims to prevent overfitting and enhance the generalization ability of data. Multi-style layer regularization refers to regularizing language content that contains multiple styles. Here, multi-style layer regularization aims to eliminate style information in the language content. The language content includes text encoding and predicted speech style. Specifically, initial layer regularization is first applied to the text encoding to obtain initial regularized text encoding, enhancing its generalization ability. Then, multi-style layer regularization is applied to the enhanced generalization initial regularized text encoding and predicted speech style to obtain style-regularized text encoding, eliminating style information in both and making the resulting style-regularized text encoding more standardized. The style-regularized text encoding is both the encoding obtained from multi-style layer regularization of the text encoding and predicted speech style, containing both the textual information of the text data to be transformed and the style information of the predicted speech style.
[0037] S50. Input the style-regularized text code into the pitch predictor to obtain the prosodic variation data of the text data to be converted.
[0038] Understandably, the pitch predictor is used to predict prosodic variations corresponding to input data. This pitch predictor consists of a two-layer ReLU (Rectified Linear Unit) activated convolutional network. Each layer is followed by an initial regularization layer and a dropout layer. Specifically, style-regularized text encoding is used as input data to the pitch predictor. The pitch predictor extracts prosodic features from the style-regularized text encoding, and then predicts the prosodic variations of the style-regularized text encoding based on these prosodic features, obtaining prosodic variation data. This prosodic variation data indicates the prosodic variations of the linguistic content contained in the style-regularized text encoding.
[0039] S60. Input the prosodic variation data and the style regularized text encoding into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and generate the target speech based on the Mel spectrogram.
[0040] Understandably, a speech synthesis decoder is used to decode input data containing linguistic content into a Mel spectrogram containing speech information. Specifically, prosodic variation data and style-regularized text encoding are input into the speech synthesis decoder, which decodes and synthesizes the prosodic variation data and style-regularized text encoding to obtain a Mel spectrogram containing speech information. Then, the obtained Mel spectrogram is input into a vocoder for reconstruction to obtain the target speech corresponding to the text data to be converted. This target speech contains both the textual information of the text data to be converted and the multi-level style information of the speech to be transferred.
[0041] In steps S10-S60, the text data to be converted is input into a speech synthesis encoder to obtain text encoding; a multi-level style encoder is used to obtain a multi-level speech style representation of the speech to be transferred; the text data to be converted and the multi-level speech style representation are input into a style predictor to predict the speech style, resulting in a predicted speech style; multi-style layer regularization is applied to the text encoding and the predicted speech style to obtain style-regularized text encoding; the style-regularized text encoding is input into a pitch predictor to obtain prosodic variation data of the text data to be converted; the prosodic variation data and the style-regularized text encoding are input into a speech synthesis decoder to obtain a Mel spectrogram corresponding to the text data to be converted, and the target speech is generated based on the Mel spectrogram. This embodiment, by obtaining a multi-level speech style representation from the speech to be transferred, fully considers the style representation at each level in the speech to be transferred, that is, it focuses on the style of words and phonemes in the sentence as well as the overall prosodic style of the speaker. Furthermore, multi-style regularization was applied to the multi-level speech style representation and text encoding, eliminating style information in the language content and making the obtained style-regularized text encoding more standardized. Further, prosodic variation prediction was performed based on the standardized style-regularized text encoding, ensuring that the synthesized target speech, based on style-regularized text encoding and prosodic variation data, considers not only the prosodic information of the text data to be converted but also the style of words and phonemes in the transferred style speech, as well as the overall prosodic style of the speaker. This results in a more fluent and expressive synthesized target speech, improving its overall speech quality.
[0042] Optionally, the multi-level style encoder includes a word-level style encoder, a phoneme-level style encoder, a discourse-level style encoder, and a global style encoder;
[0043] In step S20, namely, obtaining the multi-level speech style representation of the speech to be transferred through a multi-level style encoder, the following steps are included:
[0044] S201. Input the speech to be transferred into the word-level style encoder, the phoneme-level style encoder, the discourse-level style encoder and the global style encoder respectively to obtain the word-level style representation, the phoneme-level style representation, the discourse-level style representation and the global style representation.
[0045] S202. Based on the word-level style representation, the phoneme-level style representation, the discourse-level style representation, and the global style representation, the multi-level speech style representation is obtained.
[0046] Understandably, a word-level style encoder is used to acquire and encode the word-level style of the speech to be transferred, resulting in a word-level style representation. Here, word-level style refers to the speech style of the words contained in the speech to be transferred. A phoneme-level style encoder is used to acquire and encode the phoneme-level style of the speech to be transferred, resulting in a phoneme-level style representation. Here, phoneme-level style refers to the speech style of the phonemes contained in the speech to be transferred. A discourse-level style encoder is used to acquire and encode the discourse-level style of the speech to be transferred, resulting in a discourse-level style representation. Here, discourse-level style refers to the speech style of the discourses contained in the speech to be transferred. A global style encoder is used to acquire and encode the global style of the speech to be transferred, resulting in a global style representation. Here, global style refers to the overall speech style of the speech to be transferred. Multi-level speech style representations are obtained based on word-level style representations, phoneme-level style representations, discourse-level style representations, and global style representations.
[0047] In steps S201 and S202, the speech style of the speech to be transferred is obtained through multiple style encoders of different levels, resulting in a multi-level speech style representation that integrates the speech style of the speech to be transferred at the word level, phoneme level, discourse level and global level. The speech style in the speech to be transferred is analyzed in a comprehensive and detailed manner, which makes the accuracy of the obtained multi-level speech style representation higher.
[0048] Optionally, in step S40, i.e., performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding, the following steps are included:
[0049] S401. Perform initial layer regularization on the text encoding to obtain initial regularized text encoding;
[0050] S402. Perform multi-style layer regularization on the initial regularized text encoding and the predicted speech style to obtain the style regularized text encoding.
[0051] Understandably, regularization aims to prevent overfitting and enhance the generalization ability of data. Initial-level regularization refers to regularizing the text encoding. Multi-style-level regularization refers to regularizing language content containing multiple styles. Here, multi-style-level regularization is performed to eliminate style information in the language content. Specifically, multi-style-level regularization is applied to both the initially regularized text encoding and the predicted speech style.
[0052] In this implementation, the text encoding is first subjected to initial layer regularization, and then the text encoding that has been initially regularized is subjected to multi-style layer regularization together with the predicted speech style, in order to better eliminate style information in the language content.
[0053] Optionally, after step S401, that is, after performing initial layer regularization on the text encoding to obtain the initial regularized text encoding, the process includes:
[0054] S4011, Input the initial regularized text encoding, the predicted speech style, and the text encoding into the time predictor;
[0055] S4012. Extract the initial regularized text encoding, the predicted speech style, and the temporal features of the text encoding using the time predictor;
[0056] S4013. Based on the time characteristics, the duration of the Mel spectrogram is obtained.
[0057] Understandably, the time predictor is used to predict the duration of the generated speech based on the initial regularized text encoding, the predicted speech style, and the text encoding. Specifically, the time predictor extracts the temporal features of the initial regularized text encoding, the predicted speech style, and the text encoding, and based on the extracted temporal features, obtains the duration of the final generated Mel spectrogram.
[0058] In this embodiment, duration prediction is performed based on the initial regularized text encoding, predicted speech style, and the temporal features of the text encoding. This approach considers the temporal features of three data dimensions, resulting in more accurate predicted durations.
[0059] Optionally, in step S50, i.e., inputting the style-regularized text encoding into the pitch predictor to obtain the prosodic variation data of the text data to be converted, the following is included:
[0060] S501. Input the style-regularized text code into the pitch predictor, and extract the prosodic features of the style-regularized text code through the pitch predictor.
[0061] S502. Based on the prosodic features, predict the prosodic changes of the style-regularized text encoding to obtain the prosodic change data.
[0062] Understandably, the pitch predictor is used to predict the prosodic variation data of the text data to be converted. Specifically, the style-regularized text code of the text data to be converted is input into the pitch predictor, which extracts the prosodic features of the style-regularized text code. Then, based on the obtained prosodic features, the prosodic variation of the style-regularized text code is predicted, and finally, the prosodic variation data of the text data to be converted is output.
[0063] In this embodiment, a pitch predictor predicts prosodic variations in the text data to be converted that are unrelated to speech style, so that the final generated target speech takes into account not only speech style but also prosodic variations, resulting in higher quality target speech and improved user experience.
[0064] Optionally, in step S60, i.e., generating the target speech based on the Mel spectrogram, the following steps are included:
[0065] S601. Input the Mel spectrogram into the vocoder, and use the vocoder to perform audio reconstruction on the Mel spectrogram to obtain the target speech corresponding to the Mel spectrogram.
[0066] In essence, the Mel spectrogram is reconstructed into speech using a vocoder. Specifically, the Mel spectrogram is decoded and reconstructed using a vocoder to obtain the target speech corresponding to the Mel spectrogram. This target speech contains both the textual information of the data to be converted and the multi-level style information of the speech whose style is to be transferred.
[0067] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0068] In one embodiment, a speech synthesis style transfer apparatus is provided, which corresponds one-to-one with the speech synthesis style transfer method described in the above embodiments. For example... Figure 3 As shown, the speech synthesis style transfer device includes a text encoding module 10, a multi-level speech style representation module 20, a predicted speech style module 30, a style regularization text encoding module 40, a prosodic variation data module 50, and a target speech module 60. Detailed descriptions of each functional module are as follows:
[0069] The text encoding module 10 is used to input the text data to be converted into the speech synthesis encoder to obtain the text encoding;
[0070] The multi-level speech style representation module 20 is used to obtain the multi-level speech style representation of the speech to be transferred through a multi-level style encoder.
[0071] The speech style prediction module 30 is used to input the text data to be converted and the multi-level speech style representation into the style predictor to predict the speech style and obtain the predicted speech style.
[0072] The style regularization text encoding module 40 is used to perform multi-style layer regularization on the text encoding and the predicted speech style to obtain style regularization text encoding.
[0073] The prosody variation data module 50 is used to input the style-regularized text encoding into the pitch predictor to obtain the prosody variation data of the text data to be converted.
[0074] The target speech module 60 is used to input the prosodic variation data and the style regularized text encoding into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and to generate the target speech based on the Mel spectrogram.
[0075] Optionally, the multi-level style encoder includes a word-level style encoder, a phoneme-level style encoder, a discourse-level style encoder, and a global style encoder;
[0076] The multi-level speech style representation module 20 includes:
[0077] The style representation unit is used to input the speech with the style to be transferred into the word-level style encoder, the phoneme-level style encoder, the discourse-level style encoder and the global style encoder respectively, and obtain the word-level style representation, the phoneme-level style representation, the discourse-level style representation and the global style representation accordingly.
[0078] A multi-level speech style representation unit is used to obtain the multi-level speech style representation based on the word-level style representation, the phoneme-level style representation, the discourse-level style representation, and the global style representation.
[0079] Optionally, the style regularization text encoding module 40 includes:
[0080] An initial layer regularization unit is used to perform initial layer regularization on the text encoding to obtain an initial regularized text encoding;
[0081] A multi-style layer regularization unit is used to perform multi-style layer regularization on the initial regularized text encoding and the predicted speech style to obtain the style-regularized text encoding.
[0082] Optionally, the speech synthesis style transfer device further includes:
[0083] The data input unit is used to input the initial regularized text encoding, the predicted speech style, and the text encoding into the time predictor;
[0084] A temporal feature unit is used to extract the temporal features of the initial regularized text encoding, the predicted speech style, and the text encoding through the temporal predictor.
[0085] The duration unit is used to obtain the duration of the Mel spectrogram based on the time characteristics.
[0086] Optionally, the prosodic variation data module 50 includes:
[0087] The prosodic feature unit is used to input the style-regularized text code into the pitch predictor and extract the prosodic features of the style-regularized text code through the pitch predictor.
[0088] The prosodic variation data unit is used to predict the prosodic variation of the style-regularized text encoding based on the prosodic features, thereby obtaining the prosodic variation data.
[0089] Optionally, the target voice module 60 includes:
[0090] The target speech unit is used to input the Mel spectrogram into a vocoder, and use the vocoder to perform audio reconstruction on the Mel spectrogram to obtain the target speech corresponding to the Mel spectrogram.
[0091] Specific limitations regarding the speech synthesis style transfer device can be found in the limitations of the speech synthesis style transfer method described above, and will not be repeated here. Each module in the aforementioned speech synthesis style transfer device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0092] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes a readable storage medium and internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database stores data related to the speech synthesis style transfer method. The network interface communicates with external terminals via a network connection. When the computer-readable instructions are executed by the processor, a speech synthesis style transfer method is implemented. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.
[0093] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor performs the following steps when executing the computer-readable instructions:
[0094] The text data to be converted is input into the speech synthesis encoder to obtain the text encoding;
[0095] A multi-level speech style representation of the speech to be transferred is obtained through a multi-level style encoder.
[0096] The text data to be converted and the multi-level speech style representation are input into the style predictor to predict the speech style, thereby obtaining the predicted speech style.
[0097] The text encoding and the predicted speech style are subjected to multi-style layer regularization to obtain style-regularized text encoding.
[0098] The style-regularized text code is input into the pitch predictor to obtain the prosodic variation data of the text data to be converted;
[0099] The prosodic variation data and the style-regularized text encoding are input into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and the target speech is generated based on the Mel spectrogram.
[0100] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which, when executed by one or more processors, perform the following steps:
[0101] The text data to be converted is input into the speech synthesis encoder to obtain the text encoding;
[0102] A multi-level speech style representation of the speech to be transferred is obtained through a multi-level style encoder.
[0103] The text data to be converted and the multi-level speech style representation are input into the style predictor to predict the speech style, thereby obtaining the predicted speech style.
[0104] The text encoding and the predicted speech style are subjected to multi-style layer regularization to obtain style-regularized text encoding.
[0105] The style-regularized text code is input into the pitch predictor to obtain the prosodic variation data of the text data to be converted;
[0106] The prosodic variation data and the style-regularized text encoding are input into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and the target speech is generated based on the Mel spectrogram.
[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0109] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech synthesis style transfer method, characterized in that, include: The text data to be converted is input into the speech synthesis encoder to obtain the text encoding; A multi-level speech style representation of the speech to be transferred is obtained through a multi-level style encoder. The multi-level speech style representation is obtained based on word-level style representation, phoneme-level style representation, discourse-level style representation and global style representation; The text data to be converted and the multi-level speech style representation are input into the style predictor to predict the speech style, thereby obtaining the predicted speech style. The text encoding and the predicted speech style are subjected to multi-style layer regularization to obtain style-regularized text encoding. The style-regularized text code is input into the pitch predictor to obtain the prosodic variation data of the text data to be converted; The prosodic variation data and the style-regularized text encoding are input into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and the target speech is generated based on the Mel spectrogram. The step of inputting the style-regularized text code into the pitch predictor to obtain the prosodic variation data of the text data to be converted includes: The style-regularized text code is input into the pitch predictor, and the prosodic features of the style-regularized text code are extracted by the pitch predictor. Based on the prosodic features, the prosodic changes of the style-regularized text encoding are predicted to obtain the prosodic change data.
2. The speech synthesis style transfer method as described in claim 1, characterized in that, The multi-level style encoder includes a word-level style encoder, a phoneme-level style encoder, a speech-level style encoder, and a global style encoder; The process of obtaining a multi-level speech style representation of the speech to be transferred using a multi-level style encoder includes: The speech to be transferred in style is input into the word-level style encoder, the phoneme-level style encoder, the discourse-level style encoder and the global style encoder respectively, and word-level style representation, phoneme-level style representation, discourse-level style representation and global style representation are obtained accordingly. The multi-level speech style representation is obtained based on the word-level style representation, the phoneme-level style representation, the discourse-level style representation, and the global style representation.
3. The speech synthesis style transfer method as described in claim 1, characterized in that, The step of performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding includes: The text encoding is subjected to initial layer regularization to obtain the initial regularized text encoding; The initial regularized text encoding and the predicted speech style are subjected to multi-style layer regularization to obtain the style-regularized text encoding.
4. The speech synthesis style transfer method as described in claim 3, characterized in that, After performing initial layer regularization on the text encoding to obtain the initial regularized text encoding, the process includes: The initial regularized text encoding, the predicted speech style, and the text encoding are input into the time predictor; The time predictor extracts the initial regularized text encoding, the predicted speech style, and the temporal features of the text encoding. Based on the time characteristics, the duration of the Mel spectrogram is obtained.
5. The speech synthesis style transfer method as described in claim 1, characterized in that, The step of generating target speech based on the Mel spectrogram includes: The Mel spectrogram is input into a vocoder, and the vocoder performs audio reconstruction on the Mel spectrogram to obtain the target speech corresponding to the Mel spectrogram.
6. A speech synthesis style transfer device, characterized in that, include: The text encoding module is used to input the text data to be converted into the speech synthesis encoder to obtain the text encoding; A multi-level speech style representation module is used to obtain multi-level speech style representations of the speech to be transferred through a multi-level style encoder. The multi-level speech style representation is obtained based on word-level style representation, phoneme-level style representation, discourse-level style representation and global style representation; The speech style prediction module is used to input the text data to be converted and the multi-level speech style representation into the style predictor to predict the speech style and obtain the predicted speech style. The style-regularized text encoding module is used to perform multi-style-layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding. The prosody variation data module is used to input the style-regularized text encoding into the pitch predictor to obtain the prosody variation data of the text data to be converted. The target speech module is used to input the prosodic variation data and the style regularized text encoding into the speech synthesis decoder to obtain the Mel spectrogram corresponding to the text data to be converted, and to generate the target speech based on the Mel spectrogram; The prosodic variation data module includes: The prosodic feature unit is used to input the style-regularized text code into the pitch predictor and extract the prosodic features of the style-regularized text code through the pitch predictor. The prosodic variation data unit is used to predict the prosodic variation of the style-regularized text encoding based on the prosodic features, thereby obtaining the prosodic variation data.
7. The speech synthesis style transfer apparatus as described in claim 6, characterized in that, The multi-level style encoder includes a word-level style encoder, a phoneme-level style encoder, a speech-level style encoder, and a global style encoder; The multi-level speech style representation module includes: The style representation unit is used to input the speech with the style to be transferred into the word-level style encoder, the phoneme-level style encoder, the discourse-level style encoder and the global style encoder respectively, and obtain the word-level style representation, the phoneme-level style representation, the discourse-level style representation and the global style representation accordingly. A multi-level speech style representation unit is used to obtain the multi-level speech style representation based on the word-level style representation, the phoneme-level style representation, the discourse-level style representation, and the global style representation.
8. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the speech synthesis style transfer method as described in any one of claims 1 to 5.
9. One or more readable storage media storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by one or more processors, the one or more processors cause the speech synthesis style transfer method as described in any one of claims 1 to 5.