Speech synthesis method and device, electronic equipment and storage medium
By employing multi-scale style prediction and vectorized codebook methods, the data dependency problem in cross-linguistic style transfer is solved, achieving efficient adaptability and authentic representation of cross-linguistic styles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2023-02-16
- Publication Date
- 2026-04-14
AI Technical Summary
Existing speech synthesis technologies rely heavily on data resources when performing cross-language style transfer, and it is difficult to achieve the same style performance in different languages, resulting in high costs and poor results.
By acquiring textual features of the target language and identifiers of the original language, and utilizing multi-scale style prediction and vectorized codebooks, cross-language style transfer is achieved. Different codebooks are used to vectorize style features, reducing dependence on audio features and enhancing style adaptability.
It enables cross-language style transfer, reduces data requirements, alleviates accent issues, and improves the adaptability and expressiveness of styles in different languages.
Smart Images

Figure CN116312459B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech synthesis technology, specifically to speech synthesis methods, devices, electronic devices, and storage media. Background Technology
[0002] In speech synthesis applications, improving the expressiveness of speech synthesis often requires separate style modeling. To enable the target timbre to exhibit different styles, style transfer techniques are needed. However, current style transfer methods primarily operate within the same language, requiring corresponding style data for that language to achieve the desired style representation. This excessive reliance on data resources increases the cost of speech synthesis tasks. Furthermore, in some applications, it is necessary to represent the same style across different languages, which traditional style transfer tasks struggle to achieve. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a speech synthesis method, apparatus, electronic device, and storage medium to solve the problem of style transfer across languages during speech synthesis.
[0004] According to a first aspect, embodiments of this disclosure provide a speech synthesis method, including:
[0005] Obtain the text features of the target language and the identifier of the original language;
[0006] Based on the text features of the target language and the identifier of the original language, style prediction of the original language is performed to obtain style features. Based on the style features, the codebook corresponding to the target language is queried to obtain vectorized style features. The codebook corresponds one-to-one with the language and is used to vectorize the style features.
[0007] Encoding and decoding processes are performed based on the vectorized style features to determine the target speech of the target language.
[0008] According to a second aspect, embodiments of this disclosure provide a speech synthesis apparatus, comprising:
[0009] The acquisition module is used to acquire text features of the target language and the identifier of the original language;
[0010] The style vectorization module is used to predict the style of the original language based on the text features of the target language and the identifier of the original language, to obtain style features, and to query the codebook corresponding to the target language based on the style features to obtain vectorized style features. The codebook corresponds one-to-one with the language and is used to vectorize the style features.
[0011] The processing module is used to perform encoding and decoding processing based on the vectorized style features to determine the target speech of the target language.
[0012] According to a third aspect, this disclosure provides an electronic device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the speech synthesis method described in the first aspect or any embodiment of the first aspect.
[0013] According to a fourth aspect, embodiments of this disclosure provide a computer-readable storage medium storing computer instructions for causing the computer to perform the speech synthesis method described in the first aspect or any embodiment of the first aspect.
[0014] The speech synthesis method provided in this disclosure uses different codebooks to vectorize style features for different languages, enabling the use of the target language's codebook during cross-language style transfer to mitigate accent issues. Furthermore, style prediction for the original language is based on the target language's text features and the original language's identifiers; that is, style prediction is performed using text features as data. This removes the dependence on audio features, enhances the adaptability of style to different languages, and thus enables the transfer of style data from one language to other languages, reducing data requirements. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 This is a flowchart of a speech synthesis method according to an embodiment of the present disclosure;
[0017] Figure 2 This is a flowchart of a speech synthesis method according to an embodiment of the present disclosure;
[0018] Figure 3 This is a schematic diagram of the structure of speech synthesis according to an embodiment of the present disclosure;
[0019] Figure 4 This is a flowchart illustrating the method for determining a style model according to an embodiment of this disclosure;
[0020] Figure 5This is a schematic diagram of the structure of a style model according to an embodiment of the present disclosure;
[0021] Figure 6 This is a structural block diagram of a speech synthesis apparatus according to an embodiment of the present disclosure;
[0022] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this disclosure. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0028] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0029] In speech synthesis technologies, multi-scale reference encoders are generally used for style modeling. Vector quantization is applied in style modeling tasks. However, since the goal of vector quantization is not cross-language style transfer, and there is no strategy for style adaptation for this goal, this method of style modeling cannot achieve style transfer across different languages.
[0030] Based on this, the speech synthesis method provided in this disclosure quantifies style features based on a language-related codebook, thereby achieving style transfer across languages. To reduce the need for style data resources, the solution provided in this disclosure is a cross-language style transfer scheme, which can transfer styles from one language to another while maintaining accent authenticity and achieving high style similarity. For example, if the target language is English and the original language is Chinese, the speech synthesis method of this disclosure can achieve cross-language style transfer.
[0031] According to an embodiment of this disclosure, a speech synthesis method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0032] This embodiment provides a speech synthesis method that can be used in electronic devices such as mobile terminals, servers, and computers. Figure 1 This is a flowchart of a speech synthesis method according to an embodiment of the present disclosure, such as... Figure 1 As shown, the process includes the following steps:
[0033] S11, obtain the text features of the target language and the identifier of the original language.
[0034] The specific languages represented by the target language and the source language are set according to the actual application requirements, and no restrictions are imposed on them here. During speech synthesis, the text in the target language is converted into speech output in the target language.
[0035] Text features in the target language are obtained by extracting or preprocessing text features in the target language. For example, they can be obtained by passing the target language text through an encoder, or by using pre-trained text processing modules, etc. These pre-trained text processing modules include, but are not limited to, the BERT module.
[0036] The original language identifier serves as a unique identifier for the original language, and this identifier corresponds one-to-one with the language. This identifier includes, but is not limited to, numbers, characters, or other forms, and is set according to actual needs. For example, an electronic device maintains a table mapping languages to identifiers. After obtaining the identifier of the original language, the corresponding original language can be determined by querying this table.
[0037] S12: Based on the text features of the target language and the identifier of the original language, perform style prediction of the original language to obtain style features, and query the codebook corresponding to the target language based on the style features to obtain vectorized style features.
[0038] The codebook corresponds one-to-one with each language and is used to vectorize style features.
[0039] Style prediction is based on textual features of the target language and identifiers of the original language. The identifiers of the original language are used to determine the original language; that is, style prediction of the original language is performed using textual features of the target language.
[0040] For example, style prediction is achieved through a style predictor. Specifically, the target language text is first processed by a pre-trained BERT model to obtain a sequence of word-level semantic representations, i.e., text features. These text features are then used as input to a multi-scale style predictor. The multi-scale style predictor first utilizes a hierarchical context encoder containing a two-layer attention network to model the inter-word and inter-sentence relationships in the context, obtaining paragraph-level, sentence-level, and word-level contextual semantic representations. Next, the model predicts the corresponding level of style representations in the original language based on the different levels of contextual semantic representations, thus reconstructing the multi-scale speaking style in human speech. Furthermore, considering that higher-level speech styles, which are closer to the global scale, can influence lower-level speech styles, we model speech styles sequentially from higher to lower levels based on residual connections during the prediction process. Specifically, higher-level styles are predicted first and then used as conditional inputs to the lower-level style predictor. After processing by the style predictor, the style features of the original language are obtained.
[0041] In this embodiment, different codebooks are used for different languages. Assuming there are N languages, there are corresponding N codebooks. The codebook is used to vectorize style features. For example, the codebook is a VQ-VAE autoencoder, which can achieve discretized encoding capabilities and maintain a discretized codebook. When querying the codebook corresponding to the target language based on style features, the input style features are encoded into a vector within the codebook by performing a nearest neighbor search.
[0042] By using language-specific codebooks, the target language codebook can still be used when stylistic features differ across languages, further mitigating accent issues.
[0043] S13 performs encoding and decoding based on vectorized style features to determine the target speech of the target language.
[0044] After obtaining the vectorized style features, these features are added to the encoder to process the language and the vectorized style features, resulting in an encoded output. This encoded output is then fed into the decoder for decoding in conjunction with the target speaker, outputting the target speech in the target language. Because cross-language style transfer is implemented in speech synthesis, the target language's style features are still preserved in the target speech, mitigating accent information.
[0045] The speech synthesis method provided in this embodiment uses different codebooks to vectorize style features for different languages, enabling the use of the target language's codebook during cross-language style transfer to mitigate accent issues. Furthermore, style prediction for the original language is based on the target language's text features and the original language's identifiers; that is, style prediction is performed using text features as data. This removes the dependence on audio features, enhances the adaptability of style to different languages, and thus enables the transfer of style data from one language to other languages, reducing data requirements.
[0046] This embodiment provides a speech synthesis method that can be used in electronic devices such as mobile terminals, servers, and computers. Figure 2 This is a flowchart of a speech synthesis method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps:
[0047] S21, Obtain the text features of the target language and the identifier of the original language.
[0048] Please see details Figure 1 S11 of the illustrated embodiment will not be described again here.
[0049] S22. Based on the text features of the target language and the identifier of the original language, style prediction of the original language is performed to obtain style features. Based on the style features, the codebook corresponding to the target language is queried to obtain vectorized style features.
[0050] The codebook corresponds one-to-one with each language and is used to vectorize style features.
[0051] Specifically, S22 includes:
[0052] S221. Based on the text features of the target language and the identifiers of the original language, multi-scale style prediction of the original language is performed to obtain style features.
[0053] The style features include local style features and global style features.
[0054] In style prediction, a multi-scale style prediction approach is adopted. This involves performing multi-scale style prediction on the original language based on the text features of the target language, resulting in both local and global style features. This multi-scale style prediction can be implemented using a text prediction style module, such as Test2style Predictor. For example, if the original language is Chinese and the target language is English, Test2style Predictor can be used to predict the Chinese style in English text.
[0055] S222, using the target language to determine the local codebook and global codebook corresponding to the target language.
[0056] Corresponding to the results of multi-scale style prediction, namely local style features and global style features, there are local codebooks and global codebooks, respectively. Both the local and global codebooks are specific to the target language. The local codebook is used to vectorize local style features, and the global codebook is used to vectorize global style features.
[0057] S223, based on local style features and global style features, queries the local codebook and global codebook respectively to obtain vectorized style features.
[0058] The vectorized style features include vectorized local style features and vectorized global style features.
[0059] Different VQ-VAEs are used for different languages. Assuming there are N languages, there are N codebooks for both the vq-vae local style and the vq-vae global style. By dividing the VQ codebooks by language, the target language's style codebook can still be used when there are cross-language styles, further mitigating accent issues. If the speech synthesis method supports processing N languages, the local codebook for language 1 is VQ-vae1-local, and the global codebook is VQ-vae1-global; the local codebook for language 2 is VQ-vae2-local, and the global codebook is VQ-vae2-global; ...; the local codebook for language N is VQ-vaeN-local, and the global codebook is VQ-vaeN-global.
[0060] Continuing with the example above, if the original language is Chinese and the target language is English, that is, the input is English text features. If English is the Nth language, its corresponding local codebook is VQ-vaeN-local and its global codebook is VQ-vaeN-global. Using the predicted local style feature and global style feature, we search for the closest style features in VQ-vaeN-local and VQ-vaeN-global respectively, obtaining the vectorized local style feature VQ local style and the vectorized global style feature VQ global style.
[0061] S23 uses vectorized style features for encoding and decoding to determine the target speech of the target language.
[0062] Specifically, S23 above includes:
[0063] S231, the vectorized local style features and the vectorized global style features are concatenated to obtain the target style features.
[0064] Both vectorized local style features and vectorized global style features correspond to style features. Therefore, before further processing, these two need to be concatenated to obtain the target style feature.
[0065] In some embodiments, S231 includes:
[0066] (1) Obtain the target dimension of the vectorized local style features.
[0067] (2) Determine the number of copies of the vectorized global style features based on the target dimension.
[0068] (3) Based on the number of copies, the vectorized global style features are copied and spliced to obtain the copied features.
[0069] (4) The copied features are concatenated with the vectorized local style features to obtain the target style features.
[0070] If the target dimension of the vectorized local style feature is T, and the dimension of the vectorized global style feature is 1, in order to concatenate the two, the vectorized global style feature needs to be copied and concatenated with the vectorized local style feature in the form of a repeating feature.
[0071] Specifically, the vectorized global style features are copied T times, and then the T copies are concatenated to obtain a T-dimensional copied feature. This T-dimensional copied feature is then added to the T-dimensional vectorized local style features to obtain the T-dimensional target style feature. The dimension of the target style feature is the same as the dimension of the text features.
[0072] The vectorized global style features are concatenated with the vectorized local style features in a repeating manner to obtain the copied features. On this basis, the copied features are then concatenated with the vectorized local style features to ensure the alignment of the target style features with the text features.
[0073] S232 performs encoding and decoding processing based on target style features to determine the target speech of the target language.
[0074] For details on encoding and decoding processes, please refer to the above text. Figure 1 As described in S13 of the illustrated embodiment, it will not be repeated here.
[0075] In some implementations, the speech synthesis processing structure is as follows: Figure 3 As shown, during speech synthesis, the text features of the target language are used as input and the identifier of the original language are used as conditions. The system can predict the style of the original language in the target language and obtain the vectorized local style feature VQ localstyle and the vectorized global style feature VQ global style by searching the codebook corresponding to the target language. Furthermore, the system removes accent information from the style.
[0076] The vectorized global style feature (VQ global style) is then added to the vectorized local style feature (VQ local style) in the form of repeated features to form the target style feature, whose length dimension is consistent with the length dimension of the input target language text feature. The target style feature is added to the encoder to realize the encoder's processing of language and target style features to output the target speech in the target language.
[0077] The speech synthesis method provided in this embodiment performs multi-scale style prediction on the text features of the target language. Based on this, the multi-scale styles are quantified and combined with text feature style prediction to achieve style adaptation, solving the problem of adaptive multi-scale styles across different languages. Simultaneously, vectorized local style features are concatenated with vectorized global style features to achieve multi-scale style feature fusion. Encoding and decoding processing is then performed based on these multi-scale style features, enabling high-similarity cross-language style transfer during speech synthesis while maintaining authentic accents.
[0078] In some implementations, the processing in S22 described above is based on a style model. Therefore, S22 includes:
[0079] (1) Input the text features of the target language and the identifier of the original language into the text prediction style module of the style model to perform style prediction of the original language and obtain style features.
[0080] (2) The vectorized codebook module corresponding to the target language in the style model is determined based on the target language, and the vectorized codebook module corresponds one-to-one with the language.
[0081] (3) Input the style features into the vectorized codebook module to obtain the vectorized style features.
[0082] The style model includes a text prediction style module and a vectorized codebook module. The text prediction style module is used to predict the style of the original language and obtain style features. The vectorized codebook module is consistent with the codebook mentioned above; it corresponds one-to-one with the language and is used to vectorize the style features to obtain vectorized style features.
[0083] The text prediction style module can be the Test2style Predictor mentioned above, or a multi-scale style predictor, etc. There are no restrictions on it here. It can be set according to the actual needs.
[0084] Style models are used to predict style across languages and vectorize style features. In the process of style model processing, a vectorized codebook module that corresponds one-to-one with a language is used to vectorize style features, so that the vectorized style features are language-related and can preserve the style of the target language. While preserving the authentic accent of the target language, cross-language style transfer is achieved.
[0085] The preceding description explained that style features are derived from a style model. Based on this, the method for determining the style model is described in detail below. For example... Figure 4 As shown, the methods for determining the style model include:
[0086] S31, Obtain the reference audio features of the reference audio in the first language.
[0087] The reference audio features are obtained by extracting features from the reference audio, such as Mel features, etc.
[0088] S32, input the audio features into the multi-scale encoder of the preset style model for multi-scale feature encoding to obtain reference local features and reference global features.
[0089] A multi-scale reference encoder is used to extract local and global features from the input audio features, resulting in reference local features and reference global features.
[0090] S33, input the reference local features and reference global features into the local vectorization codebook module and the global vectorization codebook module corresponding to the first language, respectively, to obtain the vectorized reference local features and the vectorized reference global features.
[0091] The vectorized codebook module corresponds to a specific language. Since S32 above yields reference local features and reference global features, the corresponding vectorized codebook module for the first language includes a local vectorized codebook module and a global vectorized codebook module. Nearest neighbor lookup is performed in the local vectorized codebook module using the reference local features to obtain vectorized reference local features; similarly, nearest neighbor lookup is performed in the global vectorized codebook module using the reference global features to obtain vectorized reference global features.
[0092] S34. Based on the reference local features and reference global features, as well as the vectorized reference local features and vectorized reference global features, the loss is calculated to obtain the vectorized loss.
[0093] Because the vectorized codebook module uses the nearest neighbor search method, vectorization-related loss is used to train the module and make the vectorized reference local features (Reference VQ local style) closer to the reference local features. For example, to achieve the goal of the vectorized reference local features being close to the reference local features (Referecne local style), the local vectorization loss (vq_loss_local) is calculated using the following formula:
[0094] vq_loss_local=MSE(referenceVQlocalstyle,sg(referencelocalstyle))+β1*MSE(sg(referenceVQlocalstyle),referencelocalstyle)
[0095] Where β1 is a constant.
[0096] Accordingly, in order to achieve the goal of making the vectorized reference global feature (Reference VQ Global style) approximate the reference global feature (Referecne Global style), the global vectorization loss (vq_loss_global) is calculated using the following formula:
[0097] vq_loss_global=MSE(referenceVQglobalstyle,sg(referenceglobalstyle))+β2*MSE(sg(referenceVQglobalstyle),referenceglobalstyle)
[0098] Where β2 is a constant.
[0099] After obtaining the local and global vectorized losses, they are fused, for example, by weighted averaging, to obtain the vectorized loss. Of course, other methods can also be used to fuse the local and global vectorized losses. There are no restrictions on the implementation method here; the specific settings can be configured according to actual needs.
[0100] S35 updates the parameters of the preset style model based on vectorized loss to determine the style model.
[0101] Based on the vectorized loss, the parameters of the preset style model are updated; after multiple iterations, the parameters of the preset style model are fixed to determine the style model. The stopping condition for iteration can be that the number of iterations reaches a preset number, or that the vectorized loss is less than a preset loss value, etc.
[0102] When training the preset style model, the vectorized reference features (i.e., vectorized reference local features and vectorized reference global features) are made closer to the reference features (reference local features and reference global features). Vector loss is used to update the preset style model, which ensures the reliability of the trained preset style model.
[0103] In some implementations... Figure 5 A schematic diagram of the style model is shown, based on which the above S32 includes:
[0104] (1) Input the reference audio features into the multi-scale encoder of the preset style model for multi-scale feature encoding to obtain local features and reference global features.
[0105] (2) Input the local features into the attention module to obtain the reference local features. The attention module is used to align the reference local features with the reference text features.
[0106] During the training phase of the style model, audio features are first processed by a multi-scale reference encoder to obtain local features and reference global features. Based on this, an attention module is used to input the local features to obtain reference local features, thereby aligning the reference local features with the reference text features. Specifically, the attention module works as follows: it uses the reference text features as the query and the local features as the key and value, and its output features have the same length as the reference text features.
[0107] In some implementations, to remove the dependence on reference audio features during use and to enhance the style's adaptability to different languages, a Text2style Predictor is modeled, using reference text features as data and the original language identifier as a condition, to predict local style features and global style features. The targets of these two features correspond to the reference local features and the reference global features, respectively. Based on this, a style prediction loss is also introduced in the loss calculation. Figure 5 A schematic diagram of the style model is shown, based on which the above S35 includes:
[0108] (1) Obtain the reference text features of the second language and the identifier of the first language.
[0109] (2) Input the reference text features and the identifier of the first language into the text prediction style module to obtain the predicted local style features and predicted global style features of the first language.
[0110] (3) Based on the predicted local style features and predicted global style features of the first language, as well as the reference local features and reference global features, the loss is calculated to obtain the style prediction loss.
[0111] (4) Based on vectorization loss and style prediction loss, update the parameters of the preset style model to determine the style model.
[0112] like Figure 5 As shown, the reference text features and the identifier of the first language are input into the Text2style Predictor module, which outputs the predicted local style feature (localstyle) and the predicted global style feature (globalstyle) for the first language. As mentioned above, the targets of the predicted local style feature and the predicted global style feature for the first language correspond to the reference local style feature (referencelocalstyle) and the reference global style feature (referenceglobalstyle), respectively. Therefore, the style prediction loss is calculated based on these features.
[0113] For example, the style prediction loss style_predictor_loss can be expressed by the following formula:
[0114] style_predictor_loss=MSE(localstyle,referencelocalstyle)
[0115] +MSE(globalstyle,referenceglobalstyle)
[0116] Of course, the above formulas for calculating vectorization loss and style prediction loss are merely examples and do not limit the scope of protection of this disclosure. The specific loss function is set according to actual needs.
[0117] The text prediction style module, which uses reference text features as data and the identifier of the first language as a condition, obtains predicted local style features and predicted global segmentation features. The targets of these two features correspond to the reference local features and the reference global features, respectively. Based on this, a style prediction loss is introduced, which can remove the dependence on audio features when using it, and also enhance the adaptability of style to language.
[0118] As a specific application example of speech synthesis in this disclosure, the field of AI video translation presents a need to transfer a user's style in one language to another non-native language. Video translation generally refers to translating the speech in a video from the original language to the target language, ensuring consistency between the translated speech and the video. Video translation typically consists of multiple cascaded systems, including speech recognition, machine translation, and speech synthesis. To ensure the translated speech corresponds to the original video, the text length is usually controlled first in the machine translation stage, and then the length of the synthesized speech is adjusted in the speech synthesis stage. For example, to translate Chinese speech into English speech, speech recognition is first performed to convert the Chinese speech into Chinese text, and then combined with machine translation to obtain the English text. Subsequent processing of the English text is the speech synthesis processing described in this disclosure. Feature processing is performed on the English text to obtain English text features. These English text features, along with Chinese identifiers, are input into a style model, thereby utilizing the text prediction style module in the style model to predict the Chinese style in the English text. If the Nth language in the style model corresponds to English, the predicted local and global style features are searched using VQ-vaeN-local and VQ-vaeN-global methods, respectively, to obtain the closest style representations: vectorized local style (VQ local style) and vectorized global style (VQ global style). The vectorized global style feature is then added to the vectorized local style feature as a repeating feature to obtain the target style feature. This target style feature is then added to the encoder for encoding, and the encoded result is decoded to output the target English speech.
[0119] This embodiment also provides a speech synthesis device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0120] This embodiment provides a speech synthesis device, such as... Figure 6 As shown, it includes:
[0121] The acquisition module 41 is used to acquire the text features of the target language and the identifier of the original language;
[0122] The style vectorization module 42 is used to predict the style of the original language based on the text features of the target language and the identifier of the original language, to obtain style features, and to query the codebook corresponding to the target language based on the style features to obtain vectorized style features. The codebook corresponds one-to-one with the language and is used to vectorize the style features.
[0123] Processing module 43 is used to perform encoding and decoding processing based on the vectorized style features to determine the target speech of the target language.
[0124] In some implementations, the style vectorization module 42 includes:
[0125] The prediction unit is used to perform multi-scale style prediction of the original language based on the text features of the target language and the identifier of the original language, to obtain the style features, which include local style features and global style features.
[0126] The first determining unit is used to determine the local codebook and global codebook corresponding to the target language using the target language.
[0127] The query unit is used to query the local codebook and the global codebook respectively based on the local style features and the global style features to obtain the vectorized style features, wherein the vectorized style features include vectorized local style features and vectorized global style features.
[0128] In some implementations, the processing module 43 includes:
[0129] The splicing unit is used to splice the vectorized local style features and the vectorized global style features to obtain the target style features;
[0130] The processing unit is used to perform encoding and decoding processing based on the target style features to determine the target speech of the target language.
[0131] In some implementations, the splicing unit includes:
[0132] Obtain sub-units to obtain the target dimension of the vectorized local style features;
[0133] Determine a subunit for determining the number of copies of the vectorized global style feature based on the target dimension;
[0134] The first splicing subunit is used to copy and splice the vectorized global style features based on the number of copies to obtain the copied features;
[0135] The second splicing subunit is used to splice the copied features with the vectorized local style features to obtain the target style features.
[0136] In some implementations, the style vectorization module 42 includes:
[0137] The input unit is used to input the text features of the target language and the identifier of the original language into the text prediction style module of the style model to perform style prediction of the original language and obtain the style features;
[0138] The second determining unit is used to determine the vectorized codebook module in the style model corresponding to the target language based on the target language, wherein the vectorized codebook module corresponds one-to-one with the language;
[0139] The first vectorization unit is used to input the style features into the vectorization codebook module to obtain the vectorized style features.
[0140] In some implementations, the style model determination module includes:
[0141] The acquisition unit is used to acquire the reference audio features of the reference audio in the first language;
[0142] The encoding unit is used to input the audio features into a multi-scale encoder of a preset style model for multi-scale feature encoding to obtain reference local features and reference global features.
[0143] The second vectorization unit is used to input the reference local features and reference global features into the local vectorization codebook module and the global vectorization codebook module corresponding to the first language, respectively, to obtain vectorized reference local features and vectorized reference global features.
[0144] The loss unit is used to calculate the vectorized loss based on the reference local features and reference global features, as well as the vectorized reference local features and vectorized reference global features.
[0145] An update unit is used to update the parameters of the preset style model based on the vectorized loss, so as to determine the style model.
[0146] In some implementations, the updating unit includes:
[0147] The acquisition subunit is used to acquire reference text features of the second language and the identifier of the first language;
[0148] The prediction subunit is used to input the reference text features and the identifier of the first language into the text prediction style module to obtain the predicted local style features and the predicted global style features of the first language.
[0149] The loss subunit is used to calculate the style prediction loss based on the predicted local style features and predicted global style features of the first language, as well as the reference local features and reference global features.
[0150] An update subunit is used to update the parameters of the preset style model based on the vectorization loss and the style prediction loss, so as to determine the style model.
[0151] In some implementations, the encoding unit includes:
[0152] The encoding subunit is used to input the reference audio features into a multi-scale encoder of a preset style model for multi-scale feature encoding to obtain local features and reference global features.
[0153] An attention subunit is used to input the local features into the attention module to obtain the reference local features, and the attention module is used to align the reference local features with the reference text features.
[0154] In this embodiment, the speech synthesis device is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0155] Further functional descriptions of the above modules are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0156] This disclosure also provides an electronic device having the above-described features. Figure 6 The speech synthesis device shown.
[0157] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of this disclosure, such as... Figure 7As shown, the electronic device may include: at least one processor 51, such as a CPU (Central Processing Unit), at least one communication interface 53, memory 54, and at least one communication bus 52. The communication bus 52 is used to enable communication between these components. The communication interface 53 may include a display screen or a keyboard; optionally, the communication interface 53 may also include a standard wired interface or a wireless interface. The memory 54 may be high-speed RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 54 may also be at least one storage device located remotely from the aforementioned processor 51. The processor 51 may be combined with... Figure 6 The described apparatus has an application program stored in memory 54, and the processor 51 calls the program code stored in memory 54 to perform any of the above method steps.
[0158] The communication bus 52 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 52 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0159] The memory 54 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 54 may also include a combination of the above types of memory.
[0160] The processor 51 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP.
[0161] The processor 51 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0162] Optionally, the memory 54 is also used to store program instructions. The processor 51 can invoke the program instructions to implement the speech synthesis method as shown in any embodiment of this application.
[0163] This disclosure also provides a non-transitory computer storage medium storing computer-executable instructions that can execute the speech synthesis method in any of the above method embodiments. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.
[0164] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A speech synthesis method, characterized in that, include: Obtain the text features of the target language and the identifier of the original language; Based on the text features of the target language and the identifier of the original language, style prediction of the original language is performed to obtain style features. Based on the style features, the codebook corresponding to the target language is queried to obtain vectorized style features. The codebook corresponds one-to-one with the language and is used to vectorize the style features. Encoding and decoding processes are performed based on the vectorized style features to determine the target speech of the target language.
2. The method according to claim 1, characterized in that, The process of predicting the style of the original language based on the text features of the target language and the identifier of the original language to obtain style features, and querying the codebook corresponding to the target language based on the style features to obtain vectorized style features, includes: Based on the text features of the target language and the identifier of the original language, multi-scale style prediction of the original language is performed to obtain the style features, which include local style features and global style features. Determine the local codebook and global codebook corresponding to the target language using the target language; Based on the local style features and the global style features, the local codebook and the global codebook are queried respectively to obtain the vectorized style features, which include vectorized local style features and vectorized global style features.
3. The method according to claim 2, characterized in that, The encoding and decoding process based on the vectorized style features to determine the target speech of the target language includes: The vectorized local style features and the vectorized global style features are concatenated to obtain the target style features; Encoding and decoding processes are performed based on the target style features to determine the target speech of the target language.
4. The method according to claim 3, characterized in that, The concatenation of the vectorized local style features and the vectorized global style features to obtain the target style features includes: Obtain the target dimension of the vectorized local style features; The number of copies of the vectorized global style feature is determined based on the target dimension; Based on the number of copies, the vectorized global style features are copied and concatenated to obtain the copied features; The copied features are concatenated with the vectorized local style features to obtain the target style features.
5. The method according to claim 1, characterized in that, The process of predicting the style of the original language based on the text features of the target language and the identifier of the original language to obtain style features, and querying the codebook corresponding to the target language based on the style features to obtain vectorized style features, includes: The text features of the target language and the identifier of the original language are input into the text prediction style module of the style model to perform style prediction of the original language, thereby obtaining the style features; Based on the target language, a vectorized codebook module corresponding to the target language is determined in the style model, and the vectorized codebook module corresponds one-to-one with the language; The style features are input into the vectorized codebook module to obtain the vectorized style features.
6. The method according to claim 5, characterized in that, The methods for determining the style model include: Obtain the reference audio features of the reference audio in the first language; The audio features are input into a multi-scale encoder of a preset style model for multi-scale feature encoding to obtain reference local features and reference global features. The reference local features and reference global features are respectively input into the local vectorization codebook module and the global vectorization codebook module corresponding to the first language to obtain the vectorized reference local features and the vectorized reference global features. Based on the aforementioned reference local features and reference global features, as well as the vectorized reference local features and vectorized reference global features, a vectorized loss is calculated. The parameters of the preset style model are updated based on the vectorized loss to determine the style model.
7. The method according to claim 6, characterized in that, The step of updating the parameters of the preset style model based on the vectorized loss to determine the style model includes: Obtain the reference text features of the second language and the identifier of the first language; The reference text features and the identifier of the first language are input into the text prediction style module to obtain the predicted local style features and predicted global style features of the first language. Based on the predicted local style features and predicted global style features of the first language, as well as the reference local features and reference global features, a loss calculation is performed to obtain the style prediction loss. Based on the vectorization loss and the style prediction loss, the parameters of the preset style model are updated to determine the style model.
8. The method according to claim 7, characterized in that, The step of inputting the audio features into a multi-scale encoder of a preset style model for multi-scale feature encoding to obtain reference local features and reference global features includes: The reference audio features are input into a multi-scale encoder of a preset style model for multi-scale feature encoding to obtain local features and reference global features. The local features are input into the attention module to obtain the reference local features, and the attention module is used to align the reference local features with the reference text features.
9. A speech synthesis device, characterized in that, include: The acquisition module is used to acquire text features of the target language and the identifier of the original language; The style vectorization module is used to predict the style of the original language based on the text features of the target language and the identifier of the original language, to obtain style features, and to query the codebook corresponding to the target language based on the style features to obtain vectorized style features. The codebook corresponds one-to-one with the language and is used to vectorize the style features. The processing module is used to perform encoding and decoding processing based on the vectorized style features to determine the target speech of the target language.
10. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the speech synthesis method of any one of claims 1-8 by executing the computer instructions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the speech synthesis method according to any one of claims 1-8.
Citation Information
Patent Citations
Voice data synthesis method and device, electronic equipment and storage medium
CN116863910A