A speech synthesis method, apparatus, electronic device and storage medium

By determining word segmentation tags, position categories, and original pitches in a speech synthesis system, generating prosodic values ​​using a pre-defined prosodic mapping rule table, training an acoustic model, and converting it into audio, the problem of insufficient prosodic encoding in traditional systems is solved, achieving more natural and accurate speech synthesis.

CN115312026BActive Publication Date: 2025-10-31AVATAR WORKS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211046344.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-10-31
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

Traditional speech synthesis systems lack joint encoding of prosody and position when processing prosodic languages ​​such as Chinese, resulting in unnatural speech synthesis effects and even failures.

Method used

By determining the word segmentation label, position category, and original pitch of each character in the text to be processed, a preset prosodic mapping rule table is queried to map the original pitch to a prosodic value. These prosodic values ​​are then used as front-end text features to train the acoustic model, generating acoustic features, which are finally converted into audio.

Benefits of technology

It improves the accuracy, naturalness, and stability of speech synthesis, increases the MOS score of the acoustic model, and reduces the error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312026B_ABST
    Figure CN115312026B_ABST
Patent Text Reader

Abstract

This invention discloses a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: determining the word segmentation tag, position category, and original pitch of each character in the text to be processed; querying a preset prosodic mapping rule table based on the position category, word segmentation tag, and original pitch, and mapping the original pitch to a prosodic value based on the query result; training an acoustic model using each prosodic value as a front-end text feature, and obtaining a trained acoustic model; generating acoustic features corresponding to the text to be processed based on the trained acoustic model; and converting the acoustic features into audio corresponding to the text to be processed based on a vocoder. The word segmentation tag represents the position of each character in the word segmentation and whether it is a single-character word, while the position category represents the position of each character in the sentence, improving the naturalness and stability of the speech synthesis, thereby further improving the accuracy of the speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech technology, and more specifically, to a speech synthesis method, apparatus, electronic device, and storage medium. Background Technology

[0002] Speech synthesis systems generally consist of three main parts: front-end text feature extraction, acoustic model, and vocoder. How the front-end text features are represented greatly affects the synthesis effect, especially for rhythmic languages ​​like Chinese.

[0003] Traditional DNN (Deep Neural Networks) speech synthesis typically constructs numerous contextual features to represent the position of the current phoneme in a sentence, word, etc., and uses #1, #2, #3, #4 to represent pause levels, i.e., prosodic information. In end-to-end speech synthesis, the common practice is to remove contextual features while retaining pause levels. However, regardless of the approach, there is a lack of joint encoding of prosody and position, making it difficult for the model to learn the corresponding features, resulting in unnatural prosody and even speech synthesis failure.

[0004] Therefore, how to further improve the accuracy of speech synthesis is a technical problem that needs to be solved.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] This invention provides a speech synthesis method, apparatus, electronic device, and storage medium to further improve the accuracy of speech synthesis.

[0007] Firstly, a speech synthesis method is provided, the method comprising:

[0008] Determine the word segmentation tag, position category, and original tone for each character in the text to be processed;

[0009] The preset prosody mapping rule table is queried according to the location category, the word segmentation tag and the original tone, and the original tone is mapped to a prosody value according to the query result;

[0010] The prosodic values ​​are used as front-end text features to train the acoustic model, and the trained acoustic model is obtained.

[0011] Based on the trained acoustic model, acoustic features corresponding to the text to be processed are generated;

[0012] The acoustic features are converted into audio corresponding to the text to be processed based on the vocoder;

[0013] The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

[0014] In some embodiments, determining the word segmentation tag, position category, and original pitch of each character in the text to be processed includes:

[0015] Each word segmentation tag is determined based on a preset word segmentation model;

[0016] Each of the aforementioned location categories is determined based on a preset location recognition model;

[0017] The original tones of each character are determined based on a preset polyphonic character model and a non-polyphonic character tone table.

[0018] The non-polyphonic character tone table includes the original tones of all non-polyphonic characters.

[0019] In some embodiments, determining each location category based on a preset location recognition model includes:

[0020] The text to be processed is input into the preset position recognition model, and the preset position recognition model re-adds preset type punctuation marks to the text to be processed.

[0021] The location category is determined based on the output of the preset location recognition model;

[0022] The preset punctuation marks are commas and periods.

[0023] In some embodiments, the position categories include the beginning of a sentence, the end of a sentence, before a punctuation mark in the middle of a sentence, after a punctuation mark in the middle of a sentence, and other positions in the sentence; the word segmentation tags include the starting position of the word segmentation, the ending position of the word segmentation, other positions of the word segmentation, and single-character words.

[0024] In some embodiments, the preset polyphonic character model is a single model built based on all polyphonic Chinese characters, and determining each of the original tones based on the preset polyphonic character model and the tone table of non-polyphonic characters includes:

[0025] If the current character is a polyphonic character, the original tone of the current character is predicted based on the preset polyphonic character model;

[0026] If the current character is not a polyphonic character, the original tone of the current character is determined by querying the tone table of non-polyphonic characters based on the current character.

[0027] In some embodiments, the preset prosody mapping rule table is established based on the correspondence between the position category, the word segmentation tag, the original tone and the prosody value, wherein the number of prosody values ​​is greater than the number of original tones.

[0028] In some embodiments, the model structures of the preset word segmentation model, the preset position recognition model, and the preset polyphonic character model are all formed by sequentially connecting a Bert model, a 2-layer BLSTM model, a 1-layer fully connected layer, and a softmax layer.

[0029] Secondly, a speech synthesis apparatus is provided, the apparatus comprising:

[0030] The determination module is used to determine the word segmentation tag, position category, and original tone of each character in the text to be processed;

[0031] The mapping module is used to query a preset prosody mapping rule table based on the position category, the word segmentation tag and the original tone, and map the original tone to a prosody value based on the query result;

[0032] The training module is used to train the acoustic model by using the prosodic values ​​as front-end text features, and to obtain the trained acoustic model.

[0033] The generation module is used to generate acoustic features corresponding to the text to be processed based on the trained acoustic model;

[0034] A conversion module is used to convert the acoustic features into audio corresponding to the text to be processed based on a vocoder;

[0035] The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

[0036] Thirdly, an electronic device is provided, comprising:

[0037] Processor; and

[0038] Memory for storing the executable instructions of the processor;

[0039] The processor is configured to execute the speech synthesis method as described in the first aspect by executing the executable instructions.

[0040] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the speech synthesis method as described in the first aspect.

[0041] By applying the above technical solutions, the word segmentation tag, position category, and original tone of each character in the text to be processed are determined; a preset prosodic mapping rule table is queried according to the position category, the word segmentation tag, and the original tone, and the original tone is mapped to a prosodic value according to the query result; each prosodic value is used as a front-end text feature to train the acoustic model, and a trained acoustic model is obtained; acoustic features corresponding to the text to be processed are generated based on the trained acoustic model; the acoustic features are converted into audio corresponding to the text to be processed based on a vocoder; wherein, each word segmentation tag divides the text to be processed into multiple words, the word segmentation tag represents the position of each character in the word segment and whether it is a single word, and the position category represents the position of each character in the sentence, which improves the naturalness and stability of speech synthesis, thereby further improving the accuracy of speech synthesis. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 A schematic flowchart of a speech synthesis method proposed in an embodiment of the present invention is shown;

[0044] Figure 2 A flowchart illustrating a speech synthesis method according to another embodiment of the present invention is shown;

[0045] Figure 3 A schematic diagram of the structure of a speech synthesis device according to an embodiment of the present invention is shown;

[0046] Figure 4 A block diagram of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] It should be noted that other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated in the claims section.

[0049] It should be understood that this application is not limited to the precise structures described below and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0050] The following is combined Figures 1-2 This application describes a speech synthesis method according to exemplary embodiments thereof. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.

[0051] This application provides a speech synthesis method, such as... Figure 1 As shown, the method includes the following steps:

[0052] Step S101: Determine the word segmentation tag, position category, and original tone of each character in the text to be processed.

[0053] In this embodiment, the text to be processed is the text that needs to be synthesized into speech, which includes multiple characters. Word segmentation is a basic module of text analysis. Reasonable word segmentation can greatly help in the recognition of polyphonic characters, pauses, etc. Each word segmentation tag can divide the text to be processed into multiple words. The word segmentation tag represents the position of each character in the segmentation and whether it is a single-character word. Among them, a single-character word is a word with a single character, such as words like "you," "I," "he," "bucket," and "rice." The characters in the text to be processed can form one or more sentences, and the position category represents the position of each character in the sentence. Each character has its original tone. For example, if the text to be processed is Chinese, the original tone may include the first tone, second tone, third tone, fourth tone, and fifth tone (neutral tone).

[0054] Optionally, the word segmentation tag, position category, and original tone of each character can be identified and determined using machine learning algorithms. Specifically, a recognition model can be established using machine learning algorithms to simultaneously identify the word segmentation tag, position category, and original tone; alternatively, three different recognition models can be established using machine learning algorithms to identify the word segmentation tag, position category, and original tone respectively. Those skilled in the art can choose flexibly according to the actual situation.

[0055] In some embodiments of this application, the position categories include the beginning of a sentence, the end of a sentence, before punctuation marks in a sentence, after punctuation marks in a sentence, and other positions in the sentence. The word segmentation tags include the starting position of the word segmentation, the ending position of the word segmentation, other positions of the word segmentation, and single-character words, thereby more accurately determining the position of each character in the sentence and performing word segmentation on the text to be processed, which can improve the accuracy of speech synthesis.

[0056] It should be noted that those skilled in the art may use other different position categories and word segmentation tags as needed, which does not affect the scope of protection of this application.

[0057] Step S102: Query the preset prosody mapping rule table according to the position category, the word segmentation tag and the original tone, and map the original tone to a prosody value according to the query result.

[0058] In this embodiment, prosody refers to subjective linguistic categories such as pauses, stresses, and rhythms during reading or speaking. A pre-established prosody mapping rule table is created, which includes prosody mapping rules that map the original tone to prosodic values. After querying the pre-established prosody mapping rule table based on position category, word segmentation tag, and original tone, the original tone can be mapped to a prosodic value. Those skilled in the art can use different prosody mapping rules according to actual needs, which does not affect the scope of protection of this application.

[0059] In some embodiments of this application, the preset prosody mapping rule table is established based on the correspondence between the position category, the word segmentation tag, the original tone and the prosody value, and the types of prosody values ​​are greater than the types of the original tone.

[0060] In this embodiment, a preset prosodic mapping rule table is established based on the correspondence between position category, word segmentation label, and original tone and prosodic value. When people speak, they often consider information such as the position in word segmentation and sentence position, which can cause slight tone changes. By making the types of prosodic values ​​greater than the types of original tones, the final prosodic values ​​are made more consistent with the tone change rules, thereby enabling more accurate speech synthesis.

[0061] Those skilled in the art can flexibly set different numbers of prosody values, which does not affect the scope of protection of this application. In the specific application scenario of this application, the types of original tones are 1-5, the types of prosody values ​​are 1-15, and the preset prosody mapping rule table can be as shown in Table 1.

[0062] Table 1

[0063]

[0064]

[0065] Combined with Table 1, for example, the text to be processed "Good morning. Have you had your meal?" is processed as follows:

[0066] Good morning. Have you had your meal?

[0067] Word segmentation: / Good morning / , / you / / had / / your / / meal / ?

[0068] Original pitch recognition: / 3 4 3 / , / 3 / 1 4 / 1 / 5

[0069] Prosody value mapping:

[0070] Morning: At the beginning of the sentence, the starting position of word segmentation, 3->12, prosody value is 12;

[0071] Good: At the beginning of the sentence, other positions of word segmentation, prosody unchanged, prosody value is 4;

[0072] Morning: At the end of the sentence, the end position of word segmentation, 3->14, prosody value is 14;

[0073] You: After the punctuation mark in the sentence, single word, 3->12, prosody value is 12;

[0074] Had: Other positions in the sentence, 1->8, prosody value is 8;

[0075] Your: Other positions in the sentence, 4->15, prosody value is 15;

[0076] Meal: Other positions in the sentence, 1->8, prosody value is 8;

[0077] ? : At the end of the sentence, single word & end position of word segmentation, prosody unchanged, prosody value is 5.

[0078] Step S103, use each of the prosody values as front-end text features to train an acoustic model, and obtain a trained acoustic model.

[0079] In this embodiment, the acoustic model can generate acoustic features according to the front-end text features. The acoustic model can be of types such as GMM-HMM, DNN-HMM, DNN / LSTM, DNN+CTC, etc. In the specific application scenario of this application, since each prosody value takes into account the word segmentation label, position category, and original pitch of each character, the MOS (Mean Opinion Score) of the trained acoustic model is increased from 4.41 to 4.52 (with a full score of 5), and the error rate is reduced from 0.03% to 0.01%.

[0080] The specific process of using each prosodic value as a feature of the front-end text to train the acoustic model is obvious to those skilled in the art and will not be elaborated here.

[0081] Step S104: Generate acoustic features corresponding to the text to be processed based on the trained acoustic model.

[0082] In this embodiment, each prosodic value is input into the trained acoustic model, and the trained acoustic model outputs the acoustic features corresponding to the text to be processed.

[0083] In some embodiments of this application, the acoustic model is a tacotron-based model. In the acoustic model, there are multiple speakers. The prenet layer of the Decoder includes speaker information. The Encoder uses the backbone model and adds an adversarial loss function. The Attention layer uses GMMV2. The data offset between the Decoder output and the postnet output is 5 frames.

[0084] In this embodiment, the speech synthesis method adopts an end-to-end approach. The acoustic model is a tacotron-based model with improvements to the tacotron scheme, including:

[0085] Adding multiple speakers and sharing their data improves the robustness of the acoustic model. Incorporating speaker information into the decoder's prenet layer ensures the acoustic model's ability to model speakers.

[0086] The encoder uses a backbone model and adds an adversarial loss function to ensure that the encoder learns features that are independent of the speaker and language.

[0087] The attention layer uses gmmv2, which improves the stability of the model;

[0088] The Decoder output and PostNet output are offset by 5 frames to ensure that the model can see more historical information.

[0089] Step S105: Convert the acoustic features into audio corresponding to the text to be processed based on the vocoder.

[0090] In this embodiment, the vocoder can convert acoustic features into audio. The acoustic features are input into the vocoder, and the vocoder outputs the audio corresponding to the text to be processed.

[0091] In some embodiments of this application, the vocoder is based on Style GAN and incorporates a sinusoidal excitation signal for pitch.

[0092] In this embodiment, based on Style GAN, a sine excitation signal of pitch is added, which greatly improves the modeling ability of the vocoder model.

[0093] By applying the above technical solution, the word segmentation label, position category, and original pitch of each character in the text to be processed are determined; a preset prosody mapping rule table is queried according to the position category, the word segmentation label, and the original pitch, and the original pitch is mapped to a prosody value according to the query result; each prosody value is used as a front-end text feature to train an acoustic model, and a trained acoustic model is obtained; an acoustic feature corresponding to the text to be processed is generated based on the trained acoustic model; the acoustic feature is converted into an audio corresponding to the text to be processed based on a vocoder; wherein, each word segmentation label divides the text to be processed into multiple word segments, the word segmentation label represents the position of each character in the word segment and whether it is a single-word character, and the position category represents the position of each character in the text to be processed in the sentence, which improves the naturalness and stability of speech synthesis, and further improves the accuracy of speech synthesis.

[0094] An embodiment of the present application also proposes a speech synthesis method, as Figure 2 shown, including the following steps:

[0095] Step S201, determining each word segmentation label based on a preset word segmentation model.

[0096] In this embodiment, a preset word segmentation model is established in advance, the text to be processed is input into the preset word segmentation model, and each word segmentation label is determined according to the output result of the preset word segmentation model, so as to obtain a more accurate word segmentation result. In the training stage of the preset word segmentation model, corpora in the field of speech synthesis are collected and annotated. The word segmentation labels include: B, E, M, S. B represents the starting position of the word segment, E represents the ending position of the word segment, M represents other positions of the word segment, and S represents a single-word character. In the prediction stage of the preset word segmentation model, the corresponding word segmentation label is predicted for each character, and the word segmentation result can be obtained.

[0097] In addition, the preset word segmentation model adopts a semantic word segmentation strategy. For example: Good morning, everyone! In the prior art, the usual word segmentation is: everyone, morning, good. In this embodiment, considering the issue of pronunciation and semantic consistency, the word segmentation is divided into: everyone, good morning, and the corresponding word segmentation labels are: BE / BME. Thus, a more accurate word segmentation result can be obtained.

[0098] Step S202, determining each position category based on a preset position recognition model.

[0099] In this embodiment, a preset location recognition model is established in advance. The text to be processed is input into the preset location recognition model, and the location category is determined according to the output result of the preset location recognition model, so that the location recognition can be performed more accurately.

[0100] In some embodiments of this application, determining each location category based on a preset location recognition model includes:

[0101] The text to be processed is input into the preset position recognition model, and the preset position recognition model re-adds preset type punctuation marks to the text to be processed.

[0102] The location category is determined based on the output of the preset location recognition model;

[0103] The preset punctuation marks are commas and periods.

[0104] Location recognition can be achieved using simple punctuation marks, but for some sentences, the externally input punctuation marks may be inappropriate, and some training samples may lack punctuation marks altogether. However, for speech synthesis tasks, different punctuation marks and their distances result in different performance. In this embodiment, a preset location recognition model can determine whether punctuation marks need to be added for each character in the text to be processed. Only commas and periods are required; by integrating insensitive punctuation types, the accuracy of location recognition is improved.

[0105] Step S203: Determine the original tones of each character based on the preset polyphonic character model and the tone table of non-polyphonic characters.

[0106] In this embodiment, a preset polyphonic character model and a non-polyphonic character tone table are established in advance. Chinese characters can be divided into polyphonic characters and non-polyphonic characters. The preset polyphonic character model can identify polyphonic characters, and the non-polyphonic character tone table includes the original tone of all non-polyphonic characters, so as to accurately determine the original tone of each character in the text to be processed.

[0107] In some embodiments of this application, the preset polyphonic character model is a single model built based on all polyphonic Chinese characters, and the determination of each of the original tones based on the preset polyphonic character model and the tone table of non-polyphonic characters includes:

[0108] If the current character is a polyphonic character, the original tone of the current character is predicted based on the preset polyphonic character model;

[0109] If the current character is not a polyphonic character, the original tone of the current character is determined by querying the tone table of non-polyphonic characters based on the current character.

[0110] For speech synthesis, the correct pronunciation of each character is also very important. For example, for the word "重要" (important), if it is pronounced as "chong2 yao4", it may affect the listener's understanding. When the industry deals with this problem, generally a separate model for the pronunciation of each character is established through a decision tree or a finite state machine. In this embodiment, the preset polyphonic character model is a single model established based on all Chinese polyphonic characters. There are more than 100 Chinese polyphonic characters, involving more than 300 pronunciations, so that shallow features can be shared to avoid repeated calculations; sharing the same model facilitates the rapid training of the model, sharing training data, and improving the robustness of the model; the embedding (that is, the vector corresponding to each character) trained with a large amount of data can be utilized.

[0111] In this embodiment, for polyphonic characters, the original tone can be predicted based on the preset polyphonic character model. For non-polyphonic characters, the original tone can be determined by querying the non-polyphonic character tone table, so that the original tones of each character can be accurately determined.

[0112] In some embodiments of the present application, the model structures of the preset word segmentation model, the preset position recognition model, and the preset polyphonic character model are all formed by sequentially connecting a Bert (Bidirectional Encoder Representations from Transformer, bidirectional encoder representation based on Transformer) model, a 2-layer BLSTM (Bidirectional Long Short Term Memory Network, bidirectional long short-term memory network) model, a 1-layer fully connected layer, and a softmax layer.

[0113] In this embodiment, the hidden vector is extracted by a Bert model pre-trained with a large amount of data, and then the hidden vector is input into the BLSTM network. Finally, classification is performed through a 1-layer fully connected layer and softmax. Thus, the word segmentation labels, position categories, and original tones can be more accurately determined respectively. Those skilled in the art can adopt other different model structures according to actual needs, which does not affect the protection scope of the present application.

[0114] Optionally, in the preset word segmentation model, the number of nodes in the first-layer BLSTM model is 512, the number of nodes in the second-layer BLSTM model is 256, and the number of nodes in the fully connected layer is 4. In the preset position recognition model, the number of nodes in the first-layer BLSTM model is 512, the number of nodes in the second-layer BLSTM model is 256, and the number of nodes in the fully connected layer is 3. In the preset polyphonic character model, the number of nodes in the first-layer BLSTM model is 512, the number of nodes in the second-layer BLSTM model is 256, and the number of nodes in the fully connected layer is 315.

[0115] Step S204: Query the preset prosody mapping rule table according to the position category, the word segmentation tag and the original tone, and map the original tone to a prosody value according to the query result.

[0116] Step S205: Use each of the prosodic values ​​as front-end text features to train the acoustic model and obtain the trained acoustic model.

[0117] Step S206: Generate acoustic features corresponding to the text to be processed based on the trained acoustic model.

[0118] Step S207: Convert the acoustic features into audio corresponding to the text to be processed based on the vocoder.

[0119] In this embodiment, the specific implementation of steps S204-S207 can refer to the aforementioned steps S102-S105, and will not be repeated here.

[0120] By applying the above technical solutions, the word segmentation tags are determined based on a preset word segmentation model; the position categories are determined based on a preset position recognition model; the original tones are determined based on a preset polyphonic character model and a non-polyphonic character tone table; a preset prosodic mapping rule table is queried according to the position category, the word segmentation tags, and the original tones, and the original tones are mapped to prosodic values ​​according to the query results; the prosodic values ​​are used as front-end text features to train the acoustic model, and a trained acoustic model is obtained; acoustic features corresponding to the text to be processed are generated based on the trained acoustic model; and the acoustic features are converted into audio corresponding to the text to be processed using a vocoder. This allows for more accurate determination of the word segmentation tags, position categories, and original tones of each character, improving the naturalness and stability of speech synthesis, thereby further improving the accuracy of speech synthesis.

[0121] This application also proposes a speech synthesis device, such as... Figure 3 As shown, the device includes:

[0122] The determination module 301 is used to determine the word segmentation tag, position category, and original tone of each character in the text to be processed;

[0123] The mapping module 302 is used to query a preset prosody mapping rule table according to the position category, the word segmentation tag and the original tone, and map the original tone to a prosody value according to the query result;

[0124] Training module 303 is used to train the acoustic model by using the prosody values ​​as front-end text features, and to obtain the trained acoustic model.

[0125] The generation module 304 is used to generate acoustic features corresponding to the text to be processed based on the trained acoustic model;

[0126] Conversion module 305 is used to convert the acoustic features into audio corresponding to the text to be processed based on a vocoder;

[0127] The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

[0128] In specific application scenarios, module 301 is specifically used for:

[0129] Each word segmentation tag is determined based on a preset word segmentation model;

[0130] Each of the aforementioned location categories is determined based on a preset location recognition model;

[0131] The original tones of each character are determined based on a preset polyphonic character model and a non-polyphonic character tone table.

[0132] The non-polyphonic character tone table includes the original tones of all non-polyphonic characters.

[0133] In specific application scenarios, module 301 is also specifically used for:

[0134] The text to be processed is input into the preset position recognition model, and the preset position recognition model re-adds preset type punctuation marks to the text to be processed.

[0135] The location category is determined based on the output of the preset location recognition model;

[0136] The preset punctuation marks are commas and periods.

[0137] In specific application scenarios, the position categories include the beginning of a sentence, the end of a sentence, before punctuation marks in a sentence, after punctuation marks in a sentence, and other positions in the sentence. The word segmentation tags include the starting position of the word segmentation, the ending position of the word segmentation, other positions of the word segmentation, and single-character words.

[0138] In specific application scenarios, the preset polyphonic character model is a single model built based on all polyphonic Chinese characters. The determining module 301 is also specifically used for:

[0139] If the current character is a polyphonic character, the original tone of the current character is predicted based on the preset polyphonic character model;

[0140] If the current character is not a polyphonic character, the original tone of the current character is determined by querying the tone table of non-polyphonic characters based on the current character.

[0141] In specific application scenarios, the preset prosody mapping rule table is established based on the correspondence between the position category, the word segmentation tag, the original tone and the prosody value, and the types of prosody values ​​are greater than the types of original tones.

[0142] In specific application scenarios, the model structures of the preset word segmentation model, the preset position recognition model, and the preset polyphonic character model are all formed by sequentially connecting a Bert model, a 2-layer BLSTM model, a 1-layer fully connected layer, and a softmax layer.

[0143] The above-described apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the description of the method embodiments.

[0144] This invention also provides an electronic device, such as... Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.

[0145] Memory 403 is used to store the processor's executable instructions;

[0146] Processor 401 is configured to execute the following via executing the executable instructions:

[0147] Determine the word segmentation tag, position category, and original tone for each character in the text to be processed;

[0148] The preset prosody mapping rule table is queried according to the location category, the word segmentation tag and the original tone, and the original tone is mapped to a prosody value according to the query result;

[0149] The prosodic values ​​are used as front-end text features to train the acoustic model, and the trained acoustic model is obtained.

[0150] Based on the trained acoustic model, acoustic features corresponding to the text to be processed are generated;

[0151] The acoustic features are converted into audio corresponding to the text to be processed based on the vocoder;

[0152] The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

[0153] The aforementioned communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0154] The communication interface is used for communication between the aforementioned terminal and other devices.

[0155] The memory may include RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0156] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0157] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which, when executed by a processor, implements the speech synthesis method as described above.

[0158] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the speech synthesis method as described above.

[0159] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0160] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0161] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0162] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A speech synthesis method, characterized in that, The method includes: Each word segmentation tag is determined based on a preset word segmentation model; the text to be processed is input into a preset position recognition model, and the preset position recognition model adds preset type punctuation marks to the text; the position category is determined according to the output of the preset position recognition model; wherein, the preset type punctuation marks are commas and periods; each original tone is determined based on a preset polyphonic character model and a non-polyphonic character tone table; wherein, the non-polyphonic character tone table includes the original tones of all non-polyphonic characters; The system queries a preset prosodic mapping rule table based on the location category, the word segmentation tag, and the original tone, and maps the original tone to a prosodic value based on the query result. The preset prosodic mapping rule table is established based on the correspondence between the location category, the word segmentation tag, the original tone, and the prosodic value, and the types of prosodic values ​​are greater than the types of original tones. The prosodic values ​​are used as front-end text features to train the acoustic model, and the trained acoustic model is obtained. Based on the trained acoustic model, acoustic features corresponding to the text to be processed are generated; The acoustic features are converted into audio corresponding to the text to be processed based on the vocoder; The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

2. The method as described in claim 1, characterized in that, The position categories include the beginning of a sentence, the end of a sentence, before punctuation marks in the sentence, after punctuation marks in the sentence, and other positions in the sentence. The word segmentation tags include the starting position of the word segmentation, the ending position of the word segmentation, other positions of the word segmentation, and single-character words.

3. The method as described in claim 1, characterized in that, The preset polyphonic character model is a single model built based on all polyphonic Chinese characters. The determination of each original tone based on the preset polyphonic character model and the tone table of non-polyphonic characters includes: If the current character is a polyphonic character, the original tone of the current character is predicted based on the preset polyphonic character model; If the current character is not a polyphonic character, the original tone of the current character is determined by querying the tone table of non-polyphonic characters based on the current character.

4. The method as described in claim 1, characterized in that, The model structures of the preset word segmentation model, the preset position recognition model, and the preset polyphonic character model are all formed by sequentially connecting a Bert model, a 2-layer BLSTM model, a 1-layer fully connected layer, and a softmax layer.

5. A speech synthesis device, characterized in that, The device includes: The determination module is used to determine each word segmentation tag based on a preset word segmentation model; input the text to be processed into a preset position recognition model, and make the preset position recognition model re-add preset type punctuation marks to the text to be processed; determine the position category according to the output result of the preset position recognition model; wherein, the preset type punctuation marks are commas and periods; determine each original tone based on a preset polyphonic character model and a non-polyphonic character tone table; wherein, the non-polyphonic character tone table includes the original tones of all non-polyphonic characters; The mapping module is used to query a preset prosody mapping rule table based on the position category, the word segmentation tag, and the original tone, and map the original tone to a prosody value based on the query result; the preset prosody mapping rule table is established based on the correspondence between the position category, the word segmentation tag, the original tone, and the prosody value, and the types of prosody values ​​are greater than the types of original tones; The training module is used to train the acoustic model by using the prosodic values ​​as front-end text features, and to obtain the trained acoustic model. The generation module is used to generate acoustic features corresponding to the text to be processed based on the trained acoustic model; A conversion module is used to convert the acoustic features into audio corresponding to the text to be processed based on a vocoder; The segmentation tags divide the text to be processed into multiple segments. Each segmentation tag represents the position of each character in the segment and whether it is a single-character word. The position category represents the position of each character in the sentence of the text to be processed.

6. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the speech synthesis method according to any one of claims 1 to 4 by executing the executable instructions.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech synthetic method based on rhythm character

    CN101000765A

  • NLP-based mobile phone short message identification method and related equipment

    CN110377699A

  • Voice playing method and computer equipment

    CN114495896A