Speech synthesis method and device
Through the proposed speech synthesis method, through phoneme conversion, decoupling, feature extraction and fusion processing, the problems of pronunciation errors and insufficient rhythm in the prior art are solved, and high-quality speech synthesis is achieved.
Patent Information
- Application Number
- CN202510307878.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
There are problems of pronunciation errors and insufficient pronunciation in existing pronunciation synthesis technologies, especially when it comes to pronunciation synthesis for small languages.
A speech synthesis method is proposed, by obtaining the text to be synthesized, performing phoneme conversion, decoupling phonemes and tones, extracting phoneme features and semantic features, and performing fusion processing to generate speech.
Improve the accuracy and naturalness of speech synthesis to ensure that synthetic speech achieves the expected results in terms of naturalness, pronunciation and semantic consistency.
Smart Images

Figure CN120164451A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a speech synthesis method and apparatus. Background Art
[0002] Speech synthesis (Text-to-Speech, TTS) is a technology that converts text information into speech information. In related technologies, when performing speech synthesis, there are problems such as pronunciation errors and insufficient prosody. Summary of the Invention
[0003] In view of this, the present disclosure provides a speech synthesis method, apparatus, electronic device, storage medium, and computer program product.
[0004] According to one aspect of the present disclosure, there is provided a speech synthesis method, the method including:
[0005] Obtaining text to be synthesized;
[0006] Performing phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located;
[0007] Decoupling the phonemes and pitches in the first phoneme sequence, and extracting phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; where the second phoneme sequence includes the at least one phoneme;
[0008] Performing text encoding on the text to be synthesized in the syllable dimension, and extracting semantic features of the text to be synthesized;
[0009] Performing fusion processing on the semantic features and the phoneme features, and generating speech based on the fused features.
[0010] In a possible implementation, the performing phoneme conversion on the text to be synthesized to obtain a first phoneme sequence includes:
[0011] Filtering characters in the text to be synthesized that are not in the target language based on the target language corresponding to the text to be synthesized, to obtain filtered text;
[0012] Segmenting the filtered text according to a preset pause symbol corresponding to the target language, to obtain a plurality of clauses; where the preset pause symbol is used to represent the pause of a sentence;
[0013] Performing phoneme conversion processing on each of the plurality of clauses respectively, to obtain a phoneme subsequence corresponding to each clause;
[0014] Merging the phoneme subsequences corresponding to each clause, to obtain the first phoneme sequence.
[0015] In a possible implementation, the method further includes:
[0016] Based on the first phoneme sequence and the second phoneme sequence, a pitch sequence is obtained; the pitch sequence includes the pitches corresponding to the at least one phoneme; wherein, the pitch corresponding to any phoneme is determined by the pitch of the syllable where the phoneme is located;
[0017] In the process of generating speech based on the fused features, the pitch of the speech is controlled by the pitch sequence.
[0018] In a possible implementation, the text encoding of the text to be synthesized in the syllable dimension and extracting the semantic features of the text to be synthesized includes:
[0019] Performing word segmentation on the text to be synthesized according to syllables to obtain at least one syllable;
[0020] Combined with a preset vocabulary, determining the encoding corresponding to each syllable in the at least one syllable; wherein, the preset vocabulary includes the corresponding relationship between different syllables and different encodings;
[0021] Based on the encoding corresponding to each syllable in the at least one syllable, the semantic features of the text to be synthesized are extracted.
[0022] In a possible implementation, the respectively performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes:
[0023] For any one of the multiple clauses, determining whether the clause contains a preset character; and performing normalization processing in the case of containing the preset character;
[0024] Performing phoneme conversion processing on the clause after the normalization processing to obtain a phoneme subsequence corresponding to the clause.
[0025] In a possible implementation, the respectively performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes:
[0026] For any one of the multiple clauses, performing phoneme conversion processing by using a phoneme conversion model to obtain a phoneme conversion result of the clause;
[0027] Determining the phoneme to be corrected in the phoneme conversion result of the clause;
[0028] Based on a preset phoneme correction dictionary, convert the phonemes to be corrected in the phoneme conversion result of the clause into standard phonemes to obtain a phoneme subsequence corresponding to the clause; wherein, the phoneme correction dictionary includes: the corresponding relationship between different phonemes to be corrected and different standard phonemes.
[0029] In a possible implementation manner, the method further includes:
[0030] In the process of generating speech based on the fused features, predict the duration of each phoneme in the at least one phoneme based on a duration model; wherein, the duration model is trained based on the true alignment labels of phoneme-duration.
[0031] In a possible implementation manner, the method further includes:
[0032] Obtain the preset duration corresponding to each clause in the multiple clauses;
[0033] In the case that the duration of the generated speech exceeds the sum of the preset durations of the multiple clauses, perform coaxial processing on at least two adjacent clauses in the speech. After the coaxial processing, add one or more spaces between the at least two adjacent clauses so that the duration of the processed speech is not greater than the sum of the preset durations of the multiple clauses.
[0034] According to another aspect of the present disclosure, there is provided a speech synthesis device, the device includes:
[0035] An acquisition module, configured to acquire the text to be synthesized;
[0036] A phoneme conversion module, configured to perform phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located;
[0037] A phoneme feature extraction module, configured to decouple the phonemes and pitches in the first phoneme sequence, and extract the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; wherein, the second phoneme sequence includes the at least one phoneme;
[0038] A semantic feature extraction module, configured to perform text encoding on the text to be synthesized in the syllable dimension, and extract the semantic features of the text to be synthesized;
[0039] A speech synthesis module, configured to perform fusion processing on the semantic features and the phoneme features, and generate speech based on the fused features.
[0040] According to another aspect of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0041] According to another aspect of the present disclosure, there is provided a non - volatile computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above - mentioned method are implemented.
[0042] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non - volatile computer - readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above - mentioned method are implemented.
[0043] Through the above - mentioned aspects of the present disclosure, obtain the text to be synthesized; perform phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located; decouple the phonemes and pitches in the first phoneme sequence, and extract the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; wherein, the second phoneme sequence includes the at least one phoneme; perform text encoding on the text to be synthesized in the syllable dimension to extract the semantic features of the text to be synthesized; perform fusion processing on the semantic features and the phoneme features, and generate speech based on the fused features. In this way, decouple the phonemes and pitches in the first phoneme sequence, thereby isolating the action space domain of the pitch, and then extract the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence, improving the accuracy of the extracted phoneme features; at the same time, perform text encoding on the text to be synthesized in the syllable dimension, and more accurate and rich semantic features can be obtained; furthermore, fuse the phoneme features extracted from the decoupled second phoneme sequence with the semantic features, so as to ensure the quality of the fused features. The speech synthesized based on the fused features not only has accurate pronunciation but also natural prosody, ensuring that the synthesized speech achieves the expected effects in terms of naturalness, prosody, and semantic consistency.
[0044] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0045] The drawings included in the specification and constituting a part of the specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present disclosure, and are used to explain the principles of the present disclosure.
[0046] Figure 1 A flowchart showing a speech synthesis method according to an embodiment of the present disclosure;
[0047] Figure 2 A schematic structural diagram showing a VITS model according to an embodiment of the present disclosure;
[0048] Figure 3Schematic diagram showing extraction of semantic features and phoneme features according to an embodiment of the present disclosure;
[0049] Figure 4 Schematic flowchart showing speech synthesis using a speech synthesis model according to an embodiment of the present disclosure;
[0050] Figure 5 Structural diagram showing a speech synthesis device according to an embodiment of the present disclosure;
[0051] Figure 6 Block diagram showing an electronic device 1900 according to an embodiment of the present disclosure. Detailed implementation manners
[0052] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0053] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.
[0054] When an element is referred to as being "connected", "coupled", "responsive", or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.
[0055] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0056] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0057] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0058] In scenarios such as film and television drama dubbing production, the scope of speech synthesis is not limited to Chinese, but also extends to various minority languages. For example, the Thai subtitles of a Chinese film and television drama can be converted into Thai speech through speech synthesis. However, in the process of speech synthesis for different languages, especially for minority languages, there are problems such as pronunciation errors and insufficient rhythm.
[0059] To solve the above problems, the embodiments of the present disclosure provide a speech synthesis method (detailed description is as follows). According to the uniqueness of the target language, specific optimizations are carried out for the speech synthesis of the target language. Thus, the quality of the synthesized speech in the target language is further improved, ensuring that the synthesized speech not only has accurate pronunciation but also natural rhythm, providing an immersive auditory experience for the audience.
[0060] Exemplarily, the speech synthesis method can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be a desktop terminal or a mobile terminal. For example, it can be various types of electronic devices such as a laptop computer, a tablet computer, a desktop computer, a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle terminal, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0061] It should be noted that the above-described film and television drama dubbing production scenario described in the embodiments of the present disclosure is for more clearly explaining the technical solutions of the embodiments of the present disclosure, and does not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art know that for the emergence of other similar or new scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0062] Figure 1 The flowchart showing a speech synthesis method according to an embodiment of the present disclosure is as Figure 1 shown, and may include the following steps:
[0063] Step 101, obtain the text to be synthesized.
[0064] Among them, the text to be synthesized can be text in any form. For example, it can be the lines of a film and television drama, the lyrics of a song, a news report, etc. The language used for the text to be synthesized (i.e., the target language) can be Chinese, English, Thai, Vietnamese, etc., and there is no limitation thereto. As an example, the target language of the text to be synthesized is a tonal language; for example, Chinese, Thai, Vietnamese, etc.
[0065] Step 102: Perform phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the tone of the syllable where the at least one phoneme is located.
[0066] Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech; for different languages, there are corresponding phonemes; for example, in Chinese, "b", "p", "a", "o", etc. are all phonemes, and in Thai etc. are all phonemes.
[0067] The tone can also be called the pitch or sub-tone, which represents the pitch of the sound. In tone languages, when the same speech sound is made, using different lengths and different pitch values of tones will form different meanings (i.e., semantics or semantic meanings). Therefore, tones can be used to express semantic differences. At the same time, the changes in the rhythm of speech brought about by tones play an important role in aspects such as tone and expressing emotions. For different tone languages, the corresponding tones may be different. For example, in Chinese, there are four tones: the first tone - high level, the second tone - rising, the third tone - falling-rising, and the fourth tone - falling. For another example, in Thai, there are five tones, namely the first tone (low level), the second tone (mid-rising), the third tone (low falling), the fourth tone (high falling), and the fifth tone (high rising).
[0068] A syllable represents a speech unit composed of phonemes and is the smallest speech segment that people naturally feel when speaking. For different languages, there are different syllable structure rules. For example, in Chinese, "mā" is a syllable, which is composed of the phonemes "m" and "a", and the corresponding tone is "high level". For another example, in Thai, is a syllable, which is composed of phonemes and the corresponding tone is "low level".
[0069] Exemplarily, phoneme conversion can be performed through a phoneme conversion model. Among them, the phoneme conversion models corresponding to different languages are different. A pre-trained phoneme conversion model corresponding to the target language can be selected, and the text to be synthesized is input into this phoneme conversion model. This phoneme conversion model performs phoneme conversion on the text to be synthesized and outputs a first phoneme sequence. Among them, the phoneme conversion model can be trained in an existing manner, and this is not limited. As an example, the phoneme conversion model can be a Grapheme-to-Phoneme (G2P) model. The G2P model performs phoneme conversion on the text to be synthesized based on a preset phoneme dictionary of the target language and outputs a first phoneme sequence.
[0070] It is understandable that for the same sentence in the target language, different phoneme dictionaries are used for phoneme conversion, and the generated first phoneme sequence may be different; illustratively, a more fine-grained phoneme dictionary can be constructed in advance according to the characteristics of the target language, for example, the preset phoneme dictionary of the target language contains the correspondence between different syllables and phonemes of the target language; in this way, when the phoneme conversion model performs phoneme conversion on the text to be synthesized based on the phoneme dictionary, it can first determine the syllables corresponding to the text to be synthesized, and then determine the phonemes contained in each syllable in turn based on the phoneme dictionary, and at the same time, it can determine the tone corresponding to each syllable, and then, according to the order of appearance of each syllable in the text to be synthesized, generate the first phoneme sequence. In this way, based on the preset phoneme dictionary of the target language, the text to be synthesized in the target language can be efficiently converted into a standardized phoneme sequence containing phonemes and tones, so as to achieve refined phoneme expression, thereby making the pronunciation in the synthesized speech more natural.
[0071] The first phoneme sequence includes the phonemes contained in each syllable corresponding to the text to be synthesized and the tones corresponding to each syllable, wherein the phonemes of any syllable may be one or more, and the tone corresponding to the syllable is the tone of the syllable where the different phonemes in the syllable are located. As an example, the first phoneme sequence may include digital encodings of phonemes and tones. For example, taking the Chinese word "現" as an example, based on a preset Chinese phoneme dictionary, the phoneme conversion model may convert the syllable "xiàn" corresponding to "現" in "現" into the phonemes "x", "i", and "an", and digitally encode the corresponding tone "qusheng" as "4", and convert the syllable "zài" corresponding to "在" into the phonemes "z", "ai", and digitally encode the corresponding tone "qusheng" as "4", and finally obtain the phoneme sequence "xian4zai4"; for another example, taking the Thai word "現" as an example, the phoneme conversion model may convert the syllable "xiàn" corresponding to "現" in "現" into the phonemes "x", "i", and "an ... For example, based on the preset Thai phoneme dictionary, the phoneme conversion model can middle The corresponding syllable "sawa" is converted into the phoneme "sawa", and the corresponding tone "middle rising tone" is digitally encoded as "2". The corresponding syllable "t" is converted into a phoneme The corresponding tone "medium rising tone" is digitally encoded as "2". The corresponding syllable "di" is converted into the phoneme "di", and the corresponding tone "low flat tone" is digitally encoded as "1", and finally the phoneme sequence "sawa2 2di1".
[0072] In a possible implementation, the phoneme conversion of the text to be synthesized to obtain a first phoneme sequence includes: filtering the characters in the text to be synthesized that are not in the target language based on the target language corresponding to the text to be synthesized to obtain a filtered text; splitting the filtered text according to the preset pause symbols corresponding to the target language to obtain a plurality of clauses; wherein the preset pause symbols are used to represent the pauses in the sentences; performing phoneme conversion processing on each of the plurality of clauses respectively to obtain a phoneme subsequence corresponding to each clause; and merging the phoneme subsequences corresponding to each clause to obtain the first phoneme sequence. In this way, the characteristics of the target language are fully considered. First, the characters in the text to be synthesized that are not in the target language are filtered, then the filtered text is split into clauses according to the preset pause symbols corresponding to the target language, and finally, phoneme conversion processing is performed on the clauses respectively; through this step-by-step processing method, the problem that the text to be synthesized cannot be correctly phoneme-converted as a whole due to non-target language characters and specific pause symbols in the target language is effectively solved, thus ensuring the accuracy and reliability of phoneme conversion.
[0073] Among them, the characters of non-target languages can be preset according to the characteristics of the target language. For example, taking the target language as Chinese, characters that do not appear in the Chinese standard expression structure, such as Korean characters, Japanese characters, and space characters, are all characters of non-target languages. In this way, the characters of non-target languages in the text to be synthesized are filtered, thus avoiding the interference of non-target language characters on the phoneme conversion process.
[0074] Exemplarily, corresponding pause symbols can be preset according to the characteristics of different languages; for example, for Chinese, the corresponding pause symbols can include commas, periods, question marks, exclamation marks, semicolons, etc. For another example, for Thai, the corresponding pause symbol can include: space; for another example, for Vietnamese, the corresponding pause symbols can include: space, period, etc. In this way, based on the preset pause symbols corresponding to the target language, the correct splitting of the text in the target language is realized; especially for languages containing spaces (such as Thai, Vietnamese, etc.), the text of this language can be correctly split according to pause symbols such as spaces, and the semantic integrity of the independent clauses split is ensured. Furthermore, phoneme conversion processing is performed on each clause respectively to ensure that each clause can accurately generate the corresponding phoneme subsequence.
[0075] Exemplarily, a phoneme conversion model can be adopted to perform phoneme conversion processing on each of the multiple clauses segmented above. As an example, each of the multiple clauses can be separately input into a G2P model. The G2P model performs phoneme conversion on each clause based on a preset phoneme dictionary of the target language, so as to obtain a phoneme subsequence corresponding to each clause. The phoneme subsequence may include the phonemes contained in each syllable in the clause and the tones corresponding to each syllable. After completing the phoneme conversion for each clause, based on the order in which each clause appears in the text to be synthesized, the phoneme subsequences are merged into a complete phoneme sequence, that is, the first phoneme sequence.
[0076] Furthermore, considering the characteristics of different languages, during the process of performing phoneme conversion using a phoneme conversion model, potential errors in the phoneme conversion process can be corrected by means such as normalization and a phoneme correction dictionary, thereby improving the reliability and accuracy of the subsequent generated speech.
[0077] In a possible implementation manner, the separately performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes: for any one of the multiple clauses, determining whether the clause contains a preset character; and performing normalization processing when the preset character is included; performing phoneme conversion processing on the clause after normalization processing to obtain a phoneme subsequence corresponding to the clause. In this way, through normalization processing, the error rate in the phoneme conversion process is effectively reduced, and the accuracy and stability of the phoneme conversion are improved.
[0078] Among them, the preset character refers to a character or combination of characters that can appear in the expression structure specified by the target language but cannot be directly pronounced according to the pronunciation rules of the target language. For example, the preset character may include characters such as Arabic numerals, symbols, units, etc., and may also include combinations of these characters and literal characters. For any language, the preset character can be preset according to the characteristics of the language; for example, in Chinese, "2kg", "5℃", "188XXXX0000", "1999 yuan", etc.
[0079] Normalization processing refers to converting preset characters into text characters of a target language, where the text characters of the target language are in a standard and canonical representation form so that they can be read according to the pronunciation rules of the target language. In this way, when a clause contains preset characters, normalization processing is performed through preset normalization rules, and the clause containing the preset characters is converted into a standard and canonical representation to ensure the subsequent generation of correct speech; wherein the preset normalization rule is used to indicate the correspondence between the preset characters and the text characters of the target language, and the preset normalization rule can be set according to the characteristics of the target language. For example, in Chinese, the number and text combination "1999 yuan" is normalized to "one thousand nine hundred and ninety-nine yuan", the number and text combination "1999 year" is normalized to "one thousand nine hundred and ninety-nine years", the symbol "℃" is normalized to "degrees Celsius", the text and number combination "telephone number 188XXXX0000" is normalized to "telephone number one eighty-eight XXXX zero zero zero zero", the number and unit combination "2kg" is normalized to "two kilograms", etc.; for another example, in Thai and Vietnamese, for year numbers, the monosyllabic morpheme simplified reading is adopted, and the number and text combination "1999 year" is normalized to "one thousand ninety-nine years".
[0080] Considering that when the phoneme conversion model converts any clause into phonemes, it can first determine the syllable corresponding to the clause based on the correspondence between characters and syllables, and then convert any syllable into the phoneme corresponding to the syllable based on the phoneme dictionary; however, due to the existence of polyphones, that is, the same character may correspond to different syllables, the character may be mapped to the wrong syllable, resulting in the final conversion of the wrong phoneme. For example, the syllable corresponding to the character "行" in "银行" is "háng", while the syllable corresponding to the character "xíng" in "走" is "xíng". The correct phonemes corresponding to the syllables are different; or, when training the phoneme conversion model, the training samples may have deficiencies such as noise, pronunciation habits, unclear pronunciation, accent differences, and annotation errors, which cause the phoneme conversion model to learn the wrong correspondence between characters and syllables, resulting in errors in the phoneme conversion model when determining the syllables corresponding to each clause, which in turn leads to incorrect phonemes in the phoneme subsequence corresponding to the clause. For example, the syllable corresponding to the character "中" is "zhōng", which may be annotated as "zōng" in the training samples due to pronunciation habits.
[0081] In a possible implementation, the step of respectively performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes: for any one of the multiple clauses, performing phoneme conversion processing using a phoneme conversion model to obtain a phoneme conversion result of this clause; determining the phonemes to be corrected in the phoneme conversion result of this clause; based on a preset phoneme correction dictionary, converting the phonemes to be corrected in the phoneme conversion result of this clause into standard phonemes to obtain the phoneme subsequence corresponding to this clause; wherein, the phoneme correction dictionary includes: the corresponding relationships between different phonemes to be corrected and different standard phonemes. In this way, by determining the phonemes to be corrected in the phoneme conversion result of this clause, and then through the constructed phoneme correction dictionary, the correct mapping from the phonemes to be corrected to the standard phonemes is realized, thereby correcting potential incorrect phonemes in real time during the phoneme conversion process, and significantly improving the accuracy and stability of the phoneme conversion.
[0082] Exemplarily, after performing phoneme conversion processing using a phoneme conversion model to obtain a phoneme conversion result of any one clause, it is possible to detect the phoneme conversion result of the clause based on the pronunciation rules of the target language or common incorrect pronunciation problems, so as to determine whether there are errors in the phoneme conversion result. If there are errors, then correspondingly determine the incorrect phonemes or phoneme combinations in the phoneme conversion result as the phonemes to be corrected. If there are no errors, then correspondingly configure the phonemes to be corrected as empty. As an example, based on the pronunciation rules of Chinese, the phonemes 'j', 'q', 'x' do not directly combine with the phoneme 'a'. Therefore, if it is detected that 'ja', 'qa' or 'xa' exists in the phoneme conversion result, it is determined that there is an error in the phoneme conversion result, and the corresponding 'ja', 'qa' or 'xa' can be used as the phonemes to be corrected. As another example, it is possible to pre-determine the characters that are prone to errors based on common incorrect pronunciation problems in Chinese, such as polyphonic characters, easily confused characters, flat and retroflex characters, etc. Then, based on the semantics in this clause, determine the reference phonemes of the syllables corresponding to these preset characters in this clause, and then determine whether the reference phonemes are consistent with the phonemes corresponding to these preset characters in the phoneme conversion result. If they are not consistent, it is determined that there is an error in the phoneme conversion result, and correspondingly, the incorrect phonemes or phoneme combinations in the phoneme conversion result are used as the phonemes to be corrected; for example, if the preset character 'truck' is included in this clause, the reference phoneme of the corresponding syllable can be determined as 'ka' through semantic analysis. If the phoneme corresponding to this clause in the phoneme conversion result is 'qia', it is determined that there is an error, and 'qia' is used as the phoneme to be corrected; for another example, if the preset character'material' is included in this clause, the reference phoneme of the corresponding syllable can be determined as 'liao' through semantic analysis. If the phoneme corresponding to this clause in the phoneme conversion result is 'niao', it is determined that there is an error, and the phoneme 'n' is used as the phoneme to be corrected.
[0083] Exemplarily, samples can be collected for training a phoneme error correction model according to the pronunciation rules of the target language and common incorrect pronunciation problems. The output result of the phoneme conversion model can be automatically detected by the trained phoneme error correction model. Among them, the phoneme correction model has the ability to detect whether there are errors in the phoneme conversion result of the above detection clause and locate the phoneme to be corrected; the training process of the phoneme correction model can be implemented in an existing manner; in addition, the phoneme to be corrected in the phoneme conversion result of the clause can also be determined by manual judgment or manual-assisted judgment, and this is not limited herein.
[0084] Exemplarily, for any clause, after determining that there is an error in the phoneme conversion result of the clause and determining the phoneme to be corrected, the phoneme to be corrected can be converted into a standard phoneme based on a preset phoneme correction dictionary, so as to obtain the final phoneme subsequence of the clause. Among them, the phoneme correction dictionary can be preset according to the pronunciation rules of the target language, the pronunciation habits of the people using the target language, etc.; as an example, the phoneme correction dictionary can include corresponding relationships related to pronunciation rules, such as "ja" - "jia", "qa" - "qia", "xa" - "xia", etc.; it can also include corresponding relationships related to flat and retroflex sounds, such as "zh" - "z", "sh" - "s", "z" - "zh", etc.; it can also include corresponding relationships related to polyphonic characters, such as "xing" - "hang", "ka" - "qia", etc.; it can also include corresponding relationships related to accents, such as "n" - "l", etc.
[0085] Exemplarily, based on a preset phoneme correction dictionary, the phoneme to be corrected in the phoneme conversion result of the clause is converted into a standard phoneme to obtain the phoneme subsequence corresponding to the clause, and it can also include: based on the preset phoneme correction dictionary, the phoneme to be corrected in the phoneme conversion result of the clause is converted into a standard phoneme, and when the tone of the syllable where the standard phoneme is located is different from the tone of the syllable where the phoneme to be corrected is located, the tone of the syllable where the phoneme to be corrected is replaced with the tone of the syllable where the standard phoneme is located; based on the standard phoneme and the converted tone, the phoneme subsequence corresponding to the clause is obtained. For example, if the clause includes the preset character "长", it can be determined through semantic analysis that the phoneme of the corresponding syllable is "张", and the tone of the syllable is determined to be "上声". If the corresponding phoneme in the phoneme conversion result of the clause is "昌", it is determined that there is an error, and "昌" is used as the phoneme to be corrected, and the tone corresponding to the syllable where the phoneme is located is "阳平". Based on the corresponding relationship "chang"-"zhang" in the phoneme correction dictionary, the phoneme "chang" to be corrected is replaced with the standard phoneme "zhang". At the same time, the tone "yangping" of the syllable where the phoneme to be corrected is located is replaced with "shangsheng", and the phoneme subsequence corresponding to the sentence is obtained. The phoneme subsequence includes the phoneme "zhang" of the syllable corresponding to the character "chang" and the tone "shangsheng" of the syllable.
[0086] Step 103: decouple the phonemes and tones in the first phoneme sequence, and extract the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; wherein the second phoneme sequence includes the at least one phoneme.
[0087] Exemplarily, in the first phoneme sequence, all the phonemes and all the tones can be split, so as to achieve the decoupling of the phonemes and tones, and the split phonemes constitute the second phoneme sequence, that is, the decoupled second phoneme sequence, and the second phoneme sequence is a phoneme sequence that does not contain tones. In this way, the phonemes and tones in the first phoneme sequence are decoupled, and the generated second phoneme sequence includes the phonemes of each syllable corresponding to the text to be synthesized, and does not contain the tones corresponding to each syllable. Then, when the phoneme features of the text to be synthesized are extracted based on the second phoneme sequence, the spatial domain of the tones is isolated, so that the phoneme features of the text to be synthesized can be extracted more accurately, thereby improving the quality of the fusion features obtained when the phoneme features and semantic features are subsequently fused.
[0088] Among them, the phoneme feature is the feature of the phoneme, which is used to represent the pronunciation information. After the second phoneme sequence is obtained as above, the phonemes in the second phoneme sequence can be embedded (Embedding) to extract the phoneme features of the text to be synthesized. As an example, the second phoneme sequence can be processed using a model such as G2P to extract the phoneme features of the text to be synthesized.
[0089] In a possible implementation, a tone sequence can also be obtained based on the first phoneme sequence and the second phoneme sequence; the tone sequence includes the tone corresponding to the at least one phoneme; wherein the tone corresponding to any phoneme is determined by the tone of the syllable in which the phoneme is located.
[0090] Since the first phoneme sequence includes the phonemes of each syllable corresponding to the text to be synthesized and the tones corresponding to each syllable, and the second phoneme sequence includes the phonemes of each syllable, the first phoneme sequence and the second phoneme sequence include the same phonemes, and the difference between the two is whether they include tones. Therefore, based on the first phoneme sequence and the second phoneme sequence, the tones corresponding to each phoneme in the first phoneme sequence can be determined to form a tone sequence; thereby achieving the independence of the tones from the first phoneme sequence.
[0091] Exemplarily, for any syllable, the tone corresponding to the syllable can be used as the tone corresponding to each phoneme in the syllable, thereby ensuring that the dimension of the second phoneme sequence and the tone sequence are the same, that is, any phoneme in the second phoneme sequence corresponds to a tone in the tone sequence, so as to subsequently control the pronunciation of each phoneme. For example, the first phoneme sequence is "xian4zai4", the second phoneme sequence is "xianzai", and the tone sequence is "44444".
[0092] Step 104: perform text encoding on the text to be synthesized in the syllable dimension to extract semantic features of the text to be synthesized.
[0093] Among them, semantic features can provide contextual information of the text to be synthesized, and can reflect the relationship between the corresponding syllables in the text to be synthesized, the rhythm of pauses, etc., which are crucial to the prediction of phoneme duration. For example, for the text to be synthesized "I go to the station", by extracting semantic features, we can know that "station" as a whole has fewer pauses between "car" and "station", while the pause between "go" and "car" is relatively long, that is, the duration of the syllable "chē" corresponding to "car" and the syllable "zhàn" corresponding to "station" is shorter than the duration of the syllable "qù" corresponding to "go"; in this way, after extracting the semantic features of the text to be synthesized, the duration of each phoneme can be predicted based on the semantic features.
[0094] For example, the text to be synthesized can be encoded in the syllable dimension by a semantic feature extraction model to extract the semantic features of the text to be synthesized. For different languages, corresponding semantic feature extraction models can be pre-trained, and the trained semantic feature extraction models can learn the ability to capture the semantic features of the text of the language from the syllable dimension.
[0095] As an example, the semantic feature extraction model can be the BERT (Bidirectional Encoder Representations from Transformers) model. Among them, BERT is a pre-trained model of the Transformer architecture. It shows extremely high performance by pre-training on a large amount of unlabeled text and then fine-tuning on specific tasks. For example, the BGE-M3 (Bert Embedding with Multi-Linguality, Multi-Functionality, Multi-Granularity) model can be adopted. The BGE-M3 model is a model for text embedding representation based on BERT with multi-linguality, multi-functionality, and multi-granularity. Exemplarily, the output of the last layer of the BERT model can be used as the semantic features of the text to be synthesized, so as to capture the deep semantic information of the text to be synthesized.
[0096] In a possible implementation manner, the text encoding of the text to be synthesized in the syllable dimension and the extraction of the semantic features of the text to be synthesized include: performing word segmentation processing on the text to be synthesized according to syllables to obtain at least one syllable; combining a preset vocabulary, determining the encoding corresponding to each syllable in the at least one syllable; wherein, the preset vocabulary includes the corresponding relationship between different syllables and different encodings; based on the encoding corresponding to each syllable in the at least one syllable, extracting the semantic features of the text to be synthesized.
[0097] Among them, the preset vocabulary corresponds to the target language. The syllables in the vocabulary and the encodings corresponding to each syllable can be pre-configured according to the characteristics of the target language. Exemplarily, word segmentation can be performed based on the syllables in the vocabulary in a matching manner to sequentially determine the syllables corresponding to the text to be synthesized. Furthermore, the encodings corresponding to the syllables obtained after word segmentation processing can be used to query the corresponding encodings in the vocabulary to determine the encodings corresponding to each syllable, so as to realize text encoding of the text to be synthesized in the syllable dimension. Furthermore, the encodings corresponding to each syllable can be combined into an encoding sequence and input into the semantic feature extraction model, so that the semantic features of the text to be synthesized can be extracted. The encoding can be any encoding form such as Chinese character encoding, custom encoding, etc. In this way, performing word segmentation and encoding on the text to be synthesized in the syllable dimension can obtain more accurate and rich semantic features, thus providing a high-quality feature basis for the subsequent deep fusion of phoneme features and semantic features.
[0098] Step 105: Perform fusion processing on the semantic features and the phoneme features, and generate speech based on the fused features.
[0099] Exemplarily, a speech synthesis model can be adopted to fuse the semantic features and the phoneme features, and perform speech synthesis based on the fused features to generate speech. Speech synthesis models corresponding to different languages can be pre-trained. For the target language corresponding to the text to be synthesized, the trained speech synthesis model corresponding to this language is selected to fuse the semantic features and phoneme features of the text to be synthesized extracted above, and output the synthesized speech.
[0100] As an example, the speech synthesis model can be a model constructed based on the Variational Inference with adversarial learning for end-to-end Text-to-Speech (VITS). VITS is a speech synthesis architecture that combines the Variational Autoencoder (VAE) and the Generative Adversarial Network (GAN); among them, the variational autoencoder is responsible for efficiently encoding and decoding audio features, and can map audio features to a low-dimensional latent space. Furthermore, it can be restored to a high-quality speech signal through the decoder; in this way, not only the original characteristics of the sound are retained, but also by introducing randomness, it provides rich expressiveness for speech generation, making the synthesized speech more natural and diverse. In the training stage of the VITS model, the discriminator in the generative adversarial network can perform real-time quality evaluation and optimization on the generated speech signal; by distinguishing the difference between the speech generated by the decoder in the variational autoencoder and the real speech, feedback is provided to the decoder in the variational autoencoder, thereby continuously optimizing the ability of the decoder to generate speech. This process ensures that the synthesized speech meets the expected effect in terms of naturalness, prosody, and semantic consistency.
[0101] Figure 2 Fig. shows a schematic structural diagram of a VITS model according to an embodiment of the present disclosure, as Figure 2As shown, the trained voice synthesis model may include: an encoder (Text Encoder), a projection layer (Projection), a stochastic duration predictor (Stochastic Duration Predictor), a flow model (Flow), and a decoder (Decoder); through this VITS model, the refined phoneme features and semantic features can be fused and converted into high-quality speech; among them, the encoder is used to fuse the phoneme features and semantic features to generate the fused features (which can also be called intermediate features or audio features). For example, the phoneme features corresponding to each phoneme can be weighted and summed, concatenated, or multiplied with the semantic features corresponding to each phoneme to generate the fused features corresponding to each phoneme; the projection layer is used to map the fused features to the latent space; the stochastic duration predictor is used to predict the pronunciation duration of each phoneme corresponding to the text to be synthesized based on the fused features, map it to the latent space, and through the monotonic alignment model (Monotonic Alignment Search) in the latent space, align the fused features with the pronunciation duration of each phoneme; the flow model is used to generate latent hidden variables based on the above alignment results, and the decoder is used to decode the latent hidden variables to generate the speech corresponding to the text to be synthesized, that is, the synthesized speech. In this way, through the trained VITS model, high-quality speech that sounds natural and matches the text to be synthesized can be generated; among them, the training process of the voice synthesis model can be implemented through existing technologies, which is not limited herein.
[0102] Exemplarily, the decoder in VITS can be the vocoder of a high-fidelity GAN model (High-Fidelity GAN, HiFiGAN). HiFiGAN is a powerful neural network vocoder that can generate high-fidelity speech waveforms while maintaining a high generation speed. Exemplarily, the above semantic feature extraction model and phoneme conversion model can also be configured in the encoder of VITS. In this way, the text to be synthesized can be input into the encoder of VITS, and the encoder extracts the semantic features and phoneme features of the text to be synthesized and performs feature fusion.
[0103] Exemplarily, when fusing the semantic features and phoneme features of the text to be synthesized, the dimensions of the semantic features and phoneme features can be aligned first, and then the aligned semantic features and phoneme features are fused to generate the fused features. As an example, the dimension alignment of the semantic features and phoneme features can be performed by the Greedyrepeat method, which evenly distributes the semantic features to each phoneme according to the dimensions and the sorting of the phonemes in the second phoneme sequence. For the semantic features of the remaining dimensions, still according to the sorting of the phonemes, the semantic features of the remaining dimensions are preferentially assigned to the phonemes with a higher sorting. For example, if the dimension of the semantic features is 7 and the number of phonemes is 3, the dimension of the semantic features corresponding to each phoneme is 2, and then the semantic features of the remaining one dimension are assigned to the first phoneme. In this way, the accurate alignment of the phoneme features and semantic features of the text to be synthesized is achieved, the deep fusion of the semantic features and phoneme features is realized, and thus the pronunciation and prosody corresponding to each phoneme in the generated speech are more accurate.
[0104] In a possible implementation manner, the method further includes: during the process of generating speech based on the fused features, predicting the duration of each phoneme in the at least one phoneme based on a duration model; wherein, the duration model is trained based on the true alignment labels of phoneme-duration.
[0105] Among them, the duration model can be a DurationNet model, that is, a DurationNet model based on the Ground Truth. The true alignment labels of phoneme-duration are used to indicate the corresponding relationship between each phoneme in the training sentence and the true duration of each phoneme in the training sentence. Exemplarily, the random duration predictor in the above Figure 2 VITS model can be replaced with a DurationNet model based on GroundTruth, and during the training process, training sentences including phonemes and the true alignment labels of phoneme-duration are used for training.
[0106] In this way, the trained DurationNet model can better synthesize the prosody of speech. Among them, the training of the DurationNet model is carried out in an existing manner, which is not limited herein.
[0107] For example, taking Vietnamese as the target language, the DurationNet model may include: a bidirectional gated recurrent unit (GRU), a fully connected layer, a softplus activation function, etc. Compared with structures such as convolution, the GRU has significant advantages in processing sequential data, especially in capturing temporal dependencies. Through the bidirectional GRU, the forward and backward dependencies of the fused features can be captured, and at the same time, the forward and backward context information is utilized to generate high-dimensional temporal features; the fully connected layer performs a non-linear transformation on the high-dimensional temporal features output by the bidirectional GRU to complete the dimensional mapping from temporal features to durations; the softplus activation function can ensure that the predicted duration is a positive number. In this way, the features obtained by fusing semantic features and phoneme features are input into the DurationNet model, and the model can accurately predict the durations of each phoneme in the second phoneme sequence, thereby further improving the quality of speech synthesis.
[0108] Furthermore, considering that in scenarios such as dubbing for film and television dramas, there may be requirements for the duration of the generated speech. For example, for multiple consecutive lines of text of an actor, the duration of the generated speech needs to be consistent with the duration of the corresponding video frame of the line text to maintain lip-sync. However, due to differences in different languages, for the same line, the lengths of the corresponding synthesized speeches may vary. Therefore, operations such as axis merging and axis splitting need to be performed on the generated speech to make the lengths of the speeches in different languages corresponding to the same line the same; among them, axis merging means removing the pauses between adjacent clauses to merge them into one clause; axis splitting means adding pauses in a clause to split the clause into multiple clauses.
[0109] In a possible implementation manner, the method further includes: obtaining a preset duration corresponding to each of the multiple clauses; in the case where the duration of the generated speech exceeds the sum of the preset durations of the multiple clauses, performing axis merging processing on at least two adjacent clauses in the speech, and after the axis merging processing, adding one or more spaces between the at least two adjacent clauses to make the duration of the processed speech not greater than the sum of the preset durations of the multiple clauses.
[0110] Exemplarily, after the coaxial processing, one or more spaces can be added between at least two adjacent clauses in the way of blank coding, wherein the duration of each space can be set as required, and the duration of one or more spaces added between two adjacent clauses is less than the pause duration of these two adjacent clauses before coaxial processing; that is, relative to the pause duration between two adjacent clauses before coaxial processing, the pause between two adjacent clauses with one or more spaces added becomes shorter, so that while shortening the overall duration of two adjacent clauses through coaxial processing, the sense of pause between the two adjacent clauses participating in coaxial processing is strengthened, and it is ensured that the duration of the finally obtained processed speech is not greater than the sum of the preset durations of the multiple clauses.
[0111] As an example, after the coaxial processing, one or more spaces can be added between the at least two adjacent clauses to make the duration of the processed speech the same as the sum of the preset durations of the multiple clauses.
[0112] In this way, in the case where the duration of the generated synthetic speech exceeds the sum of the preset durations of the multiple clauses, coaxial processing is performed on at least two adjacent clauses in the speech, thereby shortening the duration of the synthetic speech and avoiding problems such as out-of-sync audio and video caused by the inconsistency between the duration of the synthetic speech and the preset duration in scenarios such as film and television drama dubbing. The preset duration can be determined by the duration of the corresponding pictures of the multiple clauses. In addition, considering that during the coaxial processing, the pauses between the clauses participating in coaxial processing are removed, resulting in a weakened sense of pause in the synthetic speech. Therefore, after the coaxial processing, one or more spaces are added between two adjacent clauses participating in coaxial processing, strengthening the sense of pause in speech synthesis, which can not only add a more distinct sense of rhythm to the film and television plot, but also strengthen the transmission of emotions and the presentation of scenes.
[0113] In a possible implementation manner, the method further includes: during the process of generating speech based on the fused features, controlling the pitch of the speech through the pitch sequence. Since the phoneme sequence includes: the pitches corresponding to the phonemes of each syllable of the text to be synthesized; in this way, controlling the pitch of the speech through the pitch sequence enables the pronunciation and prosody of each phoneme to be accurate enough. In addition, when generating the fused features, phoneme features are extracted from the second phoneme sequence after decoupling the phonemes and pitches in the first phoneme sequence, and then the phoneme features are fused with the semantic features of the text to be synthesized to obtain the fused features, realizing the isolation processing of the action space domain of the pitch; at the same time, controlling the pitch of the synthetic speech through the pitch sequence independently split from the phoneme sequence, thereby realizing the independent control of the pitch and phonemes and ensuring the quality of the synthesized speech.
[0114] Exemplarily, during the process of speech synthesis, the pitch is usually expressed by the fundamental frequency (F0) curve. For example, in Chinese, different pitches correspond to different F0 curves. In this way, through the pitch sequence, during the process of speech synthesis, the fundamental frequency of the audio segment corresponding to each phoneme can be determined, thereby realizing the control of the pitch of the generated speech. For example, when using the VITS model for speech synthesis, the fundamental frequency of the audio segment corresponding to each phoneme in the decoder can be controlled through the pitch sequence to achieve the control of the pitch of the synthesized speech.
[0115] In the embodiments of the present disclosure, the text to be synthesized is obtained; the text to be synthesized is subjected to phoneme conversion to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located; the phonemes and pitches in the first phoneme sequence are decoupled, and the phoneme features of the text to be synthesized are extracted based on the decoupled second phoneme sequence, where the second phoneme sequence includes the at least one phoneme; the text to be synthesized is text-encoded in the syllable dimension to extract the semantic features of the text to be synthesized; the semantic features and the phoneme features are fused, and speech is generated based on the fused features. In this way, the phonemes and pitches in the first phoneme sequence are decoupled, thereby isolating the action space domain of the pitch, and then the phoneme features of the text to be synthesized are extracted based on the decoupled second phoneme sequence, improving the accuracy of the extracted phoneme features; at the same time, text-encoding the text to be synthesized in the syllable dimension can obtain more accurate and rich semantic features; furthermore, the phoneme features extracted from the decoupled second phoneme sequence are fused with the semantic features, so that the quality of the fused features can be guaranteed. The speech synthesized based on the fused features not only has accurate pronunciation but also natural prosody, ensuring that the synthesized speech achieves the expected effects in terms of naturalness, prosody, and semantic consistency.
[0116] As an example, since the text to be synthesized is normalized and a phoneme correction dictionary is used during the process of speech synthesis, the content of the text to be synthesized can be standardized, thereby generating a standardized and correct phoneme sequence, which is beneficial for the speech synthesis model to generate a relatively accurate synthesized speech according to the standardized phoneme sequence, reducing the situations of missing words, omitting words, and pronunciation errors in the synthesized speech, and improving the accuracy of speech synthesis. In addition, for target languages such as minority languages, as determined by people whose mother tongue is the target language, the speech generated by the speech synthesis method in the embodiments of the present disclosure is indistinguishable from that of a real person in terms of pronunciation accuracy, prosody, etc.; it can be effectively applied to scenarios such as international film and television drama dubbing.
[0117] For example, the above speech synthesis method can be implemented through a trained speech synthesis model. Taking the speech synthesis model as an example of a VITS-based model, it can be in the above Figure 2Configure the trained BERT model and G2P model in the encoder of VITS as shown.
[0118] Figure 3 A schematic diagram of extracting semantic features and phoneme features according to an embodiment of the present disclosure is shown. As Figure 3 shown, the text to be synthesized can be input into the BERT model, and the BERT model can extract the semantic features of the text to be synthesized; the text to be synthesized is input into the G2P model, and the G2P model extracts the phoneme features of the text to be synthesized; furthermore, the semantic features and phoneme features are fused to obtain audio features, that is, the fused features.
[0119] Figure 4 A schematic flowchart of speech synthesis using a speech synthesis model according to an embodiment of the present disclosure is shown. As Figure 4 shown, first, the text to be synthesized is input into the VITS model configured with the BERT model and G2P model (corresponding to Figure 1 step 101 in), the G2P model in the VITS model performs phoneme conversion on the text to be synthesized to obtain a phoneme sequence (corresponding to Figure 1 step 102 in), and then the phonemes and tones in the phoneme sequence are decoupled, and phoneme features are extracted after decoupling (corresponding to Figure 1 step 103 in); the BERT model performs text encoding on the text to be synthesized to extract semantic features (corresponding to Figure 1 step 104 in); finally, the phoneme features and semantic features are fused and then speech synthesis is performed, so that high-quality, natural and fluent speech can be synthesized, achieving a speech synthesis effect with high expressiveness (corresponding to Figure 1 step 105 in).
[0120] Based on the same inventive concept of the above method embodiments, an embodiment of the present disclosure further provides a speech synthesis device, and this device can be used to execute the technical solutions described in the above method embodiments.
[0121] Figure 5 A structural diagram of a speech synthesis device according to an embodiment of the present disclosure is shown. As Figure 5As shown in the figure, the device includes: an acquisition module 501, configured to acquire a text to be synthesized; a phoneme conversion module 502, configured to perform phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located; a phoneme feature extraction module 503, configured to decouple the phonemes and pitches in the first phoneme sequence, and extract the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; where the second phoneme sequence includes the at least one phoneme; a semantic feature extraction module 504, configured to perform text encoding on the text to be synthesized in the syllable dimension to extract the semantic features of the text to be synthesized; and a speech synthesis module 505, configured to perform fusion processing on the semantic features and the phoneme features, and generate speech based on the fused features.
[0122] In an embodiment of the present disclosure, a text to be synthesized is acquired; phoneme conversion is performed on the text to be synthesized to obtain a first phoneme sequence, where the first phoneme sequence includes at least one phoneme and the pitch of the syllable where the at least one phoneme is located; the phonemes and pitches in the first phoneme sequence are decoupled, and the phoneme features of the text to be synthesized are extracted based on the decoupled second phoneme sequence; where the second phoneme sequence includes the at least one phoneme; text encoding is performed on the text to be synthesized in the syllable dimension to extract the semantic features of the text to be synthesized; fusion processing is performed on the semantic features and the phoneme features, and speech is generated based on the fused features. In this way, the phonemes and pitches in the first phoneme sequence are decoupled, so as to isolate the action space domain of the pitch, and then the phoneme features of the text to be synthesized are extracted based on the decoupled second phoneme sequence, improving the accuracy of the extracted phoneme features; at the same time, text encoding is performed on the text to be synthesized in the syllable dimension, and more accurate and rich semantic features can be obtained; furthermore, the phoneme features extracted from the decoupled second phoneme sequence are fused with the semantic features, so as to ensure the quality of the fused features, and make the finally synthesized speech achieve the expected effect in terms of naturalness, prosody, and semantic consistency.
[0123] In a possible implementation manner, the phoneme conversion module 502 is further configured to: filter the characters in the text to be synthesized that are not in the target language based on the target language corresponding to the text to be synthesized to obtain a filtered text; segment the filtered text according to the preset pause symbols corresponding to the target language to obtain a plurality of clauses; where the preset pause symbols are used to represent the pauses in the sentences; perform phoneme conversion processing on each of the plurality of clauses respectively to obtain a phoneme subsequence corresponding to each clause; and merge the phoneme subsequences corresponding to each clause to obtain the first phoneme sequence.
[0124] In a possible implementation, the speech synthesis module 505 is further configured to: obtain a pitch sequence based on the first phoneme sequence and the second phoneme sequence; the pitch sequence includes the pitches corresponding to the at least one phoneme; wherein, the pitch corresponding to any phoneme is determined by the pitch of the syllable where the phoneme is located; in the process of generating speech based on the fused features, control the pitch of the speech through the pitch sequence.
[0125] In a possible implementation, the semantic feature extraction module 504 is further configured to: perform word segmentation on the text to be synthesized according to syllables to obtain at least one syllable; combine a preset vocabulary to determine the encoding corresponding to each syllable in the at least one syllable; wherein, the preset vocabulary includes the corresponding relationship between different syllables and different encodings; extract the semantic features of the text to be synthesized based on the encoding corresponding to each syllable in the at least one syllable.
[0126] In a possible implementation, the phoneme conversion module 502 is further configured to: for any one of the multiple clauses, determine whether the clause contains a preset character; and perform normalization processing in the case of containing the preset character; perform phoneme conversion processing on the normalized clause to obtain the phoneme subsequence corresponding to the clause.
[0127] In a possible implementation, the phoneme conversion module 502 is further configured to: for any one of the multiple clauses, perform phoneme conversion processing using a phoneme conversion model to obtain the phoneme conversion result of the clause; determine the phoneme to be corrected in the phoneme conversion result of the clause; based on a preset phoneme correction dictionary, convert the phoneme to be corrected in the phoneme conversion result of the clause into a standard phoneme to obtain the phoneme subsequence corresponding to the clause; wherein, the phoneme correction dictionary includes: the corresponding relationship between different phonemes to be corrected and different standard phonemes.
[0128] In a possible implementation, the speech synthesis module 505 is further configured to: in the process of generating speech based on the fused features, predict the duration of each phoneme in the at least one phoneme based on a duration model; wherein, the duration model is trained based on the true alignment labels of phoneme-duration.
[0129] In a possible implementation, the speech synthesis module 505 is further configured to: obtain the preset duration corresponding to each of the multiple clauses; in the case that the duration of the generated speech exceeds the sum of the preset durations of the multiple clauses, perform coaxial processing on at least two adjacent clauses in the speech, and after the coaxial processing, add one or more spaces between the at least two adjacent clauses so that the duration of the processed speech is not greater than the sum of the preset durations of the multiple clauses.
[0130] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0131] The embodiments of the present disclosure also provide an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0132] The embodiments of the present disclosure also provide a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0133] The embodiments of the present disclosure also provide a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0134] Figure 6 FIG. 13 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 6 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0135] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0136] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions. The above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0137] A computer-readable storage medium can be a tangible device that can retain and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0138] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0139] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on a user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0140] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0141] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0142] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0144] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A speech synthesis method, characterized in that: The method comprises: Get the text to be synthesized; Performing phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, wherein the first phoneme sequence includes at least one phoneme and a tone of a syllable where the at least one phoneme is located; Decoupling the phonemes and tones in the first phoneme sequence, and extracting the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; wherein the second phoneme sequence includes the at least one phoneme; Performing text encoding on the text to be synthesized in the syllable dimension to extract semantic features of the text to be synthesized; The semantic features and the phoneme features are fused, and speech is generated based on the fused features.
2. The method according to claim 1, characterized in that The step of performing phoneme conversion on the text to be synthesized to obtain a first phoneme sequence includes: Based on the target language corresponding to the text to be synthesized, filtering characters of a non-target language in the text to be synthesized to obtain a filtered text; Segmenting the filtered text according to preset pause symbols corresponding to the target language to obtain a plurality of clauses; wherein the preset pause symbols are used to indicate a pause in a sentence; Performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause; The phoneme subsequences corresponding to each of the clauses are merged to obtain the first phoneme sequence.
3. The method according to claim 1, characterized in that The method further comprises: Based on the first phoneme sequence and the second phoneme sequence, a tone sequence is obtained; the tone sequence includes the tone corresponding to the at least one phoneme; wherein the tone corresponding to any phoneme is determined by the tone of the syllable in which the phoneme is located; In the process of generating speech based on the fused features, the pitch of the speech is controlled by the pitch sequence.
4. The method according to claim 1, characterized in that: The step of encoding the text to be synthesized in the syllable dimension to extract semantic features of the text to be synthesized includes: Performing word segmentation processing on the text to be synthesized according to syllables to obtain at least one syllable; Determine the code corresponding to each syllable in the at least one syllable in combination with a preset vocabulary; wherein the preset vocabulary includes a correspondence between different syllables and different codes; Based on the code corresponding to each syllable in the at least one syllable, the semantic features of the text to be synthesized are extracted.
5. The method according to claim 2, characterized in that: The performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes: For any clause among the multiple clauses, determining whether the clause contains a preset character; and performing normalization processing if the clause contains the preset character; The normalized clause is subjected to phoneme conversion to obtain a phoneme subsequence corresponding to the clause.
6. The method according to claim 2, characterized in that The performing phoneme conversion processing on each of the multiple clauses to obtain a phoneme subsequence corresponding to each clause includes: For any clause among the multiple clauses, a phoneme conversion model is used to perform phoneme conversion processing to obtain a phoneme conversion result of the clause; determining the phoneme to be corrected in the phoneme conversion result of the clause; Based on a preset phoneme correction dictionary, the phonemes to be corrected in the phoneme conversion result of the clause are converted into standard phonemes to obtain a phoneme subsequence corresponding to the clause; wherein the phoneme correction dictionary includes: corresponding relationships between different phonemes to be corrected and different standard phonemes.
7. The method according to claim 1, characterized in that The method further comprises: In the process of generating speech based on the fused features, the duration of each phoneme in the at least one phoneme is predicted based on a duration model; wherein the duration model is trained based on the real alignment label of the phoneme-duration.
8. The method according to claim 2, characterized in that: The method further comprises: Obtaining a preset duration corresponding to each of the multiple clauses; When the duration of the generated speech exceeds the sum of the preset durations of the multiple clauses, at least two adjacent clauses in the speech are subjected to axis-merging processing. After the axis-merging processing, one or more spaces are added between the at least two adjacent clauses so that the duration of the processed speech is not greater than the sum of the preset durations of the multiple clauses.
9. A speech synthesis device, characterized in that: The device comprises: An acquisition module, used to acquire the text to be synthesized; A phoneme conversion module, configured to perform phoneme conversion on the text to be synthesized to obtain a first phoneme sequence, wherein the first phoneme sequence includes at least one phoneme and a tone of a syllable where the at least one phoneme is located; A phoneme feature extraction module, used for decoupling the phonemes and tones in the first phoneme sequence, and extracting the phoneme features of the text to be synthesized based on the decoupled second phoneme sequence; wherein the second phoneme sequence includes the at least one phoneme; A semantic feature extraction module, used for performing text encoding on the text to be synthesized in the syllable dimension to extract the semantic features of the text to be synthesized; The speech synthesis module is used to fuse the semantic features and the phoneme features and generate speech based on the fused features.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
11. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product, comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Speech synthesis method and device, computer readable medium and electronic equipment
CN114495902A
Chinese and English cross-language speech synthesis method and device, electronic equipment and storage medium
CN114664282A
Audio processing method and related device
CN115862592A
Data processing method, and speech synthesis model training method and device
CN116129851A
Speech synthesis method and device, electronic equipment, storage medium and program product
CN116978353A
Cited By
Systems and methods for transposing spoken or textual input to music
US20240071343A1