Speech synthesis method, apparatus, device, storage medium and computer program product

By converting the text to be synthesized for speech into a standard pronunciation tree and adjusting the pronunciation information according to the user's pronunciation needs, the problem of difficulty in personalizing pronunciation in the prior art is solved, and an efficient and flexible pronunciation synthesis solution is achieved.

CN118471194BActive Publication Date: 2025-06-13MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410725897.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-06-13
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

The existing voice synthesis technology is difficult to adjust standard Mandarin pronunciation in real time according to the user's pronunciation needs, and cannot meet the needs of personalized voice synthesis.

Method used

By converting the synthesized text to be pronounced into a standard pronunciation tree composed of the pronunciation layer nodes, and adjusting the pronunciation information according to the user's pronunciation method needs, the target synthesized speech is generated.

Benefits of technology

It realizes real-time adjustment of voice synthesis according to the user's pronunciation needs, meeting the needs of personalized voice synthesis, which is simple and easy to operate, cheaper, and easy to deploy and update and maintain on demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118471194B_ABST
    Figure CN118471194B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech synthesis method, apparatus, device, storage medium and computer program product, relating to the technical fields of speech synthesis, tree structure and personalized speech. The method includes: obtaining a text to be speech-synthesized and a pronunciation manner; converting the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes, where the pronunciation layer nodes in the standard pronunciation tree record the standard pronunciation information of the text to be speech-synthesized; adjusting the standard pronunciation information based on the pronunciation manner to obtain a target pronunciation tree; and generating a target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree. By converting the text to be speech-synthesized into a pronunciation tree and adjusting the nodes recording pronunciation information in the pronunciation tree according to the requirements of the pronunciation manner to synthesize corresponding speech, it is simple and easy to implement, has lower costs, and is convenient for deployment on demand and update and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, specifically to technical fields such as speech synthesis, tree structures, personalized speech, etc., and particularly relates to a speech synthesis method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Speech synthesis technology (also known as text-to-speech technology, i.e., Text to Speech) is used to convert text information into speech information in real time and read or play the speech information in an appropriate manner, which is equivalent to installing an artificial mouth on a machine. The main problem to be solved is how to convert text information into audible sound information, that is, to make the machine speak like a human. Common applications include voice assistants, readers, or accessibility technologies, etc.

[0003] In some application scenarios, the speech synthesis requirement is not only limited to providing the standard Mandarin pronunciation of Chinese, but can also adjust the standard Mandarin pronunciation based on the pronunciation mode requirements proposed by the user to generate a synthesized speech with a specified pronunciation mode corresponding to the pronunciation mode requirement. Summary of the Invention

[0004] Embodiments of the present disclosure propose a speech synthesis method, apparatus, electronic device, computer-readable storage medium, and computer program product, providing a solution for converting a text to be speech-synthesized into a pronunciation tree and adjusting pronunciation information according to pronunciation mode requirements based on the information representation form of the pronunciation tree to complete corresponding speech synthesis, which is simple and easy to implement, has lower costs, and is convenient for deployment and update maintenance as needed.

[0005] In a first aspect, embodiments of the present disclosure propose a speech synthesis method, including: obtaining a text to be speech-synthesized and a pronunciation mode; converting the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes, where the pronunciation layer nodes in the standard pronunciation tree record the standard pronunciation information of the text to be speech-synthesized; adjusting the standard pronunciation information based on the pronunciation mode to obtain a target pronunciation tree; and generating a target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree.

[0006] In a second aspect, embodiments of the present disclosure propose a speech synthesis apparatus, including: an information acquisition unit configured to obtain a text to be speech-synthesized and a pronunciation mode; a standard pronunciation tree conversion unit configured to convert the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes, where the pronunciation layer nodes in the standard pronunciation tree record the standard pronunciation information of the text to be speech-synthesized; a pronunciation tree adjustment unit configured to adjust the standard pronunciation information based on the pronunciation mode to obtain a target pronunciation tree; and a speech synthesis unit configured to generate a target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree.

[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the speech synthesis method described in the first aspect.

[0008] In a fourth aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the speech synthesis method described in the first aspect when executed.

[0009] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, and when the computer program is executed by a processor, it can implement the steps of the speech synthesis method described in the first aspect.

[0010] In the speech synthesis solution provided by the present disclosure, by converting the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes for recording pronunciation information, the relevant information required for speech synthesis can be presented in the form of a tree, which is also convenient for adjusting the standard pronunciation information according to the specified pronunciation method to obtain the target pronunciation tree, and then synthesizing the corresponding synthesized speech. That is, by representing the relevant information required for speech synthesis in the form of a tree, it is convenient to directly adjust the pronunciation information recorded on the nodes of the pronunciation tree in combination with special pronunciation requirements, which is simple, has lower costs, and is convenient for on-demand deployment and update maintenance.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent:

[0013] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0014] Figure 2 is a flowchart of a speech synthesis method provided by an embodiment of the present disclosure;

[0015] Figure 3-1 is a flowchart of a method for forming a standard pronunciation tree only composed of character layer nodes and phoneme layer nodes provided by an embodiment of the present disclosure;

[0016] Figure 3-2is a schematic structural diagram of a standard pronunciation tree corresponding to Figure 3-1 ;

[0017] Figure 4-1 is a flowchart of a method for constructing a standard pronunciation tree composed of character-level nodes, word-level nodes, sentence-level nodes, and phoneme-level nodes provided by an embodiment of the present disclosure;

[0018] Figure 4-2 is a schematic structural diagram of a standard pronunciation tree corresponding to Figure 4-1 ;

[0019] Figure 5-1 is a flowchart of a method for constructing a standard pronunciation tree composed of character-level nodes, word-level nodes, sentence-level nodes, syllable-level nodes, and phoneme-level nodes provided by an embodiment of the present disclosure;

[0020] Figure 5-2 is a schematic structural diagram of a standard pronunciation tree corresponding to Figure 5-1 ;

[0021] Figure 6-1 is a flowchart of a method for adjusting the pronunciation information of a Mandarin pronunciation tree to obtain a dialect pronunciation tree provided by an embodiment of the present disclosure;

[0022] Figure 6-2 and Figure 6-3 are respectively schematic diagrams of the first replacement and the second replacement of the standard pronunciation tree shown in Figure 5-2 ;

[0023] Figure 7 is a schematic flowchart of a complete speech synthesis scheme combining specific application scenarios provided by an embodiment of the present disclosure;

[0024] Figure 8 is a structural block diagram of a speech synthesis device provided by an embodiment of the present disclosure;

[0025] Figure 9 is a schematic structural diagram of an electronic device suitable for executing a speech synthesis method provided by an embodiment of the present disclosure. Detailed Embodiments

[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0027] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0028] Figure 1 An exemplary system architecture 100 is shown that can apply the embodiments of the speech synthesis method, apparatus, electronic device, and computer-readable storage medium of the present disclosure.

[0029] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0030] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications for implementing information communication between the two can be installed on the terminal devices 101, 102, 103 and the server 105, such as speech synthesis applications, data transmission applications, instant messaging applications, etc.

[0031] The terminal devices 101, 102, 103 and the server 105 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.; when the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices, which can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and no specific limitation is made here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and no specific limitation is made here.

[0032] Server 105 can provide various services through various built-in applications. Taking the text-to-speech application that can provide text-to-speech services as an example, when Server 105 runs this text-to-speech application, the following effects can be achieved: First, it receives a text-to-speech request transmitted by the user through terminal devices 101, 102, and 103 via network 104, and then determines the text to be synthesized and the pronunciation method requirements based on this text-to-speech request; Next, it converts the text to be synthesized into a standard pronunciation tree composed of pronunciation layer nodes, and the pronunciation layer nodes in this standard pronunciation tree record the standard pronunciation information of this text to be synthesized; In the next step, it adjusts the standard pronunciation information based on this pronunciation method to obtain a target pronunciation tree; Finally, it generates a target synthesized voice based on the adjusted pronunciation information recorded in this target pronunciation tree.

[0033] Furthermore, Server 105 can also return the synthesized target synthesized voice back to terminal devices 101, 102, and 103 via network 104, so that terminal devices 101, 102, and 103 play the received audio data to complete the user's text-to-speech requirements.

[0034] It should be noted that, in addition to being obtained in real time from terminal devices 101, 102, and 103 via network 104, the text-to-speech request can also be pre-stored locally in Server 105 in various ways. Therefore, when Server 105 detects that these data have been stored locally (for example, when starting to process a pending text-to-speech task left before), it can choose to directly obtain these data from the local. In this case, the exemplary system architecture 100 may not include terminal devices 101, 102, and 103 and network 104.

[0035] Since speech synthesis requires a large amount of computing resources and strong computing power, the speech synthesis methods provided in subsequent embodiments of the present disclosure are generally executed by the server 105 with strong computing power and a large amount of computing resources. Correspondingly, the speech synthesis device is generally also set in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the above operations originally performed by the server 105 through the speech synthesis applications installed thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing capabilities at the same time, when the speech synthesis application determines that the terminal device where it is located has strong computing power and a large amount of remaining computing resources, the terminal device can be allowed to execute the above operations, thereby appropriately reducing the computing pressure on the server 105. Correspondingly, the speech synthesis device can also be set in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may not include the server 105 and the network 104.

[0036] It should be understood that Figure 1 the number of terminal devices, networks, and servers in

[0037] Please refer to Figure 2 , Figure 2 which is a flowchart of a speech synthesis method provided by an embodiment of the present disclosure. The process 200 includes the following steps:

[0038] Step 201: Obtain the text to be speech-synthesized and the pronunciation method;

[0039] The purpose of this step is for the execution entity of the speech synthesis method (such as Figure 1 the server 105 shown) to obtain the text to be speech-synthesized and the pronunciation method. Specifically, the text to be speech-synthesized and the pronunciation method included therein can be extracted by parsing the received speech synthesis request.

[0040] Among them, the text to be speech-synthesized is the content text information that serves as the basis for speech synthesis in the speech synthesis request, and it can be extracted from the specified position or specified identifier in the speech synthesis request. The pronunciation method is used to indicate the pronunciation method for speech synthesis. Usually, in the Chinese language environment, the default pronunciation method can be Mandarin pronunciation. For example, when the field content of the corresponding field in the speech synthesis request used to record or store the pronunciation method requirement is empty, it is understood that the user wants to use the default pronunciation method for speech synthesis; conversely, if the field content is not empty, the speech synthesis operation with the text to be speech-synthesized as the content is performed according to the pronunciation method corresponding to the recorded field content.

[0041] Other pronunciation methods different from the Mandarin pronunciation method usually refer to the pronunciation methods of a certain regional accent or the corresponding dialect pronunciation methods. Specifically, the pronunciation method requirement can include the target region information (such as province information like Guangdong, Shanxi, Zhejiang, etc.) used to indicate pronunciation in the target regional accent, or it can also be the dialect type information corresponding to the accent type (such as Cantonese, Hakka, etc.).

[0042] In addition, the pronunciation method can be determined in various ways. For example, when the speech synthesis request is an audio signal sent by the user in voice form, the pronunciation method can be a clear pronunciation method instruction spoken by the user (such as generating XXX in Shaanxi dialect), or when the user does not clearly state which pronunciation method instruction, the dialect or regional accent type recognized from the user's speech can be intelligently recognized as the pronunciation method, etc. There is no specific limitation here and it can be determined according to the actual situation.

[0043] Step 202: Convert the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes;

[0044] Based on Step 201, the purpose of this step is for the above-mentioned execution entity to convert the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes.

[0045] Specifically, the pronunciation layer node is used to record the pronunciation information of the text to be speech-synthesized. At the same time, the standard pronunciation tree also includes a text layer node and the text layer node is the parent node of the pronunciation layer node.

[0046] Among them, text layer nodes refer to nodes at different text levels into which the text to be speech synthesized is divided according to the text structure it contains, such as sentence layer nodes at the sentence level, word layer nodes (also called word layer nodes, vocabulary layer nodes) at the word level (also called phrase level, vocabulary level), and character layer nodes (also called Chinese character layer nodes) at the character level (also called Chinese character level). Furthermore, if the text to be speech synthesized is a paragraph containing multiple sentences or even an article containing multiple paragraphs, it can also be divided into segment-level nodes under the paragraph level and article-level nodes under the article level. At the same time, according to the relationship between the text structures, the article-level node is the parent node of the paragraph-level node (it can also be expressed as the article level is the upper level of the paragraph level), the paragraph-level node is the parent node of the sentence-level node (it can also be expressed as the paragraph level is the upper level of the sentence level), the sentence-level node is the parent node of the word-level node (it can also be expressed as the sentence level is the upper level of the word level), and the word-level node is the parent node of the character-level node (it can also be expressed as the word level is the upper level of the character level).

[0047] Among them, the pronunciation layer node is used to record the standard pronunciation information of the text at different text levels in the text to be synthesized for speech, which can specifically include syllable layer nodes and phoneme layer nodes, wherein phonemes include initial consonants and finals with tones, and the syllable of a Chinese character is composed of the phonemes of the Chinese character. Taking the Chinese character "我" as an example, its syllable information can be expressed as: "wo3", where "w" is the initial consonant and "o3" is the final with three tones. Taking the Chinese character "哦" as an example, its syllable information can be expressed as: "o4", which shows that it only contains the final with four tones "o4". Since a syllable is just an integration of the phonemes of the same Chinese character, a pronunciation layer node can only contain a phoneme layer node, or it can simultaneously contain a phoneme layer node and a syllable layer node that exists as the parent node of the phoneme layer node. It should be understood that the pronunciation level where the pronunciation layer node is located should be the lower level of the text level where the text level node is located (because the pronunciation level describes the pronunciation information of the corresponding level text recorded in the text level). In other words, the pronunciation layer node is the child node of the text layer node (it can also be expressed as the text layer node is the parent node of the pronunciation layer node).

[0048] Specifically, after determining the text layer nodes and pronunciation layer nodes to be included, they are constructed into a standard pronunciation tree in a tree diagram according to the hierarchical relationship. It should be understood that the standard pronunciation tree is usually a Mandarin pronunciation tree that follows the Mandarin pronunciation standard.

[0049] Step 203: adjusting the standard pronunciation information based on the pronunciation mode to obtain a target pronunciation tree;

[0050] On the basis of step 202, this step aims to adjust the standard pronunciation information (i.e., pronunciation information according to the pronunciation standard of standard Mandarin) recorded in the pronunciation layer node of the standard pronunciation tree based on the pronunciation method by the above-mentioned execution subject, that is, to adjust it to the pronunciation information corresponding to the pronunciation method, so as to determine the pronunciation tree after the pronunciation information is adjusted as the target pronunciation tree.

[0051] When the pronunciation method requirement is used to indicate pronunciation in a target regional accent, the target pronunciation tree may correspond to a dialect pronunciation tree corresponding to the target regional accent.

[0052] Step 204: Generate target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree.

[0053] On the basis of step 203, this step aims to generate the target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree by the above-mentioned execution subject.

[0054] An implementation method including but not limited to may be:

[0055] Considering that the pronunciation layer node is a child node of the text layer node, and the pronunciation layer node records the pronunciation information of each text content, the pronunciation layer node should be a leaf node constituting the target pronunciation tree (i.e., the bottom node of the tree), so that the pronunciation information recorded by each leaf node of the target pronunciation tree can be traversed in order to generate a target pronunciation sequence corresponding to the text to be synthesized, that is, the target pronunciation sequence is generated in the same order as the text content word order of the text to be synthesized. Then, the target synthesized speech can be generated and synthesized according to the pronunciation indicated by the target pronunciation sequence to produce the content indicated by the text to be synthesized according to the pronunciation method requirement.

[0056] Furthermore, after the execution subject obtains the target synthesized speech, the target synthesized speech can also be returned to the terminal that initiated the speech synthesis request (for example, Figure 1 The terminal devices 101, 102, 103 shown in the figure) are used to enable the terminal to respond to the speech synthesis request initiated by the user by playing the target synthesized speech.

[0057] The speech synthesis method provided by the embodiment of the present disclosure converts the text to be speech synthesized into a standard pronunciation tree composed of pronunciation layer nodes for recording pronunciation information, so that the relevant information required for speech synthesis can be presented in the form of a tree, and it is also convenient to adjust the standard pronunciation information according to the specified pronunciation method to obtain the target pronunciation tree, and then synthesize the corresponding synthesized speech. That is, by representing the relevant information required for speech synthesis in the form of a tree, it is convenient to directly adjust the pronunciation information recorded by the nodes on the pronunciation tree in combination with special pronunciation requirements, which is simple and easy, has lower cost, and is convenient for on-demand deployment and update maintenance.

[0058] In order to deepen the understanding of the standard pronunciation tree construction scheme mentioned in step 202 of process 200, the present disclosure will provide corresponding specific construction schemes through three parallel situations:

[0059] First, please refer to Figure 3-1 , Figure 3-1 A flowchart of a method for forming a standard pronunciation tree consisting of only character-level nodes and phoneme-level nodes is provided in an embodiment of the present disclosure. That is, the present embodiment is based on the situation that the text-level nodes only include character-level nodes and the pronunciation-level nodes only include phoneme-level nodes. The process 300 includes the following steps:

[0060] Step 301: taking a character-layer node represented by a text to be synthesized into speech containing a single character as a root node of a pronunciation tree;

[0061] That is, the text to be synthesized into speech at this time only contains a single Chinese character, so there is only a character level, and there is no word level, sentence level, or other higher levels. Therefore, this step is intended to require the above-mentioned execution entity to only use the character-level node represented by the single Chinese character as the root node of the pronunciation tree.

[0062] Step 302: The phoneme layer nodes represented by the phoneme information of a single word are used as child nodes of the root node to obtain a standard pronunciation tree.

[0063] On the basis of step 301, this step aims to obtain a standard pronunciation tree by having the above-mentioned execution subject use the phoneme layer node represented by the phoneme information corresponding to the text to be speech synthesized as a child node of the root node (it should be understood that the phoneme layer node at this time is not only a child node of the root node, but also a leaf node in the entire pronunciation tree).

[0064] Figure 3-2 is with Figure 3-1 The schematic diagram of the structure of the corresponding standard pronunciation tree, taking the Chinese character "累" as an example, its phoneme layer nodes are: initial consonant "l" and tone final "ei4".

[0065] Next, please refer to Figure 4-1 , Figure 4-1 A flowchart of a method for forming a standard pronunciation tree by character-level nodes, word-level nodes, sentence-level nodes and phoneme-level nodes provided in an embodiment of the present disclosure, that is, the present embodiment is based on the case where the text-level nodes include character-level nodes, word-level nodes and sentence-level nodes, and the pronunciation level includes phoneme-level nodes, and the process 400 includes the following steps:

[0066] Step 401: determining the sentence-level node represented by the text to be speech synthesized as the root node of the pronunciation tree;

[0067] and Figure 3-1For the case where the text to be synthesized by speech in the corresponding embodiment is only a single Chinese character, in this embodiment, the text to be synthesized by speech is a single sentence. This step aims to have the above-mentioned executing entity determine the sentence-level node served by the text to be synthesized by speech as the root node of the pronunciation tree.

[0068] Step 402: Determine the word-level nodes served by the respective word texts that make up the text to be synthesized by speech as the first intermediate nodes that are the children nodes of the root node;

[0069] Based on Step 401, this step aims to have the above-mentioned executing entity determine the word-level nodes served by the respective word texts that make up the text to be synthesized by speech as the first intermediate nodes that are the children nodes of the root node, that is, the word-level nodes are the children nodes of the root node and at the same time the first intermediate nodes of the entire pronunciation tree.

[0070] Step 403: Determine the character-level nodes served by the respective character texts that make up the respective word texts as the second intermediate nodes that are the children nodes of the first intermediate nodes;

[0071] Based on Step 402, this step aims to determine the character-level nodes served by the respective character texts that make up the respective word texts as the second intermediate nodes that are the children nodes of the first intermediate nodes, that is, the character-level nodes are the children nodes of the word-level nodes (i.e., the first intermediate nodes) and at the same time the second intermediate nodes of the entire pronunciation tree.

[0072] Step 404: Determine the phoneme-level nodes served by the phoneme information corresponding to the respective character texts as the children nodes of the second intermediate nodes to obtain the standard pronunciation tree.

[0073] Based on Step 403, this step aims to have the above-mentioned executing entity determine the phoneme-level nodes served by the phoneme information corresponding to the respective character texts as the children nodes of the second intermediate nodes (it should be understood that in addition to being the children nodes of the character-level nodes, the phoneme-level nodes are at the same time the leaf nodes of the entire pronunciation tree), and finally construct the standard pronunciation tree.

[0074] Figure 4-2 For Figure 4-1 is a schematic structural diagram of the corresponding standard pronunciation tree, and specifically taking the sentence "I love Beijing" as an example, its sentence-level node is: "I love Beijing", the word-level nodes are respectively: "I", "love", and "Beijing", the character-level nodes are respectively: "I", "love", "north", and "Beijing", and the phoneme-level nodes are respectively: "w", "o3", "ai4", "b", "ei3", "j", and "ing1".

[0075] Finally, please refer to Figure 5-1 , Figure 5-1A flowchart of a method for constructing a standard pronunciation tree composed of character layer nodes, word layer nodes, sentence layer nodes, syllable layer nodes, and phoneme layer nodes provided by an embodiment of the present disclosure. That is, in this embodiment, on the premise that the text layer nodes include character layer nodes, word layer nodes, and sentence layer nodes, and the pronunciation layer nodes include syllable layer nodes and phoneme layer nodes, the process 500 includes the following steps:

[0076] Step 501: Determine the sentence layer node served by the text to be speech synthesized as the root node of the pronunciation tree;

[0077] Step 502: Determine the word layer nodes served by the respective word texts constituting the text to be speech synthesized as the first intermediate nodes that are the children nodes of the root node;

[0078] Step 503: Determine the character layer nodes served by the respective character texts constituting the respective word texts as the second intermediate nodes that are the children nodes of the first intermediate nodes;

[0079] Steps 501 - 503 are the same as steps 401 - 403 in process 400, and will not be repeated here. For the same parts, please refer to the detailed description at the corresponding positions in the above embodiments.

[0080] Step 504: Determine the syllable layer nodes served by the syllable information corresponding to the respective character texts as the third intermediate nodes that are the children nodes of the second intermediate nodes;

[0081] Based on step 503, this step aims to have the above-mentioned execution entity determine the syllable layer nodes served by the syllable information corresponding to the respective character texts as the third intermediate nodes that are the children nodes of the second intermediate nodes. That is, the syllable layer nodes are the children nodes of the character layer nodes (i.e., the second intermediate nodes), and at the same time are the third intermediate nodes of the entire pronunciation tree.

[0082] Step 505: Determine the phoneme layer nodes served by the phoneme information constituting the respective syllable information as the children nodes of the third intermediate nodes to obtain a standard pronunciation tree.

[0083] Based on step 504, this step aims to have the above-mentioned execution entity determine the phoneme layer nodes served by the phoneme information constituting the respective syllable information as the children nodes of the third intermediate nodes (it should be understood that the phoneme layer nodes, in addition to being the children nodes of the syllable layer nodes, are also the leaf nodes of the entire pronunciation tree), and finally construct a standard pronunciation tree.

[0084] Figure 5-2 For Figure 5-1 the corresponding structural schematic diagram of the standard pronunciation tree, and Figure 4-2 taking the same text to be speech synthesized - "I love Beijing" as an example, Figure 4-2 the difference is that, compared with Figure 4-1In the pronunciation layer nodes of the illustrated embodiments, syllable layer nodes are added. Therefore, Figure 4-2 a layer of syllable layer nodes is further added between the character layer nodes and the phoneme layer nodes of Figure 4-2 , that is, the syllable layer nodes are the parent nodes of the phoneme layer nodes. Specifically, the syllable layer nodes include: "wo3", "ai4", "bei3", and "jing4".

[0085] Based on any of the above embodiments, to deepen the understanding of how to adjust the common pronunciation information according to the dialect pronunciation method, in this embodiment, corresponding dialect pronunciation libraries are respectively created in advance for different regions (for example, a Shanxi dialect pronunciation library is established for Shanxi, a Hunan dialect pronunciation library is established for Hunan, a Dongguan dialect pronunciation library is established for Dongguan, etc.), so as to use the corresponding relationship between the common pronunciation information and the dialect pronunciation information recorded in the dialect pronunciation library to replace the common pronunciation information recorded in the common pronunciation tree with the dialect pronunciation information determined based on the corresponding relationship. Specifically, the dialect pronunciation library may include: the first corresponding relationship between texts at different text levels (such as the character level, word level, sentence level, etc. mentioned in the above embodiments) and the pronunciation information under the corresponding regional accent, and the second corresponding relationship between the different pronunciation information of the common language and the corresponding regional accent for expressing the same semantics, so as to use the first corresponding relationship and the second corresponding relationship to complete the comprehensive and accurate replacement of the common pronunciation information recorded in the common pronunciation tree.

[0086] Considering that the text level may include the sentence level, word level, and character level, then the first corresponding relationship should actually also correspondingly include: the sentence pronunciation relationship between the sentence text and the pronunciation phonemes under the corresponding regional accent (taking "I love Beijing" as an example, the sentence pronunciation relationship can be: I love Beijing - w,o3 ai4b,ei3 j,ing1), the word pronunciation relationship between the word text and the pronunciation phonemes under the corresponding regional accent (taking "Beijing" as an example, the word pronunciation relationship can be: Beijing - b,ei3 j,ing1), and the character pronunciation relationship between the character text and the pronunciation phonemes under the corresponding regional accent (taking "love" as an example, the character pronunciation relationship can be: love - ai4). Similarly, since the pronunciation layer nodes include syllable layer nodes and phoneme layer nodes, the second corresponding relationship should also correspondingly include: the syllable-phoneme corresponding relationship between the pronunciation syllables of the same character in the common language and the pronunciation phonemes of the character under the corresponding regional accent (for example, the syllable corresponding relationship between the common language and a certain dialect for the same Chinese character "zhong": zhong1–zh,ong2), and the phoneme-phoneme corresponding relationship between the different pronunciation phonemes of the same semantics in the common language and the corresponding regional accent (for example, the corresponding relationship between initials: zh–z, and the corresponding relationship between vowels with tones: eng1–en1).

[0087] In addition, in order to ensure that the dialect pronunciation information finally presented in the dialect pronunciation tree obtained by adjusting the Mandarin pronunciation information according to the dialect pronunciation method is accurate enough, it is also necessary to determine an appropriate replacement order for the information recorded at each level of the pronunciation tree and the hierarchical order between each level.

[0088] In the case where the target dialect pronunciation method is determined according to the target regional accent corresponding to the target regional information, a replacement order that includes but is not limited to can be:

[0089] For each layer of nodes constituting the Mandarin pronunciation tree, determine the target dialect pronunciation information corresponding to the current layer of nodes in the bottom-up hierarchical order of the nodes; adjust the Mandarin pronunciation tree according to the target dialect pronunciation information to obtain the dialect pronunciation tree.

[0090] Specifically, a preset processing operation as follows can be performed on each layer of nodes constituting the Mandarin pronunciation tree in the bottom-up hierarchical processing order to obtain the target pronunciation tree for which all layers of nodes have been processed:

[0091] First, determine whether the Mandarin pronunciation information recorded by the current layer of nodes or the text to be speech-synthesized matches the target dialect pronunciation information in the target dialect pronunciation method;

[0092] In response to matching the target dialect pronunciation information and the current layer of nodes being the pronunciation layer nodes recording the Mandarin pronunciation information, replace the Mandarin pronunciation information recorded by the pronunciation layer nodes with the target dialect pronunciation information;

[0093] In response to matching the target dialect pronunciation information and the current layer of nodes being the text layer nodes recording part or all of the text to be speech-synthesized, replace the Mandarin pronunciation information of the text recorded by the pronunciation layer nodes with the target dialect pronunciation information.

[0094] The reason why the replacement order is in the bottom-up hierarchical order of the nodes (that is, starting from the bottom layer and gradually going up layer by layer until the top layer) is that considering that there may be corresponding relationships for single phonemes, single syllables, single characters, as well as words containing multiple characters and sentences containing multiple words, but the pronunciation information in the corresponding relationships when they exist alone and when they exist as part of a word or a sentence is not necessarily the same. Therefore, in combination with the actual situation, replacement needs to be carried out in the replacement order from the bottom layer to the top layer to finally obtain the dialect pronunciation tree recording the most accurate pronunciation information.

[0095] In the above two embodiments, a solution for replacing Mandarin pronunciation information by pre - constructing a dialect pronunciation library is provided, and a solution for determining the target dialect pronunciation information corresponding to the nodes of each layer of the pronunciation tree in the bottom - up hierarchical order of the nodes and then implementing the replacement is provided. In another embodiment, these two parts can also be combined to obtain the following embodiment:

[0096] When the pronunciation mode requirement is specifically the target region information used to indicate pronunciation in the target regional accent, first, the above - mentioned execution entity determines the target dialect pronunciation library corresponding to the target region information. The target dialect pronunciation library records the first correspondence relationship between texts at different text levels and pronunciation information in the target regional accent, and the second correspondence relationship between the different pronunciation information used by Mandarin and the target regional accent to express the same semantics;

[0097] Next, according to the second correspondence relationship, determine the first dialect pronunciation information corresponding to the Mandarin pronunciation information in the pronunciation layer nodes of the Mandarin pronunciation tree, and replace the Mandarin pronunciation information with the corresponding first dialect pronunciation information to obtain an adjusted pronunciation tree. That is, this step uses the second correspondence relationship to complete the replacement of the pronunciation information recorded in the pronunciation layer nodes.

[0098] Then, according to the first correspondence relationship, determine the second dialect pronunciation information corresponding to the actual text recorded in the text layer nodes of the adjusted pronunciation tree, and replace the current pronunciation information corresponding to the actual text recorded in the adjusted pronunciation tree with the second dialect pronunciation information to obtain a dialect pronunciation tree. That is, this step uses the first correspondence relationship to complete the replacement of the pronunciation corresponding to the text recorded in the text layer nodes.

[0099] That is, this implementation method first replaces the pronunciation information of the pronunciation layer nodes at the bottom layer, and then combines the correspondence relationship between the text and the pronunciation information to replace the pronunciation of the text recorded in the text nodes with the corresponding dialect pronunciation information, so as to finally obtain an accurate dialect pronunciation tree.

[0100] Combined with the multiple correspondence relationships subdivided above, please continue to refer to Figure 6-1 , Figure 6-1 which is a flowchart of a method for adjusting the pronunciation information of a Mandarin pronunciation tree to obtain a dialect pronunciation tree provided by an embodiment of the present disclosure. To fully combine and use the multiple correspondence relationships subdivided above, its process 600 includes the following steps:

[0101] Step 601: Use the phoneme - phoneme correspondence relationship to determine the first dialect phoneme information corresponding to the Mandarin phoneme information recorded in the phoneme layer nodes of the Mandarin pronunciation tree, and replace the current phoneme information recorded in the phoneme layer nodes with the first dialect phoneme information to obtain a first adjusted pronunciation tree;

[0102] This step aims to replace the Mandarin phoneme information recorded by the phoneme layer nodes of the Mandarin pronunciation tree with the corresponding first dialect phoneme information determined through the phoneme-factor correspondence relationship (i.e., corresponding to the Mandarin phoneme information), and determine the current pronunciation tree with the phoneme information in the phoneme layer nodes replaced as the first adjusted pronunciation tree.

[0103] Step 602: Use the syllable-phoneme correspondence relationship to determine the second dialect phoneme information corresponding to the Mandarin syllable information recorded by the syllable layer nodes in the Mandarin pronunciation tree, and replace the current phoneme information recorded by the phoneme layer nodes with the second dialect phoneme information to obtain the pronunciation tree being adjusted;

[0104] Based on Step 601, this step aims to use the syllable-phoneme correspondence relationship by the above-mentioned execution entity to determine the second dialect phoneme information corresponding to the Mandarin syllable information recorded by the syllable layer nodes in the Mandarin pronunciation tree, and replace the current phoneme information recorded by the phoneme layer nodes with the second dialect phoneme information, thereby determining the pronunciation tree after completing this step as the pronunciation tree being adjusted.

[0105] If the phoneme information is indeed adjusted using the syllable-phoneme correspondence relationship, since the above steps only adjust the phoneme information of the phoneme layer nodes and do not adjust the syllable information of the syllable layer nodes, which are the parent nodes of the phoneme layer nodes, for the purpose of making the pronunciation information recorded by the entire pronunciation tree correct, the syllable information currently recorded by the syllable layer nodes in the pronunciation tree being adjusted can also be corrected according to the phoneme information currently recorded by the phoneme layer nodes in the pronunciation tree being adjusted.

[0106] Step 603: Use the character pronunciation relationship to determine the third dialect phoneme information corresponding to the character text recorded by the character layer nodes in the pronunciation tree being adjusted, and replace the current phoneme information recorded by the phoneme layer nodes in the pronunciation tree being adjusted with the third dialect phoneme information to obtain the second adjusted pronunciation tree;

[0107] Based on Step 602, this step aims to use the character pronunciation relationship by the above-mentioned execution entity to determine the third dialect phoneme information corresponding to the character text recorded by the character layer nodes in the pronunciation tree being adjusted, and replace the current phoneme information recorded by the corresponding phoneme layer nodes in the pronunciation tree being adjusted with the third dialect phoneme information to obtain the second adjusted pronunciation tree.

[0108] Similarly, if the pronunciation information of a certain phoneme layer node adjusted in Step 602 is adjusted again for the pronunciation information of the same phoneme layer node due to the character pronunciation relationship in this step, then it is equivalent to the pronunciation information replaced in Step 602 being further replaced by the pronunciation information determined in this step.

[0109] Step 604: Determine the fourth dialect phoneme information corresponding to the word text recorded in the word layer node of the second adjusted pronunciation tree based on the word pronunciation relationship, and replace the current phoneme information recorded in the phoneme layer node of the second adjusted pronunciation tree with the fourth dialect phoneme information to obtain the third adjusted pronunciation tree;

[0110] Based on step 603, this step aims to enable the above-mentioned execution entity to determine the fourth dialect phoneme information corresponding to the word text recorded in the word layer node of the second adjusted pronunciation tree based on the word pronunciation relationship, and replace the current phoneme information recorded in the corresponding phoneme layer node of the second adjusted pronunciation tree with this fourth dialect phoneme information to obtain the third adjusted pronunciation tree.

[0111] Similarly, if the pronunciation information of a certain phoneme layer node adjusted in step 603 is further adjusted in this step due to the word pronunciation relationship for the pronunciation information of the same phoneme layer node, it is equivalent to the pronunciation information replaced in step 603 being further replaced by the pronunciation information determined in this step.

[0112] Step 605: Determine the fifth dialect phoneme information corresponding to the sentence text recorded in the sentence layer node of the third adjusted pronunciation tree based on the sentence pronunciation relationship, and replace the current phoneme information recorded in the phoneme layer node of the third adjusted pronunciation tree with the fifth dialect syllable information to obtain the dialect pronunciation tree.

[0113] Based on step 604, this step aims to enable the above-mentioned execution entity to determine the fifth dialect phoneme information corresponding to the sentence text recorded in the sentence layer node of the third adjusted pronunciation tree based on the sentence pronunciation relationship, and replace the current phoneme information recorded in the corresponding phoneme layer node of the third adjusted pronunciation tree with this fifth dialect phoneme information to obtain the final dialect pronunciation tree.

[0114] If the phoneme information is indeed adjusted using the character pronunciation relationship, word pronunciation relationship, or sentence pronunciation information, but since the above steps only adjust the phoneme information of the phoneme layer node and do not adjust the syllable information of the syllable layer node, which is the parent node of the phoneme layer node, it is also possible to correct the syllable information currently recorded in the syllable layer node of the dialect pronunciation tree based on the phoneme information currently recorded in the phoneme layer node of the dialect pronunciation tree for the purpose of making all the pronunciation information recorded in the entire pronunciation tree correct.

[0115] Similarly, if the pronunciation information of a certain phoneme layer node adjusted in step 604 is further adjusted in this step due to the sentence pronunciation relationship for the pronunciation information of the same phoneme layer node, it is equivalent to the pronunciation information replaced in step 604 being further replaced by the pronunciation information determined in this step.

[0116] Figure 6-2 Combined with the pronunciation requirements of a certain regional accent, it shows Figure 5-2The standard pronunciation tree shown is the first pronunciation information replacement process for the phoneme information corresponding to the character-level node, that is, according to the character pronunciation relationship recorded in the dialect pronunciation library: "北-b, ei1", the "b, ei3" recorded in the phoneme-level node under "北" in the standard pronunciation tree is replaced with "b, ei1".

[0117] Figure 6-3 Based on 6-2, it shows Figure 6-2 The second pronunciation information replacement process of the phoneme information corresponding to the word-level node in the current pronunciation tree is presented, that is, according to the word pronunciation relationship recorded in the same dialect pronunciation library: "Beijing-b,ei4 j,ing2", the "b,ei1" recorded in the phoneme-level node under "Beijing" in the current pronunciation tree is replaced with "b,ei4", and the "j,ing1" is replaced with "j,ing2", so as to finally obtain the dialect pronunciation tree.

[0118] It should be noted that Figure 6-1 In the illustrated embodiment, steps 601-602 actually provide a specific implementation method for how to use the second correspondence to complete the replacement of the corresponding pronunciation information in the standard pronunciation tree to obtain the pronunciation tree being adjusted, in combination with the phoneme-phoneme correspondence and syllable-phoneme correspondence specifically included in the second correspondence. Steps 603-605 actually provide a specific implementation method for how to use the first correspondence to complete the replacement of the pronunciation information in the pronunciation tree being adjusted to obtain the target pronunciation tree, in combination with the sentence pronunciation relationship, word pronunciation relationship and character pronunciation relationship specifically included in the first correspondence.

[0119] However, there is no causal or dependency relationship between the specific implementation methods provided by steps 601-602 and the specific implementation methods provided by steps 603-605. They can be combined with the above embodiments to form independent embodiments. This embodiment only exists as a preferred embodiment that includes the above two preferred implementation methods.

[0120] Further, the above-mentioned corresponding relationships can be specifically identified in various forms including key-value pairs, lists or arrays, sets or functions, etc. Taking the key-value pair form as an example, the syllable-phoneme corresponding relationship includes: syllable-phoneme key-value pairs with the pronunciation syllables of the same character in Mandarin as the key and the pronunciation phonemes in the corresponding regional accent as the value; the phoneme corresponding relationship includes phoneme-phoneme key-value pairs with the pronunciation phonemes of the same semantics in Mandarin as the key and the pronunciation phonemes in the corresponding regional accent as the value. The sentence pronunciation relationship can be expressed as: sentence pronunciation key-value pairs with the sentence text as the key and the pronunciation phonemes in the corresponding regional accent as the value; the word pronunciation relationship includes word pronunciation key-value pairs with the word text as the key and the pronunciation factors in the corresponding regional accent as the value; the character pronunciation relationship includes character pronunciation key-value pairs with the character text as the key and the pronunciation factors in the corresponding regional accent as the value. The remaining forms of the corresponding relationship will not be listed one by one.

[0121] It should be further noted that the sentence pronunciation corresponding relationship, word pronunciation corresponding relationship, and character pronunciation corresponding relationship mentioned in the above embodiments are all corresponding relationships between the corresponding hierarchical text and phoneme information. Therefore, when it is indeed necessary to perform pronunciation information replacement by combining the corresponding relationships in steps 603 - 605, the replacement is directly performed on the phoneme information recorded in the phoneme layer nodes. However, replacing only the phoneme information recorded in the phoneme layer nodes will cause the syllable information recorded in the syllable layer nodes of its parent node to be mismatched. Considering that only the leaf nodes of the dialect pronunciation tree need to be traversed to generate the target synthetic speech subsequently, the syllable information recorded in the syllable layer nodes can also be corrected without ensuring that all the information in the dialect pronunciation tree is correct. In some other embodiments, if the sentence pronunciation corresponding relationship, word pronunciation corresponding relationship, and character pronunciation corresponding relationship are all corresponding relationships between the corresponding hierarchical text and syllable information, then the replacement object each time is the syllable information of the syllable layer nodes. To ensure that the correct target pronunciation sequence can be obtained by traversing each leaf node subsequently for synthesizing the target synthetic speech, it is also necessary to adaptively update the lower-level phoneme layer nodes according to the replaced syllable layer nodes.

[0122] Which specific scheme to choose can be flexibly determined according to the actual situation and will not be specifically limited here.

[0123] Based on any of the above embodiments, if the number of pronunciation methods is greater than 1, the target text corresponding to each pronunciation method can also be determined in the text to be speech-synthesized, and then the standard pronunciation information corresponding to different target texts can be adjusted according to the corresponding pronunciation methods to obtain a target pronunciation tree containing pronunciation information under multiple pronunciation methods, so as to more flexibly meet the special speech synthesis requirements of users.

[0124] To deepen understanding, the present disclosure also provides a specific implementation solution in combination with a specific application scenario, see Figure 7 The complete process diagram shown is:

[0125] S1: Define the Mandarin Chinese language system, that is, define the words, Chinese characters and pinyin combinations corresponding to the Mandarin Chinese speech synthesis system. Each target speech synthesis system has an independent pinyin system. Based on the design conventions of conventional Chinese speech synthesis systems, each pronunciation unit is a Chinese character, each Chinese character corresponds to an independent syllable, and each syllable consists of an initial consonant and a final with a tone, for example: 中 (syllable zhong1: initial consonant zh, final with a tone ong1) 国 (syllable guo2: initial consonant g, final with a tone uo2).

[0126] Therefore, the combined data obtained in the current step is: 1. N Chinese word combinations, 2. M Chinese character combinations, 3. K syllables, 4. I initial consonants and J finals with tones. Among them, all Chinese characters in the N Chinese words appear in the M Chinese character combinations, each Chinese character can be pronounced in the K syllables, and each syllable is composed of one or more of I+J initial consonants or finals with tones.

[0127] It should be noted that the phonetic system of the Chinese characters selected in this step can be a standard Chinese phonetic system (as described above) or a non-standard International Phonetic Alphabet system.

[0128] S2: Initialize the regional accent pronunciation library, that is, based on the standard Mandarin speech synthesis system determined in S1, select the target regional accent A (A can be any regional accent with Chinese characters as the main pronunciation unit), and initialize the regional accent pronunciation library. The regional accent pronunciation library consists of 5 levels, namely sentences, words, Chinese characters, syllables, and phonemes. In the initial state, the five levels of information of the "dialect pronunciation library" are empty.

[0129] It should be noted that the regional accent pronunciation library A initialized in this step can be a data structure in a specific program temporarily stored in the program memory before completing S3, as a memory object of an independent pronunciation library, wherein the pronunciation information of all 5 levels is stored in a data structure of a key-value pair, based on a hash table (or programming languages ​​such as Java, C++) or a dictionary (python programming language) in the application program, the selection of specific programming language and specific data structure is not limited here. At the same time, the initialized pronunciation library A can also be based on a configuration file (yaml file or json file) mode, stored on a hard disk, usually need to be selected according to the design of the application program, which is not limited here.

[0130] S3: Manually prepare a regional accent pronunciation library, that is, add specific sentences, words, Chinese characters, syllables, and phonemes corresponding to the pronunciation library according to the pronunciation habits of dialect A. The final regional accent pronunciation library has a fixed data structure and key value type. Each key is a string, representing a specific unit of the level of an independent regional accent pronunciation library. Each key contains only one value, which is also a string, indicating the pronunciation method in the regional accent pronunciation library. For the sentence level, the pronunciation value is a syllable sequence separated by spaces (such as: I love Beijing Tiananmen-wo3 ai4 bei3 jing1 tian1 an1 men2), for the vocabulary level, the pronunciation value is a syllable sequence separated by spaces (such as: Beijing-bei3 jing1), and for the Chinese character level, the pronunciation value is an independent syllable (such as love-ai4). The key value of the syllable is composed of the pinyin syllable itself (such as zhong1-zhong2). The phonemes include two types: initials and finals with tones. Their key values ​​are all the mapping relationship when there are phonemes of the same type, representing the pronunciation of a specific accent (such as: zh-z, eng1-en1). After the regional accent pronunciation library is added, for regional accent A, there will be specific pronunciation key value mappings at different levels relative to the standard Mandarin pronunciation.

[0131] It should be noted that this step requires the pronunciation mapping rules of specific regional accents to be sorted out based on linguistic information. In practice, it can be sorted out manually by linguists or personnel in the region, or it can be sorted out automatically using pre-trained related models or tools. Taking manual sorting as an example, the pronunciation library object temporarily stored in the memory of S2 can be updated in the visual operation window provided by the application, or the relevant mapping key-value pairs can be manually added to the corresponding configuration file, which is not limited here.

[0132] S4: Import the speech synthesis system, that is, import the prepared regional accent pronunciation library corresponding to regional accent A into the front-end module of the speech synthesis system, which supports the stylized pronunciation capability of regional accent A, and then restart the speech synthesis service (of course, if the system supports hot start, the call to the new regional accent pronunciation library can also be completed without restarting the service).

[0133] It should be noted that in this step, the speech synthesis system service needs to read the regional accent pronunciation library prepared in S2 and S3, and store it in the memory of the speech synthesis system service for subsequent use.

[0134] S5: The service receives a request, that is, the speech synthesis service receives a specific external synthesis text request, which includes 1. a specific text sentence T that needs to be synthesized, and 2. a pronunciation that needs to be performed with a regional accent A.

[0135] S6: The front-end module constructs a standard Mandarin pronunciation tree, that is, through the front-end text analysis module of the speech synthesis service, the sentence T is organized into a specific pronunciation tree. The pronunciation tree consists of five levels (the same five pronunciation levels as the regional accent pronunciation library mentioned in S2: sentence, vocabulary, Chinese character, syllable, phoneme). Level 1 is the sentence T to be synthesized in the request (such as: I love Beijing Tiananmen), level 2 is the vocabulary contained in sentence T (such as: [I, love, Beijing, Tiananmen]), level 3 is the Chinese characters of the words in level 2 (such as: I), Level 4 is the Chinese syllable corresponding to the Chinese character at level 3 (e.g., wo3), and level 5 is the phoneme corresponding to level 4 (phonemes include initials and rhythms with tones, such as wo3–[w,o3]). The pronunciation tree only stores the last available pronunciation information at the phoneme node. If you need to get the last pronunciation information based on a pronunciation tree, you need to traverse the bottom-level (phoneme) nodes of the tree in order (e.g., [w,o3,ai4,b,ei3,j,ing1,t,ian1,an1,m,eng2]). Note that in the pronunciation tree initialized at the current step, the pronunciation information of Mandarin Chinese is stored.

[0136] S7: Pronunciation replacement based on regional accent A, that is, the service front-end module finds the pronunciation library of regional accent A imported in S4 according to the regional accent A specified in the S5 request, and then matches and replaces each level of the pronunciation tree prepared in S6 according to the pronunciation library hierarchy order from low to high (phoneme-syllable-Chinese character-vocabulary-sentence) according to the regional accent A pronunciation library (if the information of this level of the current pronunciation tree exists on the same level of the regional accent pronunciation library, it is directly replaced; if it does not exist, it is skipped and not processed).

[0137] Note that for the Chinese characters, vocabulary, and sentence levels in the pronunciation library, the pinyin sequence needs to be split by spaces and then all pronunciation information of the downstream sub-nodes needs to be replaced. For the syllable and phoneme levels, the current node of the pronunciation tree and the downstream sub-node information can be directly replaced according to the pronunciation information. Note that after the high-level pronunciation is replaced, the low-level pronunciation nodes also need to be replaced synchronously. According to this replacement order, the high-level pronunciation rules in the regional pronunciation library have higher priority than the low-level pronunciation rules (for example, if bei3 needs to be replaced with bei1 in regional accent A, it will be replaced in the pronunciation tree first, but "Beijing" needs to be pronounced as bei4 jing2, then bei4 will overwrite the previous replacement of bei1).

[0138] S8, S9: Extract the phoneme sequence for the back-end to complete the synthesis. That is, the service front-end module obtains the pronunciation tree replaced according to the regional accent A pronunciation library in S7, and extracts the pronunciation information of the new pronunciation tree. Similarly, to obtain the final pronunciation information based on a pronunciation tree, it is necessary to traverse the bottom layer (phoneme) nodes of this tree in order (such as [w, o3, ai4, b, ei3, j, ing1, t, ian1, an1, m, eng2]), and then the service front-end returns the pronunciation information in S8, transmits it to the back-end of the speech synthesis system for speech synthesis to obtain a specific voice, and returns it to the client.

[0139] Further reference Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speech synthesis device. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0140] As Figure 8 shown, the speech synthesis device 800 in this embodiment may include: an information acquisition unit 801 configured to acquire the text to be speech-synthesized and the pronunciation method; a standard pronunciation tree conversion unit 802 configured to convert the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes, and the pronunciation layer nodes in the standard pronunciation tree record the standard pronunciation information of the text to be speech-synthesized; a pronunciation tree adjustment unit 803 configured to adjust the standard pronunciation information based on the pronunciation method to obtain a target pronunciation tree; and a speech synthesis unit 804 configured to generate a target synthesized voice based on the adjusted pronunciation information recorded in the target pronunciation tree.

[0141] In this embodiment, in the speech synthesis device 800: the specific processing of the information acquisition unit 801, the standard pronunciation tree conversion unit 802, the pronunciation tree adjustment unit 803, and the speech synthesis unit 804 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201 - 204 in the corresponding embodiments, which will not be elaborated here.

[0142] In some optional implementation manners of this embodiment, the standard pronunciation tree conversion unit 802 may include:

[0143] a conversion subunit configured to convert the text to be speech-synthesized into a standard pronunciation tree composed of text layer nodes and pronunciation layer nodes, where the text layer nodes are the parent nodes of the pronunciation layer nodes.

[0144] In some optional implementation manners of this embodiment, the conversion subunit:

[0145] A text layer node determination module, configured to determine text layer nodes included in the text to be speech synthesized according to the text structure of the text to be speech synthesized;

[0146] A pronunciation layer node determination module, configured to determine pronunciation layer nodes for recording standard pronunciation information of the text to be speech synthesized;

[0147] A standard pronunciation tree construction module, configured to convert the text to be speech synthesized into a standard pronunciation tree composed of text layer nodes and pronunciation layer nodes.

[0148] In some optional implementation manners of this embodiment, the text layer nodes include character layer nodes; the pronunciation layer nodes include phoneme layer nodes.

[0149] Among them, the standard pronunciation tree construction module is further configured to:

[0150] Use the character layer node served by the text to be speech synthesized containing a single character as the root node of the pronunciation tree;

[0151] Use the phoneme layer node served by the phoneme information of a single character as the child node of the root node to obtain the standard pronunciation tree.

[0152] In some optional implementation manners of this embodiment, the text layer nodes include character layer nodes, word layer nodes and sentence layer nodes, and the pronunciation layer includes phoneme layer nodes.

[0153] Among them, the standard pronunciation tree construction module is further configured to:

[0154] Determine the sentence layer node served by the text to be speech synthesized as the root node of the pronunciation tree;

[0155] Determine the word layer node served by each word text constituting the text to be speech synthesized as the first intermediate node that is the child node of the root node;

[0156] Determine the character layer node served by each character text constituting each word text as the second intermediate node that is the child node of the first intermediate node;

[0157] Determine the phoneme layer node served by the phoneme information corresponding to each character text as the child node of the second intermediate node to obtain the standard pronunciation tree.

[0158] In some optional implementation manners of this embodiment, the text layer nodes include character layer nodes, word layer nodes and sentence layer nodes; the pronunciation layer nodes include syllable layer nodes and phoneme layer nodes.

[0159] Among them, the standard pronunciation tree construction module is further configured to:

[0160] Determine the sentence layer node served by the text to be speech synthesized as the root node of the pronunciation tree;

[0161] Determine the word layer nodes, which are composed of the word texts of the text to be speech synthesized, as the first intermediate nodes that are the child nodes of the root node;

[0162] Determine the character layer nodes, which are composed of the character texts of each word text, as the second intermediate nodes that are the child nodes of the first intermediate nodes;

[0163] Determine the syllable layer nodes, which are composed of the syllable information corresponding to each character text, as the third intermediate nodes that are the child nodes of the second intermediate nodes;

[0164] Determine the phoneme layer nodes, which are composed of the phoneme information that makes up each syllable information, as the child nodes of the third intermediate nodes, and obtain the standard pronunciation tree.

[0165] In some alternative implementation manners of this embodiment, the pronunciation manner includes target region information for indicating pronunciation in the target region accent, the standard pronunciation tree is a Mandarin pronunciation tree that follows the Mandarin pronunciation standard, and the target pronunciation tree is a dialect pronunciation tree corresponding to the target region accent.

[0166] In some alternative implementation manners of this embodiment, the speech synthesis device 800 may further include:

[0167] A dialect pronunciation library creation unit, configured to create corresponding dialect pronunciation libraries for different regions respectively; wherein, the dialect pronunciation library includes: a first correspondence between texts at different text levels and pronunciation information in the corresponding regional accent, and a second correspondence between different pronunciation information of Mandarin and the corresponding regional accent for expressing the same semantics.

[0168] In some alternative implementation manners of this embodiment, the pronunciation tree adjustment unit 803 may be further configured to:

[0169] Determine the target dialect pronunciation manner according to the target region accent corresponding to the target region information;

[0170] For each layer of nodes that make up the Mandarin pronunciation tree, determine the target dialect pronunciation information corresponding to the current layer of nodes in the bottom-up hierarchical order of the nodes;

[0171] Adjust the Mandarin pronunciation tree according to the target dialect pronunciation information to obtain the dialect pronunciation tree.

[0172] In some alternative implementation manners of this embodiment, the pronunciation tree adjustment unit 803 may include:

[0173] A target dialect pronunciation library determination subunit, configured to determine a target dialect pronunciation library corresponding to target regional information; wherein, corresponding dialect pronunciation libraries are respectively created in advance for different regions, and the dialect pronunciation libraries include: a first correspondence relationship between texts at different text levels and pronunciation information in the corresponding regional accent, and a second correspondence relationship between different pronunciation information used by Mandarin and the corresponding regional accent to express the same semantics;

[0174] A first adjustment subunit, configured to determine first dialect pronunciation information corresponding to the Mandarin pronunciation information in the pronunciation layer node of the Mandarin pronunciation tree according to the second correspondence relationship, and replace the Mandarin pronunciation information with the corresponding first dialect pronunciation information to obtain an adjusted pronunciation tree;

[0175] A second adjustment subunit, configured to determine second dialect pronunciation information corresponding to the text recorded in the text layer node of the adjusted pronunciation tree according to the first correspondence relationship, and replace the current pronunciation information corresponding to the text in the adjusted pronunciation tree with the second dialect pronunciation information to obtain a dialect pronunciation tree.

[0176] In some optional implementation manners of this embodiment, the pronunciation layer node includes a syllable layer node and a phoneme layer node; the second correspondence relationship includes: a syllable-phoneme correspondence relationship between the pronunciation syllables of the same character in Mandarin and the pronunciation phonemes of the character in the corresponding regional accent, and a phoneme-phoneme correspondence relationship between different pronunciation phonemes of the same semantics in Mandarin and the corresponding regional accent respectively,

[0177] wherein, the first adjustment subunit can be further configured to:

[0178] Use the phoneme-phoneme correspondence relationship to determine first dialect phoneme information corresponding to the Mandarin phoneme information recorded in the phoneme layer node of the Mandarin pronunciation tree, and replace the current phoneme information recorded in the phoneme layer node with the first dialect syllable information to obtain a first adjusted pronunciation tree;

[0179] Use the syllable-phoneme correspondence relationship to determine second dialect phoneme information corresponding to the Mandarin syllable information recorded in the syllable layer node of the Mandarin pronunciation tree, and replace the current phoneme information recorded in the phoneme layer node with the second dialect phoneme information to obtain an adjusted pronunciation tree.

[0180] In some optional implementation manners of this embodiment, the first adjustment subunit may further include:

[0181] A first correction module, configured to correct the syllable information currently recorded in the syllable layer node of the adjusted pronunciation tree according to the phoneme information currently recorded in the phoneme layer node of the adjusted pronunciation tree.

[0182] In some alternative implementation manners of this embodiment, the syllable-phoneme correspondence includes: a syllable-phoneme key-value pair with the pronunciation syllable of the same character in Mandarin as the key and the pronunciation phoneme in the corresponding regional accent as the value; the phoneme correspondence includes a phoneme-phoneme key-value pair with the pronunciation phoneme of the same semantics in Mandarin as the key and the pronunciation phoneme in the corresponding regional accent as the value.

[0183] In some alternative implementation manners of this embodiment, the text layer nodes include character layer nodes, word layer nodes, and sentence layer nodes; the first correspondence includes: the sentence pronunciation relationship between the sentence text and the pronunciation phonemes in the corresponding regional accent, the word pronunciation relationship between the word text and the pronunciation phonemes in the corresponding regional accent, and the character pronunciation relationship between the character text and the pronunciation phoneme information in the corresponding regional accent.

[0184] Wherein, the second adjustment subunit is further configured to:

[0185] Use the character pronunciation relationship to determine the third dialect phoneme information corresponding to the character text recorded in the character layer node in the pronunciation tree during adjustment, and replace the current phoneme information recorded in the phoneme layer node in the pronunciation tree during adjustment with the third dialect phoneme information to obtain a second adjusted pronunciation tree;

[0186] Use the word pronunciation relationship to determine the fourth dialect phoneme information corresponding to the word text recorded in the word layer node in the second adjusted pronunciation tree, and replace the current phoneme information recorded in the phoneme layer node in the second adjusted pronunciation tree with the fourth dialect phoneme information to obtain a third adjusted pronunciation tree;

[0187] Use the sentence pronunciation relationship to determine the fifth dialect phoneme information corresponding to the sentence text recorded in the sentence layer node in the third adjusted pronunciation tree, and replace the current phoneme information recorded in the phoneme layer node in the third adjusted pronunciation tree with the fifth dialect syllable information to obtain a dialect pronunciation tree.

[0188] In some alternative implementation manners of this embodiment, the second adjustment subunit may further include:

[0189] A second correction module, configured to correct the syllable information currently recorded in the syllable layer node in the dialect pronunciation tree according to the phoneme information currently recorded in the phoneme layer node in the dialect pronunciation tree.

[0190] In some alternative implementation manners of this embodiment, the sentence pronunciation relationship includes a sentence pronunciation key-value pair with the sentence text as the key and the pronunciation phonemes in the corresponding regional accent as the value; the word pronunciation relationship includes a word pronunciation key-value pair with the word text as the key and the pronunciation factors in the corresponding regional accent as the value; the character pronunciation relationship includes a character pronunciation key-value pair with the character text as the key and the pronunciation factors in the corresponding regional accent as the value.

[0191] In some alternative implementation manners of this embodiment, the pronunciation tree adjustment unit 803 may be further configured to:

[0192] In response to the number of pronunciation manners being greater than 1, respectively determine target texts corresponding to each pronunciation manner in the text to be speech-synthesized;

[0193] Adjust the standard pronunciation information corresponding to different target texts according to the corresponding pronunciation manners to obtain a target pronunciation tree.

[0194] In some alternative implementation manners of this embodiment, the speech synthesis unit 804 may be further configured to:

[0195] Traverse in sequence the pronunciation information recorded in each leaf node of the target pronunciation tree to generate a target pronunciation sequence corresponding to the text to be speech-synthesized;

[0196] Synthesize a target synthesized speech according to the pronunciation manners indicated by the target pronunciation sequence.

[0197] This embodiment exists as a device embodiment corresponding to the above method embodiment. The speech synthesis device provided in this embodiment converts the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes for recording pronunciation information, so that the relevant information required for speech synthesis can be presented in the form of a tree, and it is also convenient to adjust the standard pronunciation information according to the specified pronunciation manners to obtain a target pronunciation tree, and then synthesize the corresponding synthesized speech. That is, by representing the relevant information required for speech synthesis in the form of a tree, it is convenient to directly adjust the pronunciation information recorded in the nodes of the pronunciation tree according to special pronunciation requirements, which is simple and easy to implement, has lower costs, and is convenient for deployment and update maintenance as needed.

[0198] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can implement the speech synthesis method described in any of the above embodiments.

[0199] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions for enabling a computer to implement the speech synthesis method described in any of the above embodiments when executed.

[0200] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, and when the computer program is executed by a processor, it can implement the steps of the speech synthesis method described in any of the above embodiments.

[0201] Figure 9 FIG. 1 shows a schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0202] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0203] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0204] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the speech synthesis method. For example, in some embodiments, the speech synthesis method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the speech synthesis method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the speech synthesis method in any other suitable manner (e.g., by means of firmware).

[0205] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0206] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0207] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0208] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0209] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0210] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0211] According to the technical solution of the embodiment of the present disclosure, by converting the text to be speech-synthesized into a standard pronunciation tree composed of pronunciation layer nodes for recording pronunciation information, the relevant information required for speech synthesis can be presented in the form of a tree, which also facilitates adjusting the standard pronunciation information according to the specified pronunciation method to obtain the target pronunciation tree, and then synthesizing the corresponding synthesized speech. That is, by representing the relevant information required for speech synthesis in the form of a tree, it is convenient to directly adjust the pronunciation information recorded on the nodes of the pronunciation tree in combination with special pronunciation requirements, which is simple, has lower costs, and is convenient for deployment and update maintenance as needed.

[0212] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure of the present invention can be achieved, and no limitations are imposed herein.

[0213] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Acquire a text to be speech synthesized and a pronunciation method, wherein the pronunciation method includes target region information for indicating pronunciation in a target region accent; Converting the text to be synthesized into a standard pronunciation tree based on pronunciation layer nodes, wherein the pronunciation layer nodes in the standard pronunciation tree record standard pronunciation information of the text to be synthesized, and the standard pronunciation tree is a Mandarin pronunciation tree that complies with Mandarin pronunciation standards; The standard pronunciation information is adjusted based on the pronunciation mode to obtain a target pronunciation tree, including: determining a target dialect pronunciation mode according to a target regional accent corresponding to the target regional information; determining the target dialect pronunciation information corresponding to the current layer node in a bottom-up hierarchical order of the nodes for each layer of nodes constituting the Mandarin pronunciation tree; adjusting the Mandarin pronunciation tree according to the target dialect pronunciation information to obtain a dialect pronunciation tree corresponding to the target regional accent; Based on the adjusted pronunciation information recorded in the target pronunciation tree, a target synthesized speech is generated.

2. The method according to claim 1, characterized in that The step of converting the text to be synthesized into a standard pronunciation tree based on pronunciation layer nodes comprises: The text to be speech synthesized is converted into a standard pronunciation tree consisting of text layer nodes and pronunciation layer nodes, wherein the text layer nodes are parent nodes of the pronunciation layer nodes.

3. The method according to claim 2, characterized in that The step of converting the text to be synthesized into a standard pronunciation tree consisting of text layer nodes and pronunciation layer nodes comprises: Determining the text layer nodes included in the text to be synthesized into speech according to the text structure of the text to be synthesized into speech; Determine a pronunciation layer node for recording standard pronunciation information of the text to be speech synthesized; The text to be speech synthesized is converted into a standard pronunciation tree consisting of the text layer nodes and the pronunciation layer nodes.

4. The method according to claim 3, characterized in that The text layer nodes include character layer nodes; the pronunciation layer nodes include phoneme layer nodes, The step of converting the text to be synthesized into a standard pronunciation tree composed of the text layer nodes and the pronunciation layer nodes includes: Taking the character-level node represented by the text to be synthesized containing a single character as the root node of the pronunciation tree; The phoneme-level node represented by the phoneme information of the single word is used as a child node of the root node to obtain the standard pronunciation tree.

5. The method according to claim 3, characterized in that: The text layer nodes include character layer nodes, word layer nodes and sentence layer nodes, and the pronunciation layer nodes include phoneme layer nodes. The step of converting the text to be synthesized into a standard pronunciation tree composed of the text layer nodes and the pronunciation layer nodes includes: Determine the sentence-level node served by the text to be speech synthesized as the root node of the pronunciation tree; Determine the word-layer node served by each word text constituting the text to be synthesized into speech as the first intermediate node as a child node of the root node; Determine the character-layer node served by each character text constituting each of the word texts as a second intermediate node that is a child node of the first intermediate node; The phoneme layer nodes represented by the phoneme information corresponding to each of the character texts are determined as child nodes of the second intermediate node to obtain the standard pronunciation tree.

6. The method according to claim 3, characterized in that The text layer nodes include character layer nodes, word layer nodes and sentence layer nodes; the pronunciation layer nodes include syllable layer nodes and phoneme layer nodes. The step of converting the text to be synthesized into a standard pronunciation tree consisting of the text layer nodes and pronunciation layer nodes includes: Determine the sentence-level node served by the text to be speech synthesized as the root node of the pronunciation tree; Determine the word-layer node served by each word text constituting the text to be synthesized into speech as the first intermediate node as a child node of the root node; Determine the character-layer node served by each character text constituting each of the word texts as a second intermediate node that is a child node of the first intermediate node; Determine the syllable-level node represented by the syllable information corresponding to each of the character texts as a third intermediate node that is a child node of the second intermediate node; The phoneme-layer nodes represented by the phoneme information constituting each of the syllable information are determined as child nodes of the third intermediate node to obtain the standard pronunciation tree.

7. The method according to claim 1, characterized in that Also includes: Create corresponding dialect pronunciation libraries for different regions respectively; wherein the dialect pronunciation library includes: a first correspondence between texts at different text levels and pronunciation information under corresponding regional accents, and a second correspondence between different pronunciation information of the Mandarin and the corresponding regional accents used to express the same semantics.

8. The method according to claim 1, characterized in that The step of adjusting the standard pronunciation information based on the pronunciation mode to obtain a target pronunciation tree includes: Determine a target dialect pronunciation library corresponding to the target regional information; wherein corresponding dialect pronunciation libraries are pre-created for different regions, and the dialect pronunciation library includes: a first correspondence between texts at different text levels and pronunciation information under corresponding regional accents, and a second correspondence between different pronunciation information of the Mandarin and the corresponding regional accents for expressing the same semantics; Determine the first dialect pronunciation information corresponding to the Mandarin pronunciation information in the pronunciation layer node of the Mandarin pronunciation tree according to the second corresponding relationship, and replace the Mandarin pronunciation information with the corresponding first dialect pronunciation information to obtain an adjusting pronunciation tree; The second dialect pronunciation information corresponding to the text recorded in the text layer node of the pronunciation tree under adjustment is determined according to the first corresponding relationship, and the current pronunciation information corresponding to the text recorded in the pronunciation tree under adjustment is replaced with the second dialect pronunciation information to obtain the dialect pronunciation tree.

9. The method according to claim 8, characterized in that The pronunciation layer nodes include syllable layer nodes and phoneme layer nodes; the second corresponding relationship includes: a syllable-phoneme corresponding relationship between the pronunciation syllables of the same character in the Mandarin and the pronunciation phonemes of the character in the corresponding regional accent, and a phoneme-phoneme corresponding relationship between different pronunciation phonemes of the same semantics in the Mandarin and the corresponding regional accent, The step of determining the first dialect pronunciation information corresponding to the Mandarin pronunciation information in the pronunciation layer node of the Mandarin pronunciation tree according to the second corresponding relationship, and replacing the Mandarin pronunciation information with the corresponding first dialect pronunciation information to obtain the adjusting pronunciation tree includes: Determine the first dialect phoneme information corresponding to the Mandarin phoneme information recorded by the phoneme layer node in the Mandarin pronunciation tree by using the phoneme-phoneme correspondence, and replace the current phoneme information recorded by the phoneme layer node with the first dialect phoneme information to obtain a first adjusted pronunciation tree; The syllable-phoneme correspondence is used to determine the second dialect phoneme information corresponding to the Mandarin syllable information recorded by the syllable layer node in the Mandarin pronunciation tree, and the current phoneme information recorded by the phoneme layer node is replaced with the second dialect phoneme information to obtain the adjusting pronunciation tree.

10. The method according to claim 9, characterized in that Also includes: According to the phoneme information currently recorded in the phoneme-level node in the pronunciation tree under adjustment, the syllable information currently recorded in the syllable-level node in the pronunciation tree under adjustment is modified.

11. The method according to claim 9, characterized in that The syllable-phoneme correspondence includes: a syllable-phoneme key-value pair with the pronunciation syllable of the same word in the mandarin as the key and the pronunciation phoneme in the corresponding regional accent as the value; the phoneme correspondence includes a phoneme-phoneme key-value pair with the pronunciation phoneme of the same semantics in the mandarin as the key and the pronunciation phoneme in the corresponding regional accent as the value.

12. The method according to claim 8, characterized in that The text layer nodes include character layer nodes, word layer nodes and sentence layer nodes; the first corresponding relationship includes: sentence pronunciation relationship between sentence text and pronunciation phonemes under corresponding regional accents, word pronunciation relationship between word text and pronunciation phonemes under corresponding regional accents, and word pronunciation relationship between word text and pronunciation phoneme information under corresponding regional accents, The method of determining the second dialect pronunciation information corresponding to the text recorded by the text layer node of the pronunciation tree under adjustment according to the first corresponding relationship, and replacing the current pronunciation information corresponding to the text recorded in the pronunciation tree under adjustment with the second dialect pronunciation information to obtain the dialect pronunciation tree includes: Determine the third language phoneme information corresponding to the character text recorded by the character layer node in the pronunciation tree under adjustment by using the character pronunciation relationship, and replace the current phoneme information recorded by the phoneme layer node in the pronunciation tree under adjustment with the third language phoneme information to obtain a second adjusted pronunciation tree; Determine the fourth dialect phoneme information corresponding to the word text recorded by the word layer node in the second adjusted pronunciation tree by using the word pronunciation relationship, and replace the current phoneme information recorded by the phoneme layer node in the second adjusted pronunciation tree with the fourth dialect phoneme information to obtain a third adjusted pronunciation tree; The sentence pronunciation relationship is used to determine the fifth dialect phoneme information corresponding to the sentence text recorded by the sentence layer node in the third adjusted pronunciation tree, and the current phoneme information recorded by the phoneme layer node in the third adjusted pronunciation tree is replaced with the fifth dialect phoneme information to obtain the dialect pronunciation tree.

13. The method according to claim 12, characterized in that Also includes: According to the phoneme information currently recorded in the phoneme layer node in the dialect pronunciation tree, the syllable information currently recorded in the syllable layer node in the dialect pronunciation tree is corrected.

14. The method according to claim 12, characterized in that The sentence pronunciation relationship includes sentence pronunciation key-value pairs with sentence text as key and pronunciation phonemes under corresponding regional accent as value; the word pronunciation relationship includes word pronunciation key-value pairs with word text as key and pronunciation factors under corresponding regional accent as value; the character pronunciation relationship includes character pronunciation key-value pairs with character text as key and pronunciation factors under corresponding regional accent as value.

15. The method according to claim 1, characterized in that The step of adjusting the standard pronunciation information based on the pronunciation mode to obtain a target pronunciation tree includes: In response to the number of the pronunciation modes being greater than 1, determining target texts corresponding to each pronunciation mode in the text to be speech synthesized; The standard pronunciation information corresponding to different target texts is adjusted according to corresponding pronunciation methods to obtain the target pronunciation tree.

16. The method according to claim 1, characterized in that The step of generating a target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree comprises: Traversing the pronunciation information recorded in each leaf node of the target pronunciation tree in order to generate a target pronunciation sequence corresponding to the text to be synthesized into speech; The target synthesized speech is synthesized in the pronunciation manner indicated by the target pronunciation sequence.

17. A speech synthesis device, characterized in that: The device comprises: An information acquisition unit is configured to acquire a text to be speech synthesized and a pronunciation method, wherein the pronunciation method includes target region information for indicating pronunciation in a target region accent; A standard pronunciation tree conversion unit is configured to convert the text to be speech synthesized into a standard pronunciation tree based on pronunciation layer nodes, wherein the pronunciation layer nodes in the standard pronunciation tree record standard pronunciation information of the text to be speech synthesized, and the standard pronunciation tree is a Mandarin pronunciation tree that complies with Mandarin pronunciation standards; The pronunciation tree adjustment unit is configured to adjust the standard pronunciation information based on the pronunciation method to obtain a target pronunciation tree, and is further configured to: determine a target dialect pronunciation method according to a target regional accent corresponding to the target regional information; determine the target dialect pronunciation information corresponding to the current layer node in a bottom-up hierarchical order of the nodes for each layer constituting the Mandarin pronunciation tree; adjust the Mandarin pronunciation tree according to the target dialect pronunciation information to obtain a dialect pronunciation tree corresponding to the target regional accent; The speech synthesis unit is configured to generate a target synthesized speech based on the adjusted pronunciation information recorded in the target pronunciation tree.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 16.

19. A non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause the computer to execute the method of any one of claims 1-16.

20. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Voice synthesis method and apparatus, electronic device, and non-transitory computer storage medium

    CN109065016A

  • Speech synthesis method and device, computer equipment and storage medium

    CN117894293A