Pronunciation method and device, speech synthesis system, storage medium and electronic equipment

CN117711370BActive Publication Date: 2026-08-07IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-12-28
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]然而,现有语音合成系统在文本分析阶段经常输出错误的音素序列,影响语音合成效果

Benefits of technology

[0039] By applying the solution of this invention, since the target dialect engine pronunciation dictionary distinguishes between Mandarin entries and dialect entries, and at least some Mandarin entries include index information of corresponding dialect entries, when a corresponding dialect entry exists in the first word face to be transcribed, the phoneme sequence information in the corresponding dialect entry can be directly obtained and output as the phoneme sequence information assigned to the transcribed entry. Thus, when there are entries in the transcribed entry that require dialectal expression conversion, authentic phoneme sequences can be assigned to the transcribed entry, thereby improving the speech synthesis results and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117711370B_ABST
    Figure CN117711370B_ABST
Patent Text Reader

Abstract

A phonetic annotation method and device, a speech synthesis system, a storage medium, and an electronic device. The method includes: when a to-be-annotated word is a non-multi-sound word in Putonghua, querying a target dialect engine pronunciation dictionary to obtain a first word face that matches the to-be-annotated word; the target dialect engine pronunciation dictionary includes: Putonghua words and dialect words, and at least part of the Putonghua words include index information of corresponding dialect words; determining whether the first word face is in a Putonghua word that has a corresponding dialect word; when there is a corresponding dialect word, obtaining phoneme sequence information in the corresponding dialect word as phoneme sequence information assigned to the to-be-annotated word and outputting. The above scheme can improve the accuracy of the phoneme sequence output in the text analysis stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and specifically to a phonetic notation method, apparatus, speech synthesis system, storage medium, and electronic device. Background Technology

[0002] Speech synthesis, also known as text-to-speech (TTS), aims to convert input text into fluent and natural output speech, and is a key technology for realizing intelligent human-computer voice interaction.

[0003] Speech synthesis generally involves three stages: text analysis, acoustic modeling, and vocoder processing. The results of each stage will affect the quality of speech synthesis. Among them, the text analysis module is closely related to the construction of front-end resources.

[0004] A complete text analysis process generally includes five major operations: text normalization, word segmentation, part-of-speech prediction, prosodic prediction, and phonetic annotation. The specific processing steps may vary depending on the dialect in which the text is synthesized. Phonetic annotation involves assigning phoneme sequences to the words obtained after text normalization, word segmentation, and part-of-speech prediction.

[0005] For words with two or more pronunciations, i.e., polyphonic words, the speech synthesis system mainly predicts the pronunciation of the words based on a model. For non-polyphonic words, the speech synthesis system assigns phoneme sequences directly based on the pronunciation dictionary of the target language engine.

[0006] However, existing speech synthesis systems often output incorrect phoneme sequences during the text analysis stage, affecting the speech synthesis effect. Summary of the Invention

[0007] The problem this invention aims to solve is to improve the accuracy of the phoneme sequences output during the text analysis stage.

[0008] To address the above problems, embodiments of the present invention provide a phonetic notation method, the method comprising:

[0009] When the word to be annotated is a non-polyphonic word in Mandarin, the target dialect engine pronunciation dictionary is queried to obtain the first word surface that matches the word to be annotated; the target dialect engine pronunciation dictionary includes: Mandarin words and dialect words, and at least some of the Mandarin words include the index information of the corresponding dialect words;

[0010] Determine whether the Mandarin entry containing the first word exists as a corresponding dialect entry;

[0011] When a corresponding dialect entry exists, the phoneme sequence information in the corresponding dialect entry is obtained and used as the phoneme sequence information assigned to the entry to be annotated, and then output.

[0012] Optionally, the method further includes:

[0013] When there is no corresponding dialect entry, the phoneme sequence information in the Mandarin entry where the first word is located is obtained and used as the phoneme sequence information assigned to the word to be annotated and output.

[0014] Optionally, the phoneme sequence information of OOV words in the target dialect engine's pronunciation dictionary is obtained by statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the words in the entry, and is the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

[0015] Optionally, the method further includes:

[0016] When the word to be annotated is a dialect word that is not a polyphonic word, the target dialect engine pronunciation dictionary is queried to obtain a second word face that matches the word to be annotated.

[0017] Obtain the phoneme sequence information in the dialect entry where the second word is located, and use it as the phoneme sequence information assigned to the word to be annotated, and output it.

[0018] Optionally, the target dialect engine pronunciation dictionary is constructed using the following method:

[0019] Based on existing dictionary resources, the word information of words and phrases already covered by the existing dictionary resources in the target dialect is obtained;

[0020] Identify word entries that are not covered by existing dictionary resources.

[0021] Optionally, determining the word entry information for words not covered by existing dictionary resources includes: obtaining phoneme sequence information for words not covered by existing dictionary resources but whose single-character pronunciation is not missing through manual screening; obtaining phoneme sequence information for words not covered by existing dictionary resources but whose single-character pronunciation is missing through statistical analysis based on Mandarin phonetic information and dialect phonetic information; and determining other word entry information besides phoneme sequence information to obtain complete word entry information for words not covered by existing dictionary resources.

[0022] Optionally, the statistical analysis based on Mandarin phonetic information and dialect phonetic information to obtain phoneme sequence information of words not covered by existing dictionary resources and with missing pronunciations includes:

[0023] Generate a Chinese character pronunciation table for words not covered by the existing dictionary resources and whose single-character pronunciation is missing. The Chinese character pronunciation table includes dialect pronunciation information, Mandarin pronunciation information, and Middle Chinese pronunciation information for each word not covered by the existing dictionary resources; the pronunciation information includes syllable information and tone information.

[0024] Based on the syllable and tone information in the Mandarin phonetic information of words not covered by existing dictionary resources and words with missing pronunciations, as well as the syllable and tone information in the Middle Chinese phonetic information, statistical analysis was conducted to obtain the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

[0025] Based on the phoneme sequence information corresponding to the most frequently occurring syllable, the phoneme sequence information of words not covered by the existing dictionary resources and whose single-character pronunciation is missing is obtained.

[0026] Optionally, the phonetic information further includes: initial consonant information and final vowel information;

[0027] The step of obtaining phoneme sequence information for words not covered by existing dictionary resources and with missing pronunciations based on the phoneme sequence information corresponding to the most frequently occurring syllable also includes:

[0028] Statistical analysis was conducted on the initial consonant information in Mandarin phonetic information and Middle Chinese phonetic information that are not covered by existing dictionary resources to obtain the phoneme sequence information corresponding to the most frequently occurring initial consonant.

[0029] Based on the final and tone information in the Mandarin phonetic information of words not covered by existing dictionary resources, and the final and tone information in the Middle Chinese phonetic information, statistical analysis was conducted to obtain the phoneme sequence information corresponding to the most frequently occurring final.

[0030] The phoneme sequence information of the syllables, the initial consonant, and the final vowel is manually checked to determine the phoneme sequence information of words not covered by the existing dictionary resources.

[0031] This invention also provides a phonetic transcription device, the phonetic transcription device comprising:

[0032] The query unit is adapted to query the target dialect engine pronunciation dictionary when the word to be annotated is a Mandarin non-polyphonic word, and obtain a first word surface that matches the word to be annotated; the target dialect engine pronunciation dictionary includes: Mandarin words and dialect words, and at least some of the Mandarin words include index information of the corresponding dialect words;

[0033] The determining unit is adapted to determine whether there is a corresponding dialect entry for the Mandarin entry containing the first word;

[0034] The phoneme sequence output unit is adapted to, when a corresponding dialect entry exists, acquire the phoneme sequence information in the corresponding dialect entry, and output it as the phoneme sequence information assigned to the entry to be annotated.

[0035] This invention also provides a speech synthesis system, which includes the aforementioned phonetic transcription device.

[0036] This invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of any of the methods described above.

[0037] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the steps of any of the methods described above when running the computer program.

[0038] Compared with the prior art, the technical solution of the embodiments of the present invention has the following advantages:

[0039] By applying the solution of this invention, since the target dialect engine pronunciation dictionary distinguishes between Mandarin entries and dialect entries, and at least some Mandarin entries include index information of corresponding dialect entries, when a corresponding dialect entry exists in the first word face to be transcribed, the phoneme sequence information in the corresponding dialect entry can be directly obtained and output as the phoneme sequence information assigned to the transcribed entry. Thus, when there are entries in the transcribed entry that require dialectal expression conversion, authentic phoneme sequences can be assigned to the transcribed entry, thereby improving the speech synthesis results and enhancing the user experience.

[0040] Furthermore, the phoneme sequence information of OOV words in the target dialect engine's pronunciation dictionary is obtained by statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the words in the entry. The phoneme sequence information corresponding to the most frequently occurring syllable is obtained. The pronunciation is inferred based on the statistical analysis results. Compared with manually supplementing the pronunciation of OOV words, this can greatly improve the accuracy of phonetic notation, thereby further improving the speech synthesis results and enhancing the user experience. Attached Figure Description

[0041] Figure 1 This is a flowchart of an OOV word pronunciation correction method;

[0042] Figure 2 This is a flowchart of a phonetic notation method according to an embodiment of the present invention;

[0043] Figure 3 This is a flowchart illustrating how to determine the term information of OOV words in an embodiment of the present invention;

[0044] Figure 4 This is a flowchart illustrating how to determine the dialect pronunciation of the word OOV in an embodiment of the present invention;

[0045] Figure 5 This is a schematic diagram of the structure of a phonetic device according to an embodiment of the present invention;

[0046] Figure 6 This is a schematic diagram of the workflow of a speech synthesis system according to an embodiment of the present invention. Detailed Implementation

[0047] Existing dialect engine pronunciation dictionary construction generally includes the following two steps:

[0048] The first step is to comprehensively collect and organize existing dictionary resources for the target dialect;

[0049] The second step is to manually proofread the pronunciation of words and phrases (OOV, Out-of-Vocabulary) that are not covered by existing dictionary resources, based on business needs, and supplement the target dialect engine's pronunciation dictionary.

[0050] When collecting and organizing existing dictionary resources for the target dialect, all existing dictionary resources can be comprehensively collected and organized, and stored in the target dialect engine's pronunciation dictionary. The entries in this target dialect engine's pronunciation dictionary are generally in the format of "word literal: phoneme sequence: segmentation cost; syllable count; part of speech;". For example, the format of an entry in the Minnan (Hokkien) language engine's pronunciation dictionary is "eyes: gan3 tsing1: 102; 2; n;". Of course, the format of the entries will vary depending on the target dialect.

[0051] In the process of supplementing the target dialect engine's pronunciation dictionary, after collecting and organizing existing character dictionary resources, OOV words were compiled based on business needs, and their pronunciations were manually proofread to supplement the engine's pronunciation dictionary. Specifically, refer to... Figure 1 When proofreading the pronunciation of existing OOV words, the first step is to determine whether the single-character pronunciation of the OOV word exists (as shown in step 11). If the single-character pronunciation exists, the phoneme sequence of all characters under that entry is usually listed for proofreaders to filter (as shown in step 12). For entries with missing single-character pronunciations, a "blind proofreading" method is used directly (as shown in step 13), that is, the word is provided and the phoneme sequence is directly supplemented manually. Finally, the format of the engine's pronunciation dictionary is organized (as shown in step 14).

[0052] The above approach to constructing a target dialect engine pronunciation dictionary presents the following two significant problems:

[0053] 1. Dialect pronunciation proofreading for words with missing OOV characters mainly relies on manual proofreading. For entries with missing pronunciation for single characters, providing only the word and manually supplementing the pronunciation in a "blind proofreading" manner is difficult and has a low accuracy rate. However, this proofreading method is extremely common in the construction of dialect dictionary resources, especially for dialects with scarce dictionary resources.

[0054] 2. The pronunciation of entries in existing dialect engine pronunciation dictionaries is usually obtained by converting the Mandarin pronunciation of the entry into the dialect pronunciation. The dictionary itself does not distinguish between Mandarin and dialect entries, nor does it establish a connection between synonyms of the two. Therefore, according to the existing dialect synthesis text analysis process, when the input Mandarin text contains entries that need to be converted to authentic dialect expressions, the pronunciation dictionary is prone to assigning incorrect or non-authentic phoneme sequences, and the final speech synthesis result will affect the user's understanding.

[0055] For example, the Mandarin word for "eyes" is pronounced [gan3tsing1] in Hokkien, based on its literal dialect pronunciation. However, in actual Hokkien usage, "eyes" is not pronounced [gan3 tsing1], but rather as "eyeball [bag8 tsiu1]". If the input text is Mandarin: "Your eyes are very beautiful", because the pronunciation dictionary does not establish a connection between "eyes" and "eyeball", the phoneme sequence will be directly assigned as [gan3 tsing1] whenever "eyes" is encountered during text analysis, causing difficulty for the user to understand.

[0056] When constructing a target dialect engine pronunciation dictionary using the above scheme, the speech synthesis system often outputs incorrect phoneme sequences during the text analysis stage, affecting the speech synthesis effect.

[0057] To address this problem, this invention provides a phonetic annotation method. This method improves the target dialect engine's pronunciation dictionary, enabling it to distinguish between Mandarin and dialect entries. Furthermore, it ensures that at least some Mandarin entries include index information for corresponding dialect entries. Thus, when a corresponding dialect entry exists in the first word match of the word to be annotated, the phoneme sequence information from that dialect entry can be directly obtained and used as the phoneme sequence information assigned to the word to be annotated. This allows for the assignment of authentic phoneme sequences to the word to be annotated when it contains entries requiring dialectal authentic expression conversion, thereby improving the speech synthesis results.

[0058] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0059] Reference Figure 2 This invention provides a phonetic notation method, which may include the following steps:

[0060] Step 21: When the word to be annotated is a non-polyphonic word in Mandarin, query the target dialect engine pronunciation dictionary to obtain the first word face that matches the word to be annotated.

[0061] The target dialect engine pronunciation dictionary includes: Mandarin entries and dialect entries, with at least some Mandarin entries including index information of the corresponding dialect entries.

[0062] In its implementation, the target dialect engine's pronunciation dictionary distinguishes between Standard Mandarin entries and dialect entries. Furthermore, for Standard Mandarin entries with corresponding dialect entries, index information for the corresponding dialect entry can be set within that Standard Mandarin entry, allowing for retrieval of the corresponding dialect entry based on this index information. For Standard Mandarin entries without corresponding dialect entries, index information for the corresponding dialect entry does not need to be set within that Standard Mandarin entry.

[0063] In practice, the index information of the dialect entry can be the word literal of the corresponding dialect entry or other information, as long as it allows the user to jump to the corresponding dialect entry based on the index information.

[0064] In one embodiment, the format of the Mandarin entries in the target dialect engine's pronunciation dictionary can be:

[0065] Lexical literal meaning: phoneme sequence: segmentation cost; number of syllables; part of speech; source tags; (dialectal idiomatic expression of lexical literal meaning;)

[0066] The format of dialect entries in the target dialect engine's pronunciation dictionary can be:

[0067] Word literal: Phoneme sequence: segmentation cost; number of syllables; part of speech; source tags;

[0068] Here, segmentation cost refers to the segmentation cost calculated when processing the word segment. Syllable count refers to the number of syllables contained in the word segment; typically, a two-character word segment has two syllables. Part of speech indicates whether the word segment is a verb or a noun, etc. Source tag identifies whether the word segment is in Standard Mandarin or a dialect. For Standard Mandarin entries, when authentic dialect expressions exist, the authentic dialect expressions can also be included.

[0069] For example, in the Minnan dialect pronunciation dictionary, the Mandarin entry format is: “eyes: gan3 tsing1: 102; 2; n; ZhCn; eyeballs”. The dialect entry format is: “eyeballs: bag8 tsiu1: 97; 2; n; MiCn;”.

[0070] In practice, when the word to be annotated is a non-polyphonic word in Mandarin, the first matching word in the target dialect engine's pronunciation dictionary is the word in the matching Mandarin word.

[0071] Step 22: Determine whether there is a corresponding dialect entry for the Mandarin entry containing the first word.

[0072] In practice, after obtaining the first word face that matches the word to be annotated, the content of the corresponding Mandarin word can be retrieved, and the index information of the corresponding dialect word can be queried from the retrieved Mandarin word. If the index information of the corresponding dialect word exists, then the corresponding dialect word exists; otherwise, the corresponding dialect word does not exist.

[0073] If a corresponding dialect entry exists, proceed to step 23; otherwise, proceed to step 24.

[0074] Step 23: Obtain the phoneme sequence information in the corresponding dialect entry, use it as the phoneme sequence information assigned to the entry to be annotated, and output it.

[0075] In practice, if a corresponding dialect entry exists, the system jumps to that dialect entry and reads the phoneme sequence information from it, which is then used as the phoneme sequence information assigned to the entry to be annotated and output.

[0076] By distinguishing between dialect entries and standard Mandarin entries and associating the two, once a word to be transcribed has an authentic dialect expression, the corresponding authentic dialect expression can be output, thus avoiding the assignment of incorrect or non-authentic phoneme sequences and improving the speech synthesis effect.

[0077] Step 24: Obtain the phoneme sequence information in the Mandarin entry where the first word is located, and output it as the phoneme sequence information assigned to the word to be annotated.

[0078] In practice, if there is no corresponding dialect entry for the Mandarin entry containing the word to be annotated, the phoneme sequence information in the Mandarin entry is directly used as the phoneme sequence information assigned to the word to be annotated and output.

[0079] In practice, the phoneme sequence information of OOV words in the target dialect engine's pronunciation dictionary is obtained by statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the words in the entry, and is the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

[0080] Specifically, OOV words refer to characters and words not covered by existing dictionary resources. OOV words can be Mandarin or dialect words. When an OOV word is a Mandarin word, its phoneme sequence information in the target dialect engine's pronunciation dictionary is obtained through statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the word in the corresponding Mandarin entry, corresponding to the phoneme sequence information corresponding to the most frequently occurring dialect syllable. When an OOV word is a dialect word, its phoneme sequence information in the target dialect engine's pronunciation dictionary is obtained through statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the corresponding dialect entry, corresponding to the phoneme sequence information corresponding to the most frequently occurring dialect syllable.

[0081] It should be noted that, in this embodiment of the invention, the target dialect can be any dialect of Chinese. Regardless of the dialect, the phonetic notation method of this embodiment can be used for phonetic notation.

[0082] This invention also provides a method for constructing a pronunciation dictionary for a target dialect engine. Specifically, when constructing the pronunciation dictionary for a target dialect engine, one can first obtain the word information of words already covered by the existing word dictionary resources in the target dialect, and then determine the word information of words not covered by the existing word dictionary resources (i.e., OOV words).

[0083] In practice, existing dictionary resources already cover word and phrase entries, including both Mandarin and dialect entries. Specifically, this can be achieved by collecting and organizing existing dictionary resources, converting the phonetic notation of entries in those resources into a specific phonetic format. For example, the mapping between the phonetic notation of entries in existing dictionary resources and the International Phonetic Alphabet (IPA) can be established, and custom phonetic notation conversions, merging, and deduplication can be performed on the phonetic notation of each dictionary entry.

[0084] In addition, the source of each entry needs to be indicated. The source can be a dialect word, a standard Mandarin word, or a word of doubt. Typically, the collected entries can be compared with the *Modern Chinese Dictionary* to determine the source. For example, if an entry overlaps with an entry in the *Modern Chinese Dictionary*, it is a standard Mandarin word; otherwise, it is a dialect word. The sources of the collected entries can then be manually verified.

[0085] When determining the term information for OOV, refer to Figure 3 This may include the following steps:

[0086] Step 31: Determine whether words not covered by existing dictionary resources lack individual pronunciations.

[0087] If a word or phrase not covered by an existing dictionary resource lacks a single-character pronunciation, proceed to step 32; otherwise, proceed to step 33.

[0088] Step 32: Obtain phoneme sequence information from words that are not covered by the existing dictionary resources and whose pronunciations are not missing through manual screening.

[0089] In other words, the correct pronunciation of individual characters is manually selected to obtain the corresponding word information. In this case, the person selecting the pronunciation can be a native speaker of the target dialect, thus improving the accuracy of the pronunciation.

[0090] Step 33: Statistical analysis is performed based on Mandarin phonetic information and dialect phonetic information to obtain phoneme sequence information of words not covered by the existing dictionary resources and whose single-character pronunciation is missing.

[0091] By using statistical analysis to determine the phoneme sequence information of words with missing pronunciations in OOV words, it is equivalent to inferring the dialect pronunciation of words with missing pronunciations in OOV words through statistical analysis. Compared with manual blind pronunciation correction, this can improve the accuracy of pronunciation and reduce the difficulty of pronunciation determination. Subsequent speech synthesis using the target dialect engine's pronunciation dictionary will naturally yield better results.

[0092] Step 34: Determine other word information besides phoneme sequence information to obtain complete word information for words not covered by the existing dictionary resources.

[0093] In practice, after determining the phoneme sequence information of words in an OOV word whose individual pronunciations are not missing, the phoneme sequence information of the OOV word can be obtained by combining the pronunciations of each individual word in the OOV word.

[0094] In practice, after obtaining the phoneme sequence information of a word with a missing pronunciation in an OOV word, that word may combine with other words to form a Standard Mandarin entry. In this case, the phoneme sequence information can be combined with the phoneme sequences of other words to form the phoneme sequence information in the Standard Mandarin entry. The word may also combine with other words to form a dialect entry, in which case the phoneme sequence information can be combined with the phoneme sequences of other words to form the phoneme sequence information in the dialect entry.

[0095] In practice, in addition to phoneme sequence information, it is also necessary to collect and organize other information from the entry, including: segmentation cost, number of syllables, part of speech, and entry source. The collection and organization of this other information can be done before or after determining the phoneme sequence information.

[0096] In practice, the obtained phoneme sequence information can be checked by native speakers of the dialect for the OOV words. Other information in the entries can also be checked manually.

[0097] Taking the source of an entry as an example, after a single character is combined with other characters and words, the source of the entry can be determined first, and then the phoneme sequence information can be determined. When determining the source, the combined word can be compared with words in the *Modern Chinese Dictionary*. If the combined word overlaps with words in the *Modern Chinese Dictionary*, then the source of the combined word is a Standard Mandarin word; otherwise, it is a dialect word. Afterward, native speakers can be organized to conduct further random checks and verifications of the source of the entry. For Standard Mandarin entries with authentic dialect expressions, synonymous dialect words can be manually marked in the Standard Mandarin entries. After determining the source, statistical analysis can then be used to determine the phoneme sequence information of the combined word.

[0098] For example, the Hokkien OOV words "eyes" and "eyeballs" are transcribed as Mandarin words, with the former being marked as such. <zhcn>The latter is marked as a dialect word, that is <micn>During manual verification, the word "eyes" is marked as "eyes" in the Mandarin dictionary entry for that word. <zhcn>=Eyeball <micn>This links the two terms "eyes" and "eyeballs".

[0099] During subsequent proofreading, the source of the entry can be checked first, and then the pronunciation of the OOV word can be checked.

[0100] Figure 4 A flowchart for determining the dialect pronunciation of OOV words is shown. Specifically:

[0101] Reference Figure 4 In one embodiment, the step of performing statistical analysis based on Mandarin phonetic information and dialect phonetic information to obtain phoneme sequence information of words not covered by existing dictionary resources and with missing pronunciations may include:

[0102] Step 41: Generate a Chinese character pronunciation table for words not covered by the existing dictionary resources and whose single-character pronunciations are missing. The Chinese character pronunciation table includes dialect pronunciation information, Mandarin pronunciation information, and Middle Chinese pronunciation information for words not covered by the existing dictionary resources; the pronunciation information includes syllable information and tone information.

[0103] Specifically, when generating a Chinese character pronunciation table, the pronunciation information of all characters in the target dialect can be organized from four dimensions: syllable, initial consonant, final vowel, and tone. This ensures that each character in the table has complete Mandarin and Middle Chinese phonetic information. The Mandarin phonetic information includes syllable, initial consonant, final vowel, and tone, which can be obtained from the *Xinhua Dictionary* (Commercial Press), the *Hanyu Da Zidian* (Commercial Press), or other authoritative Mandarin dictionaries. The Middle Chinese phonetic information includes Middle Chinese phonological status, initial consonant category, final vowel (rhyme group + open / closed + etc.), and tone category, which can be obtained from the *Dialect Survey Character List* (Institute of Linguistics, Chinese Academy of Social Sciences, Commercial Press). Phonological information for rare characters can be obtained from the *Guangyun* dictionary or the online comprehensive rhyme dictionary search tool "Yun Dian Wang".

[0104] All collected and organized dialect pronunciation information for Chinese characters was supplemented and organized into the Chinese character pronunciation table according to four dimensions: "initial consonant", "final vowel", "tone", and "syllable". Then, based on the correspondence between the target dialect and Standard Mandarin information and Middle Chinese phonetic information, the phonetic information of the target dialect corresponding to the same phonetic information of Standard Mandarin and Middle Chinese characters was organized into the Chinese character pronunciation table.

[0105] The correspondence between the target dialect and Standard Mandarin information, as well as the Middle Chinese phonetic information, can be obtained from various dialect survey data, including but not limited to the "Modern Chinese Dialect Phonetic Database" series (often named "XX Dialect Phonetic Database" as the dialect data). For dialects outside the "Modern Chinese Dialect Phonetic Database" series, their corresponding local gazetteers, dialect volumes, etc., can be examined, and some dialect system survey data will also be appended.

[0106] Taking the Hefei dialect as an example, the generated Chinese character pronunciation table is shown in Table 1. In Table 1, Middle Chinese phonology, initial consonant, final vowel, and tone represent the Middle Chinese phonological information of the corresponding Chinese character, forming Middle Chinese phonetic information. Mandarin_ini, Mandarin_rym, Mandarin_ton, and Mandarin_syl represent the Mandarin initial consonant, final vowel, tone, and syllable information of the corresponding Chinese character, forming Mandarin phonetic information. Target_ini, Target_rym, Target_ton, and Target_syl represent the dialect initial consonant, final vowel, tone, and syllable information of the corresponding Chinese character, forming dialect phonetic information.

[0107] Table 1

[0108]

[0109] Step 42: Based on the syllable and tone information in the Mandarin phonetic information of words not covered by existing dictionary resources and words with missing pronunciations, as well as the syllable and tone information in the Middle Chinese phonetic information, statistical analysis is performed to obtain the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

[0110] After obtaining the Chinese character pronunciation table, the syllable and tone information in the Standard Mandarin phonetic information can be used to search existing resources to determine the phoneme sequence information of the dialect corresponding to the most frequently occurring syllable. Additionally, the syllable and tone information in the Middle Chinese phonetic information can be used to search existing resources to determine the phoneme sequence information of the dialect corresponding to the most frequently occurring syllable.

[0111] Step 43: Based on the phoneme sequence information corresponding to the most frequently occurring syllable, obtain the phoneme sequence information of words that are not covered by the existing dictionary resources and whose single-character pronunciation is missing.

[0112] In one embodiment, the phoneme sequence information of the dialect corresponding to the most frequent syllable can be directly used as the final phoneme sequence information of the word.

[0113] In another embodiment, the phoneme sequence information corresponding to the most frequently occurring syllable is obtained, and the final phoneme sequence information of the word can also be determined by manual verification.

[0114] Specifically, when determining the final phoneme sequence information of the single character through manual proofreading, statistical analysis can be performed based on the initial consonant information in the Mandarin phonetic notation information of words not covered by the existing character dictionary resources and the initial consonant information in the Middle Chinese phonetic notation information to obtain the phoneme sequence information corresponding to the initial consonant with the highest occurrence frequency. Additionally, statistical analysis can be carried out based on the final vowel information and tone information in the Mandarin phonetic notation information of words not covered by the existing character dictionary resources and the final vowel information and tone information in the Middle Chinese phonetic notation information to obtain the phoneme sequence information corresponding to the final vowel with the highest occurrence frequency. Finally, manual proofreading is performed based on the phoneme sequence information corresponding to the syllable, the phoneme sequence information corresponding to the initial consonant, and the phoneme sequence information corresponding to the final vowel to determine the phoneme sequence information of the entry of the word not covered by the existing character dictionary resources.

[0115] In Table 2, Mandarin_syn represents Mandarin syllable information. Middle Chinese phonology represents Middle Chinese syllable information. Target_Possible_syl represents the phoneme sequence information of the dialect corresponding to the syllable with the highest occurrence frequency. The data within () in the column where Target_Possible_syl is located indicates the frequency of occurrence of the phoneme sequence. Statistics & Sorting refers to the result of the occurrence frequencies of each syllable from high to low. Target_Possible_ini represents the phoneme sequence information of the dialect corresponding to the initial consonant with the highest occurrence frequency. Target_Possible_rym represents the phoneme sequence information of the dialect corresponding to the final vowel with the highest occurrence frequency.

[0116] For example, referring to Table 2, for the OOV character "彼" with missing pronunciation of a single character, its Mandarin syllable information is "bi3", its Middle Chinese syllable information is "止开三帮上支B", and the phoneme sequence information of the dialect corresponding to the syllable with the highest occurrence frequency is The phoneme sequence information of the dialect corresponding to the initial consonant with the highest occurrence frequency is "p(102), f(32), p‘(14), m(1)", and the phoneme sequence information of the dialect corresponding to the final vowel with the highest occurrence frequency is

[0117] Table 2

[0118]

[0119] In specific implementation, for all the inference results of the single characters with missing pronunciation in the OOV words obtained through the statistical analysis method, native speakers of the dialect can be organized for pronunciation proofreading. During proofreading, priority is given to proofreading whether the syllable inference result (i.e., the phoneme sequence information of the dialect corresponding to the syllable with the highest occurrence frequency) is correct. If there is no correct syllable phonetic notation information, it can be further determined whether the phoneme sequence information of the dialect corresponding to the initial consonant with the highest occurrence frequency is correct, and whether the phoneme sequence information of the dialect corresponding to the final vowel with the highest occurrence frequency is correct.

[0120] Using the above statistical methods, multiple inference results can be provided for manual proofreading. Compared with manual blind proofreading, this not only reduces the difficulty of proofreading but also improves pronunciation accuracy and efficiency.

[0121] In practical implementation, after obtaining the phoneme sequence information, in order to reduce the generation of erroneous or non-authentic Mandarin phoneme sequences in the dialect speech synthesis system during the text analysis stage, the word entries in the target dialect engine's pronunciation dictionary can be stored in the following format:

[0122] Lexical literal meaning: phoneme sequence: segmentation cost; number of syllables; part of speech; source tags; (dialectal idiomatic expression of lexical literal meaning;)

[0123] The collected and proofread information is organized according to the above-mentioned dictionary entry storage format. For example, the Mandarin entry format of the Minnan language engine pronunciation dictionary is "eyes: gan3 tsing1:102;2;n;ZhCn; eyeball", and the dialect entry format is "eyeball: bag8tsiu1:97;2;n;MiCn;".

[0124] By employing the phonetic annotation method in this embodiment of the invention, since the target dialect engine's pronunciation dictionary distinguishes between Mandarin entries and dialect entries, and ensures that at least some Mandarin entries include index information of corresponding dialect entries, when a corresponding dialect entry exists in the first word face to be annotated, the phoneme sequence information in the corresponding dialect entry can be directly obtained as the phoneme sequence information assigned to the word to be annotated. Thus, when there are entries in the word to be annotated that require dialectal authentic expression conversion, an authentic phoneme sequence can be assigned to the word to be annotated, thereby improving the speech synthesis results.

[0125] To enable those skilled in the art to better understand and implement the present invention, the apparatus, system, electronic device and computer-readable storage medium corresponding to the above method are described in detail below.

[0126] Reference Figure 5 This invention also provides a phonetic transcription device 50, which may include: a query unit 51, a determination unit 52, and a phoneme sequence output unit 53. Wherein:

[0127] The query unit 51 is adapted to query the target dialect engine pronunciation dictionary to obtain a first word surface that matches the word to be transcribed when the word to be transcribed is a Mandarin non-polyphonic word; the target dialect engine pronunciation dictionary includes: Mandarin words and dialect words, and at least some of the Mandarin words include index information of the corresponding dialect words;

[0128] The determining unit 52 is adapted to determine whether there is a corresponding dialect entry for the Mandarin entry containing the first word;

[0129] The phoneme sequence output unit 53 is adapted to acquire the phoneme sequence information in the corresponding dialect entry when a corresponding dialect entry exists, and output it as the phoneme sequence information assigned to the entry to be annotated.

[0130] This invention also provides a speech synthesis system, which includes the aforementioned phonetic transcription device 50.

[0131] Figure 6 This is a schematic diagram illustrating the workflow of the speech synthesis system. (Refer to...) Figure 6 The specific workflow of the speech synthesis system is as follows:

[0132] First, after receiving Mandarin text, the speech synthesis system can perform text regularization on the text. Text regularization involves converting written text words into spoken words to eliminate ambiguity in the pronunciation of non-standard words.

[0133] After text regularization, the speech synthesis system can perform word segmentation, part-of-speech prediction, and prosodic prediction on the results of text regularization.

[0134] Next, the speech synthesis system determines whether the word, after processing such as word segmentation, part-of-speech prediction, and prosody prediction, is a polyphonic word. If it is a polyphonic word, the speech synthesis system can use the model to predict the pronunciation of the polyphonic word; otherwise, it queries the target dialect engine's pronunciation dictionary to match the pronunciation corresponding to the same word.

[0135] By querying the target dialect engine's pronunciation dictionary, it can be determined whether the Mandarin text has a corresponding dialect entry. If it does, the system jumps to the corresponding dialect entry to perform a pronunciation query and extract the pronunciation. The extracted information is then used to replace the original phoneme sequence information of non-polyphonic Mandarin words, transforming the Mandarin entry into an authentic dialect expression. Otherwise, the system queries the corresponding phoneme sequence from the corresponding Mandarin entry and then outputs the phoneme sequence.

[0136] Subsequent speech synthesis systems can perform further processing based on the output phoneme sequence, such as acoustic model and vocoder processing.

[0137] This invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of any of the above methods.

[0138] In specific implementations, the computer-readable storage medium may include ROM, RAM, disk, or optical disk, etc.

[0139] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor runs the computer program, it performs the steps of any of the methods described above.

[0140] Regarding the modules / units included in the various devices and products described in the above embodiments, they can be software modules / units, hardware modules / units, or a combination of both. For example, for various devices and products applied to or integrated into a chip, all of their modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs that run on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits; for various devices and products applied to or integrated into a chip module, all of their modules / units can be implemented using hardware methods such as circuits, and different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The components can be implemented using software programs that run on the processor integrated within the chip module. The remaining (if any) modules / units can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into the terminal, each of its components / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or in different components within the terminal. Alternatively, at least some modules / units can be implemented using software programs that run on the processor integrated within the terminal, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits.

[0141] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.< / micn> < / zhcn> < / micn> < / zhcn>

Claims

1. A method for phonetic notation, characterized in that, include: When the word to be annotated is a non-polyphonic word in Mandarin, the target dialect engine pronunciation dictionary is queried to obtain the first word surface that matches the word to be annotated. The target dialect engine pronunciation dictionary includes: Mandarin entries and dialect entries, with at least some of the Mandarin entries including index information of the corresponding dialect entries; Determine whether the Mandarin entry containing the first word exists as a corresponding dialect entry; When a corresponding dialect entry exists, the phoneme sequence information in the corresponding dialect entry is obtained and used as the phoneme sequence information assigned to the entry to be annotated and output. Among them, the phoneme sequence information of OOV words in the target dialect engine pronunciation dictionary is obtained by statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the word in the entry, and the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

2. The phonetic notation method as described in claim 1, characterized in that, Also includes: When there is no corresponding dialect entry, the phoneme sequence information in the Mandarin entry where the first word is located is obtained and used as the phoneme sequence information assigned to the word to be annotated and output.

3. The phonetic notation method as described in claim 1, characterized in that, Also includes: When the word to be annotated is a dialect word that is not a polyphonic word, the target dialect engine pronunciation dictionary is queried to obtain a second word face that matches the word to be annotated. Obtain the phoneme sequence information in the dialect entry where the second word is located, and use it as the phoneme sequence information assigned to the word to be annotated, and output it.

4. The phonetic notation method according to any one of claims 1 to 3, characterized in that, The target dialect engine pronunciation dictionary was constructed using the following method: Based on existing dictionary resources, the word information of words and phrases already covered by the existing dictionary resources in the target dialect is obtained; Identify word entries that are not covered by existing dictionary resources.

5. The phonetic notation method as described in claim 4, characterized in that, The process of determining the word entry information for words not covered by existing dictionary resources includes: obtaining phoneme sequence information for words not covered by existing dictionary resources but whose single-character pronunciation is not missing through manual screening; obtaining phoneme sequence information for words not covered by existing dictionary resources but whose single-character pronunciation is missing through statistical analysis based on Mandarin phonetic information and dialect phonetic information; and determining other word entry information besides phoneme sequence information to obtain complete word entry information for words not covered by existing dictionary resources.

6. The phonetic notation method as described in claim 5, characterized in that, The statistical analysis based on Mandarin phonetic information and dialect phonetic information yields phoneme sequence information for words not covered by existing dictionary resources and with missing pronunciations, including: Generate a Chinese character pronunciation table for words not covered by the existing dictionary resources and whose single-character pronunciation is missing. The Chinese character pronunciation table includes dialect pronunciation information, Mandarin pronunciation information, and Middle Chinese pronunciation information for each word not covered by the existing dictionary resources; the pronunciation information includes syllable information and tone information. Based on the syllable and tone information in the Mandarin phonetic information of words not covered by existing dictionary resources and words with missing pronunciations, as well as the syllable and tone information in the Middle Chinese phonetic information, statistical analysis was conducted to obtain the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence. Based on the phoneme sequence information corresponding to the most frequently occurring syllable, the phoneme sequence information of words not covered by the existing dictionary resources and whose single-character pronunciation is missing is obtained.

7. The phonetic notation method as described in claim 6, characterized in that, The phonetic information also includes: initial consonant information and final vowel information; The step of obtaining phoneme sequence information for words not covered by existing dictionary resources and with missing pronunciations based on the phoneme sequence information corresponding to the most frequently occurring syllable also includes: Statistical analysis was conducted on the initial consonant information in Mandarin phonetic information and Middle Chinese phonetic information that are not covered by existing dictionary resources to obtain the phoneme sequence information corresponding to the most frequently occurring initial consonant. Based on the final and tone information in the Mandarin phonetic information of words not covered by existing dictionary resources, and the final and tone information in the Middle Chinese phonetic information, statistical analysis was conducted to obtain the phoneme sequence information corresponding to the most frequently occurring final. The phoneme sequence information of the syllables, the initial consonant, and the final vowel is manually checked to determine the phoneme sequence information of words not covered by the existing dictionary resources.

8. A phonetic notation device, characterized in that, include: The query unit is adapted to query the pronunciation dictionary of the target dialect engine when the word to be annotated is a non-polyphonic word in Mandarin, and obtain a first word surface that matches the word to be annotated. The target dialect engine pronunciation dictionary includes: Mandarin entries and dialect entries, with at least some of the Mandarin entries including index information of the corresponding dialect entries; The determining unit is adapted to determine whether there is a corresponding dialect entry for the Mandarin entry containing the first word; The phoneme sequence output unit is adapted to, when a corresponding dialect entry exists, acquire the phoneme sequence information in the corresponding dialect entry, use it as the phoneme sequence information assigned to the entry to be annotated, and output it. Among them, the phoneme sequence information of OOV words in the target dialect engine pronunciation dictionary is obtained by statistical analysis of the Mandarin syllable information and Middle Chinese phonological information of the word in the entry, and the phoneme sequence information corresponding to the syllable with the highest frequency of occurrence.

9. A speech synthesis system, characterized in that, Includes the phonetic device as described in claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the steps of the method according to any one of claims 1 to 7.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, characterized in that, When the processor runs the computer program, it performs the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Chinese language learning device

    JP1994242717A