Phonetic Code Marking Voiceprint Stitching Encoding Method and Its Phonetic Code
Through the splicing and coding method of vocal code marking, the problem that the existing pronunciation synthesis system cannot fully express the tone and tone is solved, and rich pronunciation expression and naturalness are achieved, and personalized pronunciation correction is supported.
Patent Information
- Application Number
- CN202211439181.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The existing pronunciation synthesis system cannot fully express the tone, tone, length of the original creator of the text, resulting in unnatural pronunciation, especially when rich pronunciation expression needs are needed.
The splicing and coding method of voice-print marking is adopted. By collecting phonetic materials, identifying Chinese characters, numbers, punctuation, pitch, and length, generating voice codes and establishing a voice library, combining six-key encoding and annotation, enriching the pronunciation and expression of the voice library.
It realizes the full expression of different pronunciations of each character, improves the naturalness of pronunciation, especially by marking the relationship between the front and back sounds, enhances the naturalness of pronunciation of new words and common phrases, and supports personalized pronunciation correction.
Smart Images

Figure CN115798454B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software, and particularly to speech synthesis technology. Background Art
[0002] Currently, the mainstream speech synthesis systems on the market are based on text-to-speech technology (referred to as TTS, Text To Speech), which requires preparing a piece of text first and then converting this text into speech, such as the speech synthesis technology of iFlytek. Due to the limitation of the text information volume, it is impossible to express the original intention of the text originator, that is, the author's intention, such as tone, intonation, pitch length, pitch height, etc. In other words, if you are not satisfied with some words in the speech synthesis and want to change to a more appropriate voice, the mainstream speech synthesis systems currently cannot achieve this function.
[0003] The patent applied by the applicant of this patent, with the application number 202010880919.X and the name "Speech Synthesis Method and Device", provides a three-key tone code encoding technology, which adds tone codes to make the speech more rich. However, when it is necessary to broadcast a news release or read a book, the three-key tone code encoding scheme cannot meet the rich speech expression requirements. Summary of the Invention
[0004] The purpose of the present invention is to provide a tone code marked voiceprint splicing encoding method to solve the above technical problems.
[0005] The purpose of the present invention is also to provide a tone code for the tone code marked voiceprint splicing encoding method.
[0006] The technical problems solved by the present invention can be realized by the following technical solutions:
[0007] The tone code marked voiceprint splicing encoding method includes a tone code recognition encoding method, and the tone code recognition encoding method includes the following steps:
[0008] Step 1), collect the voice materials of a person;
[0009] Step 2), recognize Chinese characters, numbers, punctuation marks, pitch height, and pitch length in the voice materials;
[0010] Step 3), generate tone codes for Chinese characters, numbers, punctuation marks, pitch height, and pitch length, and generate audio files corresponding to the tone codes after associating Chinese characters, numbers, punctuation marks with pitch height and pitch length;
[0011] Step 4), establish a voice library according to the generated tone codes and audio files.
[0012] The tone code marking voiceprint splicing coding method further includes a method for annotating tone codes for speech synthesis text. The method for annotating tone codes for speech synthesis text is to retrieve the audio file corresponding to the selected tone code.
[0013] In step 1), it is preferable to obtain the speech materials in the same recording facility environment.
[0014] In step 2), it may further include identifying the positions of Chinese characters, numbers in the speech materials. The errors or inaccuracies in Chinese characters, numbers, punctuation marks, pitch, duration, and positions can be proofread and corrected through the artificial listening and playback method.
[0015] When step 2) identifies the positions of Chinese characters, numbers, and punctuation marks, step 3) also generates tone codes for the positions of Chinese characters, numbers, and punctuation marks.
[0016] In addition, the tone code marking voiceprint splicing coding method can also automatically match the tone codes of the speech library and synthesize the corresponding speech according to the text file of a new book. The synthesized speech can also be corrected through artificial listening and playback. For the pronunciations that do not fully express the author's intention, tone codes are marked for individual words, and more appropriate speech is retrieved to change the pronunciation method.
[0017] For the new words that often appear in the new book, the tone code marking voiceprint splicing coding method uses a word-formation system to generate a large number of new word pronunciations to replace the unnatural pronunciations synthesized automatically. The tone code marking voiceprint splicing coding method can automatically learn and modify the annotated tone codes and word formation through artificial intelligence.
[0018] The tone code used in the tone code marking voiceprint splicing coding method is characterized in that it includes a first tone code for annotating initials. The first tone code includes the initials "ch", "zh", "sh". Among them, the initial "ch" is associated with the U key on the keyboard, the initial "zh" is associated with the I key on the keyboard, and the initial "sh" is associated with the V key on the keyboard.
[0019] The phonetic code marking voiceprint splicing encoding method uses a phonetic code, and further includes a second phonetic code for marking the final, where the second phonetic code includes finals "in", "ou", "ing", "ong", "iong", "ue", "ve", "uai", "uo", "ie", "iu", "ang", "ao", "eng", "ei", "ia", "ua", "ian", "iang", "uang", "un", "uan", "an", "ui", "ai", "en", "iao". Among them, the final "in" is associated with the Q key on the keyboard, the final "ou" is associated with the W key on the keyboard, the final "ing" is associated with the R key on the keyboard, the finals "ong" and "iong" are associated with the T key on the keyboard, the finals "ue", "ve", and "uai" are associated with the Y key on the keyboard, the final "uo" is associated with the O key on the keyboard, the final "ie" is associated with the P key on the keyboard, the final "iu" is associated with the S key on the keyboard, the final "ang" is associated with the D key on the keyboard, the final "ao" is associated with the F key on the keyboard, the final "eng" is associated with the G key on the keyboard, the final "ei" is associated with the H key on the keyboard, the finals "ia" and "ua" are associated with the J key on the keyboard, the final "ian" is associated with the K key on the keyboard, the finals "iang" and "uang" are associated with the L key on the keyboard, the final "un" is associated with the A key on the keyboard, the final "uan" is associated with the X key on the keyboard, the final "an" is associated with the C key on the keyboard, the final "ui" is associated with the V key on the keyboard, the final "ai" is associated with the B key on the keyboard, the final "en" is associated with the N key on the keyboard, and the final "iao" is associated with the M key on the keyboard.
[0020] The tone code used in the tone code marked voiceprint splicing coding method further includes a third tone code for marking tones. The third tone code has five zones, namely the first tone zone, the second tone zone, the third tone zone, the fourth tone zone, and the light tone zone. Each zone includes the first tone mark "ˉ", the second tone mark "ˊ", the third tone mark "ˇ", the fourth tone mark "ˋ", and the light tone mark "˙". Among them, the first tone mark "ˉ" in the first tone zone is associated with the U key on the keyboard, the second tone mark "ˊ" is associated with the I key on the keyboard, the third tone mark "ˇ" is associated with the O key on the keyboard, the fourth tone mark "ˋ" is associated with the P key on the keyboard, and the light tone mark "˙" is associated with the Y key on the keyboard; the first tone mark "ˉ" in the second tone zone is associated with the J key on the keyboard, the second tone mark "ˊ" is associated with the K key on the keyboard, the third tone mark "ˇ" is associated with the L key on the keyboard, the fourth tone mark "ˋ" is associated with the M key on the keyboard, and the light tone mark "˙" is associated with the H key on the keyboard; the first tone mark "ˉ" in the third tone zone is associated with the R key on the keyboard, the second tone mark "ˊ" is associated with the E key on the keyboard, the third tone mark "ˇ" is associated with the W key on the keyboard, the fourth tone mark "ˋ" is associated with the Q key on the keyboard, and the light tone mark "˙" is associated with the T key on the keyboard; the first tone mark "ˉ" in the fourth tone zone is associated with the F key on the keyboard, the second tone mark "ˊ" is associated with the D key on the keyboard, the third tone mark "ˇ" is associated with the S key on the keyboard, the fourth tone mark "ˋ" is associated with the A key on the keyboard, and the light tone mark "˙" is associated with the G key on the keyboard; the first tone mark "ˉ" in the light tone zone is associated with the V key on the keyboard, the second tone mark "ˊ" is associated with the C key on the keyboard, the third tone mark "ˇ" is associated with the X key on the keyboard, the fourth tone mark "ˋ" is associated with the Z key on the keyboard, and the light tone mark "˙" is associated with the B key on the keyboard.
[0021] The third tone code further includes a sentence start zone. The first tone mark "ˉ" in the sentence start zone is associated with the number 1 key on the keyboard, the second tone mark "ˊ" is associated with the number 2 key on the keyboard, the third tone mark "ˇ" is associated with the number 3 key on the keyboard, the fourth tone mark "ˋ" is associated with the number 4 key on the keyboard, and the light tone mark "˙" is associated with the number 0 key on the keyboard.
[0022] The tone code used in the tone code marked voiceprint splicing coding method further includes a fourth tone code for marking the nature of the previous tone. If the previous tone is a Chinese character, the final sound of this Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked. Commonly used punctuation marks are:,,?,!, / , :, etc.
[0023] The tone code used in the tone code marked voiceprint splicing coding method further includes a fifth tone code for marking the nature of the next tone. If the next tone is a Chinese character, the final sound of this Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked. Commonly used punctuation marks are:,,?,!, / , :, etc.
[0024] The phonetic code marking voiceprint splicing encoding method uses a phonetic code, and also includes a sixth phonetic code for marking pitch and duration. The sixth phonetic code has five regions, namely the ultra-low pitch region, the low pitch region, the middle pitch region, the high pitch region, and the ultra-high pitch region. Each region includes an ultra-short sound "-2", a short sound "-1", a middle sound "0", a long sound "+1", and an ultra-long sound "+2". Among them, the ultra-short sound "-2" in the ultra-low pitch region is associated with the Z key on the keyboard, the short sound "-1" is associated with the X key on the keyboard, the middle sound "0" is associated with the C key on the keyboard, the long sound "+1" is associated with the V key on the keyboard, and the ultra-long sound "+2" is associated with the B key on the keyboard; the ultra-short sound "-2" in the low pitch region is associated with the H key on the keyboard, the short sound "-1" is associated with the J key on the keyboard, the middle sound "0" is associated with the K key on the keyboard, the long sound "+1" is associated with the L key on the keyboard, and the ultra-long sound "+2" is associated with the M key on the keyboard; the ultra-short sound "-2" in the middle pitch region is associated with the A key on the keyboard, the short sound "-1" is associated with the S key on the keyboard, the middle sound "0" is associated with the D key on the keyboard, the long sound "+1" is associated with the F key on the keyboard, and the ultra-long sound "+2" is associated with the G key on the keyboard; the ultra-short sound "-2" in the high pitch region is associated with the Q key on the keyboard, the short sound "-1" is associated with the W key on the keyboard, the middle sound "0" is associated with the E key on the keyboard, the long sound "+1" is associated with the R key on the keyboard, and the ultra-long sound "+2" is associated with the T key on the keyboard; the ultra-short sound "-2" in the ultra-high pitch region is associated with the Y key on the keyboard, the short sound "-1" is associated with the U key on the keyboard, the middle sound "0" is associated with the I key on the keyboard, the long sound "+1" is associated with the O key on the keyboard, and the ultra-long sound "+2" is associated with the P key on the keyboard.
[0025] Beneficial effects: In mainstream speech synthesis systems, for single characters that are difficult to form phrases but are commonly used, the naturalness is not high, such as numbers (1, 2, 3), conjunctions (and, with), etc. Because these single character sounds often form liaisons with the surrounding speech in actual use, but do not form common phrases with other characters, so it is impossible to build a speech library and perform speech synthesis in the form of phrases. Basically, they are used alone and cannot form liaisons with the context, making them sound unnatural. When building the speech library in this invention, through the marking of six-key encoding, each Chinese character can be saved and used with more than 112,500 different pronunciations (5 types of the third key * more than 30 types of the fourth key * more than 30 types of the fifth key * 25 types of the sixth key = 5 * 30 * 30 * 25 = 112,500), which is sufficient to fully express the different pronunciations of each character. In particular, this invention marks the final sound of the previous pronunciation and the initial sound of the next pronunciation of these words; when in use, still retrieving the same previous final sound and next initial sound can achieve better naturalness. Description of the Drawings
[0026] Figure 1 It is a layout diagram of the position of the first phonetic code on the keyboard;
[0027] Figure 2 It is the layout diagram of the position of the second tone code on the keyboard;
[0028] Figure 3 It is the layout diagram of the position of the third tone code on the keyboard;
[0029] Figure 4 It is the layout diagram of the position of the sixth tone code on the keyboard. Specific implementation manners
[0030] In order to make the technical means, creative features, achieved purposes and functions realized by the present invention easy to understand, the present invention will be further described below in conjunction with specific drawings.
[0031] The tone code marked voiceprint splicing coding method includes a tone code recognition coding method and a voice synthesis text annotation tone code method.
[0032] The tone code recognition coding method includes the following steps:
[0033] Step 1), collect the voice materials of a person;
[0034] Step 2), recognize Chinese characters, numbers, punctuation marks, pitch and duration in the voice materials;
[0035] Step 3), generate tone codes from Chinese characters, numbers, punctuation marks, pitch and duration, and generate an audio file corresponding to the tone code after associating Chinese characters, numbers, punctuation marks with pitch and duration;
[0036] Step 4), establish a voice library according to the generated tone codes and audio files.
[0037] In step 1), it is preferable to obtain the voice materials in the same recording facility environment. In step 2), it may further include recognizing the positions of Chinese characters, numbers and punctuation marks in the voice materials. The errors or inaccurate Chinese characters, numbers, punctuation marks, pitch, duration and positions can be proofread and corrected by the artificial listening and broadcasting method.
[0038] The tone code marked voiceprint splicing coding method further includes a voice synthesis text annotation tone code method, and the voice synthesis text annotation tone code method is to retrieve the audio file corresponding to the selected tone code.
[0039] When the positions of Chinese characters, numbers and punctuation marks are recognized in step 2), the positions of Chinese characters and numbers are also generated into tone codes in step 3).
[0040] In addition, the tone code marked voiceprint splicing coding method can also automatically match the tone codes in the voice library and synthesize the corresponding voice according to the text file of a new book. The synthesized voice can also be obtained by artificial listening and broadcasting. The pronunciation that does not fully express the author's intention can be corrected, and tone codes can be marked for individual words, thereby changing the pronunciation mode.
[0041] The phonetic code marking voiceprint splicing coding method generates a large number of new word pronunciations using a word formation system for new words frequently appearing in new books, replacing the unnatural pronunciations synthesized automatically. The phonetic code marking voiceprint splicing coding method allows artificial intelligence to automatically learn and manually modify the marked phonetic codes and form new words.
[0042] The phonetic code marking voiceprint splicing coding method uses phonetic codes, including a first phonetic code for marking initials. The first phonetic code includes initials "ch", "zh", and "sh". Among them, the initial "ch" is associated with the U key on the keyboard, the initial "zh" is associated with the I key on the keyboard, and the initial "sh" is associated with the V key on the keyboard, as Figure 1 shown.
[0043] The phonetic code marking voiceprint splicing coding method also uses phonetic codes, including a second phonetic code for marking finals. The second phonetic code includes finals "in", "ou", "ing", "ong", "iong", "ue", "ve", "uai", "uo", "ie", "iu", "ang", "ao", "eng", "ei", "ia", "ua", "ian", "iang", "uang", "un", "uan", "an", "ui", "ai", "en", "iao". Among them, the final "in" is associated with the Q key on the keyboard, the final "ou" is associated with the W key on the keyboard, the final "ing" is associated with the R key on the keyboard, the finals "ong" and "iong" are associated with the T key on the keyboard, the finals "ue", "ve", and "uai" are associated with the Y key on the keyboard, the final "uo" is associated with the O key on the keyboard, the final "ie" is associated with the P key on the keyboard, the final "iu" is associated with the S key on the keyboard, the final "ang" is associated with the D key on the keyboard, the final "ao" is associated with the F key on the keyboard, the final "eng" is associated with the G key on the keyboard, the final "ei" is associated with the H key on the keyboard, the finals "ia" and "ua" are associated with the J key on the keyboard, the final "ian" is associated with the K key on the keyboard, the finals "iang" and "uang" are associated with the L key on the keyboard, the final "un" is associated with the A key on the keyboard, the final "uan" is associated with the X key on the keyboard, the final "an" is associated with the C key on the keyboard, the final "ui" is associated with the V key on the keyboard, the final "ai" is associated with the B key on the keyboard, the final "en" is associated with the N key on the keyboard, and the final "iao" is associated with the M key on the keyboard, as Figure 2As shown. The present invention optimizes position to position, and assigns one of the two compound vowels (ian / uai) represented by the K key (uai) to the Y key (ue / ve), so that the Y key contains three compound vowels (ue / ve / uai). After verification, no conflict problem will occur.
[0044] The tone code marking voiceprint splicing coding method uses tone codes, and further includes a third tone code for marking tones. Refer to Figure 3 , the third tone code has five zones, namely the first tone zone, the second tone zone, the third tone zone, the fourth tone zone, and the light tone zone. Each zone includes the first tone “ˉ”, the second tone “ˊ”, the third tone “ˇ”, the fourth tone “ˋ”, and the light tone “˙”. Among them, the first tone “ˉ” in the first tone zone is associated with the U key on the keyboard, the second tone “ˊ” is associated with the I key on the keyboard, the third tone “ˇ” is associated with the O key on the keyboard, the fourth tone “ˋ” is associated with the P key on the keyboard, and the light tone “˙” is associated with the Y key on the keyboard; the first tone “ˉ” in the second tone zone is associated with the J key on the keyboard, the second tone “ˊ” is associated with the K key on the keyboard, the third tone “ˇ” is associated with the L key on the keyboard, the fourth tone “ˋ” is associated with the M key on the keyboard, and the light tone “˙” is associated with the H key on the keyboard; the first tone “ˉ” in the third tone zone is associated with the R key on the keyboard, the second tone “ˊ” is associated with the E key on the keyboard, the third tone “ˇ” is associated with the W key on the keyboard, the fourth tone “ˋ” is associated with the Q key on the keyboard, and the light tone “˙” is associated with the T key on the keyboard; the first tone “ˉ” in the fourth tone zone is associated with the F key on the keyboard, the second tone “ˊ” is associated with the D key on the keyboard, the third tone “ˇ” is associated with the S key on the keyboard, the fourth tone “ˋ” is associated with the A key on the keyboard, and the light tone “˙” is associated with the G key on the keyboard; the first tone “ˉ” in the light tone zone is associated with the V key on the keyboard, the second tone “ˊ” is associated with the C key on the keyboard, the third tone “ˇ” is associated with the X key on the keyboard, the fourth tone “ˋ” is associated with the Z key on the keyboard, and the light tone “˙” is associated with the B key on the keyboard. The third tone code further includes a sentence-initial zone. The first tone “ˉ” in the sentence-initial zone is associated with the number 1 key on the keyboard, the second tone “ˊ” is associated with the number 2 key on the keyboard, the third tone “ˇ” is associated with the number 3 key on the keyboard, the fourth tone “ˋ” is associated with the number 4 key on the keyboard, and the light tone “˙” is associated with the number 0 key on the keyboard.
[0045] The phonetic code used in the phonetic code marked voiceprint splicing coding method further includes a fourth phonetic code for marking the nature of the previous sound. If the previous sound is a Chinese character, the final sound of this Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked. Commonly used punctuation marks are:,,?,!, / , :, etc. The phonetic code used in the phonetic code marked voiceprint splicing coding method further includes a fifth phonetic code for marking the nature of the next sound. If the next sound is a Chinese character, the final sound of this Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked. Commonly used punctuation marks are:,,?,!, / , :, etc. The fourth and fifth key coding schemes of the single-character phonetic code make the voice library of each single-character sound richer and provide more choices during voice synthesis. At the same time, recording the relationship between the previous and next sounds is also beneficial to improving the voice naturalness of creating new words.
[0046] The phonetic code used in the phonetic code marked voiceprint splicing coding method further includes a sixth phonetic code for marking pitch and duration. Refer to Figure 4 , the sixth phonetic code has five zones, namely the ultra-low pitch zone, the low pitch zone, the middle pitch zone, the high pitch zone, and the ultra-high pitch zone. Each zone includes an ultra-short sound "-2", a short sound "-1", a middle sound "0", a long sound "+1", and an ultra-long sound "+2". Among them, the ultra-short sound "-2" in the ultra-low pitch zone is associated with the Z key on the keyboard, the short sound "-1" is associated with the X key on the keyboard, the middle sound "0" is associated with the C key on the keyboard, the long sound "+1" is associated with the V key on the keyboard, and the ultra-long sound "+2" is associated with the B key on the keyboard; the ultra-short sound "-2" in the low pitch zone is associated with the H key on the keyboard, the short sound "-1" is associated with the J key on the keyboard, the middle sound "0" is associated with the K key on the keyboard, the long sound "+1" is associated with the L key on the keyboard, and the ultra-long sound "+2" is associated with the M key on the keyboard; the ultra-short sound "-2" in the middle pitch zone is associated with the A key on the keyboard, the short sound "-1" is associated with the S key on the keyboard, the middle sound "0" is associated with the D key on the keyboard, the long sound "+1" is associated with the F key on the keyboard, and the ultra-long sound "+2" is associated with the G key on the keyboard; the ultra-short sound "-2" in the high pitch zone is associated with the Q key on the keyboard, the short sound "-1" is associated with the W key on the keyboard, the middle sound "0" is associated with the E key on the keyboard, the long sound "+1" is associated with the R key on the keyboard, and the ultra-long sound "+2" is associated with the T key on the keyboard; the ultra-short sound "-2" in the ultra-high pitch zone is associated with the Y key on the keyboard, the short sound "-1" is associated with the U key on the keyboard, the middle sound "0" is associated with the I key on the keyboard, the long sound "+1" is associated with the O key on the keyboard, and the ultra-long sound "+2" is associated with the P key on the keyboard. Specific embodiments
[0048] For example, the word "I" in the sentence "I am Chinese" can be marked with the phonetic code: wo3.v(az), where the first two letters wo represent the initial consonant and final vowel of the word; 3 represents the third tone at the beginning of the sentence; . represents the beginning of the sentence; v represents the initial consonant of the following word; (az) is the letter determined by the machine based on the encoding rules of the sixth key and the pitch and length of the actual pronunciation.
[0049] The word "是" can be marked with the phonetic code: viqoi(az), where vi is the initial consonant and final vowel of the word; q is the letter determined by the third tone of the previous word and the fourth tone of the current word; o is the final vowel of the previous word; i is the initial consonant of the next word; (az) is the letter determined by the machine based on the encoding rules of the sixth key and the pitch and length of the actual pronunciation.
[0050] We won't use the middle two characters of this sentence as examples anymore. Let's take the character "人" (person) at the end of the sentence. It can be annotated with the phonetic code: rnko.(az). Here, rn represents the initial and final consonants of the character; k is the letter determined by the third tone of the previous character and the fourth tone of the current character; o is the final vowel of the previous character; . is the punctuation mark at the end of the sentence; and (az) is the letter determined by the machine based on the encoding rules of the sixth key and the actual pitch and duration of the pronunciation.
[0051] The encoding scheme for speech recognition phonetic codes for two-word and multi-word words is as follows:
[0052] Two-character words: The initial consonant of the first character + the final consonant of the second character + the final consonant of the second character + the combination of the tone of the first character and the tone of the second character (third key rule) + the final consonant of the previous sound + the initial consonant of the next sound + the combination of pitch and length (az). Still using the sentence "I am Chinese" as an example, there are three commonly used two-character words: "I am", "China", and "Guo Ren". The encoding is as follows:
[0053] "I am": woviq.i(az). Wo is the initial consonant and final vowel of the first character; vi is the initial consonant and final vowel of the second character; q is the letter determined by the third tone of the first character and the fourth tone of the second character; . represents the beginning of the sentence; i represents the initial consonant of the following character; (az) The 25 letters (except N) are determined by the pitch and duration of the two-character word.
[0054] "China": itgoiir (az). It represents the initial and final consonants of the first character; go represents the initial and final consonants of the second character; the fifth letter, i, is the letter determined by the third tone of the first character and the fourth tone of the second character; the sixth letter, i, is the final consonant of the previous character; r represents the initial consonant of the following character; and the remaining 25 letters (a, z) (except N) are determined by the pitch and duration of the two-character word.
[0055] "Chinese people.": gornkz.(a-z). Here, go is the initial and final sounds of the first character; rn is the initial and final sounds of the second character; k is the letter determined by the third tone of the first character and the fourth tone of the second character; z is the final sound of the previous character;. represents the full stop at the end of the sentence; (a-z), 25 letters (excluding N), are the letters determined by the pitch and length of the two-character word.
[0056] Three-character words: Initial sound + final sound of the first character + initial sound + final sound of the second character + initial sound + final sound of the third character + final sound of the previous sound + initial sound of the next sound + permutations and combinations of pitch value and length (a-z). Note that the tones between the three characters are not marked here because if the tones are not marked in two-character words, it is very easy to get confused, but for three-character words, not marking the tones is not likely to cause confusion. Still using the sentence "I am Chinese." above as an example, there is a commonly used three-character word: "Chinese people". The encoding is as follows:
[0057] itgorni.(a-z). Here, it is the initial and final sounds of the first character; go is the initial and final sounds of the second character; rn is the initial and final sounds of the third character; i is the final sound of the previous character;. represents the full stop at the end of the sentence; (a-z), 25 letters (excluding N), are the letters determined by the pitch and length of the two-character word.
[0058] The encoding rules for words with more than three characters are the same as those for three-character words.
[0059] Using this phonetic code annotation method can enable the machine to automatically recognize speech and annotate it as a phonetic code. When using it, that is, when synthesizing speech, the machine will also automatically match and call the appropriate speech according to the relationship between the previous and next sentences and the previous and next sounds. If a synthesized speech does not convey the meaning the author intends, it only needs to be annotated by the last code (a-z, 25 letters, excluding n) of the phonetic code of a certain word or phrase, thereby changing the machine's default pronunciation and performing manual intervention to replace it with a more appropriate pronunciation.
[0060] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for splicing and encoding a voiceprint with a phonetic code mark, including a method for identifying and encoding a phonetic code, which includes the following steps: Step 1) Collect someone's voice material; Step 2) Identify Chinese characters, numbers, punctuation marks, pitch, and duration in the speech material; identify the position of Chinese characters and numbers; Step 3) Generate Chinese characters, numbers, punctuation marks, pitch, and tone length into sound codes, and also generate the positions of Chinese characters and numbers into sound codes; and generate audio files corresponding to the sound codes after associating the Chinese characters, numbers, and punctuation marks with the pitch and tone length; Step 4) establishing a voice library based on the generated voice codes and audio files, wherein the voice codes in the voice library include a first voice code for marking initials, a second voice code for marking finals, and a third voice code for marking tones; The phonetic code also includes a fourth phonetic code that marks the nature of the previous sound. If the previous sound is a Chinese character, the vowel of the Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked; The phonetic code also includes a fifth phonetic code that marks the nature of the next phonetic sound. If the next phonetic sound is a Chinese character, the final vowel of the Chinese character is marked; if it is a number, the number is directly marked; if it is a punctuation mark, the punctuation mark is directly marked; The sound code also includes a sixth sound code that marks the pitch and length of the sound.
2. The method for concatenating and encoding phonetic code marks with voiceprints according to claim 1, characterized in that: In step 1), the voice materials are obtained in the same recording facility environment.
3. The method for concatenating and encoding phonetic code marks with voiceprints according to claim 1, characterized in that: It also includes a method for annotating phonetic codes in speech synthesis text, wherein the method is to retrieve the audio file corresponding to the selected phonetic code according to the selected phonetic code.
4. The method for concatenating and encoding phonetic code marks with voiceprints according to claim 1, characterized in that: The first sound code includes the initial consonant "ch", the initial consonant "zh", and the initial consonant "sh", wherein the initial consonant "ch" is associated with the U key on the keyboard, the initial consonant "zh" is associated with the I key on the keyboard, and the initial consonant "sh" is associated with the V key on the keyboard; The second phonetic code includes the vowel "in", the vowel "ou", the vowel "ing", the vowel "ong", the vowel "iong", the vowel "ue", the vowel "ve", the vowel "uai", the vowel "uo", the vowel "ie", the vowel "iu", the vowel "ang", the vowel "ao", the vowel "eng", the vowel "ei", the vowel "ia", the vowel "ua", the vowel "ian", the vowel "iang", the vowel "uang", the vowel "un", the vowel "uan", the vowel "an", the vowel "ui", the vowel "ai", the vowel "en", and the vowel "iao", wherein the vowel "in" is associated with the Q key on the keyboard, the vowel "ou" is associated with the W key on the keyboard, the vowel "ing" is associated with the R key on the keyboard, the vowels "ong" and "iong" are associated with the T key on the keyboard, the vowels "ue", "ve" and " ”, the final "uai" is associated with the Y key on the keyboard, the final "uo" is associated with the O key on the keyboard, the final "ie" is associated with the P key on the keyboard, the final "iu" is associated with the S key on the keyboard, the final "ang" is associated with the D key on the keyboard, the final "ao" is associated with the F key on the keyboard, the final "eng" is associated with the G key on the keyboard, the final "ei" is associated with the H key on the keyboard, the finals "ia" and "ua" are associated with the J key on the keyboard, the final "ian" is associated with the K key on the keyboard, the finals "iang" and "uang" are associated with the L key on the keyboard, the final "un" is associated with the A key on the keyboard, the final "uan" is associated with the X key on the keyboard, the final "an" is associated with the C key on the keyboard, the final "ui" is associated with the V key on the keyboard, the final "ai" is associated with the B key on the keyboard, the final "en" is associated with the N key on the keyboard, and the final "iao" is associated with the M key on the keyboard.
5. The method for encoding a voiceprint with a phonetic code mark according to claim 1, wherein: The third tone code has five areas, namely, a first tone area, a second tone area, a third tone area, a fourth tone area, and a light tone area, each area including a first tone "ˉ", a second tone "ˊ", a third tone "ˇ", a fourth tone "ˋ", and a light tone "˙". Among them, the first tone "ˉ" in the first tone area is associated with the U key on the keyboard, the second tone "ˊ" is associated with the I key on the keyboard, the third tone "ˇ" is associated with the O key on the keyboard, the fourth tone "ˋ" is associated with the P key on the keyboard, and the light tone "˙" is associated with the Y key on the keyboard; the first tone "ˉ" in the second tone area is associated with the J key on the keyboard, the second tone "ˊ" is associated with the K key on the keyboard, the third tone "ˇ" is associated with the L key on the keyboard, the fourth tone "ˋ" is associated with the M key on the keyboard, and the light tone "˙" is associated with the H key on the keyboard; The first tone "ˉ" in the three-tone zone is associated with the R key on the keyboard, the second tone "ˊ" is associated with the E key on the keyboard, the third tone "ˇ" is associated with the W key on the keyboard, the fourth tone "ˋ" is associated with the Q key on the keyboard, and the light tone "˙" is associated with the T key on the keyboard; the first tone "ˉ" in the four-tone zone is associated with the F key on the keyboard, the second tone "ˊ" is associated with the D key on the keyboard, the third tone "ˇ" is associated with the S key on the keyboard, the fourth tone "ˋ" is associated with the A key on the keyboard, and the light tone "˙" is associated with the G key on the keyboard; the first tone "ˉ" in the light tone zone is associated with the V key on the keyboard, the second tone "ˊ" is associated with the C key on the keyboard, the third tone "ˇ" is associated with the X key on the keyboard, the fourth tone "ˋ" is associated with the Z key on the keyboard, and the light tone "˙" is associated with the B key on the keyboard; The third sound code also includes a sentence start area, in which the first tone "ˉ" in the sentence start area is associated with the number 1 key on the keyboard, the second tone "ˊ" is associated with the number 2 key on the keyboard, the third tone "ˇ" is associated with the number 3 key on the keyboard, the fourth tone "ˋ" is associated with the number 4 key on the keyboard, and the light tone "˙" is associated with the number 0 key on the keyboard.
6. The method for concatenating and encoding phonetic code marks with voiceprints according to claim 1, characterized in that: The sixth tone code has five areas, namely, super bass area, bass area, middle area, treble area, and super treble area, each area includes super short tone "-2", short tone "-1", middle tone "0", long tone "+1", and super long tone "+2", among which, the super short tone "-2" in the super bass area is associated with the Z key on the keyboard, the short tone "-1" is associated with the X key on the keyboard, the middle tone "0" is associated with the C key on the keyboard, the long tone "+1" is associated with the V key on the keyboard, and the super long tone "+2" is associated with the B key on the keyboard; the super short tone "-2" in the bass area is associated with the H key on the keyboard, the short tone "-1" is associated with the J key on the keyboard, the middle tone "0" is associated with the K key on the keyboard, the long tone "+1" is associated with the L key on the keyboard, and the super long tone "+2" is associated with the M key on the keyboard. ; In the middle range, the ultra-short note "-2" is associated with the A key on the keyboard, the short note "-1" is associated with the S key on the keyboard, the middle note "0" is associated with the D key on the keyboard, the long note "+1" is associated with the F key on the keyboard, and the ultra-long note "+2" is associated with the G key on the keyboard; in the treble range, the ultra-short note "-2" is associated with the Q key on the keyboard, the short note "-1" is associated with the W key on the keyboard, the middle note "0" is associated with the E key on the keyboard, the long note "+1" is associated with the R key on the keyboard, and the ultra-long note "+2" is associated with the T key on the keyboard; in the ultra-treble range, the ultra-short note "-2" is associated with the Y key on the keyboard, the short note "-1" is associated with the U key on the keyboard, the middle note "0" is associated with the I key on the keyboard, the long note "+1" is associated with the O key on the keyboard, and the ultra-long note "+2" is associated with the P key on the keyboard.
Citation Information
Patent Citations
Speech synthesis method and device
CN112002304B
Speech synthesis method and device
CN112002304A
Imbedded voice synthesis method and system
CN1455386A