Speech synthesis methods, devices, electronic equipment and storage media

By acquiring training data for the target language and using a speech recognition system to identify phonemes and adjust the pronunciation dictionary based on tone, the problem of insufficient training data was solved, and high-quality speech synthesis for non-mainstream languages ​​was achieved.

CN115995226BActive Publication Date: 2026-04-03GUANGDONG ELECTRIC POWER COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing speech synthesis systems struggle to acquire sufficient training data for languages ​​with limited training data, such as dialects, resulting in poor speech synthesis quality that fails to meet usage requirements.

Method used

By acquiring training data of the target language, a speech recognition system is used to identify the phonemes and tones of text units, and the standard pronunciation dictionary is adjusted to generate the target pronunciation dictionary, thereby constructing a target speech synthesis system.

Benefits of technology

Rapidly construct a target speech synthesis system that meets the requirements, reduce the requirements for sample data, and improve the speech synthesis quality of non-mainstream languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115995226B_ABST
    Figure CN115995226B_ABST
Patent Text Reader

Abstract

This application provides a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: acquiring training data of a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; recognizing the training data based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes the phoneme and tone corresponding to the text unit; adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit to obtain a target pronunciation dictionary; and determining a target speech synthesis system based on the target pronunciation dictionary. This application obtains a target pronunciation dictionary by adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit, which can quickly obtain a target speech synthesis system that meets the requirements, and has low requirements for sample data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to speech synthesis methods, apparatus, electronic devices, and storage media. Background Technology

[0002] With the development of technology, speech synthesis technology has become more and more mature. Speech synthesis systems can convert text into speech, which is not only convenient and fast, but also saves human resources. Therefore, speech synthesis systems have been applied in more and more fields.

[0003] However, in order to produce high-quality speech synthesized by a speech synthesis system, a large amount of training data is required to train the system. When the speech type is a common language such as Mandarin, it is relatively easy to obtain training data. However, when the speech type is a less common language, such as various regional dialects, it is difficult to obtain enough training data to train the speech synthesis system. As a result, the quality of the speech synthesized by the system is poor and cannot meet the usage requirements. Summary of the Invention

[0004] In view of the above problems, this application proposes a speech synthesis method, apparatus, electronic device and storage medium to improve the above problems.

[0005] In a first aspect, embodiments of this application provide a speech synthesis method, the method comprising: acquiring training data of a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; recognizing the training data based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes phonemes and tones corresponding to the text units; adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text units to obtain a target pronunciation dictionary; and determining a target speech synthesis system based on the target pronunciation dictionary.

[0006] Secondly, embodiments of this application also provide a speech synthesis apparatus, comprising: an acquisition unit for acquiring training data of a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; a recognition unit for recognizing the training data based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes phonemes and tones corresponding to the text units; an adjustment unit for adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text units to obtain a target pronunciation dictionary; and a determination unit for determining a target speech synthesis system based on the target pronunciation dictionary.

[0007] Thirdly, embodiments of this application also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the speech synthesis method as described in the first aspect.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for enabling an electronic device to perform the speech synthesis method as described in the first aspect.

[0009] This application provides a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: acquiring training data of a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; recognizing the training data based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes the phoneme and tone corresponding to the text unit; adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit to obtain a target pronunciation dictionary; and determining a target speech synthesis system based on the target pronunciation dictionary. This application obtains a target pronunciation dictionary by adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit, which can quickly obtain a target speech synthesis system that meets the requirements, and has low requirements for sample data. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0011] Figure 1 This is a schematic flowchart of a speech synthesis method provided in an embodiment of this application.

[0012] Figure 2 yes Figure 1 A detailed flowchart of step 130 in the diagram.

[0013] Figure 3 yes Figure 1 A further detailed flowchart of step 130 in the process.

[0014] Figure 4 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of this application.

[0015] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0016] Figure 6 This is a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] With the development of technology, speech synthesis technology has become more and more mature. Speech synthesis systems can convert text into speech, which is not only convenient and fast, but also saves human resources. Therefore, speech synthesis systems have been applied in more and more fields.

[0019] For example, many movie commentary videos with AI (Artificial Intelligence) voice narration have appeared on various video websites, and many novels read aloud by AI have appeared in novel apps, greatly enriching people's material lives.

[0020] However, training a speech synthesis system places high demands on the training data. For example, it requires the same person to read a specified text in a native dialect, and the data must be captured using a high-definition microphone. The requirements for the speaker, the corpus content, and the acquisition equipment are far higher than those for speech recognition. The corpus needed to train a speech synthesis model, besides the speech read by a broadcaster and its corresponding text, also includes phoneme tables, pronunciation dictionaries, and pronunciation rules. This corpus is used to train the text processing front-end of the speech synthesis system, which then converts the text into speech.

[0021] However, as dialects are gradually assimilated into Standard Mandarin, it is becoming increasingly difficult to find corpus collectors with standard pronunciation for dialects with accents. Furthermore, these dialects differ from mainstream accents in their vocabulary habits, particle usage, and even have their own unique word sets. Compiling dictionaries and pronunciation rules for these dialects requires extensive research from professional linguists, but the number of scholars specializing in specific dialect accents is dwindling, especially for dialects without written characters, making the synthesis of dialects with accents extremely challenging.

[0022] Because it is difficult to obtain enough suitable training data to train speech synthesis systems, current speech synthesis systems are limited to mainstream standard languages ​​such as English, Mandarin, German, and French, resulting in a monotonous user experience and failing to meet user needs.

[0023] To address the aforementioned problems, the inventors have proposed a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: acquiring training data for a target language; wherein the training data includes speech information of the target language and corresponding text information; recognizing the training data using a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes the phonemes and tones corresponding to the text units; adjusting a standard pronunciation dictionary based on the target pronunciation sequences corresponding to the text units to obtain a target pronunciation dictionary; and determining a target speech synthesis system based on the target pronunciation dictionary. This application obtains a target pronunciation dictionary by adjusting a standard pronunciation dictionary using the target pronunciation sequences corresponding to text units. It can quickly obtain a target speech synthesis system that meets the requirements based on the target pronunciation dictionary, and has lower requirements for sample data.

[0024] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0025] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of a speech synthesis method provided in an embodiment of this application. For example... Figure 1 As shown, the method 100 includes steps 110 to 140.

[0026] Step 110: Obtain training data for the target language; wherein, the training data includes speech information of the target language and text information corresponding to the speech information.

[0027] In some implementations, the target language can be Sichuanese, Hakka, Mandarin, Cantonese, or other languages.

[0028] In some implementations, training data can be obtained from the Internet. The speech information in the training data can be relatively clear speech of the target language recorded by multiple speakers in non-specific scenarios.

[0029] In some implementations, the training data can be speech of the target language in movies and TV programs. Because the speaker and speech quality are uncontrollable, such speech cannot be directly used to train a speech synthesis system in related technologies. However, in this application, such speech can be used to obtain a target speech synthesis system. Moreover, such speech often comes with subtitles, which can quickly annotate the text content corresponding to the speech, saving the time of manually annotating the speech.

[0030] Step 120: Recognize the training data based on a speech recognition system to obtain multiple text units and the target pronunciation sequence corresponding to each text unit; wherein, the target pronunciation sequence includes the phonemes and tones corresponding to the text unit.

[0031] In some embodiments, the text information includes multiple text units, and each text unit corresponds to a certain segment of speech in the speech information. After recognizing the speech information based on the speech recognition system, the target pronunciation sequence of the speech corresponding to each text unit can be obtained.

[0032] In some embodiments, the text information can be segmented to divide text units. Exemplarily, the segmentation result of "I come to Beijing" can be "I / come / to Beijing", and the text units are respectively "I", "come", and "Beijing".

[0033] In some embodiments, the target pronunciation sequence includes the phonemes and tones corresponding to the text unit. Exemplarily, the target pronunciation sequence of the text unit "zhi" can be "zh i1", where 1 indicates the first tone.

[0034] Further, in order to improve the accuracy of speech recognition, the training data can be speech-recognized based on a standard speech recognition system.

[0035] Specifically, the standard speech recognition system is the speech recognition system corresponding to the standard language. Exemplarily, the standard language can be Mandarin. Because Mandarin has a high popularity, high-quality training data and is easy to obtain, the recognition level of the speech recognition system trained in Mandarin is high and the recognition effect is good.

[0036] However, it can be understood that if the speech recognition systems of other languages can reach a recognition level similar to or higher than that of the Mandarin speech recognition system, then the speech recognition systems of other languages can also be used. Therefore, this application does not limit the specific language of the standard language.

[0037] In some embodiments, a standard speech recognition system with a recognition accuracy higher than a preset recognition accuracy can be used for speech recognition. The preset recognition accuracy can be set according to needs, and this application does not limit it. Exemplarily, the preset recognition accuracy can be 95%.

[0038] Step 130: Adjust the standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit to obtain the target pronunciation dictionary.

[0039] Specifically, a standard pronunciation dictionary is a pronunciation dictionary corresponding to a standard language. The pronunciation dictionary includes at least multiple text units and the corresponding pronunciation sequence in the standard language for each text unit. Optionally, there are many ways to obtain a standard pronunciation dictionary, such as from open-source projects on the internet or academic institutions. It is understood that this application does not restrict the method of obtaining a standard pronunciation dictionary.

[0040] In some implementations, if a text unit obtained by recognizing training data based on a speech recognition system and the target pronunciation sequence corresponding to that text unit are different from the standard pronunciation sequence of that text unit in the standard pronunciation dictionary, then the pronunciation sequence of that text unit in the standard pronunciation dictionary is replaced by the target pronunciation sequence.

[0041] In some implementations, to improve the accuracy of pronunciation sequence replacement, step 130 includes steps 131 to 132.

[0042] Step 131: Obtain the first pronunciation probability and the second pronunciation probability of each text unit; wherein, the first pronunciation probability is the probability that the pronunciation sequence of the text unit in the target language is the target pronunciation sequence, and the second pronunciation probability is the probability that the pronunciation sequence of the text unit in the target language is the standard pronunciation sequence.

[0043] In some implementations, to avoid introducing incorrect pronunciations due to random errors in the speech recognition system of the standard language, a first pronunciation probability P(p|w) and a second pronunciation probability P(p) can be obtained for each text unit. ’ |w), where p is the pronunciation sequence of the text unit in the target language, p ’ Let w be the pronunciation sequence of a text unit in standard language, and P(p|w) represent the probability that the text unit is w and the pronunciation sequence is p in the training data. ’ |w) indicates that the text unit is w and the pronunciation sequence is p in the training data. ’ The probability of.

[0044] For example, in a standard language pronunciation dictionary, the text unit "supports" the pronunciation sequence p ’ The given value is "zh i1 chi2". However, after the speech recognition system identifies the training data of the target language, there are two possible pronunciation sequences for the text unit "support". The first pronunciation sequence is p. ’ Given the sequence "zh i1 ch i2" and the second pronunciation sequence "zh i1 ch i4", calculate the probability of the first pronunciation P(p|support) and the probability of the second pronunciation P(p|support) respectively. ’ |Support).

[0045] Step 132: If the probability of the first pronunciation is greater than the probability of the second pronunciation, then replace the standard pronunciation sequence corresponding to the text unit in the standard pronunciation dictionary with the target pronunciation sequence to obtain the target pronunciation dictionary.

[0046] Specifically, if the probability of the first pronunciation is greater than the probability of the second pronunciation, it proves that the pronunciation of the word in the target language is different from that in the standard language. The standard pronunciation sequence corresponding to the text unit in the standard pronunciation dictionary is replaced with the target pronunciation sequence. This process is repeated, and each text unit is detected and replaced.

[0047] For example, if the first pronunciation probability P(p|support) > the second pronunciation probability P(p) ’ If |support), then the pronunciation sequence of the text unit "support" in the pronunciation dictionary of the standard language will be changed from p ’ Replace with P.

[0048] By comparing the first and second pronunciation probabilities, we can avoid introducing incorrect pronunciations due to random errors in the speech recognition system of standard language, thereby improving the reliability of the obtained target pronunciation dictionary.

[0049] In some implementations, the pronunciation rules of the target language can be analyzed in depth, in which case step 130 includes steps 133 to 134.

[0050] Step 133: Obtain the mapping probability of each target pronunciation sequence; where the mapping probability is the probability that the pronunciation sequence in the standard language is a mapped text unit of the standard pronunciation sequence, and the pronunciation sequence in the target language is the target pronunciation sequence.

[0051] In some implementations, the target language has general pronunciation rules. For example, all text units with the standard pronunciation sequence "si4" have the target pronunciation sequence "si2" in the target language. Therefore, the pronunciation sequence can be replaced as a whole by mapping probabilities.

[0052] Specifically, the mapping probability is defined as P(p'|p), where p is the pronunciation sequence in the target language, and p ’ Let P(p'|p) represent the pronunciation sequence in the standard language when the pronunciation sequence of any text unit is p in the target language. ’ The probability of.

[0053] In some implementations, P(p'|p) can be calculated using Bayes' theorem, which is exemplarily shown below:

[0054] P(p'|p)=P(p|p')*P(p') / P(p)

[0055] Among them, P(p|p’) represents the probability that when the pronunciation sequence of any text unit in the target language is p’, the pronunciation sequence in the standard language is p. P(p’) represents the probability that the pronunciation sequence of the text unit in the standard language is p’, and P(p) represents the probability that the pronunciation sequence of the text unit in the target language is p.

[0056] Step 134: If the mapping probability is greater than the preset probability, replace the standard pronunciation sequence in the standard pronunciation dictionary with the target pronunciation sequence to obtain the target pronunciation dictionary.

[0057] Optionally, the preset probability can be set according to needs, and this application does not make any restrictions.

[0058] Exemplarily, if the mapping probability is P((s i2)|(s i4)) > the preset probability, replace the pronunciation sequences of all text units with the pronunciation sequence of “s i4” in the standard pronunciation dictionary with “s i2”.

[0059] Furthermore, in order to improve the credibility of the pronunciation rules, it is also possible to limit that the phonemes of the pronunciation sequences p and p’ in P(p’|p) are the same but the tones are different. It can be understood that if the phonemes are different, it proves that the difference between the target language and the standard language in this pronunciation is relatively large, and the possibility of error in replacing the entire pronunciation sequence is also relatively large. Therefore, calculate the mapping probability of the pronunciation sequences with the same phonemes but different tones to improve the replacement accuracy.

[0060] Through the above method, the coverage of the obtained target pronunciation dictionary can be improved. Exemplarily, if the text units with the pronunciation sequence of “s i2” obtained by recognizing the training data based on the speech recognition system are only “四” and “寺”, then if only the text units are used to replace the standard pronunciation dictionary, the replaced text units are only “四” and “寺”. However, if the mapping probability is obtained and it is found that the mapping probability is greater than the preset probability, and the pronunciation sequence is replaced, and all text units with the pronunciation sequence of “s i4” in the standard pronunciation dictionary are replaced with “s i2”, then the number of replaced text units is far more than just “四” and “寺” two, improving the coverage of the obtained target pronunciation dictionary.

[0061] Step 140: Determine the target speech synthesis system according to the target pronunciation dictionary. [[ID=二十]]

[0062] In some embodiments, the speech synthesis system includes a text processing front end and a speech synthesis back end. The text processing front end of the target speech synthesis system can be obtained through the target pronunciation dictionary, and then the speech synthesis back end is constructed to obtain the target speech synthesis system.

[0063] In some embodiments, in order to improve the construction speed of the target speech synthesis system, Step 140 includes:

[0064] (1) Replace the standard pronunciation dictionary in the standard speech synthesis system with the target pronunciation dictionary to obtain the target speech synthesis system.

[0065] Specifically, by using an existing standard speech synthesis system to obtain the target speech synthesis system, the time required to rebuild the speech synthesis backend can be saved, thus increasing the construction speed of the target speech synthesis system.

[0066] Furthermore, in some implementations, not only is speech recognition performed on the training data using a standard language speech recognition system, but a target pronunciation dictionary is also obtained using a standard pronunciation dictionary of the standard language. Moreover, a target speech synthesis system is obtained by replacing the standard pronunciation dictionary in the standard speech synthesis system. These technical features work together to allow for direct migration of the existing standard speech synthesis system to the target speech synthesis system without having to train the speech synthesis system from scratch. That is, by replacing the pronunciation dictionary and reusing the speech generation backend of the standard speech synthesis system, the target speech synthesis system can be obtained quickly, improving the efficiency of obtaining the target speech synthesis system. Furthermore, it eliminates the need to obtain high-standard training data, thus solving the problem of difficulty in obtaining training data for the target language that meets the requirements.

[0067] In some embodiments, the speech synthesis method provided in this application further includes:

[0068] (1) Perform word segmentation on the text information to determine all text units in the text information, and determine the target vocabulary of the target language based on the text units.

[0069] (2) Compare the target vocabulary list with the standard vocabulary list corresponding to the standard pronunciation dictionary, and identify the text units that exist in the target vocabulary list but not in the standard vocabulary list as the text units to be added.

[0070] (3) Update the target pronunciation dictionary based on the text unit to be added.

[0071] Specifically, the target language may contain words that do not exist in the standard language. In order to generate the target speech synthesis system, the text units to be added need to be added to the target pronunciation dictionary. Further, adding the text units to be added to the target pronunciation dictionary means adding the text units to be added and their corresponding pronunciation sequences to the target pronunciation dictionary.

[0072] In some implementations, the step of comparing the target vocabulary with the standard vocabulary corresponding to the standard pronunciation dictionary, and determining text units that exist in the target vocabulary but not in the standard vocabulary as text units to be added, includes:

[0073] (2.1) Determine the word frequency of each text unit based on the training data.

[0074] (2.2) Compare the target vocabulary with the standard vocabulary corresponding to the standard pronunciation dictionary, and determine the text units that exist in the target vocabulary but not in the standard vocabulary as the initial text units to be added.

[0075] (2.3) If the word frequency of the initial text unit to be added is greater than the preset word frequency, then the initial text unit to be added is determined as the text unit to be added.

[0076] In some implementations, word frequency = the number of times the text unit appears / the total number of times all text units appear.

[0077] It is understandable that a text unit may exist in the target vocabulary but not in the standard vocabulary. This could be due to a speaker's mistake in the training data or an error in the speech recognition system. Therefore, to avoid these interferences, it is necessary to determine whether the word frequency of the initial text unit to be added is greater than the preset word frequency. Only if it is greater than the preset word frequency will the initial text unit to be added be determined as the text unit to be added, thereby reducing the probability of adding the wrong text unit.

[0078] This application provides a speech synthesis method, which includes: acquiring training data of a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; recognizing the training data based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes the phonemes and tones corresponding to the text units; adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text units to obtain a target pronunciation dictionary; and determining a target speech synthesis system based on the target pronunciation dictionary. This application obtains a target pronunciation dictionary by adjusting a standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text units, which can quickly obtain a target speech synthesis system that meets the requirements, and has low requirements for sample data.

[0079] Please refer to the following: Figure 4 , Figure 4 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of this application. Figure 4 As shown, the speech synthesis device 200 includes: an acquisition unit 210, a recognition unit 220, an adjustment unit 230, and a determination unit 240.

[0080] The acquisition unit 210 is used to acquire training data of the target language; wherein, the training data includes speech information of the target language and text information corresponding to the speech information.

[0081] The recognition unit 220 is used to recognize training data based on a speech recognition system to obtain multiple text units and the target pronunciation sequence corresponding to each text unit; wherein, the target pronunciation sequence includes the phonemes and tones corresponding to the text units.

[0082] The adjustment unit 230 is used to adjust the standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit to obtain the target pronunciation dictionary.

[0083] The determination unit 240 is used to determine the target speech synthesis system based on the target pronunciation dictionary.

[0084] It should be noted that, for the device-type embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the descriptions of the method embodiments. Any processing method described in the method embodiments can be implemented in the device embodiments through corresponding processing modules, and will not be elaborated upon further in the device embodiments.

[0085] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0086] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 5 As shown, the electronic device 300 includes: one or more processors 310 and a memory 320. Figure 5 Take the 310 processor as an example.

[0087] The processor 310 and the memory 320 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0088] The processor 310 is used to acquire training data of the target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; the training data is recognized based on a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein the target pronunciation sequence includes the phonemes and tones corresponding to the text units; the standard pronunciation dictionary is adjusted according to the target pronunciation sequence corresponding to the text units to obtain a target pronunciation dictionary; and a target speech synthesis system is determined based on the target pronunciation dictionary.

[0089] The memory 320, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules of the speech synthesis method in the embodiments of this application. The processor 310 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the speech synthesis method of the above-described method embodiments.

[0090] The memory 320 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 320 may optionally include memory remotely located relative to the processor 310, and these remote memories may be connected to the controller via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0091] One or more modules are stored in memory 320. When executed by one or more processors 310, they perform the speech synthesis method in any of the above method embodiments, for example, the method described above. Figure 1 Steps 110 to 140 of the method.

[0092] Please refer to Figure 6 , Figure 6 This is a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable storage medium 400 stores program code 410, which can be called by a processor to execute the speech synthesis method described in the above method embodiments.

[0093] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code that performs any of the method steps of the above-described speech synthesis method. This program code can be read from or written to one or more computer program products. The program code may, for example, be compressed in a suitable form.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above. For the sake of brevity, they are not provided in detail; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those skilled in the art can understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

Claims

1. A speech synthesis method, characterized in that, The method includes: Acquire training data for the target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; The training data is recognized using a speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein, the target pronunciation sequence includes the phoneme and tone corresponding to the text unit; Based on the target pronunciation sequence corresponding to the text unit, the standard pronunciation dictionary is adjusted to obtain the target pronunciation dictionary: the mapping probability of each target pronunciation sequence is obtained; wherein, the mapping probability is the probability that a text unit whose pronunciation sequence in the standard language is a standard pronunciation sequence has the same phoneme but different pitch in the target language; wherein, the mapping probability restricts the pronunciation sequence in the target language and the pronunciation sequence in the standard language to have the same phoneme but different pitch; if the mapping probability is greater than a preset probability, then all standard pronunciation sequences in the standard pronunciation dictionary are replaced with the target pronunciation sequence to obtain the target pronunciation dictionary, wherein one pronunciation sequence represents one character; The target speech synthesis system is determined based on the target pronunciation dictionary.

2. The method according to claim 1, characterized in that, The step of determining the target speech synthesis system based on the target pronunciation dictionary includes: The standard pronunciation dictionary in the standard speech synthesis system is replaced with the target pronunciation dictionary to obtain the target speech synthesis system.

3. The method according to claim 2, characterized in that, The process of recognizing the training data based on the speech recognition system includes: Speech recognition is performed on the training data based on a standard speech recognition system.

4. The method according to claim 1, characterized in that, The step of adjusting the standard pronunciation dictionary according to the target pronunciation sequence corresponding to the text unit to obtain the target pronunciation dictionary includes: Obtain a first pronunciation probability and a second pronunciation probability for each text unit; wherein, the first pronunciation probability is the probability that the pronunciation sequence of the text unit in the target language is the target pronunciation sequence, and the second pronunciation probability is the probability that the pronunciation sequence of the text unit in the target language is the standard pronunciation sequence; If the first pronunciation probability is greater than the second pronunciation probability, then the standard pronunciation sequence corresponding to the text unit in the standard pronunciation dictionary is replaced with the target pronunciation sequence to obtain the target pronunciation dictionary.

5. The method according to claim 1, characterized in that, The method further includes: The text information is segmented to identify all text units in the text information, and the target vocabulary of the target language is determined based on the text units. The target vocabulary is compared with the standard vocabulary corresponding to the standard pronunciation dictionary, and text units that exist in the target vocabulary but do not exist in the standard vocabulary are identified as text units to be added. The target pronunciation dictionary is updated based on the text unit to be added.

6. The method according to claim 5, characterized in that, The step of comparing the target vocabulary with the standard vocabulary corresponding to the standard pronunciation dictionary, and determining text units that exist in the target vocabulary but not in the standard vocabulary as text units to be added, includes: The word frequency of each text unit is determined based on the training data; The target vocabulary is compared with the standard vocabulary corresponding to the standard pronunciation dictionary, and text units that exist in the target vocabulary but not in the standard vocabulary are identified as initial text units to be added. If the word frequency of the initial text unit to be added is greater than the preset word frequency, then the initial text unit to be added is determined as the text unit to be added.

7. A speech synthesis device, characterized in that, The device includes: An acquisition unit is used to acquire training data for a target language; wherein the training data includes speech information of the target language and text information corresponding to the speech information; The recognition unit is used to recognize the training data based on the speech recognition system to obtain multiple text units and a target pronunciation sequence corresponding to each text unit; wherein, the target pronunciation sequence includes the phoneme and tone corresponding to the text unit; An adjustment unit is configured to adjust a standard pronunciation dictionary based on the target pronunciation sequence corresponding to the text unit to obtain a target pronunciation dictionary: obtaining a mapping probability for each target pronunciation sequence; wherein, the mapping probability is the probability that a text unit whose pronunciation sequence in the standard language is a standard pronunciation sequence has the same phoneme but different pitch in the target language; wherein, the mapping probability restricts the pronunciation sequence in the target language and the pronunciation sequence in the standard language to have the same phoneme but different pitch; if the mapping probability is greater than a preset probability, then all standard pronunciation sequences in the standard pronunciation dictionary are replaced with the target pronunciation sequence to obtain a target pronunciation dictionary, wherein one pronunciation sequence represents one character; A determining unit is used to determine a target speech synthesis system based on the target pronunciation dictionary.

8. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the speech synthesis method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that enable an electronic device to perform the speech synthesis method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Pronunciation dictionary generation method and device, storage medium and electronic device

    CN107767858A

  • Speech synthesis method and device, electronic equipment and computer readable storage medium

    CN113450757A