A speech synthesis method and related device

By adding tone words and rhythm information to pronunciation, the problem that the pronunciation synthesis method is difficult to match the user's real voice characteristics is solved, and the authenticity and user experience of voice interaction are improved.

CN116778903BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210223272.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-07-11
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

In the prior art, the speech synthesis method is difficult to conform to the user's real voice characteristics, resulting in poor voice interaction experience.

Method used

During the speech synthesis process, modal word information and rhythm information are added to simulate the user's speaking style, generate comprehensive information to be synthesized, and improve the authenticity of the speech information.

Benefits of technology

By simulating the user's tone words and rhythm, the authenticity of voice synthesis and the user's voice interaction experience are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778903B_ABST
    Figure CN116778903B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a speech synthesis method and related devices. When performing voice interaction, if voice information needs to be synthesized, the text information to be synthesized can be obtained first, and then the modal particle information corresponding to the text information to be synthesized can be determined. The modal particle information is used to simulate the modal particles involved in expressing the text information to be synthesized by voice, that is, the modal particles that users usually emit when speaking the text information. According to the text information to be synthesized and the modal particle information, comprehensive information to be synthesized can be generated, thereby integrating the modal particle information into the text information to be synthesized. Voice synthesis can be performed according to the comprehensive information to be synthesized to synthesize the voice information corresponding to the text information to be synthesized. Since the comprehensive information to be synthesized contains modal particle information, the synthesized voice information can better simulate the vocal characteristics during actual voice communication between users, improve the realism of the synthesized voice, and enhance the user's voice interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technologies, and particularly to a speech synthesis method and related devices. Background Art

[0002] With the continuous progress of human-computer interaction technologies, the interaction between humans and computers is no longer limited to using traditional devices such as keyboards or mice, but more through interaction methods such as speech and gestures.

[0003] In related technologies, the computer first determines the text information to be fed back to the user, and then automatically synthesizes the corresponding speech information based on the text information to conduct speech interaction with the user. However, the speech information determined in related technologies is difficult to fit the real speech characteristics of users, resulting in a poor speech interaction experience for users. Summary of the Invention

[0004] To solve the above technical problems, this application provides a speech synthesis method, which can add the modal particle information corresponding to the text to the text for speech synthesis together, so that the obtained speech information is more in line with the real speaking manner of users and improves the speech interaction experience of users.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] In a first aspect, an embodiment of this application provides a speech synthesis method, and the method includes:

[0007] Obtain the text information to be synthesized;

[0008] Determine the modal particle information corresponding to the text information to be synthesized, where the modal particle information is used to simulate the modal particles involved in expressing the text information to be synthesized by voice;

[0009] Generate comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information;

[0010] Synthesize the speech information corresponding to the text information to be synthesized according to the comprehensive information to be synthesized.

[0011] In a second aspect, an embodiment of this application provides a speech synthesis device, and the device includes an obtaining unit, a first determining unit, a generating unit, and a synthesizing unit:

[0012] The obtaining unit is used to obtain the text information to be synthesized;

[0013] The first determining unit is used to determine the modal particle information corresponding to the text information to be synthesized, where the modal particle information is used to simulate the modal particles involved in expressing the text information to be synthesized by voice;

[0014] The generating unit is configured to generate comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information;

[0015] The synthesizing unit is configured to synthesize voice information corresponding to the text information to be synthesized according to the comprehensive information to be synthesized.

[0016] In a possible implementation, the modal particle information includes modal particle identification information and modal particle position information. The modal particle identification information is used to identify the target modal particle corresponding to the text information to be synthesized, and the modal particle position information is used to identify the addition position of the target modal particle in the text information to be synthesized;

[0017] Specifically, the generating unit is configured to:

[0018] Add the modal particle identification information at the addition position in the text information to be synthesized according to the modal particle position information to generate the comprehensive information to be synthesized.

[0019] In a possible implementation, specifically, the synthesizing unit is configured to:

[0020] Determine the pinyin information corresponding to the text information to be synthesized;

[0021] Determine the first phoneme information corresponding to the pinyin information and the second phoneme information corresponding to the modal particle identification information;

[0022] Synthesize the voice information corresponding to the text information to be synthesized according to the first phoneme information and the second phoneme information.

[0023] In a possible implementation, the apparatus further includes a second determination unit:

[0024] The second determination unit is configured to determine the prosody information corresponding to the text information to be synthesized, where the prosody information is used to simulate the rhythm rule when expressing the text information to be synthesized in a voice manner;

[0025] Specifically, the generating unit is configured to:

[0026] Generate the comprehensive information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information.

[0027] In a possible implementation, the prosody information includes prosody identification information and prosody boundary information. The prosody identification information is used to identify the prosody type corresponding to the text information to be synthesized, and the prosody boundary information is used to identify the text position corresponding to the prosody identification information in the text information to be synthesized:

[0028] Specifically, the generating unit is configured to:

[0029] Determine the initial comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information.

[0030] Add the prosody identification information to the text positions in the initial comprehensive information to be synthesized according to the prosody boundary information, and generate the comprehensive information to be synthesized.

[0031] In a possible implementation manner, the synthesis unit is specifically configured to:

[0032] Determine N phoneme information corresponding to the comprehensive information to be synthesized, where the N phoneme information includes the phoneme information corresponding to the text information to be synthesized and the phoneme information corresponding to the modal particle information.

[0033] Determine N sub-prosody identification information corresponding to the N phoneme information according to the prosody identification information added in the comprehensive information to be synthesized.

[0034] Synthesize the speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information.

[0035] In a possible implementation manner, the synthesis unit is specifically configured to:

[0036] Synthesize the speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information through an acoustic model.

[0037] In a possible implementation manner, the acoustic model is trained in the following manner:

[0038] Obtain sample text information, where the sample text information has corresponding sample speech information.

[0039] Determine the sample modal particle information and sample prosody information corresponding to the sample text information.

[0040] Generate sample comprehensive information to be synthesized based on the sample modal particle information, sample text information, and the sample prosody information.

[0041] Determine M phoneme information corresponding to the sample comprehensive information to be synthesized, where the M phoneme information includes the phoneme information corresponding to the sample text information and the phoneme information corresponding to the sample modal particle information.

[0042] Determine M sub-prosody identification information corresponding to the M phoneme information according to the sample prosody information.

[0043] Concatenate the M phoneme information with the corresponding sub-prosody identification information in the M sub-prosody identification information respectively to obtain M input information.

[0044] Determine the to-be-determined speech information corresponding to the sample text information according to the M input information through the initial acoustic model;

[0045] Adjust the initial acoustic model according to the difference between the to-be-determined speech information and the sample speech information to obtain the acoustic model.

[0046] In a third aspect, an embodiment of the present application provides a computer device for speech synthesis. The computer device includes a processor and a memory:

[0047] The memory is used to store program codes and transmit the program codes to the processor;

[0048] The processor is used to execute the speech synthesis method described in the foregoing aspect according to the instructions in the program codes.

[0049] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium is used to store program codes, and the program codes are used to execute the speech synthesis method described in the foregoing aspect.

[0050] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the speech synthesis method described in the foregoing aspect when executed by a processor.

[0051] It can be seen from the above technical solutions that when performing speech interaction, if speech information needs to be synthesized, the text information to be synthesized can be obtained first, and then the modal particle information corresponding to the text information to be synthesized can be determined. The modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized by voice, that is, the modal particles that users usually emit when speaking the text information. According to the text information to be synthesized and the modal particle information, comprehensive information to be synthesized can be generated, so as to integrate the modal particle information into the text information to be synthesized. According to the comprehensive information to be synthesized, speech synthesis can be performed to synthesize the speech information corresponding to the text information to be synthesized. Since the comprehensive information to be synthesized contains modal particle information, the synthesized speech information can better simulate the vocal characteristics during actual speech communication between users, improve the realism of the synthesized speech, and improve the user's speech interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 Schematic diagram of a speech synthesis method in an actual application scenario provided by an embodiment of the present application;

[0054] Figure 2 Flowchart of a speech synthesis method provided by an embodiment of the present application;

[0055] Figure 3 Schematic diagram of a speech information synthesis provided by an embodiment of the present application;

[0056] Figure 4 Schematic diagram of a speech synthesis method provided by an embodiment of the present application;

[0057] Figure 5 Schematic diagram of a model training provided by an embodiment of the present application;

[0058] Figure 6 Schematic diagram of a model training provided by an embodiment of the present application;

[0059] Figure 7 Architecture diagram of a Tacotron2 model provided by an embodiment of the present application;

[0060] Figure 8 Architecture diagram inside a model provided by an embodiment of the present application;

[0061] Figure 9 Structural block diagram of a speech synthesis device provided by an embodiment of the present application;

[0062] Figure 10 Structural diagram of a terminal provided by an embodiment of the present application;

[0063] Figure 11 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0064] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0065] Speech synthesis technology is one of the widely used technologies at present. In the related technologies, speech synthesis only converts the text to be synthesized into its corresponding speech, and the speech information and text information are exactly the same without adding any extra information other than the text. In fact, when users communicate with each other by voice, they usually add some modal particles, such as "eh", "oh", "ah", etc. When these modal particles are lacking, the synthesized speech will lack a sense of reality, be difficult to conform to the real voice communication habits of users, and bring users a poor voice interaction experience.

[0066] To solve the above technical problems, the present application provides a speech synthesis method. The processing device can add the modal particle information corresponding to the text to the text for speech synthesis together, so that the obtained speech information is more in line with the user's real speaking manner and improves the user's speech interaction experience.

[0067] It can be understood that this method can be applied to a processing device, which is a processing device capable of performing speech synthesis. For example, it can be a terminal device or a server with speech synthesis function. This method can be independently executed by a terminal device or a server, or can be applied to a network scenario where a terminal device and a server communicate, and is executed in cooperation with the terminal device and the server. Among them, the terminal device can be a device such as a computer or a mobile phone. The server can be understood as an application server or a Web server. In actual deployment, the server can be an independent server or a cluster server.

[0068] To facilitate understanding of the technical solution provided by the embodiments of the present application, next, a speech synthesis method provided by the present application will be introduced in combination with an actual application scenario.

[0069] See Figure 1 , Figure 1 is a schematic diagram of a speech synthesis method in an actual application scenario provided by the embodiments of the present application. In this actual application scenario, the processing device is server 101, and the text information to be synthesized is "The weather is really nice today".

[0070] After obtaining the text information to be synthesized, server 101 can first determine the modal particle information corresponding to the text information to be synthesized. For example, the modal particle information corresponding to the text information to be synthesized can be predicted through a modal particle information prediction model. In this actual application scenario, the determined modal particle information can be "ah", that is, it is predicted that when the user says the sentence "The weather is really nice today", the user will habitually add the modal particle "ah" at the end. Based on this, in order to simulate the speech characteristics during real user speech communication, a comprehensive information to be synthesized can be generated based on the text information to be synthesized and the modal particle information, that is, "The weather is really nice today ah". Finally, server 101 can synthesize the speech information corresponding to the text information to be synthesized according to the comprehensive information to be synthesized, so that not only can the speech information corresponding to "The weather is really nice today" be synthesized, but also the real speech expression manner can be simulated, and "ah" is added as a modal particle in the speech information, improving the authenticity of the speech information.

[0071] Next, a speech synthesis method provided by the present application will be introduced in combination with the accompanying drawings.

[0072] See Figure 2 , Figure 2The flowchart of a speech synthesis method provided by an embodiment of the present application. The method includes:

[0073] S201: Obtain the text information to be synthesized.

[0074] The text information to be synthesized refers to the text information that needs to be synthesized into speech. The manner in which the processing device obtains the text information to be synthesized can include various ways, for example, it can be the text information determined for reply based on the voice information sent by the user, etc.

[0075] S202: Determine the modal particle information corresponding to the text information to be synthesized.

[0076] In order to make the synthesized speech information sound more real and natural to the user, when performing speech synthesis, the processing device can first add information for simulating the characteristics of voice communication between real users to the text information to be synthesized.

[0077] It can be understood that when a user speaks, they usually subconsciously emit modal particles corresponding to the speech content. For example, the "ah" in "The weather is really nice today" and the "ei" in "Hey, where did I park my car", etc. Based on this, in order to improve the realism of the synthesized speech information, the processing device can determine the modal particle information corresponding to the text information to be synthesized. This modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized by voice. For example, this modal particle information can identify the modal particles that need to be added to the text information to be synthesized and the positions where the modal particles are added, so that the modal particles can be reasonably added to the text information to be synthesized. As Figure 1 shown, this modal particle information can identify that the modal particle "ah" should be added after "The weather is really nice today".

[0078] S203: Generate comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information.

[0079] The processing device can add the modal particle information to the text information to be synthesized based on the content identified by the modal particle information to generate comprehensive information to be synthesized. This comprehensive information to be synthesized can be more in line with the characteristics of how the user actually expresses the text information to be synthesized by voice. For example, based on the text information to be synthesized "The weather is really nice today" and the modal particle information "ah", comprehensive information to be synthesized "The weather is really nice today ah" can be generated.

[0080] Among them, the determination of modal particle information in this application mainly includes two situations. One is that the text information to be synthesized itself has its corresponding modal particle. The processing device can determine the modal particle in the text to be synthesized through this step and convert it into the corresponding modal particle information when generating the comprehensive text information to be synthesized, so that the processing device can perform speech synthesis targeted based on this modal particle information. The other is that the text information to be synthesized does not have the corresponding modal particle information of the text information to be synthesized. The processing device can determine its corresponding modal particle information and add it to the text information to be synthesized to generate the comprehensive text information to be synthesized.

[0081] S204: Synthesize the voice information corresponding to the text information to be synthesized according to the comprehensive text information to be synthesized.

[0082] Since the comprehensive text information to be synthesized is added with modal particle information for simulating the voice tone characteristics of the user, the processing device can synthesize voice information according to the comprehensive text information to be synthesized, so that the voice information can conform to the actual voice characteristics of the user. Thus, when the voice information is sent to the user, the voice interaction experience of the user can be made more real.

[0083] It can be seen from the above technical solution that when performing voice interaction, if voice information needs to be synthesized, the text information to be synthesized can be obtained first, and then the corresponding modal particle information of the text information to be synthesized can be determined. The modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized by voice, that is, the modal particles that the user usually emits when saying the text information. According to the text information to be synthesized and the modal particle information, a comprehensive text information to be synthesized can be generated, so that the modal particle information is integrated into the text information to be synthesized. Voice synthesis can be performed according to the comprehensive text information to be synthesized to synthesize the voice information corresponding to the text information to be synthesized. Since the comprehensive text information to be synthesized contains modal particle information, the synthesized voice information can better simulate the vocal characteristics during actual voice communication between users, improve the realism of the synthesized voice, and improve the voice interaction experience of users.

[0084] In order to make the addition of modal particle information more reasonable, in a possible implementation manner, the modal particle information may include modal particle identification information and modal particle position information. The modal particle identification information is used to identify the target modal particle corresponding to the text information to be synthesized. The target modal particle is the modal particle that needs to be added to the text information to be synthesized in order to simulate the real voice habit of the user. The modal particle position information is used to identify the addition position of the target modal particle in the text information to be synthesized. Among them, the modal particle identification information can directly be the target modal particle itself, or a special identifier corresponding to the target modal particle, such as "S1", "S2", etc.

[0085] Thus, when generating the comprehensive information to be synthesized, the processing device can add the modal particle identification information at the addition position in the text information to be synthesized according to the modal particle position information, so as to generate the comprehensive information to be synthesized, realizing the addition of the target modal particle to the appropriate position in the text information to be synthesized. Among them, the determination method of the modal particle information can include various types. In one possible implementation, the processing device can first obtain the training samples for training the modal particle prediction model. The training samples include the sample text information, and the sample text information has the labeled text information with the added modal particle information. By training the model with the training samples, a modal particle prediction model can be obtained. The modal particle prediction model can determine whether it is necessary to add a modal particle at some specific positions (such as the head and tail positions of the text information) in the text information. For example, it can output 0 to indicate that no addition is required, output 1 to indicate that addition is required, and when addition is required, it can determine the modal particle to be added.

[0086] Specifically, in one possible implementation, when synthesizing the speech information according to the comprehensive information to be synthesized, the processing device can first determine the pinyin information corresponding to the text information to be synthesized. The pinyin information is the pinyin that constitutes the pronunciation corresponding to the text information to be synthesized. For example, the pinyin information corresponding to "speech synthesis" can be "Yv3 / Yin1 / He2 / Cheng2", where the numbers represent tones. Subsequently, the processing device can determine the first phoneme information corresponding to the pinyin information and the second phoneme information corresponding to the modal particle identification information. The first phoneme information is the phoneme information that constitutes the pronunciation of the pinyin information, and the second phoneme information is the phoneme information that constitutes the pronunciation of the modal particle identified by the modal particle identification information. Since modal particles are relatively complex and diverse, some modal particles are difficult to accurately express in the form of pinyin, such as relatively abstract and colloquial modal particles like "emmm". In addition, if modal particles are also represented in the form of pinyin, it may cause the processing device to be unable to distinguish modal particles from ordinary text information in terms of speech expression, and it is difficult to express the speech characteristics of modal particles. Therefore, in the embodiments of the present application, the processing device can separately set the corresponding phoneme information for the modal particle identification information, rather than determining the phoneme information through the pinyin information corresponding to the modal particle, so as to fully analyze the pronunciation characteristics of the modal particle and make the modal particle information better in terms of speech expression.

[0087] As mentioned above, there are two situations when determining the modal particle information. For the situation that the corresponding modal particle already exists in the text information to be synthesized, in order to avoid the situation that the phoneme information of the modal particle cannot be accurately expressed through the pinyin information, the pinyin information corresponding to the text information to be synthesized determined by the processing device does not include the pinyin information corresponding to the modal particle, and the second phoneme information corresponding to the modal particle is determined by the second phoneme information corresponding to the modal particle identification information. For example, when the text information to be synthesized is "The sun is so big today", in the modal particle information corresponding to the text information to be synthesized, the target modal particle can be "ah", and the corresponding modal particle identification information can be "S2". At this time, the processing device can determine the pinyin information corresponding to the text content "The sun is so big today" other than the target modal particle in the text information to be synthesized, and then determine the corresponding first phoneme information based on the pinyin information. Subsequently, the processing device can directly determine the corresponding second phoneme information based on the identification S2, thereby avoiding determining the phoneme information corresponding to the pinyin information of the target modal particle itself as the second phoneme information.

[0088] According to the first phoneme information and the second phoneme information, the processing device can determine the pronunciation method when expressing the comprehensive information to be synthesized in a voice manner, and then synthesize the voice information corresponding to the text information to be synthesized.

[0089] For example, Figure 3 As shown, Figure 3 A schematic diagram of speech information synthesis provided for an embodiment of the present application, wherein the text information to be synthesized in the comprehensive information to be synthesized can be first converted into corresponding pinyin information, and then the corresponding phoneme information is determined, wherein the "∧" in the phoneme information represents an unpronounced initial consonant factor. The modal particle identification information is a special token S1, and its corresponding pinyin is also S1, and the corresponding phoneme information is preset [s11, s12, s13, s14]. Using 4 phoneme information to refer to the modal particle "eh~" can enable the pronunciation characteristics of the modal particles to be learned in a more fine-grained manner when training a speech synthesis model for synthesizing speech information.

[0090] Each phoneme information has a corresponding phoneme id. The processing device can map the phoneme information to the corresponding phoneme id. The phoneme id is determined by the serial number of the phoneme information in the pre-set phoneme dictionary. The phoneme dictionary is a collection of phoneme components. Various phoneme information can be formed through various combinations in the phoneme dictionary. Therefore, by mapping the phoneme id according to the phoneme information, the processing device can determine the phoneme composition of the comprehensive information to be synthesized, and then synthesize the voice information. The voice information is the voice information corresponding to the text information to be synthesized.

[0091] It can be understood that in the way of improving the authenticity of voice information, in addition to simulating the modal particles that users may add, the processing device can also simulate the prosodic features during users' voice communication. Prosody refers to the rhythm and rules in voice information, such as the cadence when speaking and the pauses between different words. Usually, when a user says a text, they will pause after saying one or several words. For example, in the sentence "The sun is so big. The weather today is really nice", the user may express it with a prosody like "The sun is so big (pause), today's (pause) weather (pause) is really nice".

[0092] Based on this, in a possible implementation, in order to further improve the authenticity and naturalness of the synthesized voice and make it more conform to the actual voice characteristics of users, the processing device can also determine the prosody information corresponding to the text information to be synthesized. This prosody information is used to simulate the rhythm rules when expressing the text information to be synthesized by voice. When generating the comprehensive text information to be synthesized, the processing device can generate the comprehensive text information to be synthesized according to the text information to be synthesized, the modal particle information and the prosody information. Thus, the synthesized voice information based on this comprehensive text information to be synthesized can not only simulate the usage of modal particles by users, but also simulate the prosodic features of users' voices, bringing users a more real voice interaction experience.

[0093] Specifically, in a possible implementation, in order to add the prosody information to the text information to be synthesized more reasonably, the prosody information determined by the processing device can include prosody identification information and prosody boundary information. The prosody identification information is used to identify the prosody type corresponding to the text information to be synthesized. Different prosody types correspond to different rhythm rules. For example, the prosody type can include syllable, prosodic word, prosodic phrase and intonation phrase. A prosodic word means a pause after a user says a word. The concept of a prosodic word is similar to word segmentation. A prosodic phrase contains prosodic words and is generally composed of prosodic words and modal particles. An intonation phrase contains prosodic phrases and generally refers to a large pause. The prosody boundary information is used to identify the text position in the text information to be synthesized corresponding to the prosody identification information.

[0094] The processing device can first determine the initial comprehensive text information to be synthesized according to the text information to be synthesized and the modal particle information. The initial comprehensive text information to be synthesized is the text information to be synthesized with the modal particle information added. Then, according to the prosody boundary information, the processing device can add the prosody identification information at the text position in the initial comprehensive text information to be synthesized to generate the comprehensive text information to be synthesized.

[0095] Similar to the method for determining modal particle information, the processing device can also determine prosody information through a model. First, the processing device can obtain sample text information, which has corresponding prosody information. The processing device can use the sample text information as a training sample and the prosody information as a training label to train a prosody prediction model, which can be used to determine the prosody information corresponding to the text information to be synthesized.

[0096] For example, as Figure 4 shown, the text information to be synthesized can be "The sun is so big. The weather today is really nice". Through the modal particle prediction model, it can be determined that the target modal particle corresponding to this text information to be synthesized is "ah", the modal particle identification information corresponding to this target modal particle can be S2, and this target modal particle is added to the end of this text information to be synthesized to obtain the initial comprehensive text information to be synthesized "The sun is so big. The weather today is really niceS2". Then, the prosody information corresponding to this text information to be synthesized can be determined through the prosody prediction model. After combining this prosody information with the initial comprehensive text information to be synthesized, the comprehensive text information to be synthesized "Tai#0yang#0hao#0da#3, jin#0tian#0de#1tian#0qi#2zhen#0bu#0cuo#0S2#0" is obtained. Among them, #0 represents a syllable, #1 represents a prosodic word, #2 represents a prosodic phrase, and #3 represents an intonation phrase. The processing device can perform speech synthesis according to this comprehensive text information to be synthesized to obtain the speech information corresponding to this text information to be synthesized. Of course, various identification methods can be adopted for the modal particle identification in this application, such as X1, X2, etc., which are not limited here.

[0097] In order to fully incorporate prosody information into the text information to be synthesized and obtain highly realistic speech information, in a possible implementation, the processing device may first determine N phoneme information corresponding to the comprehensive text information to be synthesized. The N phoneme information includes the phoneme information corresponding to the text information to be synthesized and the phoneme information corresponding to the filler word information. The processing device may determine N sub-prosody identification information corresponding to the N phoneme information according to the prosody identification information added to the comprehensive text information to be synthesized, that is, according to the correspondence between the text information and the prosody identification information in the comprehensive text information, expand the prosody identification information to each phoneme information corresponding to the text information. For example, before expansion, each character in the comprehensive text information to be synthesized has a corresponding prosody identification information, such as "lan#0yin#1he#2cheng#3". After expansion, according to the phoneme information and prosody identification information corresponding to each character, the prosody identification information may be determined as the sub-prosody identification information corresponding to each phoneme information under the character. For example, it may become "∧#0v3#0∧#1in1#1h#2e#2ch#3eng#3". Thus, when synthesizing the speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information, the processing device may fully analyze the expression of each phoneme in the speech information, enhance the combination degree between the phoneme information and the prosody information, and obtain more realistic prosody-rich speech information.

[0098] Among them, the way for the processing device to perform speech synthesis may include various methods. In a possible implementation, in order to improve the accuracy of speech synthesis, the processing device may use an acoustic model to synthesize the speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information.

[0099] When training the acoustic model, two model training methods may be adopted. The difference between the two model training methods lies in the combination method of the sub-prosody identification information and the phoneme information. One model training method is to interleave and combine the phoneme information and the sub-prosody identification information, expanding the number of phoneme information. As Figure 5 shown, 4 phoneme information corresponds to 4 sub-prosody identification information. After interleaving and combining, 8 pieces of information will be obtained. Through the 8 pieces of information for model training, after the model learns the prosody representation, the prosody representation is removed. In this way, the phoneme representation obtains the information of the prosody representation. However, this method changes the number of phoneme information, which may cause the model to not fully learn the characteristics of the prosody information based on the phoneme information.

[0100] In order to train an acoustic model that is more suitable for speech synthesis, in a possible implementation, the processing device can train the model in a way that does not change the number of phoneme information. The processing device can obtain sample text information, which has corresponding sample speech information, and then determine the corresponding sample filler word information and sample prosody information of the sample text information. Based on the sample text information, the sample filler word information, and the sample prosody information, the processing device can generate sample comprehensive information to be synthesized, and then determine M phoneme information corresponding to the sample comprehensive information to be synthesized. The M phoneme information includes the phoneme information corresponding to the sample text information and the phoneme information corresponding to the sample filler word information. According to the sample prosody information, M sub-prosody identification information corresponding to the M phoneme information can be determined, and then the M phoneme information is respectively concatenated with the corresponding sub-prosody identification information in the M sub-prosody identification information to obtain M input information. That is, after concatenating 1 phoneme information with 1 sub-prosody identification information, 1 input information is generated, rather than 2 input information, so that after the M phoneme information is fused with the prosody information, M input information is still obtained. The processing device can determine the pending speech information corresponding to the sample text information through the initial acoustic model, and then based on the difference between the pending speech information and the sample speech information, the processing device can analyze the accuracy of the initial acoustic model when synthesizing speech information, so as to adjust the initial acoustic model to obtain an acoustic model for speech synthesis.

[0101] As Figure 6 shown, Figure 6 FIG. is a schematic diagram of a model training provided by an embodiment of the present application. The characters in the text information correspond one-to-one with the phoneme tokens (i.e., prosody identification information). After converting the text information into phoneme information and expanding it, the phoneme tokens corresponding to each phoneme information can be obtained, which are 8 phoneme information and 8 sub-prosody identification information. After concatenation, 8 input information is still obtained, and the input information is input into the acoustic model for training.

[0102] There are various acoustic models that can be used in the present application. For example, an end-to-end speech synthesis model tacotron2 based on deep learning can be used, but it is not limited to this model. For example, it can also be a neural network model lstm-crf, cnn, etc. The vocoder can use melgan. Next, the training process on the tacotron2 model will be mainly introduced. As Figure 7 shown, Figure 7Disclosed is an architecture diagram of a Tacotron2 model, where the encoding layer can perform encoding in a manner of concatenating phoneme information and sub-prosody identification information as described above, without changing the number of information. When recording voice information, filler words can be recorded simultaneously, and the filler words are specially marked in the corresponding text information, such as S1, S2…SN, etc. During the training process, the filler words can be mapped to specific phoneme information. Through Figure 7 the model architecture in

[0103] the Tacotron2 model remains unchanged. As Figure 8 shown, for the encoding layer, it is replaced with the special scheme described above that can encode filler word information and prosody information to enhance its filler word synthesis ability and its performance in prosody. During its training process, the results of prosody encoding, filler word encoding, and phoneme encoding are updated together according to the gradients backpropagated during the training process to obtain better representation results.

[0104] Experiments prove that the voice information synthesized by the voice synthesis method of the present application is clearer in the mel spectrogram performance, and the user's voice interaction experience is also better.

[0105] Based on the voice synthesis method provided in the above embodiments, the embodiments of the present application also provide a voice synthesis device. Refer to Figure 9 , Figure 9 which is a structural block diagram of a voice synthesis device 900 provided by the embodiments of the present application. The device 900 includes an acquisition unit 901, a first determination unit 902, a generation unit 903, and a synthesis unit 904:

[0106] The acquisition unit 901 is configured to acquire text information to be synthesized;

[0107] The first determination unit 902 is configured to determine filler word information corresponding to the text information to be synthesized, where the filler word information is used to simulate filler words involved when expressing the text information to be synthesized by voice;

[0108] The generation unit 903 is configured to generate comprehensive information to be synthesized according to the text information to be synthesized and the filler word information;

[0109] The synthesis unit 904 is configured to synthesize voice information corresponding to the text information to be synthesized according to the comprehensive information to be synthesized.

[0110] In a possible implementation, the modal particle information includes modal particle identification information and modal particle position information. The modal particle identification information is used to identify the target modal particle corresponding to the text information to be synthesized, and the modal particle position information is used to identify the addition position of the target modal particle in the text information to be synthesized;

[0111] The generating unit 903 is specifically configured to:

[0112] Add the modal particle identification information at the addition position in the text information to be synthesized according to the modal particle position information, and generate the comprehensive text information to be synthesized.

[0113] In a possible implementation, the synthesizing unit 904 is specifically configured to:

[0114] Determine the pinyin information corresponding to the text information to be synthesized;

[0115] Determine the first phoneme information corresponding to the pinyin information and the second phoneme information corresponding to the modal particle identification information;

[0116] Synthesize the voice information corresponding to the text information to be synthesized according to the first phoneme information and the second phoneme information.

[0117] In a possible implementation, the device further includes a second determination unit:

[0118] The second determination unit is configured to determine the prosody information corresponding to the text information to be synthesized, where the prosody information is used to simulate the rhythm rule when expressing the text information to be synthesized in a voice manner;

[0119] The generating unit 903 is specifically configured to:

[0120] Generate the comprehensive text information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information.

[0121] In a possible implementation, the prosody information includes prosody identification information and prosody boundary information. The prosody identification information is used to identify the prosody type corresponding to the text information to be synthesized, and the prosody boundary information is used to identify the text position corresponding to the prosody identification information in the text information to be synthesized:

[0122] The generating unit 903 is specifically configured to:

[0123] Determine the initial comprehensive text information to be synthesized according to the text information to be synthesized and the modal particle information;

[0124] Add the prosody identification information at the text position in the initial comprehensive information to be synthesized according to the prosody boundary information, and generate the comprehensive information to be synthesized.

[0125] In a possible implementation manner, the synthesis unit 904 is specifically configured to:

[0126] Determine N phoneme information corresponding to the comprehensive information to be synthesized, where the N phoneme information includes phoneme information corresponding to the text information to be synthesized and phoneme information corresponding to the filler word information;

[0127] Determine N sub-prosody identification information corresponding to the N phoneme information according to the added prosody identification information in the comprehensive information to be synthesized;

[0128] Synthesize speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information.

[0129] In a possible implementation manner, the synthesis unit 904 is specifically configured to:

[0130] Synthesize speech information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information through an acoustic model.

[0131] In a possible implementation manner, the acoustic model is trained in the following manner:

[0132] Obtain sample text information, where the sample text information has corresponding sample speech information;

[0133] Determine sample filler word information and sample prosody information corresponding to the sample text information;

[0134] Generate sample comprehensive information to be synthesized based on the sample filler word information, sample text information, and the sample prosody information;

[0135] Determine M phoneme information corresponding to the sample comprehensive information to be synthesized, where the M phoneme information includes phoneme information corresponding to the sample text information and phoneme information corresponding to the sample filler word information;

[0136] Determine M sub-prosody identification information corresponding to the M phoneme information according to the sample prosody information;

[0137] Concatenate the M phoneme information with the corresponding sub-prosody identification information in the M sub-prosody identification information to obtain M input information;

[0138] Determine the pending speech information corresponding to the sample text information according to the M input information through an initial acoustic model;

[0139] Adjust the initial acoustic model according to the difference between the to-be-determined voice information and the sample voice information to obtain the acoustic model.

[0140] An embodiment of the present application also provides a computer device, which will be introduced below with reference to the accompanying drawings. Please refer to Figure 10 As shown, an embodiment of the present application provides a device, which may also be a terminal device. The terminal device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), a point of sale (POS for short), an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:

[0141] Figure 10 Shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiment of the present application. Refer to Figure 10 , the mobile phone includes: a radio frequency (RF for short) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless fidelity (WiFi for short) module 770, a processor 780, and a power supply 790 and other components. Those skilled in the art can understand that Figure 10 the structure of the mobile phone shown in [[ ]] does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0142] Next, with reference to [[ ]] Figure 10 specific introduction to each component of the mobile phone will be given:

[0143] The RF circuit 710 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is sent to the processor 780 for processing. Additionally, the uplink data designed is sent to the base station. Generally, the RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. Moreover, the RF circuit 710 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0144] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 720 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0145] The input unit 730 can be used to receive input numeric or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 can include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 731), and drive the corresponding connection device according to a preset program. Optionally, the touch panel 731 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute the commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 can also include other input devices 732. Specifically, the other input devices 732 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0146] The display unit 740 can be used to display the information input by the user or the information provided to the user and various menus of the mobile phone. The display unit 740 can include a display panel 741. Optionally, the display panel 741 can be configured in forms such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED). Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation thereon or nearby, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 10 it, the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0147] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer attitude calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0148] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output. On the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output to the processor 780 for processing, it is sent through the RF circuit 710 to, for example, another mobile phone, or the audio data is output to the memory 720 for further processing.

[0149] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770, which provides users with wireless broadband Internet access. Although Figure 10 the WiFi module 770 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0150] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 720, and by calling the data stored in the memory 720, it executes various functions of the mobile phone and processes data, thereby performing an overall detection of the mobile phone. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 780.

[0151] The mobile phone further includes a power supply 790 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.

[0152] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0153] In this embodiment, the processor 780 included in the terminal device further has the following functions:

[0154] Obtain the text information to be synthesized;

[0155] Determine the modal particle information corresponding to the text information to be synthesized, where the modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized in a voice manner;

[0156] Generate comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information;

[0157] Synthesize the voice information corresponding to the text information to be synthesized according to the comprehensive information to be synthesized.

[0158] The embodiment of the present application further provides a server. Please refer to Figure 11 as shown Figure 11 is a structural diagram of the server 800 provided by the embodiment of the present application. The server 800 may vary greatly due to configuration or performance differences, and may include one or more central processing units (Central Processing Units, abbreviated as CPU) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 for storing application programs 842 or data 844 (for example, one or more mass storage devices). Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not marked in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 822 may be set to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.

[0159] The server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM , FreeBSD TM and so on.

[0160] In the above embodiments, the steps executed by the server can be based on Figure 11 the server structure shown.

[0161] According to one aspect of the present application, there is provided a computer-readable storage medium for storing program code for executing the speech synthesis method described in each of the foregoing embodiments.

[0162] According to one aspect of the present application, there is provided a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the speech synthesis method provided in various alternative implementations of the above embodiments.

[0163] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disk, etc., various media that can store program code.

[0164] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0165] As described above, it is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice synthesis method, characterized in that, The method includes: Obtaining text information to be synthesized; Determining the modal particle information corresponding to the text information to be synthesized, where the modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized in a voice manner; the modal particle information includes modal particle identification information, and the modal particle identification information is used to identify the target modal particle corresponding to the text information to be synthesized; Determining the prosody information corresponding to the text information to be synthesized, where the prosody information is used to simulate the rhythm rule when expressing the text information to be synthesized in a voice manner; the prosody information includes prosody identification information; the prosody identification information is used to identify the prosody type corresponding to the text information to be synthesized; Generating comprehensive information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information; Determining the pinyin information corresponding to the text information to be synthesized, where the pinyin information corresponding to the text information to be synthesized does not include the pinyin information corresponding to the modal particle; Determining the first phoneme information corresponding to the pinyin information, and the second phoneme information corresponding to the modal particle identification information; the second phoneme information corresponding to the modal particle identification information is a pre-set phoneme information combination composed of multiple phoneme information, and is different from the phoneme information corresponding to the pinyin information of the modal particle itself; Determining N phoneme information corresponding to the comprehensive information to be synthesized, where the N phoneme information includes the first phoneme information and the second phoneme information; Determining N sub-prosody identification information corresponding to the N phoneme information according to the prosody identification information added in the comprehensive information to be synthesized; Synthesizing the voice information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information, where the N sub-prosody identification information is extended into the corresponding phoneme information in the N phoneme information.

2. The method according to claim 1, wherein The modal particle information further includes modal particle position information, and the modal particle position information is used to identify the addition position of the target modal particle in the text information to be synthesized; The generating comprehensive information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information includes: Adding the modal particle identification information at the addition position in the text information to be synthesized according to the modal particle position information to generate the comprehensive information to be synthesized.

3. The method according to claim 1, wherein The prosody information further includes prosody boundary information, and the prosody boundary information is used to identify the text position corresponding to the prosody identification information in the text information to be synthesized; The generating comprehensive information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information includes: Determining initial comprehensive information to be synthesized according to the text information to be synthesized and the modal particle information; Adding the prosody identification information at the text position in the initial comprehensive information to be synthesized according to the prosody boundary information to generate the comprehensive information to be synthesized.

4. The method according to claim 1, wherein The synthesizing the voice information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information includes: Through an acoustic model, according to the N phoneme information and the N sub-prosody identification information, synthesize the speech information corresponding to the text to be synthesized.

5. The method according to claim 4, characterized in that, The acoustic model is trained in the following manner: Obtain sample text information, and the sample text information has corresponding sample speech information; Determine the sample modal particle information and sample prosody information corresponding to the sample text information; Based on the sample modal particle information, sample text information, and the sample prosody information, generate sample comprehensive information to be synthesized; Determine M phoneme information corresponding to the sample comprehensive information to be synthesized, where the M phoneme information includes the phoneme information corresponding to the sample text information and the phoneme information corresponding to the sample modal particle information; According to the sample prosody information, determine M sub-prosody identification information corresponding to the M phoneme information; Concatenate each of the M phoneme information with the corresponding sub-prosody identification information in the M sub-prosody identification information to obtain M input information; Through an initial acoustic model, determine the pending speech information corresponding to the sample text information according to the M input information; Adjust the initial acoustic model according to the difference between the pending speech information and the sample speech information to obtain the acoustic model.

6. A voice synthesis device, characterized in that, The device includes an acquisition unit, a first determination unit, a second determination unit, a generation unit, and a synthesis unit: The acquisition unit is used to acquire text information to be synthesized; The first determination unit is used to determine the modal particle information corresponding to the text information to be synthesized, and the modal particle information is used to simulate the modal particles involved when expressing the text information to be synthesized by voice; the modal particle information includes modal particle identification information, and the modal particle identification information is used to identify the target modal particle corresponding to the text information to be synthesized; The second determination unit is used to determine the prosody information corresponding to the text information to be synthesized, and the prosody information is used to simulate the rhythm rule when expressing the text information to be synthesized by voice; the prosody information includes prosody identification information; the prosody identification information is used to identify the prosody type corresponding to the text information to be synthesized; The generation unit is used to generate comprehensive information to be synthesized according to the text information to be synthesized, the modal particle information, and the prosody information; The synthesis unit is used for: Determine the pinyin information corresponding to the text information to be synthesized, and the pinyin information corresponding to the text information to be synthesized does not include the pinyin information corresponding to the modal particle; Determine the first phoneme information corresponding to the pinyin information and the second phoneme information corresponding to the modal particle identification information; the second phoneme information corresponding to the modal particle identification information is a preset phoneme information combination composed of multiple phoneme information and is different from the phoneme information corresponding to the pinyin of the modal particle itself; Determine N phoneme information corresponding to the comprehensive information to be synthesized, where the N phoneme information includes the first phoneme information and the second phoneme information; According to the prosody identification information added in the comprehensive information to be synthesized, determine N sub-prosody identification information corresponding to the N phoneme information; Synthesize the voice information corresponding to the text to be synthesized according to the N phoneme information and the N sub-prosody identification information, wherein the N sub-prosody identification information is extended into the corresponding phoneme information in the N phoneme information.

7. The device according to claim 6, characterized in that, The modal particle information further includes modal particle position information, and the modal particle position information is used to identify the addition position of the target modal particle in the text information to be synthesized; The generating unit is specifically configured to: Add the modal particle identification information at the addition position in the text information to be synthesized according to the modal particle position information to generate the comprehensive text information to be synthesized.

8. The device according to claim 6, wherein The prosody information further includes prosody boundary information, and the prosody boundary information is used to identify the text position corresponding to the prosody identification information in the text information to be synthesized; The generating unit is specifically configured to: Determine the initial comprehensive text information to be synthesized according to the text information to be synthesized and the modal particle information; Add the prosody identification information at the text position in the initial comprehensive text information to be synthesized according to the prosody boundary information to generate the comprehensive text information to be synthesized.

9. The device according to claim 6, characterized in that, The synthesizing unit is specifically configured to: Synthesize the voice information corresponding to the text to be synthesized through an acoustic model according to the N phoneme information and the N sub-prosody identification information.

10. The device according to claim 9, characterized in that, The acoustic model is trained in the following manner: Obtain sample text information, and the sample text information has corresponding sample voice information; Determine the sample modal particle information and sample prosody information corresponding to the sample text information; Generate sample comprehensive text information to be synthesized based on the sample modal particle information, sample text information, and the sample prosody information; Determine M phoneme information corresponding to the sample comprehensive text information to be synthesized, where the M phoneme information includes the phoneme information corresponding to the sample text information and the phoneme information corresponding to the sample modal particle information; Determine M sub-prosody identification information corresponding to the M phoneme information according to the sample prosody information; Splice the M phoneme information with the corresponding sub-prosody identification information in the M sub-prosody identification information respectively to obtain M input information; Determine the pending voice information corresponding to the sample text information through an initial acoustic model according to the M input information; Adjust the initial acoustic model according to the difference between the pending voice information and the sample voice information to obtain the acoustic model.

11. A computer device, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the voice synthesis method according to any one of claims 1-5 according to the instructions in the program code.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the voice synthesis method according to any one of claims 1-5.

13. A computer program product comprising instructions, characterized in that, When it runs on a computer, it causes the computer to execute the voice synthesis method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and computer readable storage medium

    CN113838448A