Speech synthesis method and apparatus, readable medium, and electronic device
By introducing liaison tone sandhi features into the speech synthesis process, the problem of unnatural speech tones in existing technologies is solved, achieving a more natural speech synthesis effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2026-03-20
AI Technical Summary
Current speech synthesis technology lacks effective learning of tone sandhi, resulting in unnatural tones in synthesized speech.
By determining the tone and prosody information of the target text, a synthesized audio corresponding to the target text is generated, and the tone sandhi feature is introduced to control the speech synthesis process.
It improves the controllability of tone sandhi in speech synthesis and enhances the naturalness of speech synthesis.
Smart Images

Figure CN115148186B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to a speech synthesis method, device, readable medium and electronic equipment. BACKGROUND
[0002] Speech synthesis technology can convert any text into corresponding audio, which usually includes two parts, one part is to analyze the text to obtain linguistic information, and the other part is to generate sound waveform based on the analysis result. In related technologies, there is usually a lack of learning of the feature of connected tone change, so that the tone of the synthesized speech cannot be effectively controlled, resulting in unnatural synthesized audio. SUMMARY
[0003] This part is provided to briefly introduce the concepts, which will be described in detail in the following specific embodiments. This part is not intended to identify key features or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0004] In a first aspect, the present disclosure provides a speech synthesis method, comprising:
[0005] determining tone annotation information of a target text to be processed, wherein the tone annotation information comprises a connected tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate the pitch change trend of the text unit;
[0006] determining prosodic annotation information of the target text and a phoneme sequence corresponding to the target text;
[0007] generating synthesized audio corresponding to the target text according to the tone annotation information, the prosodic annotation information and the phoneme sequence.
[0008] In a second aspect, the present disclosure provides a speech synthesis device, comprising:
[0009] a first determining module configured to determine tone annotation information of a target text to be processed, wherein the tone annotation information comprises a connected tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate the pitch change trend of the text unit;
[0010] a second determining module configured to determine prosodic annotation information of the target text and a phoneme sequence corresponding to the target text;
[0011] The generating module is configured to generate synthesized audio corresponding to the target text according to the tone labeling information, the prosody labeling information, and the phoneme sequence.
[0012] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program which, when executed by a processing device, implements the steps of the method of the first aspect of the present disclosure.
[0013] In a fourth aspect, the present disclosure provides an electronic device comprising:
[0014] a storage device having stored thereon at least one computer program;
[0015] at least one processing device configured to execute the at least one computer program stored in the storage device to implement the steps of the method of the first aspect of the present disclosure.
[0016] By the above technical solution, the tone labeling information of the target text to be processed is determined, wherein the tone labeling information comprises a connected tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate a pitch change trend of the text unit. The prosody labeling information of the target text and the phoneme sequence corresponding to the target text are determined. According to the tone labeling information, the prosody labeling information, and the phoneme sequence, the synthesized audio corresponding to the target text is generated. Thus, in the speech synthesis for the target text, in addition to using the prosody features, the connected tone change features are further introduced, so that the connected tone change mode can be directly controlled in the speech synthesis, the controllability of the connected tone change phenomenon in the speech synthesis is improved, and the naturalness of the synthesized speech is further improved.
[0017] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0019] Figure 1 is a flowchart of a speech synthesis method according to an embodiment of the present disclosure;
[0020] Figure 2 is a block diagram of a speech synthesis device according to an embodiment of the present disclosure;
[0021] Figure 3 shows a structural schematic diagram of an electronic device suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0022] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted in a limited sense as set forth in the embodiments set forth herein, but rather, the embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] It should be understood that each of the steps described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0024] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.
[0025] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0026] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0028] All actions of obtaining signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.
[0029] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means in accordance with relevant laws and regulations.
[0030] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application program, a server or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.
[0031] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0032] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0033] At the same time, it can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0034] Figure 1 A flowchart of a voice synthesis method according to an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method provided by the present disclosure can include steps 11-13. Figure 1
[0035] In step 11, the tone marking information of the target text to be processed is determined.
[0036] In the present disclosure, the tone marking information includes the connected tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate the pitch change trend of the text unit.
[0037] The unit text is the smallest unit to constitute the target text. For example, if the target text is a Chinese text, the unit text is a single character, that is, a single syllable, and accordingly, the text unit is composed of at least one character.
[0038] The connected tone change type of the text unit can refer to the following definition:
[0039] If the text unit is composed of one unit text, the connected tone change type of the text unit is one of the following: a first type used to indicate that the pitch is the original tone of the text unit, and a second type used to indicate that the pitch is changed from the pitch starting point of the original tone of the text unit to a flat tone.
[0040] If the text unit is composed of two unit texts, the connected reading tone type of the text unit is one of the following: a third type for indicating that the pitch is lowered from high level tone to low level tone, a fourth type for indicating that the pitch is lowered from high level tone to medium level tone, a fifth type for indicating that the pitch is raised from medium level tone to high level tone, a sixth type for indicating that the pitch is raised from low level tone to high level tone, a seventh type for indicating that the pitch is raised from low level tone to low rising tone, and an eighth type for indicating that the pitch is raised from low level tone to medium level tone.
[0041] If the text unit is composed of more than two unit texts, the connected reading tone type of the text unit is one of the following: the third type, the fourth type, the seventh type, the eighth type, a ninth type for indicating that the pitch is raised from medium level tone to high level tone and then lowered to low level tone, a tenth type for indicating that the pitch is raised from medium level tone to high level tone and then lowered to medium level tone, an eleventh type for indicating that the pitch is raised from low level tone to high level tone and then lowered to low level tone, and a twelfth type for indicating that the pitch is raised from low level tone to high level tone and then lowered to medium level tone.
[0042] For example, the connected reading tone type of the text unit and its pitch change trend can refer to the following table:
[0043]
[0044] Wherein, H represents high level tone, M represents medium level tone, L represents low level tone, and R represents low rising tone. The symbol “-” is used to separate the two adjacent unit texts. Taking the target text as a Chinese text as an example, the unit text is a single character, that is, a single syllable. Accordingly, H-L represents a two-syllable word tone, the first syllable of which is high level tone and the second syllable of which is low level tone, and the overall pitch curve changes from high level tone to low level tone.
[0045] In the above table, [M]{n} means that n Ms are connected, that is, n medium level tones are connected, such as M{2} which is M-M, M{3} which is M-M-M, and the rest are the same.
[0046] It should be noted that, taking the third type as an example, if the text unit is composed of four unit texts, the pitch change trend is H-M-M-L, and the pitch can gradually decrease from high level tone to low level tone, that is, the pitch of the former can be slightly higher than that of the latter although both of them are medium level tones.
[0047] In step 12, the prosodic annotation information of the target text and the phoneme sequence corresponding to the target text are determined.
[0048] In the present disclosure, the phoneme sequence corresponding to the text to be synthesized can be obtained by a Grapheme-to-Phoneme (G2P) model.
[0049] For example, the G2P model can employ a recurrent neural network (RNN) and a long short-term memory (LSTM) to implement the conversion from a character to a phoneme.
[0050] The prosody annotation information is used to reflect the content related to prosody, which can include but is not limited to prosody boundary information. Among them, the prosody boundary (break index, which can be abbreviated as BRK) is used to describe the form of information organization, sentence division in the speech flow. Optionally, the prosody boundary information can include but is not limited to sentence boundary, intonation phrase boundary, prosodic phrase boundary, and prosodic word boundary.
[0051] In step 13, according to the tone annotation information, the prosody annotation information and the phoneme sequence, the synthesized audio corresponding to the target text is generated.
[0052] Through the above technical solution, the tone annotation information of the target text to be processed is determined, wherein the tone annotation information includes the connected reading tone type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected reading tone type of the text unit is used to indicate the pitch change trend of the text unit. The prosody annotation information of the target text and the phoneme sequence corresponding to the target text are determined, and according to the tone annotation information, the prosody annotation information and the phoneme sequence, the synthesized audio corresponding to the target text is generated. Therefore, in the speech synthesis for the target text, in addition to using the prosody feature, the connected reading tone feature is further introduced, so that the connected reading tone mode can be directly controlled in the speech synthesis, the controllability of the connected reading tone phenomenon in the speech synthesis is improved, and the naturalness of the synthesized speech is improved.
[0053] In order to make those skilled in the art more understand the speech synthesis method provided by the present disclosure, the above steps are illustrated in detail as follows.
[0054] First, the related content of the tone annotation information used by the present disclosure is explained.
[0055] As described above, the tone annotation information includes the connected reading tone type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected reading tone type of the text unit is used to indicate the pitch change trend of the text unit. The determination method of the tone annotation information of the target text is described in detail as follows.
[0056] In a possible implementation, the tone labeling information of the target text can be determined by manual labeling. That is, a tone labeling operation for the target text can be received, and the tone labeling information of the target text can be generated according to the tone labeling operation.
[0057] That is, the labeling personnel can directly perform a labeling operation on the tone of the target text, and label the elision feature expected to be heard from the synthesized audio into the target text. The labeling operation on the tone can be performed in two steps.
[0058] In the first step, the text units can be divided based on the content of the target text, that is, the target text is divided into multiple text units by generating boundaries.
[0059] In the second step, the elision feature type of each text unit is labeled.
[0060] In another possible implementation, the determination of the tone labeling information of the target text can be implemented by a pre-trained tone labeling model. Accordingly, the tone labeling information of the target text can be obtained in the following manner:
[0061] The target text is input into the tone labeling model to obtain an output result of the tone labeling model;
[0062] According to the output result, the tone labeling information of the target text is determined.
[0063] The tone labeling model is trained based on a second training text with tone labeling information, and the output result includes boundary information for dividing the target text into multiple text units and an elision feature type corresponding to each text unit.
[0064] The second training text can be a text extracted from a real existing voice. For such a voice, the labeling personnel can mark the appropriate position in the text by listening to the voice to obtain the tone labeling information of the second training text. Such labeling mainly depends on the listening of the labeling personnel, and the labeling idea can refer to the description of the elision feature type in the foregoing.
[0065] In a possible implementation, the output result can be directly used as the tone labeling information of the target text.
[0066] In another possible implementation, according to the output result, the determination of the tone labeling information of the target text can include the following steps:
[0067] It is determined whether a correction instruction for the output result is received;
[0068] According to the correction instruction, the output result is corrected, and the corrected output result is determined as the tonal marking information of the target text.
[0069] The correction instruction is used to indicate at least one of the boundary information and the connected tone change type in the output result.
[0070] That is, the annotator can also correct the output result of the model through the correction instruction, which can improve the efficiency of determining the tonal marking information and ensure high accuracy of the annotation.
[0071] In the above manner, the tonal marking information corresponding to the real voice of the second training text can be obtained. Thus, the second training text is taken as the input of the neural network model, and the tonal marking information corresponding to the second training text is taken as the target output of the model, the neural network model is trained, and after the training is completed, the tonal marking model capable of automatically generating tonal marking information for a text is obtained. In this way, a piece of text is input into the tonal marking model, and the tonal marking information corresponding to the piece of text output by the tonal marking model can be automatically obtained, without the need for manual annotation, which is beneficial to improving the efficiency of determining the tonal marking information.
[0072] Back to Figure 1 In step 13, the synthesized audio corresponding to the target text is generated according to the tonal marking information, the prosodic marking information, and the phoneme sequence.
[0073] In one possible implementation, step 13 can include the following steps:
[0074] According to the tonal marking information, a connected tone change label sequence corresponding to the target text is determined;
[0075] According to the prosodic marking information, a prosodic label sequence corresponding to the target text is determined;
[0076] According to the connected tone change label sequence, the prosodic label sequence, and the phoneme sequence, a pre-trained speech synthesis model is used to generate acoustic feature information corresponding to the target text;
[0077] The acoustic feature information is synthesized by using a vocoder to generate synthesized audio corresponding to the target text.
[0078] The idea of determining the connected tone change label is that the unit texts of the same text unit share the same connected tone change label, that is, the connected tone change labels of the unit texts constituting a text unit are consistent with the connected tone change label of the text unit. Further, according to the appearance order of the text unit in the target text, the connected tone change label sequence can be obtained.
[0079] As mentioned above, the prosody annotation information can comprise prosody boundary information, and correspondingly, the prosody label can comprise prosody boundary label.
[0080] The prosody annotation information generally annotates a certain position of the text, for example, a certain position of the text is a prosodic phrase boundary. In order to facilitate subsequent speech synthesis and ensure that the prosody annotation information can correspond to each phoneme of the text to be synthesized, the prosody label at the phoneme level can be further determined based on the prosody annotation information.
[0081] The idea of determining the prosody label is that, for the phoneme position with prosody annotation information, the label content is generated according to the annotation information, and for the phoneme position without prosody annotation information, the specified replacement content is used for replacement. For example, for the phoneme sequence {A1, A2, A3, A4, A5, A6}, it is assumed that the prosody annotation information comprises prosody boundary information, and the annotation content is that there is a prosodic phrase boundary at A2, and there is a intonational phrase boundary at A5, and it is specified that the prosodic phrase boundary is represented by 3, the intonational phrase boundary is represented by 4, and no label is represented by N2. Then the determined prosody boundary label is {N2, 3, N2, N2, 4, N2}.
[0082] In generating the acoustic feature information corresponding to the target text, the coarticulation label sequence, the prosody label sequence and the phoneme sequence can be input into a pre-trained speech synthesis model to obtain the acoustic feature information corresponding to the target text. For example, the acoustic feature information can be Mel spectrum, linear spectrum, etc.
[0083] The speech synthesis model can comprise an encoding network, an attention network and a decoding network. The encoding network is used to generate a text representation sequence according to the concatenation vector corresponding to the coarticulation label sequence, the prosody label sequence and the phoneme sequence; the attention network is used to generate a semantic representation according to the text representation sequence; and the decoding network is used to output the acoustic feature information corresponding to the target text according to the semantic representation.
[0084] The input of the encoding network of the speech synthesis model is the vector representation of the target text, which can comprise a first vector obtained by vectorizing the phoneme sequence, a second vector obtained by vectorizing the coarticulation label sequence and a third vector obtained by vectorizing the prosody label sequence. The above-mentioned vectors are concatenated to form a concatenation vector as the input of the encoding network. Then, the encoding network outputs a text representation sequence (TE, text embedding) of the target text. The text representation sequence output by the encoding network is input into the attention network to generate a context vector C as the semantic representation of the target text. The semantic representation generated by the attention network is input into the decoding network, and the decoding network outputs the acoustic feature information corresponding to the target text.
[0085] For example, the speech synthesis model is trained in the following manner:
[0086] The first training sample is obtained, wherein each first training sample includes a training phoneme sequence corresponding to the first training text, a training connected tone label sequence, a training prosody label sequence, and training acoustic feature information corresponding to the first training text;
[0087] The model is trained by taking the concatenation vector corresponding to the training phoneme sequence, the training connected tone label sequence, and the training prosody label sequence as the input of the model, and taking the training acoustic feature information as the target output of the model, to obtain the trained speech synthesis model.
[0088] For example, the speech synthesis model described above can use a Tacotron model.
[0089] The first training text corresponds to an audio, and the acoustic feature information of the audio is determined as the training acoustic feature information.
[0090] In the present disclosure, the training phoneme sequence of the first training text can be determined in a manner similar to the determination of the phoneme sequence of the target text in step 12, and the training connected tone label sequence and the training prosody label sequence of the first training text can be determined in a manner similar to the determination of the connected tone label sequence and the prosody label sequence of the target text described above. The above content will not be repeated here.
[0091] The purpose of training the speech synthesis model is to make the synthesized audio output by the model infinitely close to the actual audio of the first training sample, i.e., to make the acoustic feature information output by the model infinitely close to the training acoustic feature information. Therefore, the loss value of the model can be calculated based on the training acoustic feature information and the acoustic feature information output by the model during training, and the internal parameters of the current model can be adjusted using the loss value. After that, the adjusted model is used for the next training, and the cycle is repeated until the condition for stopping training is met, and the trained speech synthesis model can be obtained.
[0092] The trained speech synthesis model obtained through the above training steps can be used in a speech synthesis scenario. That is:
[0093] According to the connected tone label sequence, the prosody label sequence, and the phoneme sequence, the pre-trained speech synthesis model is used to generate acoustic feature information corresponding to the target text;
[0094] The vocoder is used to synthesize the acoustic feature information to generate synthesized audio corresponding to the target text.
[0095] After obtaining the acoustic feature information of the target text, the acoustic feature information can be input into a vocoder (for example, a Wavenet vocoder, a Griffin-Lim vocoder) for speech synthesis, so as to obtain the synthesized audio corresponding to the text to be synthesized.
[0096] Figure 2 is a block diagram of a speech synthesis device provided according to an embodiment of the present disclosure. As shown in Figure 2 the device 20 includes:
[0097] A first determination module 21 is configured to determine tonal annotation information of a target text to be processed, wherein the tonal annotation information includes a connected tone change type of each text unit in the target text, and the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate the pitch change trend of the text unit.
[0098] A second determination module 22 is configured to determine prosodic annotation information of the target text and a phoneme sequence corresponding to the target text.
[0099] A generation module 23 is configured to generate synthesized audio corresponding to the target text according to the tonal annotation information, the prosodic annotation information and the phoneme sequence.
[0100] Optionally, the generation module 23 includes:
[0101] A first determination sub-module is configured to determine a connected tone change label sequence corresponding to the target text according to the tonal annotation information.
[0102] A second determination sub-module is configured to determine a prosodic label sequence corresponding to the target text according to the prosodic annotation information.
[0103] A first generation sub-module is configured to generate acoustic feature information corresponding to the target text by using a pre-trained speech synthesis model according to the connected tone change label sequence, the prosodic label sequence and the phoneme sequence.
[0104] A synthesis sub-module is configured to perform speech synthesis on the acoustic feature information by using a vocoder to generate synthesized audio corresponding to the target text.
[0105] Optionally, the speech synthesis model includes an encoding network, an attention network and a decoding network; wherein:
[0106] The encoding network is configured to generate a text representation sequence according to a splicing vector corresponding to the connected tone change label sequence, the prosodic label sequence and the phoneme sequence.
[0107] The attention network is configured to generate a semantic representation according to the text representation sequence.
[0108] The decoding network is configured to output acoustic feature information corresponding to the target text according to the semantic representation.
[0109] Optionally, the speech synthesis model is obtained by the following modules.
[0110] The obtaining module is configured to obtain first training samples, wherein each of the first training samples comprises a training phoneme sequence, a training connected tone label sequence, a training prosody label sequence corresponding to a first training text, and training acoustic feature information corresponding to the first training text.
[0111] The training module is configured to train the model by taking the concatenation vector corresponding to the training phoneme sequence, the training connected tone label sequence and the training prosody label sequence as the input of the model and taking the training acoustic feature information as the target output of the model, so as to obtain the trained speech synthesis model.
[0112] Optionally, the first determining module 21 comprises:
[0113] The processing submodule is configured to input the target text into a tone labeling model to obtain an output result of the tone labeling model, wherein the tone labeling model is trained based on a second training text with tone labeling information, and the output result comprises boundary information for dividing the target text into a plurality of text units and a connected tone type corresponding to each text unit.
[0114] The third determining submodule is configured to determine the tone labeling information of the target text according to the output result.
[0115] Optionally, the third determining submodule comprises:
[0116] The fourth determining submodule is configured to determine whether a correction instruction for the output result is received, wherein the correction instruction is used to indicate a change in at least one of the boundary information and the connected tone type in the output result.
[0117] The correction submodule is configured to correct the output result according to the correction instruction, and determine the corrected output result as the tone labeling information of the target text.
[0118] Optionally, if a text unit is composed of one unit text, the connected tone type of the text unit is one of the following: a first type indicating that the pitch is the original tone of the text unit, and a second type indicating that the pitch is changed from the pitch starting point of the original tone of the text unit to a flat tone.
[0119] If the text unit is composed of two unit texts, the connected reading inflection type of the text unit is one of the following: a third type for indicating a decrease in pitch from high level to low level, a fourth type for indicating a decrease in pitch from high level to medium level, a fifth type for indicating an increase in pitch from medium level to high level, a sixth type for indicating an increase in pitch from low level to high level, a seventh type for indicating an increase in pitch from low level to low rising tone, an eighth type for indicating an increase in pitch from low level to medium level;
[0120] If the text unit is composed of more than two unit texts, the connected reading inflection type of the text unit is one of the following: the third type, the fourth type, the seventh type, the eighth type, a ninth type for indicating an increase in pitch from medium level to high level and then a decrease to low level, a tenth type for indicating an increase in pitch from medium level to high level and then a decrease to medium level, an eleventh type for indicating an increase in pitch from low level to high level and then a decrease to low level, a twelfth type for indicating an increase in pitch from low level to high level and then a decrease to medium level.
[0121] As to the apparatus in the above-described embodiments, the specific manners in which the respective modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.
[0122] Reference is made below to Figure 3 which shows a structure of an electronic device (600) suitable for use in implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of embodiments of the present disclosure.
[0123] As shown in Figure 3 , the electronic device 600 can include a processing device (e.g., a central processor, a graphics processor, or the like) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage device 608. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0124] In general, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 600 is shown with various devices, but it is understood that all of the illustrated devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.
[0125] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing devices 601, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0126] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0127] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0128] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.
[0129] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: determine tone labeling information of a target text to be processed, wherein the tone labeling information comprises a connected tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate a pitch change trend of the text unit; determine prosody labeling information of the target text and a phoneme sequence corresponding to the target text; and generate synthesized audio corresponding to the target text according to the tone labeling information, the prosody labeling information, and the phoneme sequence.
[0130] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0131] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified functions. It should also be noted that, in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.
[0132] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the first obtaining module can also be described as "a module for obtaining at least two internet protocol addresses".
[0133] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0134] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0135] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, the method comprising:
[0136] determining tone labeling information of a target text to be processed, wherein the tone labeling information comprises a connected reading tone change type of each text unit in the target text, wherein the text unit is composed of at least one unit text, and the connected reading tone change type of the text unit is used to indicate a pitch change trend of the text unit;
[0137] determining prosody labeling information of the target text and a phoneme sequence corresponding to the target text;
[0138] generating synthesized audio corresponding to the target text according to the tone labeling information, the prosody labeling information, and the phoneme sequence.
[0139] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, the method comprising: generating synthesized audio corresponding to the target text according to the tone labeling information, the prosody labeling information, and the phoneme sequence.
[0140] determine a sequence of connected tone change labels corresponding to the target text according to the tone marking information;
[0141] determine a sequence of prosody labels corresponding to the target text according to the prosody marking information;
[0142] generate acoustic feature information corresponding to the target text by using a pre-trained speech synthesis model according to the sequence of connected tone change labels, the sequence of prosody labels, and the phoneme sequence;
[0143] perform speech synthesis on the acoustic feature information by using a vocoder to generate synthesized audio corresponding to the target text.
[0144] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, and the speech synthesis model includes an encoding network, an attention network, and a decoding network; wherein:
[0145] The encoding network is configured to generate a text representation sequence according to a splicing vector corresponding to the sequence of connected tone change labels, the sequence of prosody labels, and the phoneme sequence;
[0146] The attention network is configured to generate a semantic representation according to the text representation sequence;
[0147] The decoding network is configured to output acoustic feature information corresponding to the target text according to the semantic representation.
[0148] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, and the speech synthesis model is obtained by the following method:
[0149] obtain first training samples, wherein each of the first training samples includes a training phoneme sequence corresponding to a first training text, a training sequence of connected tone change labels, a training sequence of prosody labels, and training acoustic feature information corresponding to the first training text;
[0150] perform model training by taking a splicing vector corresponding to the training phoneme sequence, the training sequence of connected tone change labels, and the training sequence of prosody labels as an input of the model, and taking the training acoustic feature information as a target output of the model, to obtain the trained speech synthesis model.
[0151] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, and the determination of the tone marking information of the target text to be processed includes:
[0152] input the target text into a tone labeling model to obtain an output result of the tone labeling model, wherein the tone labeling model is trained based on second training texts with tone labeling information, and the output result includes boundary information for dividing the target text into a plurality of text units and a respective continuous tone change type of each text unit;
[0153] According to the output result, the tone labeling information of the target text is determined.
[0154] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, and the tone labeling information of the target text is determined according to the output result, including:
[0155] It is determined whether a correction instruction for the output result is received, the correction instruction being used to indicate a change in at least one of the boundary information and the continuous tone change type in the output result;
[0156] According to the correction instruction, the output result is corrected, and the corrected output result is determined as the tone labeling information of the target text.
[0157] According to one or more embodiments of the present disclosure, a speech synthesis method is provided, and if a text unit is composed of one unit text, the continuous tone change type of the text unit is one of the following: a first type indicating that the pitch is the original tone of the text unit, a second type indicating that the pitch changes from the starting point of the original tone of the text unit to a flat tone;
[0158] If a text unit is composed of two unit texts, the continuous tone change type of the text unit is one of the following: a third type indicating that the pitch decreases from a high flat tone to a low flat tone, a fourth type indicating that the pitch decreases from a high flat tone to a medium flat tone, a fifth type indicating that the pitch increases from a medium flat tone to a high flat tone, a sixth type indicating that the pitch increases from a low flat tone to a high flat tone, a seventh type indicating that the pitch increases from a low flat tone to a low rising tone, and an eighth type indicating that the pitch increases from a low flat tone to a medium flat tone;
[0159] If a text unit is composed of more than two unit texts, the continuous tone change type of the text unit is one of the following: the third type, the fourth type, the seventh type, the eighth type, a ninth type indicating that the pitch increases from a medium flat tone to a high flat tone and then decreases to a low flat tone, a tenth type indicating that the pitch increases from a medium flat tone to a high flat tone and then decreases to a medium flat tone, an eleventh type indicating that the pitch increases from a low flat tone to a high flat tone and then decreases to a low flat tone, and a twelfth type indicating that the pitch increases from a low flat tone to a high flat tone and then decreases to a medium flat tone.
[0160] According to one or more embodiments of the present disclosure, a speech synthesis device is provided, the device comprising:
[0161] a first determining module configured to determine tone labeling information of a target text to be processed, wherein the tone labeling information comprises a connected tone change type of each text unit in the target text, and wherein the text unit is composed of at least one unit text, and the connected tone change type of the text unit is used to indicate a pitch variation trend of the text unit;
[0162] a second determining module configured to determine prosody labeling information of the target text and a phoneme sequence corresponding to the target text;
[0163] a generating module configured to generate a synthesized audio corresponding to the target text according to the tone labeling information, the prosody labeling information and the phoneme sequence.
[0164] According to one or more embodiments of the present disclosure, a computer readable medium is provided, and the computer readable medium has a computer program stored thereon, wherein the computer program is executed by a processing device to implement steps of the speech synthesis method provided by any of the embodiments of the present disclosure.
[0165] According to one or more embodiments of the present disclosure, an electronic device is provided, the device comprising:
[0166] a storage device having at least one computer program stored thereon;
[0167] at least one processing device configured to execute the at least one computer program in the storage device to implement steps of the speech synthesis method provided by any of the embodiments of the present disclosure.
[0168] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions of the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0169] Moreover, while operations have been depicted in a particular order, this should not be understood as requiring such an order nor limiting it to only those operations shown and described. One of ordinary skill in the art will recognize that many of the operations can be performed in a differing order, or be performed concurrently, that some operations can be performed in any order or omitted, and that some operations can be performed in parallel. Similarly, while several specific implementation details have been discussed in the context of the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.
Claims
1. A speech synthesis method, characterized in that, The method includes: The target text to be processed is input into the tone annotation model to obtain the output result of the tone annotation model. The tone annotation model is trained based on a second training text with tone annotation information. The output result includes boundary information for dividing the target text into multiple text units and the tone sandhi type corresponding to each text unit. Based on the output results, the tone marking information of the target text is determined. The tone marking information includes the tone sandhi type of each text unit in the target text. The text unit consists of at least one unit text, which is a single syllable. The tone sandhi type of the text unit is used to indicate the pitch change trend of the text unit. Determine the prosodic annotation information of the target text and the phoneme sequence corresponding to the target text; Based on the tone annotation information, the prosody annotation information, and the phoneme sequence, a synthesized audio corresponding to the target text is generated.
2. The method according to claim 1, characterized in that, The step of generating a synthesized audio corresponding to the target text based on the tone annotation information, the prosodic annotation information, and the phoneme sequence includes: Based on the tone marking information, determine the tone sandhi tag sequence corresponding to the target text; Based on the prosodic annotation information, determine the prosodic tag sequence corresponding to the target text; Based on the tone sandhi tag sequence, the prosody tag sequence, and the phoneme sequence, an acoustic feature information corresponding to the target text is generated using a pre-trained speech synthesis model. The acoustic feature information is synthesized using a vocoder to generate synthesized audio corresponding to the target text.
3. The method according to claim 2, characterized in that, The speech synthesis model includes an encoding network, an attention network, and a decoding network; wherein: The encoding network is used to generate a text representation sequence based on the concatenation vectors corresponding to the tone sandhi tag sequence, the prosody tag sequence, and the phoneme sequence; The attention network is used to generate semantic representations based on the text representation sequence; The decoding network is used to output acoustic feature information corresponding to the target text based on the semantic representation.
4. The method according to claim 2, characterized in that, The speech synthesis model is obtained through the following methods: Obtain a first training sample, wherein each first training sample includes a training phoneme sequence, a training tone sandhi tag sequence, a training prosody tag sequence corresponding to the first training text, and training acoustic feature information corresponding to the first training text. The speech synthesis model is trained by using the concatenated vectors corresponding to the training phoneme sequence, the training tone sandhi tag sequence, and the training prosodic tag sequence as inputs to the model, and the training acoustic feature information as the target output of the model.
5. The method according to claim 1, characterized in that, The step of determining the tone marking information of the target text based on the output result includes: Determine whether a correction instruction for the output result has been received, the correction instruction being used to indicate a change in at least one of boundary information and tone sandhi type in the output result; According to the correction instruction, the output result is corrected, and the corrected output result is determined as the tone annotation information of the target text.
6. The method according to any one of claims 1-5, characterized in that, If a text unit consists of a single text unit, the tone sandhi type of that text unit is one of the following: the first type, which indicates that the pitch is the original tone of the text unit; and the second type, which indicates that the pitch changes from the pitch starting point of the original tone of the text unit to a level tone. If a text unit consists of two text units, the tone sandhi type of the text unit is one of the following: the third type used to indicate that the pitch drops from a high level tone to a low level tone; the fourth type used to indicate that the pitch drops from a high level tone to a middle level tone; the fifth type used to indicate that the pitch rises from a middle level tone to a high level tone; the sixth type used to indicate that the pitch rises from a low level tone to a high level tone; the seventh type used to indicate that the pitch rises from a low level tone to a low rising tone; and the eighth type used to indicate that the pitch rises from a low level tone to a middle level tone. If a text unit consists of more than two text units, the tone sandhi type of the text unit is one of the following: the third type, the fourth type, the seventh type, the eighth type, the ninth type used to indicate that the pitch rises from the middle level tone to the high level tone and then falls back to the low level tone, the tenth type used to indicate that the pitch rises from the middle level tone to the high level tone and then falls back to the middle level tone, the eleventh type used to indicate that the pitch rises from the low level tone to the high level tone and then falls back to the low level tone, and the twelfth type used to indicate that the pitch rises from the low level tone to the high level tone and then falls back to the middle level tone.
7. A speech synthesis device, characterized in that, The device includes: The first determining module is used to input the target text to be processed into the tone annotation model and obtain the output result of the tone annotation model. The tone annotation model is trained based on a second training text with tone annotation information. The output result includes boundary information for dividing the target text into multiple text units and the tone sandhi type corresponding to each text unit. Based on the output result, the tone annotation information of the target text is determined. The tone annotation information includes the tone sandhi type of each text unit in the target text. The text unit consists of at least one unit text, the unit text is a single syllable, and the tone sandhi type of the text unit is used to indicate the pitch change trend of the text unit. The second determining module is used to determine the prosodic annotation information of the target text and the phoneme sequence corresponding to the target text; The generation module is used to generate a synthesized audio corresponding to the target text based on the tone annotation information, the prosody annotation information and the phoneme sequence.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.
9. An electronic device, characterized in that, include: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Speech synthesis method and device, readable storage medium and electronic equipment
CN114155829A
Pitch model production device, method and pitch model production program
CN1664922A