Prosody labeling method, acoustic model training method, speech synthesis method and device

By clustering and annotating phoneme-level prosodic features, the problem that acoustic models in existing technologies cannot learn the speaker's speech style and emotions is solved, and highly realistic speech synthesis is achieved.

CN116129859BActive Publication Date: 2025-10-24MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211435105.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-10-24
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

In existing technologies, it is impossible to obtain an acoustic model with high human-like speech effect by simply using prosodic annotation at the pause level, and it is impossible to effectively learn the speaker's speech style and emotional characteristics.

Method used

By dividing and clustering phonemes and audio data at the phoneme level, prosodic features of pitch, volume, and duration are obtained, forming rich prosodic markers for training acoustic models.

Benefits of technology

It improves the acoustic model's ability to learn the speaker's speech style and emotions, and synthesizes highly realistic and human-like speech audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129859B_ABST
    Figure CN116129859B_ABST
Patent Text Reader

Abstract

The present disclosure provides a prosody labeling method, an acoustic model training method, a speech synthesis method and device, and relates to the technical field of speech synthesis. The method comprises: dividing first audio data into a plurality of second audio data according to the correspondence between a plurality of phonemes in text data and the first audio data corresponding to the text data; clustering prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters; and determining prosodic labels based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters to obtain prosodic labels of the plurality of phonemes respectively. The embodiment of the present disclosure clusters the frame-level prosodic features of the phonemes to obtain the prosodic labels. The training text used for training the acoustic model is labeled by the prosodic labels at the phoneme level. Compared with the traditional word / sentence-level prosodic labeling method, the prosodic labels can better assist the acoustic model to learn the emotions, speech styles and other characteristics of the speaker, so as to synthesize high-simulation-degree speech audio.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of speech synthesis, and in particular, to a prosody labeling method, an acoustic model training method, a speech synthesis method and device. BACKGROUND

[0002] An acoustic model is one of the important components of speech synthesis (TTS) technology. In the training process of the acoustic model, a large amount of training text with prosody labels is used to ensure that the trained acoustic model can predict the prosody in the text, so as to synthesize synthesized speech with prosody and not stiff. Therefore, it is very important to ensure the accuracy of the prosody labels in the text.

[0003] In the related art, the prosody labeling of the text is mainly based on pause prosody, that is, a simple prosody symbol is used to label the pause level when reading, so that the whole speech is interrupted. However, only the pause level is used to label the prosody of the training text of the acoustic model, and the acoustic model with high human-like voice effect cannot be obtained. SUMMARY

[0004] Therefore, the present disclosure provides a prosody labeling method, an acoustic model training method, a speech synthesis method and device.

[0005] In a first aspect, a prosody labeling method is provided, comprising:

[0006] According to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, the first audio data is divided into a plurality of second audio data, and the plurality of second audio data has a corresponding relationship with the plurality of phonemes; the prosody features of the plurality of second audio data are clustered to obtain a plurality of clustering clusters; the prosody features include pitch, volume and duration; each clustering cluster in the plurality of clustering clusters is used to represent a prosody label, and the prosody label is used to reflect the prosody features including a pitch, a volume and a duration; based on the prosody features of the plurality of second audio data and the plurality of clustering clusters, the prosody label determination processing is performed to obtain the prosody label of each of the plurality of phonemes.

[0007] In a second aspect, an acoustic model training method is provided, comprising:

[0008] A training set is constructed, the training set comprising text data and audio data corresponding to the text data; the plurality of phonemes in the text data are labeled by the method of the first aspect to obtain the prosody label of each of the plurality of phonemes; the acoustic model is trained using the training set and the prosody label of each of the plurality of phonemes, and the trained acoustic model is used for speech synthesis processing of the text to be synthesized to obtain synthesized speech.

[0009] In a third aspect, a speech synthesis method is provided, including: inputting to-be-synthesized text data into a pre-trained speech synthesis model to obtain a mel spectrum of the to-be-synthesized text data, the speech synthesis model being trained based on a training set and prosodic labels of a plurality of phonemes corresponding to text data of the training set, the prosodic labels of the plurality of phonemes corresponding to the text data being obtained by the method of the first aspect; and synthesizing synthesized speech of the to-be-synthesized text data based on the mel spectrum of the to-be-synthesized text data.

[0010] In a fourth aspect, a prosodic labeling device is provided, including:

[0011] A division module is configured to divide first audio data into a plurality of second audio data according to a correspondence between a plurality of phonemes in text data and the first audio data, the plurality of second audio data having a correspondence with the plurality of phonemes;

[0012] A clustering module is configured to cluster prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters, the prosodic features including pitch, volume, and duration, and each of the plurality of clustering clusters being used to represent a prosodic label, the prosodic label being used to reflect prosodic features including a pitch, a volume, and a duration.

[0013] A determination module is configured to determine the prosodic labels based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters to obtain the prosodic labels of the plurality of phonemes.

[0014] In a fifth aspect, an acoustic model training device is provided, including:

[0015] A construction module is configured to construct a training set, the training set including text data and audio data corresponding to the text data.

[0016] A labeling module is configured to label a plurality of phonemes in the text data by the method of the first aspect to obtain the prosodic labels of the plurality of phonemes.

[0017] A training module is configured to train an acoustic model using the training set and the prosodic labels of the plurality of phonemes, the trained acoustic model being used to perform speech synthesis processing on to-be-synthesized text to obtain synthesized speech.

[0018] In a sixth aspect, a speech synthesis device is provided, including:

[0019] An acquisition module is configured to input to-be-synthesized text data into a pre-trained speech synthesis model to obtain a mel spectrum of the to-be-synthesized text data, the speech synthesis model being trained based on a training set and prosodic labels of a plurality of phonemes corresponding to text data of the training set, the prosodic labels of the plurality of phonemes corresponding to the text data being obtained by the method of the first aspect.

[0020] a synthesis module configured to synthesize speech for the text data to be synthesized based on the mel-spectrogram of the text data to be synthesized.

[0021] In a seventh aspect, an electronic device is provided, comprising: a processor; and a memory storing executable instructions of the processor; wherein the processor is configured to perform the method of the first aspect; or the method of the second aspect; or the method of the third aspect, via execution of the executable instructions.

[0022] In an eighth aspect, a computer-readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, implements the method of the first aspect; or the method of the second aspect; or the method of the third aspect.

[0023] The prosody labeling method provided in the embodiments of the present disclosure can cluster the prosodic features of the plurality of second audio data after dividing the first audio data into the plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, and then determine the prosodic labels based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters, so as to obtain the respective prosodic labels of the plurality of phonemes. The embodiments of the present disclosure cluster the frame-level prosodic features of the phonemes to obtain the prosodic labels, and label the training text for training the acoustic model by using the prosodic labels at the phoneme level. Compared with the traditional prosodic labeling method at the word or sentence level, the embodiments of the present disclosure can better assist the acoustic model to learn the characteristics of the speaker's emotion, speech style, and the like, so as to synthesize a high-simulation-degree speech audio. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 A system architecture schematic diagram of a prosody labeling method in the embodiments of the present disclosure is shown.

[0025] Figure 2 A flowchart schematic diagram of a prosody labeling method in the embodiments of the present disclosure is shown.

[0026] Figure 3 A flowchart schematic diagram of a sound model training method in the embodiments of the present disclosure is shown.

[0027] Figure 4 A network structure schematic diagram of FastPitch in the related art is shown.

[0028] Figure 5 A network structure schematic diagram of an acoustic model with embedded prosody prediction function in the embodiments of the present disclosure is shown.

[0029] Figure 6 A flowchart schematic diagram of a speech synthesis method in the embodiments of the present disclosure is shown.

[0030] Figure 7 FIG. 1 shows a structural schematic diagram of a prosody labeling device according to an embodiment of the present disclosure.

[0031] Figure 8 FIG. 3 shows a structural schematic diagram of a sound model training device according to an embodiment of the present disclosure.

[0032] Figure 9 FIG. 4 shows a structural schematic diagram of a speech synthesis device according to an embodiment of the present disclosure.

[0033] Figure 10 FIG. 5 shows a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] Example implementations will now be described with reference to the drawings. However, example implementations can be implemented in various forms and should not be construed as being limited to the examples set forth herein; rather, these implementations are provided so that the present disclosure will be more thorough and complete, and will fully convey the concept of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations in any suitable manner.

[0035] In addition, the accompanying drawings are only schematic and are not necessarily drawn to scale. Identical components have been given the same reference numerals in the various drawings and the same reference designators have been maintained where possible in the drawings and the specification. Some of the blocks in the drawings are functional blocks that can be implemented in software, hardware, or a combination of software and hardware. Some of the blocks in the drawings can be implemented as software modules or code on a machine-readable medium.

[0036] An acoustic model is one of the important components of a speech synthesis (TTS) technology. In the training process of the acoustic model, a large amount of training text with prosody labels is used to ensure that the trained speech synthesis model can predict the prosody in the text, and therefore it is very important to ensure the accuracy of the prosody labels in the text.

[0037] In the related art, prosody labeling is usually based on a syntactic structure, and the content of the labeling is mainly pause prosody. The labeling result is at the granularity of words and phrases, and the granularity is coarse. That is, the related art can only mark sentence pauses and cannot help the acoustic model learn complete speech prosody information.

[0038] As an acoustic phenomenon, complete prosody features include pitch, duration, and volume. Prosody labeling that can reflect complete prosody features is necessary for an acoustic model to more effectively learn the prosody features and speaking style of a speaker, so as to synthesize speech with high anthropomorphism.

[0039] Therefore, the scheme provided by the present disclosure can divide the first audio data into a plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data. By clustering the prosodic features of the plurality of second audio data, the prosodic labels of the plurality of phonemes can be obtained. The prosodic labeling method provided by the present disclosure considers complete prosodic features, so that when the training text labeled by the method of the present disclosure is applied to acoustic model training, the acoustic model can more accurately learn the voice style and prosodic features of the speaker, and thus can synthesize speech audio with better emotional fidelity and style fidelity.

[0040] The following will be combined with Figure 1 The system architecture of the prosodic labeling method in the embodiments of the present disclosure is described.

[0041] Figure 1 An exemplary system architecture schematic diagram of the prosodic labeling method or prosodic labeling device applied to the embodiments of the present disclosure is shown. As Figure 1 shown, the system architecture 100 can include a sample collection device 101 and at least one special-purpose or general-purpose computer processing module.

[0042] It should be noted that the sample collection device 101 can be a microphone assembly, which can include a microphone, a microphone sleeve, a mounting rod, a connecting line, etc., or a wireless microphone or a microphone circuit, which is configured to collect sound information in a scene and convert the sound information into text data by any voice conversion method, thereby obtaining text data and audio data corresponding to the text data. The present disclosure does not limit this.

[0043] It should be noted that the at least one special-purpose or general-purpose computer processing module can be any electronic device capable of executing a computer program, which can include a processor 102 and a memory 103.

[0044] The processor 102 is configured to execute program instructions, for example, the sound model training method provided by the present disclosure can be executed. The memory 103 can exist in the system architecture 100 in different forms of program storage units or data storage units, such as hard disks, read-only memories (ROM), random access memories (RAM), which can be used to store various data files used in the process of processor processing and / or executing the prosody labeling method, and possible program instructions executed by the processor. Although not shown in the figure, the system architecture 100 can also include an input / output component to support the input / output of data streams between the prosody labeling device applying the system architecture 100 and other components (such as a screen display device). In addition, the prosody labeling device applying the system architecture 100 can also send and receive information and data from the network through the communication port.

[0045] Although in Figure 1 the sample collection device 101, the processor 102 and the memory 103 are shown as separate modules, those skilled in the art can understand that the above device modules can be realized as separate hardware devices, or integrated into one or more hardware devices, such as integrated into a smart watch or other smart device. As long as the principles described in the present disclosure can be realized, the specific implementation of different hardware devices should not be considered as a limitation on the protection scope of the present disclosure.

[0046] The present example embodiment will be described in detail below with reference to the accompanying drawings and examples.

[0047] First, a prosody labeling method is provided in the embodiment of the present disclosure, which can be executed by any electronic device with computing processing capability. The electronic device mentioned herein can include a terminal or a server, wherein the terminal can include a personal computer, a notebook computer, a tablet computer, a mobile phone, a personal digital assistant (PDA), smart glasses, a smart watch, a smart ring, a smart helmet, a vehicle-mounted terminal, and the like, and the server can include a standalone physical server, a server cluster composed of multiple servers, or a cloud server capable of cloud computing. Figure 2 A flowchart of a prosody labeling method in the embodiment of the present disclosure is shown, as Figure 2 shown, the prosody labeling method provided in the embodiment of the present disclosure includes the following steps:

[0048] S201, according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, the first audio data is divided into a plurality of second audio data.

[0049] It should be noted that a phoneme is the smallest unit of speech divided according to the natural properties of speech, and is analyzed according to the pronunciation action in a syllable. For example, for Chinese language, the initial and final consonants in the Chinese pinyin corresponding to the text data can be taken as the phonemes of the text data.

[0050] It should be noted that the text data is training text applicable to acoustic model training. The first audio data refers to an audio file corresponding to the text data, which can be obtained by a microphone component. The second audio data is an audio segment corresponding to each phoneme in the text data divided from the first audio data. Therefore, the number of second audio data is the same as the number of phonemes.

[0051] In some embodiments, dividing the first audio data into a plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data can include: first aligning the text data and the first audio data corresponding to the text data to obtain the time boundary of each phoneme in the first audio data. Then, the first audio data is divided into a plurality of second audio data according to the time boundary of each phoneme in the first audio data.

[0052] For example, by using a preset speech-to-text alignment tool, such as MFA (Montreal Forced Aligner, MFA) tool, the plurality of phonemes in the text data and the first audio data are aligned to obtain the start and end positions of the second audio data corresponding to each phoneme (initial or final consonant) in the first audio data. Then, each phoneme corresponding segment can be cut from the first audio data, i.e., the second audio data corresponding to each phoneme is obtained.

[0053] By dividing the first audio data at the phoneme level, the embodiments of the present disclosure can improve the granularity of the prosody labels obtained, so that the prosody labels labeled by the method provided by the present disclosure can represent more rich prosodic features when used for speech synthesis.

[0054] S202, clustering the prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters.

[0055] It should be noted that the prosodic features refer to the features such as duration, pitch and volume in speech in addition to voice quality. When we listen to speech, we will pay attention to the features such as duration, pitch and volume while listening to individual phonemes. These features can better help us understand the semantics.

[0056] In some embodiments, before clustering the prosodic features of the plurality of second audio data, the prosodic features of each second audio data can be calculated, wherein the prosodic features include duration, pitch and volume.

[0057] For example, assuming that the second audio data is 16K sampling rate audio, the number of frames of the second audio data corresponding to each phoneme can be calculated as the duration L, with 15 milliseconds as a frame unit and 5 milliseconds as a frame step value, while the pitch sequence P is obtained by obtaining the pitch of each frame of the L frames, and the volume sequence V is obtained by further calculating the average energy value of each frame of the L frames representing the volume. Thus, the prosodic features F corresponding to each phoneme can be represented as F=(L, V, P).

[0058] In some embodiments, when clustering the prosodic features of the plurality of second audio data, the plurality of second audio data can be divided into a plurality of audio data subsets according to the pronunciation duration of the phonemes corresponding to the plurality of second audio data. Then, the plurality of audio data subsets are clustered respectively to obtain a plurality of clustering clusters. The pronunciation duration of the phonemes corresponding to each second audio data in the same audio data subset is within the preset pronunciation duration range corresponding to the audio data subset.

[0059] For example, for Chinese language, since the pronunciation duration of initial consonants is often significantly smaller than that of vowels, the plurality of second audio data can be divided into two audio data subsets according to initial consonant phonemes and vowel phonemes, and the two audio data subsets are clustered respectively to obtain two independent prosodic labels, so that more abundant prosodic labels can be obtained, and the training text labeled by the prosodic labels can represent more abundant emotions.

[0060] For example, the clustering method of the prosodic features of the plurality of second audio data can use the clustering algorithm in related technologies, such as the density-based spatial clustering of applications with noise (DBSCAN).

[0061] Specifically, in the clustering process, the similarity calculation formula of adjacent points is as follows:

[0062] Sim(x1, x2) = sqrt(L(x1) / L(x2)) * DTW(P(x1), P(x2)) * DTW(V(x1), V(x2)).

[0063] That is, the acoustic feature similarity of any two phonemes (x1 and x2) is equal to the square root of the ratio of their durations, multiplied by the dynamic time warping (DTW) similarity of their pitch sequences, multiplied by the DTW similarity of their volume sequences. Wherein, L(x1)≤L(x2), that is, the audio feature of the phoneme with longer duration is taken as the second operand. Exemplarily, for two elements with too large a difference in duration, their similarity is set to a near-zero value, that is, two phonemes with too large a difference in duration are forced to be dissimilar.

[0064] It should be noted that the specific manner of clustering by using the DBSCAN algorithm is known to those skilled in the art, and the present disclosure will not repeat it.

[0065] In S203, the prosodic features of the plurality of second audio data are determined based on the plurality of clustering clusters to obtain the prosodic labels of the plurality of phonemes.

[0066] In some embodiments, the plurality of clustering clusters is obtained by clustering the prosodic features of the plurality of second audio data. The prosodic labels of the plurality of phonemes are obtained according to the relationship between the prosodic features of the plurality of second audio data and the plurality of clustering clusters. Each clustering cluster obtained by clustering is used to show a prosodic label, and each prosodic label is used to reflect a prosodic feature including a pitch, a volume and a duration. The embodiments of the present disclosure can enrich the types of prosodic labels, improve the number of prosodic features that can be expressed by the prosodic labels, and thus enable the acoustic model based on the prosodic labels to be used to synthesize speech with higher fidelity and naturalness.

[0067] In some embodiments, the relationship between the prosodic features of the plurality of second audio data and the plurality of clustering clusters can be represented by the distance between the prosodic features and the core points of the clustering clusters. Specifically, for each second audio data in the plurality of second audio data, the distance between the prosodic feature of the second audio data and the core point of each clustering cluster in the plurality of clustering clusters can be calculated respectively to obtain a distance calculation result. Subsequently, according to the distance calculation result, a target clustering cluster corresponding to the second audio data can be determined from the plurality of clustering clusters, so as to obtain the prosodic label of the phoneme corresponding to the second audio data.

[0068] Exemplarily, based on the relationship between the prosodic features of the plurality of second audio data and the plurality of clusters, a prosodic label for each of the plurality of phonemes is obtained by traversing each cluster and calculating the distance between the prosodic feature of the second audio data and each core point of each cluster. If the distance between the prosodic feature and a core point of a cluster is less than 1 / 10 of the cluster radius ε used in the clustering algorithm, the prosodic feature is considered to be close to the core point. The phoneme corresponding to the prosodic feature is then labeled as the prosodic label of the cluster in which the core point resides.

[0069] In some embodiments, if the prosodic feature of a certain second audio data cannot be approximated to any core point using the above method, each cluster is traversed, and the distance from the prosodic feature to each core point of the cluster is calculated. If the distance from the prosodic feature to a core point is less than the cluster radius ε, the effective distance is incremented by 1, and then the ratio X of the effective distance to the total number of core points in the cluster is calculated. The ratios X obtained for each cluster are sorted, and the prosodic signature of the cluster with the largest ratio X greater than 0 is used as the prosodic signature of the phoneme corresponding to the second audio data.

[0070] In some embodiments, if the prosodic features of a certain second audio data cannot be effectively marked using the above method, the phoneme corresponding to the second audio data is marked as arrhythmic and represented by a special prosodic mark.

[0071] The prosody annotation method provided by the embodiment of the present disclosure can, after dividing the first audio data into multiple second audio data based on the correspondence between multiple phonemes in the text data and the first audio data corresponding to the text data, cluster the prosody features of the multiple second audio data to obtain multiple clusters, and then determine the prosody tags based on the prosody features of the multiple second audio data and the multiple clusters, so as to obtain the prosody tags of the multiple phonemes. The embodiment of the present disclosure obtains prosody tags by clustering the frame-level prosody features of the phonemes. By annotating the training text for training the acoustic model with such phoneme-level prosody tags, compared with the traditional word-level prosody annotation method, it can better assist the acoustic model in learning the speaker's emotions, voice style and other characteristics, thereby synthesizing highly realistic speech audio.

[0072] Based on the same inventive concept, in an application scenario of the present disclosure, an acoustic model training method is also provided. Figure 3 , shows a flow chart of an acoustic model training method provided in an embodiment of the present disclosure, such as Figure 3 As shown, the method includes the following steps.

[0073] S301, construct a training set.

[0074] It should be noted that the training set in the embodiments of the present disclosure includes text data and audio data corresponding to the text data. The source of the audio data can be obtained by recording or crawling from the network. The recording environment can use a professional recording studio or a quiet room, conference room, and the recording equipment can use professional recording equipment or simple recording equipment such as a mobile phone, and the reference voice recognition data recording conditions; network crawling needs to use noise reduction and other means to ensure the quality of the audio. The text data can be obtained by manual annotation, for example, phoneme annotation of the transcribed text of the audio data.

[0075] In some embodiments, in order to make the model have better training effect, the training set can select text data and audio data similar to the application field of the trained acoustic model.

[0076] S302, the prosody of each of the plurality of phonemes in the text data is labeled to obtain the prosodic label of each of the plurality of phonemes.

[0077] It should be noted that the prosody labeling method of the plurality of phonemes in the text data in the embodiments of the present disclosure is based on Figure 2 the prosody labeling method shown, and the present disclosure will not be repeated here.

[0078] S303, training the acoustic model using the training set and the prosodic label of each of the plurality of phonemes.

[0079] It should be noted that the acoustic model in the embodiments of the present disclosure can use the acoustic model in the related art, for example, FastPitch. For example, Figure 4 a network structure diagram of FastPitch in the related art is shown.

[0080] As Figure 4 shown, the network structure 400 of FastPitch can include an encoder 410, an adapter 420 and a decoder 430 connected in sequence. The encoder 410 is used to convert the text data to be synthesized into a feature vector that the adapter 420 can recognize. The adapter 420 includes a duration prediction module 421, a pitch prediction module 422 and a volume prediction module 423, which are used to predict the duration, pitch, volume and other hidden variable features of the text to be synthesized through the feature vector of the text to be synthesized. The decoder 430 is used to decode the prediction result output by the adapter 420, and convert it into a mel spectrum.

[0081] In some embodiments, in order to enable the acoustic model to better learn the prosody of the phonemes labeled by the Figure 2 prosody labeling method provided by the embodiments, a prosody prediction module and a prosody embedding module can be added between the above-mentioned encoder 410 and the adapter 420, thereby obtaining an acoustic model with embedded prosody prediction function.

[0082] Figure 5 FIG. 1 shows a schematic diagram of the network structure of an acoustic model with an embedded rhythm prediction function in an embodiment of the present disclosure. Figure 5 As shown, the network structure 500 of the acoustic model may include an encoder 510, an adapter 520, and a decoder 530 connected in sequence, wherein a prosody prediction module 540 and a prosody embedding module 550 are further connected between the encoder 510 and the adapter 520.

[0083] It should be noted that the prosody prediction module 540 is used to predict the prosody at the phoneme level in the text data to be synthesized. The prosody embedding module 550 is used to convert the prosody into a feature representation that can be recognized in subsequent steps. The specific functions of the encoder 510, adapter 520 and decoder 530 are similar to those of the decoder 530. Figure 4 The encoder 410 , adapter 420 , and decoder 430 shown are the same and will not be described in detail in this disclosure.

[0084] It should be noted that, during the training process of the acoustic model, the prosodic mark at the phoneme level obtained in S302 can be used to form a prosodic representation through the prosodic embedding module 550, and then added to the output of the encoder 510, thereby forming a feature representation with phoneme-level prosodic information, and then enters the adapter 520 and subsequent modules for calculation. At the same time, the output of the encoder 510 is input into the prosodic predictor module 540 shown in the figure. By performing loss calculation on the prediction result of the prosodic prediction module 540 and the phoneme-level prosodic mark, the model parameters of the prosodic prediction module 540 can be learned at the same time during the training process of the acoustic model.

[0085] In some embodiments, the prosody prediction module 540 may be composed of a recurrent neural network (RNN), a ReLU activation function, a convolutional layer (Conv), a layer normalization layer (LayerNorm), and a linear layer (Linear) connected in sequence.

[0086] Specifically, the feature vector output by the encoder 510 is first input into a recurrent neural network, which can use a long short-term memory (LSTM) operator in the embodiment of the present disclosure, so as to learn a temporal representation of context at each time step. Then, a ReLU activation function is used to increase the nonlinearity of the temporal representation. Then, a convolution layer, for example, a one-dimensional convolution layer (Conv1D), is used to perform convolution along the time direction, which can extract key information at each time step to obtain a high-quality temporal sequence feature representation. A layer normalization layer after the convolution layer is used to normalize the range of the representation value, facilitating neural network training. Finally, a linear layer is used to map the representation at each time step to a certain prosody label. The prosody label is obtained by Figure 2 The prosody labeling method shown in the figure labels the prosody label at the phoneme level.

[0087] The embodiment of the present disclosure adopts Figure 2 The prosody labeling method shown in the figure labels the prosody of each phoneme in the text data, and combines the labeled prosody label with the text data as training text of the acoustic model. Compared with the method in the related art, the prosody of the speaker, the emotion, the speech style, and the like can be better assisted to be learned by the acoustic model, so as to synthesize a high-simulation-degree speech audio.

[0088] Based on the same inventive concept, in one application scenario of the present disclosure, a speech synthesis method is further provided. Referring to Figure 6 , a flowchart of a speech synthesis method provided in the embodiment of the present disclosure is shown, as shown in Figure 6 , the method includes the following steps.

[0089] S601, input the text data to be synthesized into a pre-trained speech synthesis model to obtain the mel spectrum of the text data to be synthesized.

[0090] It should be noted that the speech synthesis model is trained based on the training set and the prosody labels of the plurality of phonemes corresponding to the text data of the training set, and the prosody labels of the plurality of phonemes corresponding to the text data are obtained by Figure 2 the prosody labeling method shown in the embodiment.

[0091] Exemplarily, as shown in the foregoing Figure 5 , the speech synthesis model in the embodiment of the present disclosure can include an encoder, an adapter, a decoder connected in sequence, wherein the encoder and the adapter are further connected with a prosody prediction module and a prosody embedding module.

[0092] Exemplarily, after the text data to be synthesized is input into the speech synthesis model, the text data to be synthesized is first converted into a feature vector recognizable by the prosody prediction module and the adapter by the encoder, then the prosody prediction module predicts the prosodic features of the text data to be synthesized through the feature vector of the text data to be synthesized, and the adapter predicts the latent variable features such as pitch, volume, and length of the text data to be synthesized through the feature vector of the text data to be synthesized. Further, the prosody predicted by the prosody prediction module is processed by the prosody embedding module, and after being combined with the prediction result of the adapter, the mel spectrum of the text data to be synthesized can be obtained through the decoding of the decoder.

[0093] S602, synthesizing the synthesized speech of the text data to be synthesized based on the mel spectrum of the text data to be synthesized.

[0094] Exemplarily, after the mel spectrum is obtained, the mel spectrum can be converted into a time-domain waveform of sound by a neural vocoder, such as WaveRNN, and then the synthesized speech of the text data to be synthesized is synthesized.

[0095] It should be noted that the main inventive concept and the effects achieved by the embodiments of the present disclosure are similar to Figure 3 The speech synthesis model training method embodiment shown is similar, and therefore the specific implementation details can be referred to Figure 3 The speech synthesis model training method embodiment shown, and the embodiments of the present disclosure will not be described here.

[0096] In some application scenarios, based on the speech synthesis method provided by the embodiments of the present disclosure, a speech synthesis scheme (which can be referred to as a multi-branch multi-style emotion speech synthesis service) having multiple branches and applicable to multiple styles (such as including narration, dialogue, etc.) and multiple emotions (such as including joy, anger, sadness, and happiness, etc.) can be configured as a cloud service as a kind of basic technology to empower users using the cloud service, and the scheme can also be used in personalized scenarios in vertical fields. For example, it can be applied to reading APP (application) intelligent reading, intelligent customer service, news broadcasting, intelligent device interaction, and other scenarios to realize intelligent speech synthesis in various scenarios.

[0097] Figure 7 A structure schematic diagram of a prosody labeling device in the embodiments of the present disclosure is shown, as Figure 7 The prosody labeling device 700 includes a division module 701, a clustering module 702, and a determination module 703.

[0098] Specifically, the division module 701 is configured to divide the first audio data into a plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, and the plurality of second audio data have a correspondence with the plurality of phonemes.

[0099] The clustering module 702 is configured to cluster the prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters; the prosodic features comprise pitch, volume and duration; each clustering cluster in the plurality of clustering clusters is configured to represent a prosodic label, and the prosodic label is configured to reflect the prosodic features comprising a pitch, a volume and a duration.

[0100] The determining module 703 is configured to determine the prosodic labels of the plurality of phonemes based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters.

[0101] In some embodiments, the determining module 703 is further configured to calculate the distance between the prosodic features of each second audio data and the core point of each clustering cluster in the plurality of clustering clusters respectively to obtain distance calculation results; determine the target clustering cluster corresponding to each second audio data from the plurality of clustering clusters according to the distance calculation results; wherein the distance between the prosodic features of each second audio data and the core point of the corresponding target clustering cluster satisfies a preset distance condition; and take the prosodic label shown by the target clustering cluster corresponding to each second audio data as the prosodic label of the phoneme corresponding to the corresponding second audio data.

[0102] In some embodiments, the clustering module 702 is further configured to divide the plurality of second audio data into a plurality of audio data subsets according to the pronunciation duration of the phonemes corresponding to the plurality of second audio data; wherein the pronunciation duration of the phonemes corresponding to each second audio data in the same audio data subset is within a preset pronunciation duration range corresponding to the audio data subset; and cluster the plurality of audio data subsets respectively to obtain the plurality of clustering clusters.

[0103] In some embodiments, the dividing module 701 is further configured to perform alignment processing on the text data and the first audio data corresponding to the text data to obtain the time boundary of each phoneme in the first audio data; and divide the first audio data into a plurality of second audio data according to the time boundary of each phoneme in the first audio data.

[0104] It should be noted that the prosodic labeling device provided in the above embodiments is used for prosodic labeling, and the division of the above functional modules is only used as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the prosodic labeling device and the prosodic labeling method provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.

[0105] Based on the same inventive concept, the embodiment of the present disclosure also provides an acoustic model training device, as follows. Since the principle of solving the problem of the embodiment is similar to the above-mentioned prosody labeling method embodiment, the implementation of the embodiment can be referred to the implementation of the above-mentioned prosody labeling method embodiment, and the repeated parts will not be described here.

[0106] Figure 8 The structure diagram of an acoustic model training device in the embodiment of the present disclosure is shown, as shown in Figure 8 The acoustic model training device 800 includes a construction module 801, a marking module 802, and a training module 803.

[0107] Specifically, the construction module 801 is configured to construct a training set, the training set including text data and audio data corresponding to the text data. The marking module 802 is configured to label the prosody of each of a plurality of phonemes in the text data by the prosody labeling method of the present disclosure to obtain the prosody label of each of the plurality of phonemes. The training module 803 is configured to train an acoustic model using the training set and the prosody label of each of the plurality of phonemes, and the trained acoustic model is configured to perform speech synthesis processing on text to be synthesized to obtain synthesized speech.

[0108] Based on the same inventive concept, the embodiment of the present disclosure also provides a speech synthesis device, as follows. Since the principle of solving the problem of the embodiment is similar to the above-mentioned prosody labeling method embodiment, the implementation of the embodiment can be referred to the implementation of the above-mentioned prosody labeling method embodiment, and the repeated parts will not be described here.

[0109] Figure 9 The structure diagram of a speech synthesis device in the embodiment of the present disclosure is shown, as shown in Figure 9 The speech synthesis device 900 includes an acquisition module 901 and a synthesis module 902.

[0110] Specifically, the acquisition module 901 is configured to input text data to be synthesized into a pre-trained speech synthesis model to obtain a mel spectrum of the text data to be synthesized, the speech synthesis model being trained based on a training set and prosody labels of a plurality of phonemes corresponding to text data of the training set, the prosody labels of the plurality of phonemes corresponding to the text data being obtained by the prosody labeling method of the present disclosure. The synthesis module 902 is configured to synthesize synthesized speech of the text data to be synthesized based on the mel spectrum of the text data to be synthesized.

[0111] Those skilled in the art can understand that the various aspects of the present disclosure can be implemented as a system, a method or a program product. Therefore, the various aspects of the present disclosure can be embodied as a whole hardware implementation, a whole software implementation (including firmware, microcode, etc.), or an implementation combined with hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.

[0112] The electronic device 1000 according to this embodiment of the present disclosure will be described below with reference to Figure 10 Figure 10 The electronic device 1000 shown is merely an example and should not limit the function and use range of the embodiments of the present disclosure.

[0113] As Figure 10 shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 can include, but are not limited to, the at least one processing unit 1010 described above, the at least one storage unit 1020 described above, and a bus 1030 connecting different system components, including the storage unit 1020 and the processing unit 1010.

[0114] The storage unit stores program code that can be executed by the processing unit 1010, so that the processing unit 1010 performs the steps described in the "Exemplary Method" section of the present specification according to various exemplary embodiments of the present disclosure. The processing unit 1010 can be a processor.

[0115] In some embodiments, the processing unit 1010 can perform the following steps of the prosody labeling method embodiments described above: dividing the first audio data into a plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, the plurality of second audio data having a correspondence between the plurality of phonemes; clustering the prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters; the prosodic features include pitch, volume and duration; each clustering cluster in the plurality of clustering clusters is used to represent a prosodic label, and the prosodic label is used to reflect the prosodic features including a pitch, a volume and a duration; determining the prosodic label based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters to obtain the prosodic label of each of the plurality of phonemes.

[0116] ​In some embodiments, the processing unit 1010 can further perform the following steps of the above-mentioned acoustic model training method embodiments: constructing a training set, the training set comprising text data and audio data corresponding to the text data; performing prosodic labeling on the plurality of phonemes in the text data respectively by the prosodic labeling method of the present disclosure to obtain the prosodic labels of the plurality of phonemes respectively; and training the acoustic model using the training set and the prosodic labels of the plurality of phonemes respectively, the trained acoustic model being used for speech synthesis processing on the text to be synthesized to obtain synthesized speech.

[0117] In some embodiments, the processing unit 1010 can further perform the following steps of the above-mentioned speech synthesis method embodiments: inputting the text data to be synthesized into the pre-trained speech synthesis model to obtain the mel spectrum of the text data to be synthesized, the speech synthesis model being trained based on a training set and prosodic labels of a plurality of phonemes corresponding to text data of the training set, the prosodic labels of the plurality of phonemes corresponding to the text data being obtained by the prosodic labeling method of the present disclosure; and synthesizing the synthesized speech of the text data to be synthesized based on the mel spectrum of the text data to be synthesized.

[0118] The storage unit 1020 can include a readable medium in the form of volatile storage unit, such as a random access memory (RAM) 10201 and / or a cache memory 10202, and can further include a read-only memory (ROM) 10203.

[0119] The storage unit 1020 can further include program / utility 10204 having a set of programs / modules 10205, including without limitation, an operating system, one or more application programs, other programs, and programmatic data, each or some combination thereof, likely implementing aspects of the network environment.

[0120] The bus 1030 can be representative of one or more of several types of bus structures, including a storage bus or bus controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0121] The electronic device 1000 can also communicate with one or more external devices 1040 such as a keyboard or pointing device, a Bluetooth device, or a database, and / or can communicate with one or more devices that enable a user to interact with the electronic device 1000 and / or one or more devices (e.g., a router, a modem, a server, etc.) that enable the electronic device 1000 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface(s) 1050. Still yet, the electronic device 1000 can communicate with one or more networks, such as one or more local area networks (LANs), wide area networks (WANs), and / or the Internet, through network adapter 1060. As depicted, network adapter 1060 communicates with the other components of the electronic device 1000 via bus 1030. It should be appreciated that although not shown, other hardware and / or software components could be used in conjunction with the electronic device 1000. These components, as well as the software components of the electronic device 1000, are meant to be illustrative only and the scope of the application is not limited to any particular software configuration or hardware configuration.

[0122] Those skilled in the art will readily understand that the example embodiments described herein can be implemented by software and / or by hardware coupled with software, as described above. Thus, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash disk, a mobile hard disk, etc.) or a network, and includes a number of instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.

[0123] In the example embodiments of the present disclosure, a computer readable storage medium is also provided, which can be a readable signal medium or a readable storage medium. A program product is stored on the computer readable storage medium, and the program product can implement the method of the present disclosure. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a terminal device to execute the steps according to various example embodiments of the present disclosure described in the above “Example Method” section of the specification when the program product is run on the terminal device.

[0124] More specific examples of the computer readable storage medium in the present disclosure can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination of the foregoing.

[0125] In this disclosure, a computer readable storage medium can include a data signal transported over a carrier wave and can take the form of any appropriate medium including but not limited to phase or frequency modulation techniques, as well as other techniques. Still yet, a computer readable medium can be any medium that can be read by a computer including magnetic storage media (e.g., floppy disks); optical storage media (e.g., CD-ROMs); and tangible storage media that is impregnated with a substance such as gold and / or silver; and others as will occur to those skilled in the art. A computer readable medium can also be any medium that can be used to store or transfer the desired program code in the form of instructions or data structures and that can be accessed by a computer.

[0126] Optionally, program code embodied on a computer readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0127] In one or more embodiments, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over a computer-readable medium, such as a non-transitory computer-readable medium, as described above. The term "computer-readable medium" includes both computer storage media and communication media, including but not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store or transfer the desired information or instructions, and which can be accessed by a computer.

[0128] It should be noted that, although several modules or units for device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. Indeed, features and functionalities of two or more modules or units described above can be embodied in one module or unit according to embodiments of the present disclosure. Conversely, features and functionalities of one module or unit described above can be further divided into several modules or units.

[0129] Moreover, although the various steps of the methods in the present disclosure are described in a particular order in the figures, this is not required or implied in any way as to the order of the steps or that all illustrated steps be performed to achieve desirable results. Additionally or alternatively, certain steps can be omitted, combined into fewer steps, separated into multiple steps, and / or performed in a different order than that shown.

[0130] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.

[0131] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure that come within known use or custom in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A prosody annotation method characterized by, The method comprises the following steps: According to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data, the first audio data is divided into a plurality of second audio data, and the plurality of second audio data has a corresponding relationship with the plurality of phonemes; The prosodic features of the plurality of second audio data are clustered to obtain a plurality of clustering clusters;The prosodic feature corresponding to each second audio data is determined by pitch, volume and length;Each clustering cluster in the plurality of clustering clusters is used to represent a prosodic label, and a prosodic label is used to reflect a prosodic feature containing a pitch, a volume and a length; Based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters, the processing of determining the prosodic label is carried out to obtain the prosodic label of each phoneme in the plurality of phonemes.

2. The method of claim 1, wherein, The processing of determining the prosodic label based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters to obtain the prosodic label of each phoneme in the plurality of phonemes comprises: The distance between the prosodic feature of each second audio data and the core point of each clustering cluster in the plurality of clustering clusters is calculated respectively to obtain the distance calculation result; According to the distance calculation result, the target clustering cluster corresponding to each second audio data is determined from the plurality of clustering clusters;Wherein, the distance between the prosodic feature of each second audio data and the core point of the corresponding target clustering cluster satisfies the preset distance condition; The prosodic label shown by the target clustering cluster corresponding to each second audio data is taken as the prosodic label of the phoneme corresponding to the corresponding second audio data.

3. The method of claim 1, wherein, The clustering of the prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters comprises: According to the pronunciation time length of the phonemes corresponding to the plurality of second audio data, the plurality of second audio data is divided into a plurality of audio data subsets;Wherein, the pronunciation time length of each second audio data corresponding to each phoneme in the same audio data subset is within the preset pronunciation time length range corresponding to the audio data subset; The plurality of audio data subsets are respectively clustered to obtain a plurality of clustering clusters.

4. The method of claim 1, wherein, The division of the first audio data into a plurality of second audio data according to the correspondence between the plurality of phonemes in the text data and the first audio data corresponding to the text data comprises: The text data and the first audio data corresponding to the text data are aligned to obtain the time boundary of each phoneme in the plurality of phonemes in the first audio data; According to the time boundary of each phoneme in the plurality of phonemes in the first audio data, the first audio data is divided into a plurality of second audio data.

5. An acoustic model training method, comprising: The method comprises the following steps: Constructing a training set, the training set comprising text data and audio data corresponding to the text data; According to any one of claims 1 to 4, the plurality of phonemes in the text data are respectively prosodically labeled to obtain the prosodic label of each phoneme in the plurality of phonemes; Using the training set and the prosodic label of each phoneme, an acoustic model is trained, and the trained acoustic model is used for speech synthesis processing of the text to be synthesized to obtain synthesized speech.

6. A speech synthesis method characterized by, The method comprises the following steps: inputting the text data to be synthesized into a pre-trained speech synthesis model to obtain a mel spectrum of the text data to be synthesized, the speech synthesis model being trained based on a training set and prosodic labels of phonemes corresponding to text data in the training set, the prosodic labels of the phonemes corresponding to the text data being obtained by the method in any one of claims 1 to 4; synthesizing synthesized speech of the text data to be synthesized based on the mel spectrum of the text data to be synthesized.

7. A prosody labeling apparatus characterized by comprising: The method comprises: a dividing module configured to divide first audio data into a plurality of second audio data according to a correspondence between a plurality of phonemes in text data and the first audio data, the plurality of second audio data having a correspondence with the plurality of phonemes; a clustering module configured to cluster prosodic features of the plurality of second audio data to obtain a plurality of clustering clusters, each second audio data corresponding to prosodic features determined by pitch, volume and duration, and each clustering cluster in the plurality of clustering clusters being used to represent a prosodic label, a prosodic label being used to reflect prosodic features including a pitch, a volume and a duration; a determining module configured to determine prosodic labels based on the prosodic features of the plurality of second audio data and the plurality of clustering clusters to obtain respective prosodic labels of the plurality of phonemes.

8. An acoustic model training apparatus, comprising: The method comprises: a constructing module configured to construct a training set, the training set comprising text data and audio data corresponding to the text data; a labeling module configured to label a plurality of phonemes in the text data with respective prosodic labels by the method in any one of claims 1 to 4 to obtain the respective prosodic labels of the plurality of phonemes; a training module configured to train an acoustic model using the training set and the respective prosodic labels of the plurality of phonemes, the trained acoustic model being used to perform speech synthesis processing on text to be synthesized to obtain synthesized speech.

9. A speech synthesis apparatus characterized by comprising: The method comprises: an obtaining module configured to input text data to be synthesized into a pre-trained speech synthesis model to obtain a mel spectrum of the text data to be synthesized, the speech synthesis model being trained based on a training set and prosodic labels of phonemes corresponding to text data in the training set, the prosodic labels of the phonemes corresponding to the text data being obtained by the method in any one of claims 1 to 4; a synthesizing module configured to synthesize synthesized speech of the text data to be synthesized based on the mel spectrum of the text data to be synthesized.

10. An electronic device, comprising: The method comprises: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the method in any one of claims 1 to 4, or the method in claim 5, or the method in claim 6, via execution of the executable instructions.

11. A computer readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the method in any one of claims 1 to 4, or the method in claim 5, or the method in claim 6. The computer program, when executed by a processor, implements the method in any one of claims 1 to 4, or the method in claim 5, or the method in claim 6.

Citation Information

Patent Citations

  • Metrical structure predicting method and metrical structure predicting device

    CN104867490A

  • Rhythm labeling method and system

    CN114255736A