A method, device and apparatus for multi-scale text prosody analysis at the paragraph level
Through the multi-scale text rhythm analysis method at the chapter level, combined with the multi-scale text rhythm analysis model at the discourse level and long-term short-term memory network, the problem of insufficient rhythm and emotional control in the existing speech synthesis technology is solved, and the speech synthesis effect with coherence in context and consistent style in long text synthesis is achieved.
Patent Information
- Application Number
- CN202310347958.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-04-03
AI Technical Summary
Existing pronunciation synthesis technologies are difficult to effectively control the pronunciation and emotions of pronunciation, especially in long text synthesis, where the context is incoherent, resulting in inconsistent styles of synthetic pronunciations between sentences.
A multi-scale text pronunciation analysis method is proposed. By splitting the text to be analyzed into multiple sentences, using the discourse level multi-scale text pronunciation analysis model to process each sentence, the local pronunciation embedded sequence characteristics and sentence-level discourse characteristics are obtained, and context information is fused through long-term and short-term memory networks to generate local pronunciation embedded sequence characteristics and global style embedding features with context information.
It realizes fine control of phonological rhythm and emotion, ensuring that the synthetic pronunciation is diverse in style within the sentence and highly consistent with the text characteristics, while also being more coherent between sentences.
Smart Images

Figure CN116386595B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a method, device and equipment for multi-scale text prosody analysis at a paragraph level. Background Art
[0002] Speech synthesis technology, also known as Text To Speech (TTS) technology, can convert any text information into standard and fluent speech. Although the existing end-to-end speech synthesis technology can achieve relatively accurate pronunciation, it lacks emotional information compared with real human pronunciation.
[0003] In related technologies, emotion tags are used to enhance emotion information, or style transfer of reference audio is used to achieve emotion control. However, the number of emotion tags is limited, and each sentence generally has only one simple tag to control emotion, which makes the expression of emotion insufficient and the control weak. Although the style transfer method of reference audio has stronger control than the simple emotion tag method, this method requires reference audio as input, which may be missing under realistic conditions. In addition, the style transfer method of reference audio focuses more on global style embedding, while information such as rhythm and emotion is constantly changing. Simple global control cannot meet actual emotional needs. Summary of the invention
[0004] In view of the above problems, embodiments of the present invention provide a method, apparatus and device for paragraph-level multi-scale text prosody analysis to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of an embodiment of the present invention discloses a method for analyzing the prosody of a multi-scale text at a paragraph level, the method comprising:
[0006] Split the text to be analyzed into multiple sentences;
[0007] Processing the multiple sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence;
[0008] Inputting the sentence-level discourse features of the plurality of sentences into a long short-term memory network for processing to obtain a global style embedding feature at the paragraph level and a sentence-level discourse feature with context information corresponding to each sentence;
[0009] Mapping the sentence-level discourse features with context information to the phoneme level to obtain the phoneme-level discourse features with context information;
[0010] The phoneme-level discourse features with contextual information and the local prosody embedding sequence features are fused to obtain local prosody embedding sequence features with contextual information. The local prosody embedding sequence features with contextual information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
[0011] Optionally, the method further comprises:
[0012] Extracting features from the text to be analyzed to obtain phoneme embedding features;
[0013] Speech synthesis is performed based on the phoneme embedding features, the local prosody embedding sequence features with contextual information, and the global style embedding features at the chapter level to obtain speech corresponding to the text to be analyzed.
[0014] Optionally, the multiple sentences are processed using a discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence, including:
[0015] Extracting text features from the sentence to obtain word-level features and sentence-level discourse features;
[0016] After fusing the word-level feature with other word-level features, the word-level feature is copied and extended using a length adjuster to obtain a phoneme-level feature;
[0017] Using other phoneme-level features and the phoneme-level features to perform multimodal feature fusion to obtain multi-scale fused text features;
[0018] Perform pitch and energy prediction based on the multi-scale fusion text features to obtain pitch features and energy features;
[0019] The pitch feature, the energy feature and the multi-scale fusion text feature are concatenated and then feature prediction is performed to obtain the local prosody embedding sequence feature.
[0020] Optionally, mapping the sentence-level discourse feature with context information to a phoneme level to obtain the phoneme-level discourse feature with context information includes:
[0021] Constructing a parameter learnable matrix according to the dimension of the sentence-level discourse feature with context information and the dimension of the local prosodic embedding sequence feature;
[0022] Based on the parameter learnable matrix, the sentence-level discourse features with context information are mapped using the Einstein summation convention to obtain the phoneme-level discourse features with context information.
[0023] Optionally, the passage-level multi-scale text prosody analysis method is implemented by a pre-trained passage-level multi-scale text prosody analysis model, and the passage-level multi-scale text prosody analysis model is trained in the following manner:
[0024] Acquire a prosodic feature training data set, wherein each training data in the prosodic feature training data set includes: a training text and a true value local prosodic embedding sequence feature and a true value global style embedding feature corresponding to the training text;
[0025] Inputting the training data into the paragraph-level multi-scale text prosody analysis model for training, so that the paragraph-level multi-scale text prosody analysis model learns the mapping relationship from text to prosody features;
[0026] After the training end condition is met, a trained chapter-level multi-scale text prosody analysis model is obtained, and the trained chapter-level multi-scale text prosody analysis model has the ability to predict prosody features according to the text.
[0027] Optionally, the discourse-level multi-scale text prosody analysis model is a sub-model in the paragraph-level multi-scale text prosody analysis model; the discourse-level multi-scale text prosody analysis model training process includes:
[0028] The discourse-level multi-scale text prosody analysis model processes the input training sentence to obtain a multi-scale fusion text feature;
[0029] Perform pitch and energy prediction based on the multi-scale fusion text features to obtain predicted pitch features and predicted energy features;
[0030] The true value pitch feature, the true value energy feature, and the multi-scale fusion text feature are spliced together to perform feature prediction to obtain a predicted local prosody embedding sequence feature;
[0031] updating the parameters of the discourse-level multi-scale text prosody analysis model according to the comparison result of the predicted pitch feature, the predicted energy feature and the predicted local prosody embedding sequence feature with the true value pitch feature, the true value energy feature and the true value local prosody embedding sequence feature;
[0032] After the training end conditions are met, the trained discourse-level multi-scale text prosody analysis model has the ability to predict pitch features, predict energy features, and predict local prosody embedding sequence features.
[0033] Optionally, the training text also corresponds to an original audio; the true value local prosody embedding sequence features and the true value global style embedding features corresponding to the training text are obtained in the following manner:
[0034] Using a style transfer model to extract features from the original audio, to obtain the true value global style embedding features, local rhythm embedding sequence features, pitch features and energy features;
[0035] The local prosody embedding sequence feature, the pitch feature and the energy feature are fused to obtain a true value local prosody embedding sequence feature.
[0036] Optionally, before extracting features from the original audio using the style transfer model, the method further includes:
[0037] Converting the training text into phoneme features;
[0038] Using a forced alignment tool to align the original audio with the phoneme feature to obtain an aligned reference audio;
[0039] The style transfer model is used to extract features from the original audio, including:
[0040] A style transfer model is used to extract features from the aligned reference audio.
[0041] A second aspect of the embodiments of the present invention discloses a paragraph-level multi-scale text prosody analysis device, the device comprising:
[0042] A text splitting module is used to split the text to be analyzed into multiple sentences;
[0043] A discourse analysis module, used to process the multiple sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosodic embedding sequence features and sentence-level discourse features corresponding to each sentence;
[0044] A context fusion module, used for inputting the sentence-level discourse features of the plurality of sentences into a long short-term memory network for processing, to obtain a global style embedding feature at the paragraph level and a sentence-level discourse feature with context information corresponding to each sentence;
[0045] A feature mapping module, used for mapping the sentence-level speech features with context information to the phoneme level to obtain the phoneme-level speech features with context information;
[0046] The local feature module is used to fuse the phoneme-level discourse features with contextual information and the local prosody embedding sequence features to obtain local prosody embedding sequence features with contextual information. The local prosody embedding sequence features with contextual information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
[0047] According to a third aspect of an embodiment of the present invention, an electronic device is disclosed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when executed, the processor implements the paragraph-level multi-scale text prosody analysis method as described in the first aspect of the embodiment of the present invention.
[0048] The embodiments of the present invention include the following advantages:
[0049] In an embodiment of the present invention, in order to solve the problem of insufficient rhythmic emotion control ability of synthetic speech, a paragraph-level multi-scale text rhythm analysis method is proposed. First, the text to be analyzed is split into multiple sentences, and the multiple sentences are processed using a discourse-level multi-scale text rhythm analysis model to obtain local rhythm embedding sequence features and sentence-level discourse features corresponding to each sentence. Then, the sentence-level discourse features of the multiple sentences are input into a long short-term memory network for processing to obtain a paragraph-level global style embedding feature and a sentence-level discourse feature with context information corresponding to each sentence. Then, the sentence-level discourse feature with context information is mapped to the phoneme level to obtain the phoneme-level discourse feature with context information. Finally, the phoneme-level discourse feature with context information and the local rhythm embedding sequence feature are fused to obtain the local rhythm embedding sequence feature with context information.
[0050] Since the local prosody embedding sequence feature with context information is a more fine-grained prosody emotion control sequence, more refined prosody emotion control can be achieved based on this feature; and the global style embedding feature at the chapter level can achieve global style control; and by analyzing the data at the chapter level, the problem of context incoherence in long-form speech synthesis is solved, so that the prosody feature-synthesized speech based on this method is not only diverse in style within the sentence and highly consistent with the text features, but also more coherent between sentences. Therefore, in the embodiment of the present invention, the local prosody embedding sequence feature with context information and the global style embedding feature at the chapter level are used to represent the prosody features of the text to be analyzed, so as to automatically obtain speech that conforms to the prosody emotion expression of the text features through pure text. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0052] Figure 1 It is a flowchart of the steps of a method for analyzing the prosody of a multi-scale text at the paragraph level provided by an embodiment of the present invention;
[0053] Figure 2 is a structural diagram of a discourse-level multi-scale text prosody analysis model provided by an embodiment of the present invention;
[0054] Figure 3 It is a structural schematic diagram of a paragraph-level multi-scale text prosody analysis model provided by an embodiment of the present invention;
[0055] Figure 4 This is an overall application architecture diagram of a multi-scale text prosody analysis at the chapter level provided by an embodiment of the present invention;
[0056] Figure 5 It is a structural schematic diagram of a paragraph-level multi-scale text prosody analysis device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] The embodiment of the present invention provides a method for analyzing the prosody of a multi-scale text at the chapter level. Figure 1 As shown, Figure 1 A flowchart of a method for analyzing the prosody of a multi-scale text at the paragraph level provided by an embodiment of the present invention includes steps S101 to S105:
[0059] Step S101: split the text to be analyzed into multiple sentences.
[0060] The text to be analyzed is a chapter-level text content composed of multiple sentences, such as novels, news reports, etc. This embodiment uses plain text as input and splits the text to be analyzed into multiple sentences according to language logic, so that each sentence can be processed separately in subsequent steps.
[0061] Step S102: Process the multiple sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence.
[0062] In this embodiment, the local prosody embedding sequence feature is a phoneme-level fine-grained sentence prosody control feature, which is used to control the stress, pause, etc. of phonemes; among them, a phoneme refers to information such as the initial consonant and final vowel of the pinyin corresponding to each character. For example, in the sentence "Hello, world", the pause between "好 (hǎo)" and "世 (shì)" is controlled by the local prosody embedding sequence feature of this sentence. And the sentence-level discourse feature is used to represent the overall feature information of the sentence. In this embodiment, a discourse-level multi-scale text prosody analysis model is used to process the multiple sentences split in step S101 in sequence, so as to obtain the local prosody embedding sequence feature and the sentence-level discourse feature corresponding to each sentence.
[0063] In an alternative embodiment, the use of the discourse-level multi-scale text prosody analysis model to process the multiple sentences to obtain the local prosody embedding sequence feature and the sentence-level discourse feature corresponding to each sentence includes steps S102-1 to S102-5:
[0064] Step S102-1: Extract text features from the sentence to obtain word-level features and the sentence-level discourse feature.
[0065] In this step, the word-level features represent the feature information of each word in the sentence. For example, for the sentence "你好吗 (Nǐ hǎo ma)", there are the feature information of the three words "你 (Nǐ)", "好 (hǎo)", and "吗 (ma)". Considering that text prosody emotion prediction is essentially a natural language understanding (NLU)-type task, and in the speech synthesis task, in order to better learn pronunciation-related information, the text input used is usually at the phoneme level, but this is not conducive to the learning of NLU-type tasks. Therefore, in step S101-1, a text feature extraction model (such as RoBERTa, BERT, CNN, etc. models) is used to extract text features from the sentence, so as to obtain the word-level features and the sentence-level discourse feature corresponding to the sentence.
[0066] Step S102-2: After fusing the word-level features with other word-level features, use a length regulator to perform replication and expansion to obtain phoneme-level features.
[0067] In this step, other word-level features refer to the features of the speaker (for example, the speaker is male, female, old, young, etc.) and the overall style features (for example, the style is fairy tale, martial arts, etc.). After fusing the word-level features with other word-level features, word-level features with a preliminary feature style are obtained, that is, according to the word-level features with a preliminary feature style, it is possible to know the voice information such as whether the pronunciation of this sentence is male or female, fairy tale or martial arts.
[0068] In order to analyze the fine-grained phoneme-level prosodic sentiment information, considering that each word may correspond to multiple phonemes, the length regulator is used to copy and expand the word-level features with the preliminary feature style to align the word-level features with the phonemes to obtain the phoneme-level features. The length of the phoneme-level features is consistent with the number of phonemes. For example, for the word "you", "you" corresponds to two phonemes "n" and "i", then the length regulator is used to copy the features corresponding to "you" to obtain features aligned with the phonemes "n" and "i".
[0069] Step S102 - 3 : performing multimodal feature fusion with other phoneme-level features and the phoneme-level features to obtain multi-scale fused text features.
[0070] In this step, other phoneme-level features include: phoneme features (e.g., which phoneme it specifically corresponds to), tone features (e.g., which tone a phoneme specifically corresponds to), position features (e.g., distinguishing initials and finals), etc. In order to further refine different phonemes and achieve more detailed control, other phoneme-level features are added to the phoneme-level features, and multi-scale fused text features are obtained through multimodal feature fusion.
[0071] Step S102 - 4 : performing pitch and energy prediction based on the multi-scale fusion text features to obtain pitch features and energy features.
[0072] Among them, pitch is the frequency of sound, that is, the speed of sound wave vibration. The higher the frequency, the higher the pitch; the lower the frequency, the lower the pitch. The pitch here refers specifically to the fundamental frequency in human audio. Energy is the amplitude of sound, that is, the amplitude of sound wave vibration. The larger the amplitude, the greater the energy of the sound; the smaller the amplitude, the smaller the energy of the sound.
[0073] In this step, the discourse-level multi-scale text prosody analysis model also has the ability to predict pitch features and energy features. By processing the multi-scale fusion text features, the corresponding pitch features and energy features are predicted.
[0074] Step S102-5: concatenating the pitch feature, the energy feature and the multi-scale fusion text feature to perform feature prediction to obtain the local prosody embedding sequence feature.
[0075] In this step, the pitch feature, energy feature and multi-scale fusion text feature are used to jointly predict the local prosody embedding sequence feature. Compared with directly predicting the five-dimensional features corresponding to each phoneme (i.e., the three-dimensional local prosody embedding sequence feature, pitch feature and energy feature), this step can see the pitch feature and energy feature when predicting the local prosody embedding sequence feature, so that the local prosody embedding sequence feature learned by the model has better synergy with the pitch feature and energy feature, and the synthesized speech is more coordinated.
[0076] Figure 2 The structure diagram of the discourse-level multi-scale text prosody analysis model is shown. Specifically, the discourse-level multi-scale text prosody analysis model includes: a text feature extraction model, a length regulator, a multimodal feature fusion module, a pitch and energy prediction module, and a local prosody embedding sequence feature prediction module. Among them, the text feature extraction model is used to extract features from sentences to obtain word-level features and sentence-level discourse features; the length regulator is used to copy and expand the input features and align them with the length of the phoneme to obtain the phoneme-level features; the multimodal feature fusion module is used to perform multimodal feature fusion on other phoneme-level features and phoneme-level features to obtain multi-scale fused text features; the pitch and energy prediction module is used to predict pitch features and energy features based on the multi-scale fused text features; the local prosody embedding sequence feature prediction module is used to predict local prosody embedding sequence features.
[0077] Step S103: Input the sentence-level discourse features of the multiple sentences into the long short-term memory network for processing to obtain the global style embedding features at the paragraph level and the sentence-level discourse features with context information corresponding to each sentence.
[0078] In this embodiment, in order to further introduce the context of the text, the sentence-level discourse features corresponding to the multiple sentences corresponding to the text to be analyzed are input into the long short-term memory network for interactive processing. Since all the sentence-level discourse features are interactive, on the one hand, the output of the last unit (i.e., the attention mechanism) in the long short-term memory network is used as the global feature at the chapter level, that is, the global style embedding feature at the chapter level is predicted, and the global style embedding feature at the chapter level can control the overall style of the entire text to be analyzed; on the other hand, each sentence-level discourse feature will obtain a new sentence-level discourse feature with context information.
[0079] Step S104: Mapping the sentence-level discourse features with context information to the phoneme level to obtain the phoneme-level discourse features with context information.
[0080] In this embodiment, in order to allow the context information to act on more fine-grained local prosody embedding sequence features, the sentence-level discourse features with context information are mapped to phoneme-level discourse features with context information, so that the phoneme-level discourse features with context information can be fused with the local prosody embedding sequence features in subsequent steps.
[0081] In an optional embodiment, mapping the sentence-level speech features with context information to the phoneme level to obtain the phoneme-level speech features with context information includes step S104-1 and step S104-2:
[0082] Step S104-1: construct a parameter learnable matrix according to the dimension of the sentence-level discourse feature with context information and the dimension of the local prosody embedding sequence feature.
[0083] Step S104 - 2 : Based on the parameter learnable matrix, the sentence-level discourse features with context information are mapped using the Einstein summation convention to obtain the phoneme-level discourse features with context information.
[0084] For example, suppose there are m sentences in a text to be analyzed, each sentence has N phonemes, and the phoneme-level discourse feature with context information corresponds to an r-dimensional feature for each sentence. Then the local prosody embedding sequence feature degree A lpe The dimension of is (m, N, d), where d = 3; the dimension of the phoneme-level discourse feature with contextual information is (m, r), then the dimension of the constructed parameter learnable D is (r, d, d). The sentence-level discourse feature with contextual information is mapped to the phoneme level using the Einstein summation convention (einsum), that is:
[0085] ΔA lpe =einsum('rdd,mr,mNd→mNd',D,U,A lpe )
[0086] Where ΔA lpe Represents phoneme-level utterance features with contextual information.
[0087] Step S105: The phoneme-level discourse features with context information and the local prosody embedding sequence features are merged to obtain local prosody embedding sequence features with context information. The local prosody embedding sequence features with context information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
[0088] In this embodiment, the phoneme-level speech features with context information and the local prosody embedding sequence features are fused, that is, the local prosody embedding sequence features are adjusted using the phoneme-level speech features with context information, so that the local prosody embedding sequence features have context information. Furthermore, the local prosody embedding sequence features with context information can make the transition between the synthesized sentences of the text to be analyzed smoother and the style more consistent, thereby bringing a better experience to the listener.
[0089] Finally, the prosodic features of the text to be analyzed are characterized by using the local prosodic embedding sequence features with contextual information and the global style embedding features at the chapter level. Since the local prosodic embedding sequence features with contextual information are a more fine-grained prosodic emotion control sequence, more sophisticated prosodic emotion control can be achieved based on this feature; and the global style embedding features at the chapter level can achieve global style control; and, by analyzing the data at the chapter level, the problem of context incoherence in long-form speech synthesis is solved, so that the prosodic feature-synthesized speech based on this method is not only diverse in style within the sentence and highly consistent with the text features, but also more coherent between sentences. Therefore, in this embodiment, the prosodic features of the text to be analyzed are characterized by using the local prosodic embedding sequence features with contextual information and the global style embedding features at the chapter level, so that speech that conforms to the prosodic emotion expression of the text features can be automatically obtained through pure text.
[0090] In practical applications, the paragraph-level multi-scale text prosody analysis method is implemented by a pre-trained paragraph-level multi-scale text prosody analysis model, such as Figure 3 As shown, it is a schematic diagram of the structure of the paragraph-level multi-scale text prosody analysis model provided by this implementation. Specifically, the paragraph-level multi-scale text prosody analysis model is composed of m discourse-level multi-scale text prosody analysis models, m feedback modules, and a long short-term memory network, where m is a positive integer, and the value of m is related to the number of sentences in the text to be analyzed.
[0091] The discourse-level multi-scale text prosody analysis model processes the sentences of the text to be analyzed, obtains the local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence, and then inputs the m sentence-level discourse features into the long short-term memory network for processing. The long short-term memory network outputs a chapter-level global style embedding feature and a sentence-level discourse feature with contextual information corresponding to each sentence. The sentence-level discourse features with contextual information are mapped to phoneme-level discourse features with contextual information through a feedback module. Finally, the phoneme-level discourse features with contextual information and the local prosody embedding sequence features are fused to obtain the local prosody embedding sequence features with contextual information.
[0092] In an optional embodiment, the text to be analyzed is converted into corresponding speech by the following method, specifically including:
[0093] Extracting features from the text to be analyzed to obtain phoneme embedding features;
[0094] Speech synthesis is performed based on the phoneme embedding features, the local prosody embedding sequence features with contextual information, and the global style embedding features at the chapter level to obtain speech corresponding to the text to be analyzed.
[0095] In this embodiment, the phoneme embedding feature refers to the phoneme information corresponding to the text to be analyzed, which is used to control the pronunciation of each word. Specifically, the feature extraction of the text to be analyzed includes: firstly processing the text to be analyzed to obtain the pinyin corresponding to the text to be analyzed, and then splitting the pinyin into corresponding phonemes, that is, obtaining the phoneme embedding feature.
[0096] Then, the local prosody embedding sequence features with context information and the global style embedding features at the chapter level obtained in the above steps S101 to S105 are used to control the pronunciation of the phoneme embedding features, thereby obtaining the semantic expression corresponding to the text to be analyzed. Specifically, the local prosody embedding sequence features with context information are used to embed the prosodic structure of the sentence (e.g., pauses, stresses, emotions, etc.), and the global style embedding features at the chapter level control the overall style of the text to be analyzed (e.g., martial arts, fairy tales, etc.).
[0097] In an optional embodiment, the paragraph-level multi-scale text prosody analysis model is trained in the following manner:
[0098] Acquire a prosodic feature training data set, wherein each training data in the prosodic feature training data set includes: a training text and a true value local prosodic embedding sequence feature and a true value global style embedding feature corresponding to the training text;
[0099] Inputting the training data into the paragraph-level multi-scale text prosody analysis model for training, so that the paragraph-level multi-scale text prosody analysis model learns the mapping relationship from text to prosody features;
[0100] After the training end condition is met, a trained chapter-level multi-scale text prosody analysis model is obtained, and the trained chapter-level multi-scale text prosody analysis model has the ability to predict prosody features according to the text.
[0101] In this embodiment, the paragraph-level multi-scale text prosody analysis model learns a mapping from text to prosody features from the paragraph-level text through training. This enables the paragraph-level multi-scale text prosody analysis model to predict the local prosody embedding sequence features with contextual information and the paragraph-level global style embedding features of the text to be analyzed based on the pure text content. However, prosody features can only be extracted from speech, so it is necessary to use a style transfer model to obtain the true value local prosody embedding sequence features and true value global style embedding features corresponding to the training text.
[0102] Specifically, the training text also corresponds to an original audio, and the true value local rhythm embedding sequence features and the true value global style embedding features corresponding to the training text are obtained in the following way:
[0103] The style transfer model is used to extract features of the original audio to obtain the true value global style embedding features, local rhythm embedding sequence features, pitch features and energy features; the local rhythm embedding sequence features, the pitch features and the energy features are fused to obtain the true value local rhythm embedding sequence features.
[0104] In an optional embodiment, before extracting features from the original audio using the style transfer model, the method further includes:
[0105] Converting the training text into phoneme features;
[0106] Using a forced alignment tool to align the original audio with the phoneme feature to obtain an aligned reference audio;
[0107] Using a style transfer model to extract features from the original audio includes: using a style transfer model to extract features from the aligned reference audio.
[0108] In this embodiment, a " / " separator is added between every two words in the sentence as an anchor point for explicitly learning pauses between words. For example, the sentence "Hello, world!" corresponds to the pinyin "ni3 / hao3 / shi4 / jie4!", which is then broken down into phonemes. In order to facilitate the extraction of local prosody embedding sequence features, the original audio needs to be aligned with the phoneme features. For example, if the pronunciation of the first phoneme "n" of "you" is 0 seconds to 0.1 seconds, the audio from 0 to 0.1 seconds in the original audio is aligned with the phoneme "n".
[0109] In actual use, the MFA forced alignment tool is used to align the original audio with the phoneme feature. If the MFA forced alignment tool extracts a pause between two words, the separator " / " between the two words will be treated as a real phoneme to facilitate the subsequent steps to extract its local prosodic embedding sequence features. On the contrary, the separator between the two words cannot correspond to a specific audio, so the corresponding local prosodic embedding sequence features cannot be extracted. At this time, its features are set to an all-0 vector. Then, in the subsequent steps, through the learning of the style transfer model, the model can learn this explicit pause feature very well. For punctuation marks, since punctuation marks do not correspond to any pronunciation, the number of corresponding phonemes is 0. For the separator " / ", the number of corresponding phonemes is 1.
[0110] In an optional embodiment, the discourse-level multi-scale text prosody analysis model is a sub-model in the paragraph-level multi-scale text prosody analysis model; the discourse-level multi-scale text prosody analysis model training process includes:
[0111] The discourse-level multi-scale text prosody analysis model processes the input training sentence to obtain a multi-scale fusion text feature;
[0112] Perform pitch and energy prediction based on the multi-scale fusion text features to obtain predicted pitch features and predicted energy features;
[0113] The true value pitch feature, the true value energy feature, and the multi-scale fusion text feature are spliced together to perform feature prediction to obtain a predicted local prosody embedding sequence feature;
[0114] updating the parameters of the discourse-level multi-scale text prosody analysis model according to the comparison result of the predicted pitch feature, the predicted energy feature and the predicted local prosody embedding sequence feature with the true value pitch feature, the true value energy feature and the true value local prosody embedding sequence feature;
[0115] After the training end conditions are met, the trained discourse-level multi-scale text prosody analysis model has the ability to predict pitch features, predict energy features, and predict local prosody embedding sequence features.
[0116] In this embodiment, the discourse-level multi-scale text prosody analysis model is trained by using the true value pitch feature, the true value energy feature, and the true value local prosody embedding sequence feature, so that the local prosody embedding sequence has a dependency relationship with the interpretable local prosody embedding sequence. Then, when the discourse-level multi-scale text prosody analysis model is actually used to process multiple sentences, the corresponding pitch features and energy features can be predicted, so as to predict the local prosody embedding sequence of the sentence based on the pitch features and energy features.
[0117] Figure 4 The overall application architecture diagram of the present embodiment is illustrated, which specifically includes three parts: style transfer, text analysis, and speech synthesis. In the training stage, a style transfer model is used to extract features from the original audio corresponding to the training text to obtain true global style embedding features and true local rhythm embedding sequence features. Then, the true global style embedding features and the true local rhythm embedding sequence features are used to train the chapter-level multi-scale text rhythm analysis model, so that the chapter-level multi-scale text rhythm analysis model learns a mapping from text to rhythm features from the chapter-level text. In actual use, the trained chapter-level multi-scale text rhythm analysis model is directly used for feature prediction to obtain local rhythm embedding sequence features with contextual information of the text to be analyzed and global style embedding features at the chapter level, and then combined with the phoneme embedding features, the automatic conversion of pure text content to speech that conforms to the rhythmic emotional expression of the text features is realized.
[0118] In the embodiment of the present invention, since the local prosody embedding sequence feature with context information is a more fine-grained prosody emotion control sequence, more refined prosody emotion control can be achieved based on this feature; and the global style embedding feature at the chapter level can achieve global style control; and by analyzing the data at the chapter level, the problem of context incoherence in long-form speech synthesis is solved, so that the prosody feature synthesized speech based on this method is not only diverse in style within the sentence and highly consistent with the text features, but also more coherent between sentences. Therefore, in the embodiment of the present invention, the local prosody embedding sequence feature with context information and the global style embedding feature at the chapter level are used to represent the prosody features of the text to be analyzed, so as to automatically obtain speech that conforms to the prosody emotion expression of the text features through pure text.
[0119] The embodiment of the present invention also provides a paragraph-level multi-scale text prosody analysis device, such as Figure 5 As shown, Figure 5 A schematic diagram of the structure of a paragraph-level multi-scale text prosody analysis device provided by an embodiment of the present invention, the device comprising:
[0120] A text splitting module 51 is used to split the text to be analyzed into multiple sentences;
[0121] A discourse analysis module 52, configured to process the plurality of sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence;
[0122] A context fusion module 53, used for inputting the sentence-level discourse features of the plurality of sentences into a long short-term memory network for processing, to obtain a global style embedding feature at the chapter level and a sentence-level discourse feature with context information corresponding to each sentence;
[0123] A feature mapping module 54, configured to map the sentence-level speech features with context information to a phoneme level to obtain a phoneme-level speech feature with context information;
[0124] The local feature module 55 is used to fuse the phoneme-level discourse features with contextual information and the local prosody embedding sequence features to obtain local prosody embedding sequence features with contextual information. The local prosody embedding sequence features with contextual information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
[0125] In an optional embodiment, the device further includes:
[0126] A phoneme extraction module, used to extract features from the text to be analyzed to obtain phoneme embedding features;
[0127] The speech synthesis module is used to perform speech synthesis based on the phoneme embedding feature, the local prosody embedding sequence feature with context information and the global style embedding feature at the paragraph level to obtain the speech corresponding to the text to be analyzed.
[0128] In an optional embodiment, the discourse analysis module includes:
[0129] A first discourse analysis submodule is used to extract text features from the sentence to obtain word-level features and sentence-level discourse features;
[0130] The second discourse analysis submodule is used to fuse the word-level features with other word-level features, and then copy and expand them using a length regulator to obtain phoneme-level features;
[0131] A third discourse analysis submodule is used to perform multimodal feature fusion using other phoneme-level features and the phoneme-level features to obtain multi-scale fused text features;
[0132] A fourth discourse analysis submodule, configured to perform pitch and energy prediction based on the multi-scale fusion text features to obtain pitch features and energy features;
[0133] The fifth discourse analysis submodule is used to perform feature prediction after splicing the pitch feature, the energy feature and the multi-scale fusion text feature to obtain the local prosody embedding sequence feature.
[0134] In an optional embodiment, the feature mapping module includes:
[0135] A first feature mapping submodule, configured to construct a parameter learnable matrix according to the dimension of the sentence-level discourse feature with context information and the dimension of the local prosodic embedding sequence feature;
[0136] The second feature mapping submodule is used to map the sentence-level discourse features with context information based on the parameter learnable matrix using the Einstein summation convention to obtain the phoneme-level discourse features with context information.
[0137] In an optional embodiment, the paragraph-level multi-scale text prosody analysis method is implemented by a pre-trained paragraph-level multi-scale text prosody analysis model, and the device further includes a training module, which includes:
[0138] A data acquisition module is used to acquire a prosodic feature training data set, wherein each training data in the prosodic feature training data set includes: a training text and a true value local prosodic embedding sequence feature and a true value global style embedding feature corresponding to the training text;
[0139] A model training module, used for inputting the training data into the paragraph-level multi-scale text prosody analysis model for training, so that the paragraph-level multi-scale text prosody analysis model learns the mapping relationship from text to prosody features;
[0140] The model output module is used to obtain a trained chapter-level multi-scale text rhythm analysis model after the training end condition is met. The trained chapter-level multi-scale text rhythm analysis model has the ability to predict rhythmic features based on the text.
[0141] In an optional embodiment, the discourse-level multi-scale text prosody analysis model is a sub-model in the paragraph-level multi-scale text prosody analysis model; the model training module includes a sub-model training module, and the sub-model training module is used to train the discourse-level multi-scale text prosody analysis model. The discourse-level multi-scale text prosody analysis model training process includes:
[0142] The discourse-level multi-scale text prosody analysis model processes the input training sentence to obtain a multi-scale fusion text feature;
[0143] Perform pitch and energy prediction based on the multi-scale fusion text features to obtain predicted pitch features and predicted energy features;
[0144] The true value pitch feature, the true value energy feature, and the multi-scale fusion text feature are spliced together to perform feature prediction to obtain a predicted local prosody embedding sequence feature;
[0145] updating the parameters of the discourse-level multi-scale text prosody analysis model according to the comparison result of the predicted pitch feature, the predicted energy feature and the predicted local prosody embedding sequence feature with the true value pitch feature, the true value energy feature and the true value local prosody embedding sequence feature;
[0146] After the training end conditions are met, the trained discourse-level multi-scale text prosody analysis model has the ability to predict pitch features, predict energy features, and predict local prosody embedding sequence features.
[0147] In an optional embodiment, the training text further corresponds to an original audio, and the data acquisition module includes:
[0148] A first data acquisition submodule is used to extract features from the original audio using a style transfer model to obtain the true value global style embedding features, local rhythm embedding sequence features, pitch features, and energy features;
[0149] The second data acquisition submodule is used to fuse the local prosody embedding sequence feature, the pitch feature and the energy feature to obtain a true value local prosody embedding sequence feature.
[0150] In an optional embodiment, the data acquisition module further includes:
[0151] A third data acquisition submodule is used to convert the training text into phoneme features;
[0152] Using a forced alignment tool to align the original audio with the phoneme feature to obtain an aligned reference audio;
[0153] The fourth data acquisition submodule is used to extract features of the aligned reference audio using a style transfer model.
[0154] An embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the paragraph-level multi-scale text prosody analysis method described in the embodiment of the present invention.
[0155] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0156] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses and devices according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0157] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0159] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0160] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0161] The above is a detailed introduction to a method, device and equipment for multi-scale text rhythm analysis at the chapter level provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A multi-scale text prosody analysis method at the paragraph level, It is characterized in that The method comprises: Split the text to be analyzed into multiple sentences; Processing the multiple sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence; Inputting the sentence-level discourse features of the plurality of sentences into a long short-term memory network for processing to obtain a global style embedding feature at the paragraph level and a sentence-level discourse feature with context information corresponding to each sentence; Mapping the sentence-level discourse features with context information to the phoneme level to obtain the phoneme-level discourse features with context information; The phoneme-level discourse features with contextual information and the local prosody embedding sequence features are fused to obtain local prosody embedding sequence features with contextual information. The local prosody embedding sequence features with contextual information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
2. The method according to claim 1, It is characterized in that The method further comprises: Extracting features from the text to be analyzed to obtain phoneme embedding features; Speech synthesis is performed based on the phoneme embedding features, the local prosody embedding sequence features with contextual information, and the global style embedding features at the chapter level to obtain speech corresponding to the text to be analyzed.
3. The method according to claim 1, It is characterized in that The method of processing the multiple sentences using the discourse-level multi-scale text prosody analysis model to obtain local prosody embedding sequence features and sentence-level discourse features corresponding to each sentence includes: Extracting text features from the sentence to obtain word-level features and sentence-level discourse features; After fusing the word-level feature with other word-level features, the word-level feature is copied and extended using a length adjuster to obtain a phoneme-level feature; Using other phoneme-level features and the phoneme-level features to perform multimodal feature fusion to obtain multi-scale fused text features; Perform pitch and energy prediction based on the multi-scale fusion text features to obtain pitch features and energy features; The pitch feature, the energy feature and the multi-scale fusion text feature are concatenated and then feature prediction is performed to obtain the local prosody embedding sequence feature.
4. The method according to claim 1, It is characterized in that Mapping the sentence-level discourse features with context information to the phoneme level to obtain the phoneme-level discourse features with context information includes: Constructing a parameter learnable matrix according to the dimension of the sentence-level discourse feature with context information and the dimension of the local prosodic embedding sequence feature; Based on the parameter learnable matrix, the sentence-level discourse features with context information are mapped using the Einstein summation convention to obtain the phoneme-level discourse features with context information.
5. The method according to any one of claims 1 to 4, It is characterized in that The passage-level multi-scale text prosody analysis method is implemented by a pre-trained passage-level multi-scale text prosody analysis model, and the passage-level multi-scale text prosody analysis model is trained in the following manner: Acquire a prosodic feature training data set, wherein each training data in the prosodic feature training data set includes: a training text and a true value local prosodic embedding sequence feature and a true value global style embedding feature corresponding to the training text; Inputting the training data into the paragraph-level multi-scale text prosody analysis model for training, so that the paragraph-level multi-scale text prosody analysis model learns the mapping relationship from text to prosody features; After the training end condition is met, a trained chapter-level multi-scale text prosody analysis model is obtained, and the trained chapter-level multi-scale text prosody analysis model has the ability to predict prosody features according to the text.
6. The method according to claim 5, It is characterized in that The discourse-level multi-scale text prosody analysis model is a sub-model in the paragraph-level multi-scale text prosody analysis model; the discourse-level multi-scale text prosody analysis model training process includes: The discourse-level multi-scale text prosody analysis model processes the input training sentence to obtain a multi-scale fusion text feature; Perform pitch and energy prediction based on the multi-scale fusion text features to obtain predicted pitch features and predicted energy features; The true value pitch feature, the true value energy feature, and the multi-scale fusion text feature are spliced together to perform feature prediction to obtain a predicted local prosody embedding sequence feature; updating the parameters of the discourse-level multi-scale text prosody analysis model according to the comparison result of the predicted pitch feature, the predicted energy feature and the predicted local prosody embedding sequence feature with the true value pitch feature, the true value energy feature and the true value local prosody embedding sequence feature; After the training end conditions are met, the trained discourse-level multi-scale text prosody analysis model has the ability to predict pitch features, predict energy features, and predict local prosody embedding sequence features.
7. The method according to claim 5, It is characterized in that The training text also corresponds to an original audio; the true value local rhythm embedding sequence feature and the true value global style embedding feature corresponding to the training text are obtained in the following way: Using a style transfer model to extract features from the original audio, to obtain the true value global style embedding features, local rhythm embedding sequence features, pitch features and energy features; The local prosody embedding sequence feature, the pitch feature and the energy feature are fused to obtain a true value local prosody embedding sequence feature.
8. The method according to claim 7, It is characterized in that Before extracting features from the original audio using the style transfer model, the method further includes: Converting the training text into phoneme features; Using a forced alignment tool to align the original audio with the phoneme feature to obtain an aligned reference audio; The style transfer model is used to extract features from the original audio, including: A style transfer model is used to extract features from the aligned reference audio.
9. A multi-scale text prosody analysis device at the chapter level, It is characterized in that The device comprises: A text splitting module is used to split the text to be analyzed into multiple sentences; A discourse analysis module, used to process the multiple sentences using a discourse-level multi-scale text prosody analysis model to obtain local prosodic embedding sequence features and sentence-level discourse features corresponding to each sentence; A context fusion module, used for inputting the sentence-level discourse features of the plurality of sentences into a long short-term memory network for processing, to obtain a global style embedding feature at the paragraph level and a sentence-level discourse feature with context information corresponding to each sentence; A feature mapping module, used for mapping the sentence-level speech features with context information to the phoneme level to obtain the phoneme-level speech features with context information; The local feature module is used to fuse the phoneme-level discourse features with contextual information and the local prosody embedding sequence features to obtain local prosody embedding sequence features with contextual information. The local prosody embedding sequence features with contextual information and the global style embedding features at the paragraph level represent the prosody features of the text to be analyzed.
10. An electronic device, It is characterized in that The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes, the method for analyzing multi-scale text rhythm at the paragraph level as claimed in any one of claims 1 to 8 is implemented.