Nucleus position prediction method, device, equipment and storage medium

By integrating phrase pauses, phonemes, and text features from Japanese text to predict the pitch kernel position, the problem of insufficient accuracy in pitch kernel position in Japanese speech synthesis is solved, thereby improving the naturalness and intelligibility of speech synthesis.

CN119541456BActive Publication Date: 2025-12-12合肥智能语音创新发展有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411462046.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-12-12
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

The accuracy of current Japanese intonation core position prediction is insufficient, which affects the naturalness and intelligibility of Japanese speech synthesis.

Method used

By acquiring phrase pause features, phoneme features, and text features from the target Japanese text, feature fusion is performed to predict the kernel position. Features are extracted using a pre-trained text processing network and a prosody prediction network, and prediction accuracy is improved by combining polyphonic word prediction.

Benefits of technology

It improves the accuracy of core position prediction and enhances the naturalness and intelligibility of Japanese speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541456B_ABST
    Figure CN119541456B_ABST
Patent Text Reader

Abstract

The application discloses a kind of to adjust the position prediction method, device and storage medium of core, and the position prediction method of core includes: obtaining the target phrase pause feature corresponding to target Japanese text, target phoneme feature and target text feature;Target phrase pause feature, target phoneme feature and target text feature are fused, and target fusion feature is obtained;Adjust the position prediction of core using target fusion feature, and obtain the position prediction result of core corresponding to target Japanese text.The accuracy of the position prediction result of core corresponding to target Japanese text can be improved by the above manner.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech synthesis, in particular to a pitch nucleus position prediction method and device, equipment and a storage medium. BACKGROUND

[0002] Speech synthesis, also known as TTS (Text To Speech), includes a front-end module and a back-end module. The front-end module is used to extract linguistic and acoustic features related to pronunciation from input text. The quality of the front-end module will directly affect various accuracy indicators of the synthesized audio, and thus affect the overall performance, so the front-end module is an important part of the whole process of speech synthesis.

[0003] Japanese is a pitch accent language, and the pitch nucleus position (i.e. the intonation core or accent position) of Japanese not only determines the high-low change of the syllable in the sentence, but also reflects the semantic emphasis and emotional expression of the sentence. In the process of Japanese speech synthesis, the front-end module is usually required to predict the pitch nucleus position of the Japanese text. For Japanese speech synthesis, the accuracy of the pitch nucleus position prediction directly affects the naturalness and intelligibility of the synthesized speech.

[0004] Therefore, how to improve the accuracy of Japanese pitch nucleus position prediction has become a technical problem to be solved. SUMMARY

[0005] The technical problem solved by the present application is to provide a pitch nucleus position prediction method, device, equipment and computer readable storage medium, which can improve the accuracy of the pitch nucleus position prediction result of the target Japanese text.

[0006] To solve the above technical problem, one technical solution adopted by the present application is to provide a pitch nucleus position prediction method, which comprises: obtaining target phrase pause features, target phoneme features and target text features corresponding to a target Japanese text; fusing the target phrase pause features, the target phoneme features and the target text features to obtain target fusion features; and predicting the pitch nucleus position by using the target fusion features to obtain a pitch nucleus position prediction result corresponding to the target Japanese text.

[0007] To solve the above technical problem, another technical solution adopted by the present application is to provide a pitch nucleus position prediction device, which comprises: an acquisition module configured to obtain target phrase pause features, target phoneme features and target text features corresponding to a target Japanese text; a fusion module configured to fuse the target phrase pause features, the target phoneme features and the target text features to obtain target fusion features; and a prediction module configured to predict the pitch nucleus position by using the target fusion features to obtain a pitch nucleus position prediction result corresponding to the target Japanese text.

[0008] To solve the above technical problems, another technical solution adopted by the present application is to provide an electronic device, comprising a memory and a processor coupled with each other, the memory storing program instructions; the processor is used to execute the program instructions stored in the memory to realize the above-mentioned core position prediction method.

[0009] To solve the above technical problems, another technical solution adopted by the present application is to provide a computer readable storage medium, the computer readable storage medium is used to store program instructions, the program instructions can be executed by the processor to realize the above-mentioned core position prediction method.

[0010] The above scheme, by fusing the target phrase pause feature, the target phoneme feature and the target text feature corresponding to the target Japanese text, obtains the target fusion feature, and uses the target fusion feature to predict the core position, obtains the core position prediction result corresponding to the target Japanese text. Because the Japanese core position and the text and pronunciation within the phrase pause boundary are related, and the target phrase pause feature, the target phoneme feature and the target text feature can reflect the phrase pause information, pronunciation information and text semantic information of the target Japanese text respectively, therefore, using the target fusion feature to predict the core position can comprehensively integrate the phrase pause, pronunciation and text semantic information of the target Japanese text in three different dimensions, so that the core position prediction result obtained by prediction has higher accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a flowchart of an embodiment of the core position prediction method provided by the present application;

[0012] Figure 2 is a flowchart of another embodiment of the core position prediction method provided by the present application;

[0013] Figure 3 is a flowchart of an embodiment of determining the target phrase pause feature provided by the present application;

[0014] Figure 4 is a flowchart of an embodiment of obtaining the target phoneme feature provided by the present application;

[0015] Figure 5 is a flowchart of an embodiment of obtaining the target fusion feature provided by the present application;

[0016] Figure 6 is a structural diagram of the target prediction model provided by the present application;

[0017] Figure 7 is a framework diagram of an embodiment of the core position prediction device provided by the present application;

[0018] Figure 8is a frame schematic diagram of an embodiment of the electronic device provided in the present application.

[0019] Figure 9 is a frame schematic diagram of an embodiment of the computer readable storage medium provided in the present application. DETAILED DESCRIPTION

[0020] For the purpose, technical solutions and effects of the present application to be clearer, more explicit, the following refers to the drawings and embodiments of the present application are further described in detail.

[0021] In the following description, for the purpose of explanation and not for the purpose of limitation, specific details such as specific system architecture, interface, technology, etc. are presented in order to thoroughly understand the present application.

[0022] It should be noted that the term "and / or" herein is only a description of the associated relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. The term "several" means at least one. In addition, the terms "first", "second" and the like in the specification and claims herein and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0023] Please refer to Figure 1 , Figure 1 is a flow schematic diagram of an embodiment of the method for predicting the position of the present application. It should be noted that the method of the present application is not limited to the order of the flow shown in Figure 1 . As shown in Figure 1 , the method comprises the following steps:

[0024] S11: obtaining target Japanese text corresponding to the target phrase pause feature, target phoneme feature and target text feature.

[0025] The target Japanese text is the Japanese text to be synthesized by voice. The target phrase pause feature, the target phoneme feature and the target text feature are all extracted from the target Japanese text. The target phrase pause feature includes the phrase pause information corresponding to the target Japanese text, the target phoneme feature includes the pronunciation information corresponding to the target Japanese text, and the target text feature includes the text semantic information corresponding to the target Japanese text.

[0026] S12: fusing the target phrase pause feature, the target phoneme feature and the target text feature to obtain the target fusion feature.

[0027] The target phrase pause feature, the target phoneme feature and the target text feature can not be features of the same granularity, for example, the target phrase pause feature is a feature of word granularity, the target phoneme feature is a feature of phoneme granularity and the target text feature is a feature of word granularity. In order to facilitate feature fusion and subsequent prediction of the position of the adjustment, the target phrase pause feature, the target phoneme feature and the target text feature are all converted into features of syllable granularity, and then the target phrase pause feature of syllable granularity, the target phoneme feature of syllable granularity and the target text feature of syllable granularity are fused to obtain the target fusion feature of syllable granularity. The target fusion feature of syllable granularity is in units of syllables, and each syllable contains phrase pause information, pronunciation information and text semantic information corresponding to the syllable.

[0028] In an embodiment, the target phrase pause feature of syllable granularity, the target phoneme feature of syllable granularity and the target text feature of syllable granularity can be directly spliced by syllable to obtain the target fusion feature of syllable granularity.

[0029] In another embodiment, the target phrase pause feature and the target phoneme feature can only reflect the phrase pause information and the pronunciation information of the target Japanese text respectively, and do not contain deep information. Therefore, in order to further improve the accuracy of the subsequent prediction of the position of the adjustment, in this embodiment, the target phrase pause feature of syllable granularity and the target phoneme feature of syllable granularity are first extracted to obtain the first extracted feature of syllable granularity, and then the target phoneme feature of syllable granularity and the first extracted feature of syllable granularity are spliced by syllable to obtain the target fusion feature.

[0030] S13: predicting the position of the adjustment by using the target fusion feature to obtain the prediction result of the position of the adjustment corresponding to the target Japanese text.

[0031] In this embodiment, the target fusion feature is obtained by fusing the target phrase pause feature, the target phoneme feature and the target text feature corresponding to the target Japanese text, and the prediction result of the position of the adjustment corresponding to the target Japanese text is obtained by predicting the position of the adjustment by using the target fusion feature. Since the Japanese adjustment position is related to the text and pronunciation within the phrase pause boundary, and the target phrase pause feature, the target phoneme feature and the target text feature can reflect the phrase pause information, the pronunciation information and the text semantic information of the target Japanese text respectively, the prediction of the position of the adjustment by using the target fusion feature can comprehensively integrate the information of the phrase pause, the pronunciation and the text semantic of the target Japanese text in three different dimensions, so that the prediction result of the position of the adjustment is more accurate.

[0032] Please refer to Figure 2 , Figure 2 is a flowchart of another embodiment of the method for predicting the position of the adjustment provided by the present application. As shown in Figure 2As shown, the method comprises the following steps:

[0033] S21: processing the target Japanese text by using a pre-trained text processing network to obtain a character granularity text feature corresponding to the target Japanese text.

[0034] The character granularity text feature is in units of characters, and the character granularity text feature comprises subtext features respectively corresponding to each character of the target Japanese text, and each character subtext feature is a multi-dimensional feature.

[0035] In an embodiment, the pre-trained text processing network is a character granularity Bert (Bidirectional Encoder Representations from Transformers) to obtain a character granularity text feature with context information. Illustratively, the character granularity Bert is a standard 12-layer 768-node structure. The structure comprises 12 layers of Transformer encoders, each of which is responsible for processing different aspects of the input data and capturing context information through a self-attention mechanism, and the size of the hidden layer of each layer is 768 to map each character of the target Japanese text to a 768-dimensional feature.

[0036] Optionally, the pre-trained text processing network is not limited to the character granularity Bert, and in other embodiments, other similar text processing networks can also be used to obtain the character granularity text feature of the target Japanese text.

[0037] Considering that Japanese contains a large number of Chinese characters and hiragana, and the segmentation is relatively ambiguous, directly processing each character in the target Japanese text, rather than first converting it into a higher-level language unit (such as a word or phrase) and then processing it, helps the pre-trained text processing network to more accurately capture the detailed information in the target Japanese text, thereby improving the accuracy of the subsequently obtained target phrase pause feature, target phoneme feature and target text feature.

[0038] S22: determining the target phrase pause feature, the target phoneme feature and the target text feature of the target Japanese text based on the character granularity text feature.

[0039] In an embodiment, the step of determining the target phrase pause feature is performed by a prosody prediction network. It should be noted that in this embodiment, the prosody of the target Japanese text mainly comprises phrase pauses (i.e. L1 pauses) and long pauses (i.e. L3 pauses), the phrase pause is a pause at the level of a stress phrase in the target Japanese text, and the long pause is a pause at the level of a sentence or paragraph in the target Japanese text.

[0040] The prosody prediction network is a cascaded MLP (Multilayer Perceptron) structure. Specifically, the prosody prediction network comprises a first feature extraction subnetwork, a second feature extraction subnetwork, a long pause prediction subnetwork, and a phrase break prediction subnetwork.

[0041] Figure 3 is a flowchart of an embodiment of determining target phrase break features provided by the present application. As shown in Figure 3 the method comprises the following steps:

[0042] S301: performing feature extraction on the word granularity text features by using the first feature extraction subnetwork to obtain word granularity long pause extraction features.

[0043] The first feature extraction subnetwork is used to further abstract and transform the word granularity text features output by the pre-trained text processing network, so as to extract features conducive to long pause prediction, i.e., word granularity long pause extraction features. The word granularity long pause extraction features are in units of characters, and the word granularity long pause extraction features comprise long pause sub-extraction features corresponding to each character of the target Japanese text. Each character corresponding long pause sub-extraction feature is a multi-dimensional feature (such as a 768-dimensional feature). Exemplarily, the first feature extraction subnetwork is a 768-node fully connected layer.

[0044] S302: predicting the word granularity long pause extraction features by using the long pause prediction subnetwork to obtain target long pause features of word granularity.

[0045] In a specific embodiment, the target long pause features comprise first classification results corresponding to each character of the target Japanese text. The first classification result of each character comprises probabilities that each character belongs to a non-long pause, a long pause, and a padding, respectively. Further, the first classification result of each character can also be represented using a one-hot label. For example, the first classification result of a certain character is [0.1, 0.6, 0.3], wherein 0.1, 0.6, and 0.3 are probabilities that the character belongs to a non-long pause, a long pause, and a padding, respectively, and the corresponding one-hot label is represented as [0, 1, 0].

[0046] S303: performing feature extraction on the word granularity long pause extraction features by using the second feature extraction subnetwork to obtain word granularity phrase break extraction features.

[0047] The second feature extraction sub-network is configured to further abstract and transform the word granularity long pause features output by the first feature extraction sub-network to extract features beneficial to the phrase pause prediction, i.e., word granularity phrase pause extraction features. The word granularity phrase pause extraction features are in units of characters, and each character of the target Japanese text corresponds to a phrase pause sub-extraction feature, each of which is a multi-dimensional feature. Exemplarily, the second feature extraction sub-network is also a fully connected layer with 768 nodes.

[0048] S304: predicting the word granularity phrase pause extraction features by using the phrase pause prediction sub-network to obtain target word granularity phrase pause features.

[0049] In an embodiment, the target pause features include second classification results corresponding to each character of the target Japanese text. Each character's second classification result includes the probability of each character belonging to non-phrase pause, phrase pause, and padding. Further, the second classification result of each character can also be represented using one-hot labels.

[0050] Exemplarily, the long pause prediction sub-network and the phrase pause prediction sub-network are both fully connected layers.

[0051] In steps S301 to S304, considering that the long pause and phrase pause in the target Japanese text have little to do with the word pronunciation, and mainly depend on the meaning of the text itself, the prosody prediction of the target Japanese text is closer to an NLP (Natural Language Processing) task, and the level of long pause in Japanese grammar is higher than that of phrase pause, i.e., long pause is necessarily phrase pause, but phrase pause is not necessarily long pause, therefore, a two-level cascade structure prosody prediction network is used as a downstream task of the pre-trained text processing network to predict long pause and phrase pause layer by layer. The first layer of the prosody prediction network includes the first feature extraction sub-network and the long pause prediction sub-network, which are configured to extract word granularity long pause extraction features and predict target word granularity long pause features; the second layer of the prosody prediction network includes the second feature extraction sub-network and the phrase pause prediction sub-network, which are configured to extract word granularity phrase pause extraction features and predict target word granularity phrase pause features. Extracting word granularity long pause extraction features first can provide more valuable context information for the prediction of target phrase pause features, which helps the network better utilize the hierarchical information of long pause and phrase pause and improve the accuracy of target phrase pause feature prediction.

[0052] In an embodiment, an initial phoneme feature of a target Japanese text and a multiple-reading word prediction result of the target Japanese text are obtained first, and then the initial phoneme feature of the target Japanese text is updated based on the multiple-reading word prediction result of the target Japanese text to obtain a target phoneme feature of the target Japanese text.

[0053] Figure 4 is a flowchart of an embodiment of obtaining a target phoneme feature provided in the present application. As shown in the figure, the method comprises the following steps: Figure 4

[0054] S401: Obtain an initial phoneme feature of a target Japanese text.

[0055] Specifically, step S401 comprises: performing a word segmentation processing on the target Japanese text to obtain a plurality of segmented Japanese words; and determining an initial phoneme feature of the target Japanese text. The initial phoneme feature of the target Japanese text comprises initial phoneme sub-features of each Japanese word determined based on a Japanese dictionary (i.e. an initial phoneme sequence of each Japanese word). After the word segmentation processing on the target Japanese text, the initial phoneme sub-features of each Japanese word can be obtained based on a query of the Japanese dictionary.

[0056] For example, the Mecab word segmenter can be used to perform the word segmentation processing on the target Japanese text, and the Japanese dictionary can be the unidic dictionary.

[0057] S402: Convert the character-granularity phrase pause extraction feature into a word-granularity phrase pause extraction feature.

[0058] The character-granularity phrase pause extraction feature comprises a plurality of phrase pause sub-extraction features corresponding to each character of the target Japanese text. Since each Japanese word comprises a plurality of characters, each Japanese word comprises a plurality of phrase pause sub-extraction features corresponding to the corresponding plurality of characters. The plurality of phrase pause sub-extraction features of each Japanese word can be integrated respectively to obtain a phrase pause extraction feature capable of representing each Japanese word, i.e. a word-granularity phrase pause extraction feature.

[0059] Specifically, a plurality of phrase pause sub-extraction features corresponding to each Japanese word are obtained; and the plurality of phrase pause sub-extraction features of each Japanese word are subjected to average pooling processing respectively to obtain a word-granularity phrase pause extraction feature. For each Japanese word, the average pooling processing of the plurality of phrase pause sub-extraction features of the Japanese word means taking the average value of the plurality of phrase pause sub-extraction features of the Japanese word.

[0060] S403: Use a multiple-reading word prediction network to predict the word-granularity phrase pause extraction feature to obtain a multiple-reading word prediction result of the target Japanese text.

[0061] ​The multi-reading word prediction result includes the multi-reading word category of each Japanese word. The multi-reading word prediction result can also be represented by a one-hot label. A multi-reading word label with a class plus 1 can be designed according to the supported Japanese multi-reading word categories. For example, 15 words and 36 categories can be designed 37 multi-reading word labels. The 0 dimension represents a non-multi-reading word, and the other dimensions correspond to the remaining 36 multi-reading word labels. Exemplarily, the multi-reading word prediction network is a fully connected layer.

[0062] In this embodiment, considering that most Japanese multi-reading words appear at the word granularity, Japanese multi-reading word prediction is performed at the word granularity. Since the multi-reading word prediction task is mainly determined by the text semantic information of the target Japanese text, and the upstream pre-training text processing network is at the word granularity and the word granularity is lower than the phrase pause level, the multi-reading word prediction in this embodiment is performed after the phrase pause prediction. The multi-reading word prediction is performed after converting the word granularity phrase pause extraction feature to the word granularity phrase pause extraction feature. In this way, the multi-reading word prediction can be performed using the phrase pause information, thereby improving the accuracy of the multi-reading word prediction result.

[0063] S404: Based on the multi-reading word prediction result, determine the multi-reading word in the plurality of Japanese words of the target Japanese text.

[0064] After obtaining the multi-reading word prediction result of the target Japanese text, the multi-reading word in the plurality of Japanese words can be further determined based on the multi-reading word category of each Japanese word.

[0065] S405: Determine the target phoneme sub-feature corresponding to each multi-reading word.

[0066] Specifically, for each multi-reading word, the target phoneme sub-feature (i.e. the phoneme sequence corresponding to the multi-reading word) of the multi-reading word can be obtained based on the specific multi-reading word category of the multi-reading word.

[0067] S406: Update the initial phoneme sub-feature of each multi-reading word in the initial phoneme feature to the corresponding target phoneme sub-feature to obtain the target phoneme feature.

[0068] In steps S401 to S406, considering that the initial phoneme sub-feature of the Japanese word obtained by querying the Japanese dictionary can be the default pronunciation of the Japanese word, i.e. the accuracy of the initial phoneme feature of the target Japanese text is not high. Therefore, based on the multi-reading word prediction result of the target Japanese text, the accurate pronunciation of the multi-reading word of the target Japanese text can be obtained, and the initial phoneme feature of the target Japanese text is updated to obtain the target phoneme feature, which can improve the accuracy of the obtained target phoneme feature.

[0069] In an embodiment, the step of determining the target text feature comprises: obtaining any one of the following features as the target text feature: the character granularity text feature, the character granularity long pause extraction feature, the character granularity phrase pause extraction feature, and the word granularity phrase pause extraction feature converted from the character granularity phrase pause extraction feature.

[0070] Since the character granularity text feature, the character granularity long pause extraction feature, the character granularity phrase pause extraction feature, and the word granularity phrase pause extraction feature converted from the character granularity phrase pause extraction feature can all reflect the text semantic information of the target Japanese text, the character granularity text feature, the character granularity long pause extraction feature, the character granularity phrase pause extraction feature, and the word granularity phrase pause extraction feature converted from the character granularity phrase pause extraction feature can all be used as the target text feature of the target Japanese text.

[0071] Exemplarily, the word granularity phrase pause extraction feature converted from the character granularity phrase pause extraction feature is used as the target text feature.

[0072] S23: Fusing the target phrase pause feature, the target phoneme feature, and the target text feature to obtain a target fusion feature.

[0073] Figure 5 is a flowchart of an embodiment of obtaining the target fusion feature provided by the present application. As shown in Figure 5 the method comprises the following steps:

[0074] S501: Converting the target phrase pause feature, the target phoneme feature, and the target text feature into features of the syllable granularity.

[0075] Since the accent nucleus position of Japanese is a feature of controlling syllable accent on the accent phrase pause boundary of Japanese, determining the accent nucleus position of Japanese needs to follow: locating the accent phrase boundary; locating the syllable boundary; and combining the text to predict the accent nucleus position on the syllable. That is, the accent nucleus position needs to be predicted on the syllable granularity. Therefore, the obtained target phrase pause feature, the target phoneme feature, and the target text feature all need to be converted into features of the syllable granularity.

[0076] In a specific embodiment, the obtained target phrase pause feature is a feature of the character granularity, the target phoneme feature is a feature of the phoneme granularity, and the target text feature is the word granularity phrase pause extraction feature, i.e., the target text feature is a feature of the word granularity.

[0077] Due to the writing system of Japanese (including Hiragana, Katakana, Kanji, etc.) and the syllable, there is no direct one-to-one mapping relationship, especially when it comes to Kanji, a Kanji may correspond to multiple syllables, depending on its specific reading in the sentence. Therefore, in Japanese, there is no certain mapping relationship between single characters and syllables, and only through the mapping relationship between characters, words and syllables, the character granularity features can be converted into word granularity features, and then into syllable granularity features at the cost of part of the character granularity information.

[0078] Specifically, the target phrase pause features of the character granularity can be mapped into the target phrase pause features of the syllable granularity according to the pre-acquired mapping relationship of character-word-syllable.

[0079] Specifically, the target text features of the word granularity (i.e., the phrase pause extraction features of the word granularity) can be further mapped into the target text features of the syllable granularity according to the pre-acquired mapping relationship of word-syllable.

[0080] Illustratively, the corresponding syllables of each Japanese word can be determined based on the pre-acquired mapping relationship of word-syllable; and the phrase pause extraction features of the Japanese word are subjected to average reverse pooling processing to obtain the phrase pause extraction features corresponding to each syllable of the Japanese word. Wherein, the phrase pause extraction features of the Japanese word are subjected to average reverse pooling processing, which means that the phrase pause extraction features of the Japanese word are copied to each syllable of the Japanese word, that is, the phrase pause extraction features of each syllable of the Japanese word are the same as the phrase pause extraction features of the corresponding Japanese word.

[0081] Specifically, the target phoneme features of the phoneme granularity can be mapped into the target phoneme features of the syllable granularity according to the pre-acquired mapping relationship of phoneme-syllable.

[0082] S502: using a third feature extraction sub-network of the kernel position prediction network to perform feature extraction on the target phrase pause features of the syllable granularity and the target phoneme features of the syllable granularity, to obtain first extraction features.

[0083] Wherein, the third feature extraction sub-network includes a first embedding sub-network, a second embedding sub-network and a recurrent neural sub-network. Specifically, step S502 includes the following sub-steps:

[0084] Sub-step one, using the first embedding sub-network to process the target phrase pause features of the syllable granularity to obtain first embedding features of the syllable granularity.

[0085] Sub-step two, using the second embedding sub-network to encode the target phoneme features of the syllable granularity to obtain second embedding features of the syllable granularity.

[0086] The first embedding sub-network and the second embedding sub-network are both embedding sub-networks.

[0087] Exemplarily, the first embedding sub-network and the second embedding sub-network are both 256 nodes, the first embedding sub-network is used to map the features of each syllable in the syllable-granularity target phrase pause feature into 256-dimensional features respectively, and the second embedding sub-network is used to map the features of each syllable in the syllable-granularity target phoneme feature into 256-dimensional features respectively, so as to better capture the semantic and contextual relationship between the features of each syllable subsequently.

[0088] In substep three, the syllable-granularity first embedding feature and the syllable-granularity second embedding feature are spliced.

[0089] Specifically, the syllable-granularity first embedding feature and the syllable-granularity second embedding feature are spliced by syllable to obtain the spliced embedding feature.

[0090] In substep four, the spliced embedding feature is extracted by a recurrent neural sub-network to obtain the syllable-granularity first extraction feature.

[0091] Exemplarily, the recurrent neural sub-network is a bidirectional LSTM, which is used to extract the contextual information on the syllable sequence from the spliced embedding feature, i.e., the first extraction feature. The first extraction feature obtained in this substep includes a deep-level comprehensive representation of the phrase pause and phoneme information.

[0092] S503: The syllable-granularity target text feature and the syllable-granularity first extraction feature are spliced to obtain a target fusion feature.

[0093] Specifically, the syllable-granularity target text feature and the syllable-granularity first extraction feature are spliced by syllable to obtain the target fusion feature.

[0094] S24: The target fusion feature is used for pitch position prediction to obtain a pitch position prediction result corresponding to the target Japanese text.

[0095] In an embodiment, step 24 includes: extracting the target fusion feature by a fourth feature extraction sub-network of the pitch position prediction network to obtain a second extraction feature; and predicting the second extraction feature by a pitch position prediction sub-network of the pitch position prediction network to obtain the pitch position prediction result.

[0096] The second extraction feature extracted by the fourth feature extraction sub-network includes a deep-level comprehensive representation of the phrase pause, phoneme, and text semantic information, which is beneficial to subsequent pitch position prediction.

[0097] The tonal nucleus position prediction result is also of syllable granularity, and specifically includes third classification results corresponding to each syllable. The third classification result of each syllable includes a probability of being a tonal nucleus position and a probability of not being a tonal nucleus position. Further, the third classification result of each syllable can also be represented using a one-hot label.

[0098] Exemplarily, the fourth feature extraction subnetwork includes a preset number of convolutional layers. For example, the preset number can be 2.

[0099] Exemplarily, the tonal nucleus position prediction subnetwork is an LSTM autoregressive prediction subnetwork, which includes a bidirectional LSTM and a fully connected layer.

[0100] Further, after obtaining the tonal nucleus position prediction result of the target Japanese text, the tonal nucleus position prediction result is input into a backend speech synthesis model to assist the backend speech synthesis.

[0101] Further, in addition to inputting the obtained tonal nucleus position prediction result into the backend speech synthesis model, the aforementioned obtained target long pause feature, target phrase pause feature, and multi-syllable word prediction result can also be input into the backend speech synthesis model to assist the backend speech synthesis.

[0102] In the related art, tonal nucleus position prediction is performed based on a CRF algorithm. For a longer and more complex target Japanese text, the tonal nucleus position prediction accuracy using this method will decrease significantly, and the generalization ability is relatively general, and relies on relatively complete data reserves.

[0103] In this embodiment, the pre-trained text processing network is used to obtain the character granularity text features corresponding to the target Japanese text, and the target phrase pause feature, the target phoneme feature, and the target text feature used for tonal nucleus position prediction are determined based on the character granularity text features. Since the character granularity text features obtained by the pre-trained text processing network contain rich context semantic information, the accuracy of the target phrase pause feature, the target phoneme feature, and the target text feature can be improved, thereby further improving the accuracy of tonal nucleus position prediction. Moreover, this embodiment integrates multi-task prediction, and can jointly predict the tonal nucleus position, prosody (including long pause and phrase pause), and multi-syllable word features required for backend speech synthesis, which can greatly improve the prediction performance of each front-end feature, and the performance on complex long sentences is more prominent, having the advantages of high prediction accuracy and low deployment cost.

[0104] Please refer to Figure 6 , Figure 6 is a structural diagram of a target prediction model provided by the present application. As shown in Figure 6 , the target prediction model includes a pre-trained text processing network, a prosody prediction network, a multi-syllable word prediction network, and a tonal nucleus position prediction network.

[0105] The pre-trained text processing network is configured to obtain word granularity text features of the target Japanese text.

[0106] The prosody prediction network is configured to extract word granularity long pause extraction features and word granularity phrase pause extraction features based on the word granularity text features, and predict target long pause features at word granularity and target phrase pause features at word granularity. The prosody prediction network specifically includes a first feature extraction subnetwork, a second feature extraction subnetwork, a long pause prediction subnetwork, and a phrase pause prediction subnetwork. The functions of the subnetworks of the prosody prediction network can be referred to the embodiments shown in Figure 3 , and will not be described here again.

[0107] The multi-phonetic word prediction network is configured to obtain a multi-phonetic word prediction result of the target Japanese text based on the word granularity phrase pause extraction features.

[0108] The core position prediction network is configured to perform core position prediction based on the target phrase pause features, target phoneme features obtained from the multi-phonetic word prediction result, and target text features, to obtain a core position prediction result of the target Japanese text. The core position prediction network specifically includes a third feature extraction subnetwork, a fourth feature extraction subnetwork, and a core position prediction subnetwork. The third feature extraction subnetwork further includes a first embedding subnetwork, a second embedding subnetwork, and a recurrent neural subnetwork (not shown in Figure 6 ). The functions of the subnetworks of the core position prediction network can be referred to the related content in the foregoing steps S23 to S24, and will not be described here again.

[0109] In this embodiment, before applying the target prediction model, a sample data set is obtained, and then the sample data set is used to train the target prediction model. The sample data set includes a plurality of sample Japanese texts, and each sample Japanese text is annotated with word segmentation information, phoneme information, sample long pause annotation features, sample phrase pause annotation features, sample multi-phonetic word annotation results, and sample core position annotation results.

[0110] To ensure the reliability of the sample data set, before the sample data set is used to train the target prediction model, the sample Japanese texts in the sample data set can also be preprocessed, for example, written text normalization, processing of abnormal character text symbols, etc.

[0111] In this embodiment, the target prediction model can be first trained. The first training is used to train the adjustment position prediction network in the target prediction model by using the first sample Japanese text. The steps of the first training include: obtaining the sample phrase pause feature, the sample phoneme feature and the sample text feature of the first sample Japanese text; obtaining the sample adjustment position prediction result of the first sample Japanese text based on the sample phrase pause feature, the sample phoneme feature and the sample text feature of the first sample Japanese text by using the adjustment position prediction network; and adjusting the network parameters of the adjustment position prediction network based on the sample adjustment position prediction result of the first sample Japanese text.

[0112] For example, the sample phrase pause feature of the first sample Japanese text is the sample phrase pause annotation feature of the first sample Japanese text. The sample phoneme feature of the first sample Japanese text can be the sample phoneme annotation feature of the first sample Japanese text, or the sample phoneme feature of the first sample Japanese text can be determined based on the sample multi-phonological word annotation result of the first sample Japanese text. The sample text feature of the first sample Japanese text can be obtained based on the sample character granularity text feature of the first sample Japanese text, and the sample character granularity text feature of the first sample Japanese text can be obtained by processing the first sample Japanese text by using the pre-trained text processing network. For example, the sample text feature of the first sample Japanese text is the sample character granularity text feature, which can also be the sample character granularity long pause extraction feature, the sample character granularity phrase pause extraction feature or the word granularity phrase pause extraction feature obtained based on the sample character granularity text feature.

[0113] Specifically, in the process of the first training, the first loss can be calculated based on the difference between the first sample adjustment position prediction result of the first sample Japanese text and the sample adjustment position annotation result of the first sample Japanese text; and the network parameters of the adjustment position prediction network are adjusted based on the first loss.

[0114] Before the first training of the target prediction model, the second training of the target prediction model is performed. The second training is used to train the pre-trained text processing network and the prosody prediction network in the target prediction model by using the second sample Japanese text. The steps of the second training include: obtaining the character granularity text feature of the second sample Japanese text by using the pre-trained text processing network; obtaining the sample long pause feature and the sample phrase pause feature of the second sample Japanese text by using the prosody prediction network based on the character granularity text feature of the second sample Japanese text; and adjusting the network parameters of the pre-trained text processing network and the prosody prediction network based on the sample long pause feature and the sample phrase pause feature of the second sample Japanese text.

[0115] Specifically, a second loss is calculated based on a difference between the sample long pause feature of the second sample Japanese text and the sample long pause labeled feature of the second sample Japanese text, and a third loss is calculated based on a difference between the sample phrase pause feature of the second sample Japanese text and the sample phrase pause labeled feature of the second sample Japanese text; based on the second loss and the third loss, the network parameters of the pre-trained text processing network and the prosody prediction network are adjusted.

[0116] After the second training and the first training of the target prediction model, the target prediction model can be further trained in a third training, and the sample multi-phonetic word labeled result of the third sample Japanese text is loaded to participate in the downstream nucleus position prediction in the third training. The steps of the third training include: obtaining the word granularity text feature of the third sample Japanese text by using the pre-trained text processing network; predicting the sample long pause feature and the sample phrase pause feature of the third sample Japanese text based on the word granularity text feature of the third sample Japanese text by using the prosody prediction network; determining the sample text feature of the third sample Japanese text based on the word granularity text feature of the third sample Japanese text; determining the sample phoneme feature of the third sample Japanese text based on the sample multi-phonetic word labeled result of the third sample Japanese text; obtaining the sample nucleus position prediction result of the third sample Japanese text based on the sample phrase pause feature, the sample phoneme feature and the sample text feature of the third sample Japanese text by using the nucleus position prediction network; and adjusting the network parameters of the pre-trained text processing network, the prosody prediction network and the nucleus position prediction network based on the sample long pause feature, the sample phrase pause feature and the sample nucleus position prediction result of the third sample Japanese text.

[0117] Specifically, a fourth loss, a fifth loss and a sixth loss are respectively calculated based on a difference between the sample long pause feature of the third sample Japanese text and the sample long pause labeled feature of the third sample Japanese text, a difference between the sample phrase pause feature of the third sample Japanese text and the sample phrase pause labeled feature of the third sample Japanese text, and a difference between the sample nucleus position prediction result of the third sample Japanese text and the sample nucleus position labeled result of the third sample Japanese text; and the network parameters of the pre-trained text processing network, the prosody prediction network and the nucleus position prediction network are adjusted based on the loss sums corresponding to the fourth loss, the fifth loss and the sixth loss.

[0118] After the third training of the target prediction model, the target prediction model can be further trained for a fourth training, in which a multi-phonetic word prediction task is added, but the sample multi-phonetic word annotation results of the fourth sample Japanese text are loaded to participate in the downstream adjustment of the position prediction. The steps of the fourth training include: obtaining the character granularity text features of the fourth sample Japanese text by using the pre-trained text processing network; obtaining the sample character granularity long pause extraction features and the sample character granularity phrase pause extraction features of the fourth sample Japanese text based on the character granularity text features of the fourth sample Japanese text by using the prosody prediction network, and predicting the sample long pause features and the sample phrase pause features of the fourth sample Japanese text based on the sample character granularity long pause extraction features and the sample character granularity phrase pause extraction features of the fourth sample Japanese text respectively; determining the sample text features of the fourth sample Japanese text based on the character granularity text features of the fourth sample Japanese text; determining the sample word granularity phrase pause extraction features of the fourth sample Japanese text based on the sample character granularity phrase pause extraction features of the fourth sample Japanese text; obtaining the sample multi-phonetic word prediction results of the fourth sample Japanese text based on the sample word granularity phrase pause extraction features of the fourth sample Japanese text by using the multi-phonetic word prediction network; determining the sample phoneme features of the fourth sample Japanese text based on the sample multi-phonetic word annotation results of the fourth sample Japanese text; obtaining the sample adjustment position prediction results of the fourth sample Japanese text based on the sample phrase pause features, the sample phoneme features and the sample text features of the fourth sample Japanese text by using the adjustment position prediction network; and adjusting the overall network parameters of the target prediction model based on the sample long pause features, the sample phrase pause features, the sample multi-phonetic word prediction results and the adjustment position prediction results of the fourth sample Japanese text.

[0119] Specifically, based on the differences between the sample long pause features of the fourth sample Japanese text and the sample long pause annotation features of the fourth sample Japanese text, the differences between the sample phrase pause features of the fourth sample Japanese text and the sample phrase pause annotation features of the fourth sample Japanese text, the differences between the sample multi-phonetic word prediction results of the fourth sample Japanese text and the sample multi-phonetic word annotation results of the fourth sample Japanese text, and the differences between the sample adjustment position prediction results of the fourth sample Japanese text and the sample adjustment position annotation results of the fourth sample Japanese text, the seventh loss, the eighth loss, the ninth loss and the tenth loss are calculated respectively; and the overall network parameters of the target prediction model are adjusted based on the loss sums corresponding to the seventh loss, the eighth loss, the ninth loss and the tenth loss.

[0120] After the fourth training of the target prediction model, the target prediction model can be further trained for a fifth training, and in the fifth training, no annotation information of the fifth sample Japanese text is loaded to participate in the downstream adjustment position prediction. The steps of the fifth training include: obtaining the word granularity text features of the fifth sample Japanese text by using the pre-training text processing network; obtaining the sample word granularity long pause extraction features and the sample word granularity phrase pause extraction features of the fifth sample Japanese text based on the word granularity text features of the fifth sample Japanese text by using the prosody prediction network, and predicting the sample long pause features and the sample phrase pause features of the fifth sample Japanese text based on the sample word granularity long pause extraction features and the sample word granularity phrase pause extraction features of the fifth sample Japanese text respectively; determining the sample text features of the fifth sample Japanese text based on the word granularity text features of the fifth sample Japanese text; determining the sample word granularity phrase pause extraction features of the fifth sample Japanese text based on the sample word granularity phrase pause extraction features of the fifth sample Japanese text; obtaining the sample multi-phonetic word prediction results of the fifth sample Japanese text based on the sample word granularity phrase pause extraction features of the fifth sample Japanese text by using the multi-phonetic word prediction network; determining the sample phoneme features of the fifth sample Japanese text based on the sample multi-phonetic word prediction results of the fifth sample Japanese text; obtaining the sample adjustment position prediction results of the fifth sample Japanese text based on the sample long pause features, the sample phrase pause features and the sample phoneme features of the fifth sample Japanese text by using the adjustment position prediction network; and adjusting the overall network parameters of the target prediction model based on the sample long pause features, the sample phrase pause features, the sample multi-phonetic word prediction results and the sample adjustment position prediction results of the fifth sample Japanese text.

[0121] Specifically, based on the differences between the sample long pause features of the fifth sample Japanese text and the sample long pause annotation features of the fifth sample Japanese text, the differences between the sample phrase pause features of the fifth sample Japanese text and the sample phrase pause annotation features of the fifth sample Japanese text, the differences between the sample multi-phonetic word prediction results of the fifth sample Japanese text and the sample multi-phonetic word annotation results of the fifth sample Japanese text, and the differences between the sample adjustment position prediction results of the fifth sample Japanese text and the sample adjustment position annotation results of the fifth sample Japanese text, the eleventh loss, the twelfth loss, the thirteenth loss and the fourteenth loss are calculated respectively; and the overall network parameters of the target prediction model are adjusted based on the loss sums corresponding to the eleventh loss, the twelfth loss, the thirteenth loss and the fourteenth loss.

[0122] Exemplarily, the first loss to the fourteenth loss described above are all cross-entropy losses. The first sample Japanese text, the second sample Japanese text, the third sample Japanese text, the fourth sample Japanese text and the fifth sample Japanese text can be the same sample Japanese text or different sample Japanese texts.

[0123] In this embodiment, a multi-task teacher-force training strategy is adopted to gradually train the target prediction model, and the target prediction model is obtained through second training, first training, third training, fourth training and fifth training in turn. Through this training method, the target prediction model can simultaneously have the ability to predict prosody, multi-sound words and tone nucleus position, and can reduce error propagation and speed up the training process, thereby improving the prediction performance of the target prediction model.

[0124] Please refer to Figure 7 , Figure 7 is a framework schematic diagram of an embodiment of the tone nucleus position prediction device provided in the present application. In this embodiment, the tone nucleus position prediction device 70 comprises an acquisition module 71, a fusion module 72 and a prediction module 73.

[0125] The acquisition module 71 is configured to acquire target phrase pause features, target phoneme features and target text features corresponding to a target Japanese text; the fusion module 72 is configured to fuse the target phrase pause features, the target phoneme features and the target text features to obtain target fusion features; and the prediction module 73 is configured to perform tone nucleus position prediction by using the target fusion features to obtain a tone nucleus position prediction result corresponding to the target Japanese text.

[0126] Optionally, the acquisition module 71 is configured to process the target Japanese text by using a pre-trained text processing network to obtain text features of a word granularity corresponding to the target Japanese text; and the target phrase pause features, the target phoneme features and the target text features are determined based on the text features of the word granularity.

[0127] Optionally, the step of determining the target phrase pause features is performed by a prosody prediction network, and the prosody prediction network comprises a first feature extraction sub-network, a second feature extraction sub-network and a phrase pause prediction sub-network. The acquisition module 71 is configured to perform feature extraction on the text features of the word granularity by using the first feature extraction sub-network to obtain long pause extraction features of the word granularity; perform feature extraction on the long pause extraction features of the word granularity by using the second feature extraction sub-network to obtain phrase pause extraction features of the word granularity; and perform prediction on the phrase pause extraction features of the word granularity by using the phrase pause prediction sub-network to obtain the target phrase pause features of the word granularity.

[0128] Optionally, the prosody prediction network further comprises a long pause prediction sub-network, and the acquisition module 71 is further configured to perform prediction on the long pause extraction features of the word granularity by using the long pause prediction sub-network to obtain target long pause features of the word granularity; wherein the tone nucleus position prediction result, the target long pause features and the target phrase pause features are all used as inputs of a backend speech synthesis model.

[0129] Optionally, before determining the target phoneme feature, the obtaining module 71 is further configured to perform word segmentation on the target Japanese text to obtain a plurality of segmented Japanese words; determine an initial phoneme feature of the target Japanese text; wherein the initial phoneme feature includes initial phoneme sub-features of the Japanese words determined based on a Japanese dictionary. The obtaining module 71 is configured to convert the character-granularity phrase pause extraction feature into a word-granularity phrase pause extraction feature; predict the word-granularity phrase pause extraction feature using a multi-phonetic word prediction network to obtain a multi-phonetic word prediction result of the target Japanese text; wherein the multi-phonetic word prediction result includes multi-phonetic word categories of the Japanese words; determine multi-phonetic words in the Japanese words based on the multi-phonetic word prediction result; determine target phoneme sub-features corresponding to each multi-phonetic word, respectively; and update the initial phoneme sub-features of the multi-phonetic words in the initial phoneme feature to the corresponding target phoneme sub-features, respectively, to obtain the target phoneme feature.

[0130] Optionally, the character-granularity phrase pause extraction feature includes phrase pause sub-extraction features corresponding to each character of the target Japanese text, respectively, and the obtaining module 71 is configured to obtain a plurality of phrase pause sub-extraction features corresponding to each Japanese word, respectively; perform average pooling processing on the plurality of phrase pause sub-extraction features of each Japanese word, respectively, to obtain the word-granularity phrase pause extraction feature; and / or the multi-phonetic word prediction result is used as input of a back-end speech synthesis model.

[0131] Optionally, the obtaining module 71 is configured to obtain any one of the following features as the target text feature: the character-granularity text feature, the character-granularity long pause extraction feature, the character-granularity phrase pause extraction feature, and the word-granularity phrase pause extraction feature obtained by converting the character-granularity phrase pause extraction feature.

[0132] Optionally, the fusion module 72 is configured to convert the target phrase pause feature, the target phoneme feature, and the target text feature into syllable-granularity features; perform feature extraction on the syllable-granularity target phrase pause feature and the syllable-granularity target phoneme feature using a third feature extraction sub-network of the syllable position prediction network to obtain a first extraction feature; and perform feature concatenation on the syllable-granularity target text feature and the first extraction feature to obtain a target fusion feature.

[0133] Optionally, the third feature extraction sub-network includes a first embedding sub-network, a second embedding sub-network, and a recurrent neural sub-network. The fusion module 72 is configured to process the syllable-granularity target phrase pause feature using the first embedding sub-network to obtain a first embedding feature of the syllable-granularity; encode the syllable-granularity target phoneme feature using the second embedding sub-network to obtain a second embedding feature of the syllable-granularity; perform feature concatenation on the first embedding feature and the second embedding feature; and perform feature extraction on the concatenated embedding feature using the recurrent neural sub-network to obtain the first extraction feature.

[0134] Optionally, the prediction module 73 is configured to perform feature extraction on the target fusion feature by using a fourth feature extraction subnetwork of the core position prediction network to obtain second extracted features; and perform prediction on the second extracted features by using a core position prediction subnetwork of the core position prediction network to obtain a core position prediction result.

[0135] Optionally, the target prediction model comprises a pre-trained text processing network, a prosody prediction network, a multi-phoneme prediction network, and a pitch location prediction network. The pitch location prediction device 70 further comprises a training module 74. The training module 74 is configured to sequentially perform first training, second training, third training, fourth training, and fifth training on the target prediction model. The first training comprises the following steps: obtaining first sample long pause prediction features and first sample phrase pause prediction features of a first sample Japanese text by using the pre-trained text processing network and the prosody prediction network; and adjusting network parameters of the pre-trained text processing network and the prosody prediction network based on the first sample long pause prediction features and the first sample phrase pause prediction features. The second training comprises the following steps: determining first sample phoneme features of a second sample Japanese text based on sample multi-phoneme annotation results of the second sample Japanese text; obtaining first sample pitch location prediction results of the second sample Japanese text by using the pitch location prediction network based on sample long pause annotation features, sample phrase pause annotation features, and the first sample phoneme features of the second sample Japanese text; and adjusting network parameters of the pitch location prediction network based on the first sample pitch location prediction results. The third training comprises the following steps: obtaining second sample long pause prediction features and second sample phrase pause prediction features of a third sample Japanese text by using the pre-trained text processing network and the prosody prediction network; determining second sample phoneme features of the third sample Japanese text based on sample multi-phoneme annotation results of the third sample Japanese text; obtaining second sample pitch location prediction results of the third sample Japanese text by using the pitch location prediction network based on the second sample long pause prediction features, the second sample phrase pause prediction features, and the second sample phoneme features; and adjusting network parameters of the pre-trained text processing network, the prosody prediction network, and the pitch location prediction network based on the second sample long pause prediction features, the second sample phrase pause prediction features, and the second sample pitch location prediction results. The fourth training comprises the following steps: obtaining third sample long pause prediction features and third sample phrase pause prediction features of a fourth sample Japanese text by using the pre-trained text processing network and the prosody prediction network; obtaining first sample multi-phoneme prediction results of the fourth sample Japanese text by using the multi-phoneme prediction network based on the third sample phrase pause prediction features; determining third sample phoneme features of the fourth sample Japanese text based on sample multi-phoneme annotation results of the fourth sample Japanese text; obtaining third sample pitch location prediction results of the fourth sample Japanese text by using the pitch location prediction network based on the third sample long pause prediction features, the third sample phrase pause prediction features, and the third sample phoneme features; and adjusting overall network parameters of the target prediction model based on the third sample long pause prediction features, the third sample phrase pause prediction features, the first sample multi-phoneme prediction results, and the third sample pitch location prediction results. The fifth training comprises the following steps: obtaining fourth sample long pause prediction features and fourth sample phrase pause prediction features of a fifth sample Japanese text by using the pre-trained text processing network and the prosody prediction network;The multi-phonetic word prediction network predicts a second sample multi-phonetic word prediction result of the fifth sample Japanese text based on the fourth sample phrase pause prediction feature; the second sample multi-phonetic word prediction result is used to determine a fourth sample phoneme feature of the fifth sample Japanese text; the adjustment position prediction network predicts a fourth sample adjustment position prediction result of the fifth sample Japanese text based on the fourth sample long pause prediction feature, the fourth sample phrase pause prediction feature and the fourth sample phoneme feature; and the fourth sample long pause prediction feature, the fourth sample phrase pause prediction feature, the second sample multi-phonetic word prediction result and the fourth sample adjustment position prediction result are used to adjust the overall network parameters of the target prediction model.

[0136] It should be noted that the device of the embodiment can execute the steps in the above method, and the details of the related content will be described in the method part, which will not be repeated here.

[0137] Please refer to Figure 8 , Figure 8 is a frame schematic diagram of an embodiment of an electronic device provided by the present application. In the embodiment, the electronic device 80 comprises a memory 81 and a processor 82.

[0138] The processor 82 can also be referred to as a CPU (Central Processing Unit). The processor 82 can be an integrated circuit chip with processing capability. The processor 82 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 82 can also be any conventional processor 82 and the like.

[0139] The memory 81 in the electronic device 80 is used to store the program instructions required by the processor 82 to run.

[0140] The processor 82 is used to execute the program instructions to realize the adjustment position prediction method in the present application.

[0141] Please refer to Figure 9 , Figure 9is a framework schematic diagram of an embodiment of the computer readable storage medium provided in the present application. The computer readable storage medium 90 of the embodiment of the present application stores program instructions 91, which, when executed, implement the pitch location prediction method provided in the present application. Wherein the program instructions 91 can form a program file and be stored in the above computer readable storage medium 90 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a sub-network device, etc.) executes all or part of the steps of the method of each embodiment of the present application. And the aforementioned computer readable storage medium 90 includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes, or a terminal device such as a computer, a server, a mobile phone, and a tablet.

[0142] The above scheme, by fusing the target phrase pause feature, the target phoneme feature and the target text feature corresponding to the target Japanese text, obtains the target fusion feature, and uses the target fusion feature to predict the pitch location, to obtain the pitch location prediction result corresponding to the target Japanese text. Since the Japanese pitch location and the text and pronunciation within the phrase pause boundary are related, and the target phrase pause feature, the target phoneme feature and the target text feature can reflect the phrase pause information, the pronunciation information and the text semantic information of the target Japanese text respectively, therefore, using the target fusion feature to predict the pitch location can comprehensively integrate the phrase pause, the pronunciation and the text semantic information of the target Japanese text in three different dimensions, so that the pitch location prediction result obtained by prediction has higher accuracy.

[0143] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiment descriptions, and the specific implementation can refer to the descriptions of the above method embodiments. For the sake of brevity, they will not be described here.

[0144] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For the sake of brevity, they will not be described here.

[0145] In several embodiments provided in the present application, it should be understood that the disclosed methods, devices and systems can be implemented in other manners. For example, the division of the apparatus embodiments is merely an example, and the division of the modules or units can be different, for example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0146] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed to multiple sub-network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0147] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0148] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a sub-network device, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0149] The above description is merely an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. A nucleosome positioning prediction method, characterized by, The method comprises: obtaining target phrase pause features, target phoneme features and target text features corresponding to the target Japanese text, wherein the target phrase pause features, the target phoneme features and the target text features are determined by using word granularity text features corresponding to the target Japanese text; fusing the target phrase pause features, the target phoneme features and the target text features to obtain target fusion features; using the target fusion features to predict the position of the target Japanese text to obtain a prediction result of the position of the target Japanese text.

2. The method of claim 1, wherein, The method comprises: processing the target Japanese text by using a pre-trained text processing network to obtain word granularity text features corresponding to the target Japanese text; determining the target phrase pause features, the target phoneme features and the target text features based on the word granularity text features.

3. The method of claim 2, wherein, The step of determining the target phrase pause features is performed by a prosody prediction network, wherein the prosody prediction network comprises a first feature extraction subnetwork, a second feature extraction subnetwork and a phrase pause prediction subnetwork. The step of determining the target phrase pause features comprises: extracting features of the word granularity text features by using the first feature extraction subnetwork to obtain word granularity long pause extraction features; extracting features of the word granularity long pause extraction features by using the second feature extraction subnetwork to obtain word granularity phrase pause extraction features; predicting the word granularity phrase pause extraction features by using the phrase pause prediction subnetwork to obtain the target phrase pause features in word granularity.

4. The method of claim 3, wherein, The prosody prediction network further comprises a long pause prediction subnetwork, and the method further comprises: predicting the word granularity long pause extraction features by using the long pause prediction subnetwork to obtain target long pause features in word granularity; wherein the prediction result of the position, the target long pause features and the target phrase pause features are all used for inputting a backend speech synthesis model.

5. The method of claim 3, wherein, before determining the target phoneme features, the method further comprises: performing word segmentation processing on the target Japanese text to obtain a plurality of segmented Japanese words; determining initial phoneme features of the target Japanese text; wherein the initial phoneme features comprise initial phoneme sub-features of each Japanese word determined based on a Japanese dictionary; the step of determining the target phoneme features comprises: converting the word granularity phrase pause extraction features into word granularity phrase pause extraction features; predicting the word granularity phrase pause extraction features by using a multi-phonetic word prediction network to obtain a multi-phonetic word prediction result of the target Japanese text; wherein the multi-phonetic word prediction result comprises a multi-phonetic word category of each Japanese word; based on the multi-phonetic word prediction result, determining multi-phonetic words in the plurality of Japanese words; determining target phoneme sub-features corresponding to each multi-phonetic word respectively; The initial phoneme sub-feature of each multi-sound word in the initial phoneme feature is updated to the corresponding target phoneme sub-feature to obtain the target phoneme feature.

6. The method of claim 5, wherein, The word granularity phrase pause extraction feature includes phrase pause sub-extraction features corresponding to each character of the target Japanese text, and converting the word granularity phrase pause extraction feature into a word granularity phrase pause extraction feature includes: obtaining a plurality of phrase pause sub-extraction features corresponding to each Japanese word; performing average pooling processing on the plurality of phrase pause sub-extraction features of each Japanese word to obtain the word granularity phrase pause extraction feature; and / or, the multi-sound word prediction result is used to input a back-end speech synthesis model.

7. The method of claim 3, wherein, The step of determining the target text feature includes: obtaining any one of the following features as the target text feature: the word granularity text feature, the word granularity long pause extraction feature, the word granularity phrase pause extraction feature, and the word granularity phrase pause extraction feature obtained by converting the word granularity phrase pause extraction feature.

8. The method of claim 1, wherein, The step of fusing the target phrase pause feature, the target phoneme feature, and the target text feature to obtain a target fusion feature includes: converting the target phrase pause feature, the target phoneme feature, and the target text feature into syllable granularity features; performing feature extraction on the syllable granularity target phrase pause feature and the syllable granularity target phoneme feature using a third feature extraction sub-network of the syllable position prediction network to obtain a first extraction feature; performing feature concatenation on the syllable granularity target text feature and the first extraction feature to obtain the target fusion feature.

9. The method of claim 8, wherein, The third feature extraction sub-network includes a first embedding sub-network, a second embedding sub-network, and a recurrent neural sub-network; performing feature extraction on the syllable granularity target phrase pause feature and the syllable granularity target phoneme feature using the third feature extraction sub-network to obtain the first extraction feature includes: processing the syllable granularity target phrase pause feature using the first embedding sub-network to obtain a syllable granularity first embedding feature; encoding the syllable granularity target phoneme feature using the second embedding sub-network to obtain a syllable granularity second embedding feature; performing feature concatenation on the first embedding feature and the second embedding feature; performing feature extraction on the concatenated embedding feature using the recurrent neural sub-network to obtain the first extraction feature.

10. The method of claim 1, wherein, The step of performing syllable position prediction using the target fusion feature to obtain a syllable position prediction result corresponding to the target Japanese text includes: performing feature extraction on the target fusion feature using a fourth feature extraction sub-network of the syllable position prediction network to obtain a second extraction feature; performing prediction on the second extraction feature using a syllable position prediction sub-network of the syllable position prediction network to obtain the syllable position prediction result.

11. The method of claim 1, wherein, The step of obtaining the target fusion feature and the step of performing the adjustment position prediction by using the target fusion feature are performed by an adjustment position prediction network, and a training step of the adjustment position prediction network comprises: obtaining sample phrase pause features, sample phoneme features and sample text features of a first sample Japanese text; obtaining sample adjustment position prediction results of the first sample Japanese text based on the sample phrase pause features, the sample phoneme features and the sample text features by using the adjustment position prediction network; adjusting network parameters of the adjustment position prediction network based on the sample adjustment position prediction results.

12. A nucleosome positioning prediction apparatus characterized by comprising: The device comprises: an obtaining module configured to obtain target phrase pause features, target phoneme features and target text features corresponding to a target Japanese text, the target phrase pause features, the target phoneme features and the target text features being determined by using word granularity text features corresponding to the target Japanese text; a fusion module configured to fuse the target phrase pause features, the target phoneme features and the target text features to obtain target fusion features; a prediction module configured to perform adjustment position prediction by using the target fusion features to obtain adjustment position prediction results corresponding to the target Japanese text.

13. An electronic device, comprising: comprise a memory and a processor coupled to each other, the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the method of any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store program instructions executable by a processor to implement the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and storage medium

    CN112735373A

  • Construction method of information prediction module, information prediction method and related equipment

    CN114333760A