A method for predicting the prosody structure of Chinese and related devices

Through the method based on pre-training Chinese BERT and combined with a multi-head attention classifier, the Chinese pronunciation structure is predicted, which solves the problem of complex feature design and poor prediction effect in the existing technology, and achieves more efficient and universal pronunciation structure prediction.

CN115221273BActive Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210713393.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-06-27
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The existing Chinese rhythm prediction method requires professionals to carry out complex and detailed feature designs. The prediction effect is poor and lacks universality, making it difficult to migrate between different applicable scenarios.

Method used

The pre-trained Chinese BERT method is used to obtain the feature sequence of the input text, and classify the rhythmic pauses of each Chinese character in the input text through a preset multi-head attention classifier, and finally predict the Chinese rhythmic structure based on the output text and feature sequence.

Benefits of technology

This avoids complex text feature design processes, improves prediction accuracy and universality, simplifies migration in different scenarios, and saves feature design and data labeling costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115221273B_ABST
    Figure CN115221273B_ABST
Patent Text Reader

Abstract

The embodiments of this application belong to the field of artificial intelligence and are applied to the field of Chinese prosody structure prediction. It involves a method for predicting the Chinese prosody structure, including obtaining the input text; obtaining the feature sequence of the input text based on the pre-trained Chinese BERT; classifying the prosodic pauses of each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier to obtain the classified output text; predicting the Chinese prosody structure of the input text based on the output text, the feature sequence, and a preset prosody structure category. This application also provides a device, a computer device, and a storage medium for predicting the Chinese prosody structure. Compared with traditional methods, this application uses Chinese BERT to avoid the complex and delicate text feature design process, and is convenient to be migrated to other scenarios, saving the high feature design cost and prosody data annotation cost. At the same time, using the multi-head attention classifier can make more effective use of context information and has higher classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence and Chinese prosodic structure prediction, and in particular, to a method for predicting Chinese prosodic structure and related devices thereof. Background Art

[0002] The standard Chinese front-end module at least includes two major functions: prosodic structure prediction and pronunciation conversion. Prosodic structure prediction is mainly used to obtain context-related prosodic information of the synthesized text. Accurate prosodic structure prediction plays a key role in the rhythm and realism of the synthesized speech. According to phonetic knowledge, Chinese prosodic features have a hierarchical structure. Prosodic structure prediction mainly predicts three levels of structures: prosodic words, prosodic phrases, and intonation phrases.

[0003] Existing Chinese prosody prediction mainly uses methods such as syntactic tree rules, conditional random fields, and neural networks. These methods all require professionals to perform complex and delicate feature design for applicable scenarios, but the prediction effect is poor, and at the same time, they lack universality and are difficult to be migrated between different applicable scenarios. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a method, device, computer device, and storage medium for predicting Chinese prosodic structure, so as to solve the problem that in the existing technology for predicting Chinese prosodic structure, professionals are required to perform complex and delicate feature design for applicable scenarios, the prediction effect is poor, and at the same time, it lacks universality and is difficult to be migrated between different applicable scenarios.

[0005] To solve the above technical problems, the embodiments of this application provide a method for predicting Chinese prosodic structure, which adopts the following technical solutions:

[0006] A method for predicting Chinese prosodic structure includes the following steps:

[0007] Obtain the input text;

[0008] Based on the pre-trained Chinese BERT, obtain the feature sequence of the input text;

[0009] Based on the feature sequence and a preset multi-head attention classifier, perform prosodic pause classification on each Chinese character in the input text to obtain the classified output text;

[0010] Based on the output text, feature sequence, and preset prosodic structure categories, predict the Chinese prosodic structure of the input text.

[0011] Further, before the step of obtaining the feature sequence of the input text based on the pre-trained Chinese BERT, it further includes:

[0012] Identify the composition type of the input text;

[0013] If the input text only contains a single Chinese sentence, obtain the word vectors and position vectors of each Chinese character in the Chinese sentence, and use the word vectors and position vectors as the feature sequence;

[0014] If the input text contains multiple Chinese sentences, obtain the sentence vectors of each Chinese sentence, the word vectors and position vectors of each Chinese character in each Chinese sentence, and use the sentence vectors, word vectors and position vectors as the feature sequence.

[0015] Further, the preset multi-head attention classifier includes a multi-head attention layer and a one-dimensional convolutional layer. The step of performing prosodic pause classification on each Chinese character in the input text based on the feature sequence and the preset multi-head attention classifier specifically includes:

[0016] Previously, based on the prosodic pause classification categories, perform a first differential naming on different prosodic pause classifications;

[0017] Use the word vectors in the feature sequence as query values, use the word vectors of each Chinese character in the reference character sets corresponding to different prosodic pause classifications as key values and attribute values, and perform a first linear transformation on the query values, key values and attribute values respectively based on preset different parameters;

[0018] Obtain the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value after the first linear transformation;

[0019] Use the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value as parameters, and respectively obtain the feature vectors obtained by the word vectors of each Chinese character in the input text in different attention layers through the Attention(Q, K, V) function of the attention layer, where Q represents the linear sequence corresponding to the query value, K represents the linear sequence corresponding to the key value, and V represents the linear sequence corresponding to the attribute value;

[0020] Calculate the vector dot product between the feature vectors of each Chinese character in the input text and the word vectors of each reference character in different prosodic pause classifications through the multi-head attention layer, splice the vector dot products corresponding to the preset different prosodic pause classifications, and obtain a splicing result;

[0021] Perform a second linear transformation on the splicing result through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the vector dot product after splicing as the output linear sequence;

[0022] Based on the first differential naming and each vector dot product in the output linear sequence, identify the prosodic pause classification corresponding to each Chinese character in the input text.

[0023] Further, before the step of performing the first differential naming on the different prosodic pause classifications, the following steps are also included:

[0024] Based on the Chinese BERT, obtain the word vectors of each Chinese character and the corresponding Chinese characters in a batch of Chinese sentences for which the Chinese prosodic structure prediction has been completed.

[0025] Use each Chinese character in the batch of Chinese sentences as a data source, and pre-classify the data source based on prosodic pause classification to obtain a reference character set corresponding to different prosodic pause classifications. Among them, when performing the pre-classification, at the same time, use each Chinese character as a label name and its corresponding word vector as an attribute value to construct a key-value pair.

[0026] Further, after the one-dimensional convolutional layer performs the second linear transformation step on the splicing result, the following steps are also included:

[0027] Based on the sigmoid activation function, perform numerical compression on the dot product of each vector in the output linear sequence, and compress it to a value range between [0, 1].

[0028] Further, the step of identifying the prosodic pause classification corresponding to each Chinese character in the input text based on the first differential naming and the dot product of each vector in the output linear sequence specifically includes:

[0029] Step A: Obtain multiple dot products corresponding to a single Chinese character in the input text in the output linear sequence, and select the maximum dot product among the multiple dot products by comparison.

[0030] Step B: Identify the reference character set corresponding to the Chinese character based on the maximum dot product, and identify the prosodic pause classification corresponding to the Chinese character through the reference character set.

[0031] Step C: Repeat the above Step A and Step B to identify the prosodic pause classification corresponding to each Chinese character in the input text.

[0032] Step D: Use the first differential naming of each Chinese character in the input text and its corresponding prosodic pause classification as the output text.

[0033] Further, after the step of classifying each Chinese character in the input text based on the feature sequence and the preset multi-head attention classifier, the following steps are also included:

[0034] Based on the cross-entropy loss function, judge the difference degree between the classification result of each Chinese character in the input text and the pre-classification result of the corresponding Chinese character in the data source.

[0035] If the difference degree does not meet the preset difference degree threshold, the configuration parameters of the classifier and the pre-trained Chinese BERT are updated until the difference degree meets the preset difference degree threshold, and then the model update is completed.

[0036] Further, after the step of obtaining the classified output text, it further includes:

[0037] If the input text is a single Chinese sentence, based on the position vectors of each Chinese character in the feature sequence, the Chinese characters in the output text are reversely spliced to generate a spliced text;

[0038] If the input text contains multiple Chinese sentences, based on the sentence vectors of each Chinese sentence in the feature sequence and the position vectors of each Chinese character in each Chinese sentence, the Chinese characters in the output text are reversely spliced to generate a spliced text.

[0039] Further, the predicting the Chinese prosodic structure of the input text based on the output text and the preset prosodic structure category specifically includes:

[0040] Pre-set second distinct names for different Chinese prosodic structures;

[0041] According to the association relationship between the preset Chinese prosodic structure and the prosodic pause classification category, the second distinct name is associated with the first distinct name;

[0042] Based on the first distinct name, position vector corresponding to each Chinese character in the output text, and the association relationship between the second distinct name and the first distinct name, identify the Chinese prosodic structure of the spliced text, that is, the Chinese prosodic structure of the input text.

[0043] To solve the above technical problems, an embodiment of the present application further provides a device for predicting the Chinese prosodic structure, and adopts the following technical solutions:

[0044] A device for predicting the Chinese prosodic structure includes:

[0045] An input text acquisition module, configured to acquire an input text;

[0046] A feature sequence acquisition module, configured to acquire a feature sequence of the input text based on a pre-trained Chinese BERT;

[0047] A prosodic pause classification module, configured to perform prosodic pause classification on each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier, and obtain a classified output text;

[0048] A Chinese prosodic structure prediction module, configured to predict the Chinese prosodic structure of the input text based on the output text, the feature sequence, and a preset prosodic structure category.

[0049] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0050] In the method for predicting the Chinese prosodic structure according to the embodiments of the present application, by obtaining the input text, based on the pre-trained Chinese BERT to obtain the feature sequence of the input text, based on the feature sequence and the preset multi-head attention classifier to classify the prosodic pauses of each Chinese character in the input text, obtaining the output text after classification, based on the output text, the feature sequence and the preset prosodic structure categories to predict the Chinese prosodic structure of the input text, using the Chinese pre-trained BERT to build a Chinese prosodic structure prediction model, avoiding the complex and delicate text feature design process, and being convenient to be migrated to other scenarios, saving the high feature design cost and prosodic data annotation cost; at the same time, using the multi-head attention classifier can make more effective use of the context information, and the classification accuracy is higher. The embodiments of the present application also monitor the accuracy of the Chinese prosodic structure prediction by introducing the cross-entropy loss function, so as to facilitate the tester to optimize the parameters of the pre-trained Chinese BERT and the classifier in the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0053] Figure 2 A flowchart of an embodiment of the method for predicting the Chinese prosodic structure according to the present application;

[0054] Figure 3 is Figure 2 a flowchart of a specific implementation manner of step 203 in;

[0055] Figure 4 is a schematic diagram of the one-to-one correspondence between the multi-head attention layer in the classifier and the prosodic pause classification categories;

[0056] Figure 5 is a schematic structural diagram of an embodiment of the device for predicting the Chinese prosodic structure according to the present application;

[0057] Figure 6 is Figure 5 a schematic structural diagram of a specific implementation manner of the prosodic pause classification module shown;

[0058] Figure 7 It is a schematic structural diagram of an embodiment of a computer device according to the present application. Detailed implementation manners

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0060] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0061] In order to enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings.

[0062] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0063] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0064] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0065] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.

[0066] It should be noted that the method for predicting the Chinese prosodic structure provided in the embodiments of the present application is generally executed by a server / terminal device. Correspondingly, the device for predicting the Chinese prosodic structure is generally set in the server / terminal device.

[0067] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0068] Continuing to refer to Figure 2 , a flowchart of an embodiment of the method for predicting the Chinese prosodic structure according to the present application is shown. The method for predicting the Chinese prosodic structure includes the following steps:

[0069] Continuing to refer to Figure 2 , a flowchart of an embodiment of the method for predicting the Chinese prosodic structure according to the present application is shown. The method for predicting the Chinese prosodic structure includes the following steps:

[0070] Step 201, obtain the input text.

[0071] In this embodiment, the electronic device on which the method for predicting the Chinese prosodic structure runs (such as Figure 1 shown in Server / Terminal Device)The input text can be obtained through a wired connection or a wireless connection. A request for Chinese prosody structure prediction is received, and the input text is passed into a preset Chinese prosody structure prediction model for Chinese prosody structure prediction. It should be noted that the above wireless connection methods can include but are not limited to 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0072] Step 202: Obtain the feature sequence of the input text based on the pre-trained Chinese BERT.

[0073] In this embodiment, obtaining the feature sequence of the input text based on the pre-trained Chinese BERT includes the steps of: presetting configuration parameters for the Chinese BERT to perform Chinese prosody structure recognition, and completing the parameter configuration of the Chinese BERT based on the configuration parameters; after completing the parameter configuration of the Chinese BERT, using the input text obtained in step 201 as the text for Chinese prosody structure recognition, and obtaining its feature sequence.

[0074] In some optional implementation manners of this embodiment, before step 202 obtains the feature sequence of the input text based on the pre-trained Chinese BERT, the above electronic device may further perform the following steps: identify the composition type of the input text; if the input text only contains a single Chinese sentence, obtain the word vectors and position vectors of each Chinese character in the Chinese sentence, and use the word vectors and position vectors as the feature sequence; if the input text contains multiple Chinese sentences, obtain the sentence vectors of each Chinese sentence in the input text, the word vectors and position vectors of each Chinese character in each Chinese sentence, and use the sentence vectors, word vectors, and position vectors as the feature sequence.

[0075] Step 203: Classify the prosody pauses of each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier, and obtain the classified output text.

[0076] In some optional implementation manners of this embodiment, after step 203 classifies the prosody pauses of each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier, the above electronic device may further perform the following steps: based on the cross-entropy loss function, judge the difference degree between the classification results of each Chinese character in the input text and the pre-classification results of the corresponding Chinese characters in the data source; if the difference degree does not meet the preset difference degree threshold, update the configuration parameters of the classifier and the pre-trained Chinese BERT until the difference degree meets the preset difference degree threshold, and then the model update is completed.

[0077] cross-entropy loss function y i represents the expected classification result of each Chinese prosodic structure in the input text, and x i represents the prediction accuracy of each Chinese prosodic structure in the input text. i represents the total number of structures that can be represented by Chinese prosodic structures in the input text. For example, if an input text can be divided into 3 prosodic words, 4 prosodic phrases, and 5 intonation phrases, then i = 12,

[0078] Obtain the corresponding loss functions when the model predicts each prosodic word, prosodic phrase, and intonation phrase respectively Then sum the values of the i loss functions obtained, which is the loss function of the input text. This application monitors the accuracy of Chinese prosodic structure prediction through the cross-entropy loss function, which is convenient for guiding testers to optimize the parameters of the prediction model and also ensures the prediction accuracy of the prediction model.

[0079] In this embodiment, using the cross-entropy loss function is both a calibration of the classifier and a calibration of the pre-trained BERT. Chinese BERT itself can perform part-of-speech classification on Chinese content. The part-of-speech classification generally includes 12 categories such as nouns, verbs, adjectives, pronouns, numeral-classifiers, auxiliary words, adverbs, prepositions, etc. And this application classifies the input text through a classifier, setting the categories to four. In essence, it is two different classification sub-models that jointly construct a model. When classifying the input text, it is necessary to continuously adjust the parameters of the two classification sub-models to obtain the final model with appropriate classification.

[0080] In some optional implementation manners of this embodiment, after step 203 obtains the classified output text, the above electronic device may further perform the following steps: If the input text is a single Chinese sentence, then based on the position vectors of each Chinese character in the feature sequence, perform inverse splicing on each Chinese character in the output text to generate a spliced text; if the input text includes multiple Chinese sentences, then based on the sentence vectors of each Chinese sentence in the feature sequence and the position vectors of each Chinese character in each Chinese sentence, perform inverse splicing on each Chinese character in the output text to generate a spliced text.

[0081] In this embodiment, the preset multi-head attention classifier includes a multi-head attention layer and a one-dimensional convolutional layer.

[0082] Continue to refer to Figure 3 , which shows the steps of classifying each Chinese character in the input text based on the preset multi-head attention classifier, specifically including:

[0083] Step 301: Based on the prosodic pause classification categories in advance, perform a first differential naming for different prosodic pause classifications. Among them, the prosodic pause classification categories include: no prosodic pause category, prosodic word category, prosodic phrase category, and intonation phrase category. Use I to represent the no prosodic pause category, B1 to represent the prosodic word category, B2 to represent the prosodic phrase category, and B3 to represent the intonation phrase category.

[0084] In some alternative implementation manners of this embodiment, before performing the first differential naming for different prosodic pause classifications, the above electronic device may further perform the following steps: Based on the Chinese BERT, obtain the word vectors of each Chinese character and each Chinese character in a batch of Chinese sentences for which Chinese prosodic structure prediction has been completed; Use each Chinese character in the batch of Chinese sentences as a data source, and perform pre-classification on the data source based on prosodic pause classification to obtain a reference character set corresponding to different prosodic pause classifications. Among them, when performing the pre-classification, at the same time, use each Chinese character as a label name and its corresponding word vector as an attribute value to construct a key-value pair.

[0085] Step 302: Use the word vectors in the feature sequence as query values, use the word vectors of each Chinese character in the reference character sets corresponding to different prosodic pause classifications as key values and attribute values, and perform a first linear transformation on the query values, key values, and attribute values respectively based on different preset parameters.

[0086] Step 303: Obtain the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value after the first linear transformation.

[0087] Step 304: Use the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value as parameters, and respectively obtain the feature vectors obtained by the word vectors of each Chinese character in the input text at different attention layers through the Attention(Q, K, V) function of the attention layer, where Q represents the linear sequence corresponding to the query value, K represents the linear sequence corresponding to the key value, and V represents the linear sequence corresponding to the attribute value.

[0088] Step 305: Calculate the vector dot product between the feature vectors of each Chinese character in the input text and the word vectors of each reference character in different prosodic pause classifications through the multi-head attention layer, splice the vector dot products corresponding to different preset prosodic pause classifications, and obtain a splicing result.

[0089] Step 306: Perform a second linear transformation on the splicing result through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the vector dot product after splicing as the output linear sequence.

[0090] In some alternative implementation manners of this embodiment, after the one-dimensional convolutional layer performs a second linear transformation on the splicing result, the electronic device may further perform the following steps: numerically compress the dot product of each vector in the output linear sequence based on the sigmoid activation function, and compress it to a value range between [0, 1].

[0091] Step 307, based on the first differential naming and the dot product of each vector in the output linear sequence, identify the prosodic pause classification corresponding to each Chinese character in the input text.

[0092] In this embodiment, the step of identifying the prosodic pause classification corresponding to each Chinese character in the input text based on the first differential naming and the dot product of each vector in the output linear sequence specifically includes: Step A: Obtain multiple dot products corresponding to a single Chinese character in the output linear sequence of the input text, and select the maximum dot product among the multiple dot products by comparison; Step B: Identify the reference character set corresponding to the Chinese character based on the maximum dot product, and identify the prosodic pause classification corresponding to the Chinese character through the reference character set; Step C: Repeat Steps A and B to identify the prosodic pause classification corresponding to each Chinese character in the input text; Step D: Use the first differential naming of each Chinese character in the input text and its corresponding prosodic pause classification as the output text.

[0093] In this application, the multi-head attention separator includes a multi-head attention layer and a one-dimensional convolutional layer. First, different data sources are set for the multi-head attention layer, that is, the data sources are pre-classified based on different prosodic pause classifications, and reference character sets corresponding to different prosodic pause classifications are obtained. The prosodic pause classification categories are divided into non-prosodic pause class, prosodic word class, prosodic phrase class, and intonation phrase class. That is, the attention layer corresponding to the multi-head attention separator has four layers, and different attention layers correspond to different prosodic pause classifications. After obtaining the word vectors of each Chinese character corresponding to the input text, use them as query values, and obtain the word vectors of each Chinese character in the reference character sets corresponding to different attention layers as key values and attribute values. The query values, key values, and attribute values are all full vector matrices. That is, the query value is a matrix constructed by the word vectors of all Chinese characters in the input text, and the key value and attribute value are matrices constructed by the word vectors of all Chinese characters in the reference character sets corresponding to different prosodic pause classifications. Therefore, the above four-head attention layer has one same matrix, that is, the query value, and has two different matrices among each other, that is, the key value and the attribute value. The key value and the attribute value of the same attention layer are the same.

[0094] When the multi-head attention classifier performs classification operations, each head attention layer generates three matrices that are multiplied by the query value, key value, and attribute value respectively. Assume them to be the first matrix, the second matrix, and the third matrix. At this time, each head attention layer has six matrices, namely the query value, key value, attribute value, the first matrix, the second matrix, and the third matrix. Multiply the query value, key value, and attribute value with the first matrix, the second matrix, and the third matrix respectively to obtain the dot product of vectors, and then add them up. Perform a first linear transformation on the added result. After performing the first linear transformation on the operation results of the four head attention layers, splice the operation results corresponding to the four head attention layers in a one-dimensional convolutional layer to obtain a splicing result. Then, perform a second linear transformation to finally obtain a linear sequence.

[0095] Continue to refer to Figure 4 , which shows a schematic diagram of the one-to-one correspondence between the multi-head attention layer in the classifier and the prosodic pause classification categories, including: classifying the prosodic pause classification categories into non-prosodic pause category, prosodic word category, prosodic phrase category, and intonation phrase category. That is, the attention layer corresponding to the multi-head attention separator is four layers. Different attention layers correspond to different prosodic pause classifications respectively. That is, the number of types of the attention layer of the multi-head attention classifier is the same as that of the prosodic pause classification categories.

[0096] Step 204, predict the Chinese prosodic structure of the input text based on the output text, feature sequence, and preset prosodic structure categories.

[0097] In this embodiment, the predicting the Chinese prosodic structure of the input text based on the output text and preset prosodic structure categories specifically includes: pre-setting second distinct names for different Chinese prosodic structures. Among them, different Chinese prosodic structures include prosodic word structure, prosodic phrase structure, and intonation phrase structure. Use #1 to represent the prosodic word structure, #2 to represent the prosodic phrase structure, and #3 to represent the intonation phrase structure; according to the association relationship between the preset Chinese prosodic structure and prosodic pause classification categories, associate the second distinct name with the first distinct name. The specific association is: pre-set that the prosodic word structure is associated with the prosodic word category, the prosodic phrase structure is associated with the prosodic phrase category, and the intonation phrase structure is associated with the intonation phrase category. That is, B1 is associated with #1, B2 is associated with #2, and B3 is associated with #3; based on the first distinct name corresponding to each Chinese character in the output text, position vector, and the association relationship between the second distinct name and the first distinct name, identify the Chinese prosodic structure of the spliced text, that is, the Chinese prosodic structure of the input text.

[0098] Based on the association relationship between the second differential naming and the first differential naming, determining the Chinese prosody structure corresponding to the output text. The specific implementation method is as follows: Assume that the first differential naming is set as I, B1, B2, B3, and the second differential naming is set as #1, #2, #3. It is preset that B1 corresponds to #1, B2 corresponds to #2, and B3 corresponds to #3. Among them, I corresponds to the prosody pause classification, B1 corresponds to the prosody word classification, B2 corresponds to the prosody phrase classification, B3 corresponds to the intonation phrase classification, #1 corresponds to the prosody word structure, #2 corresponds to the prosody phrase structure, and #3 corresponds to the intonation phrase structure. Identify the Chinese characters in the output text whose prosody classification is B1, B2, or B3, and split the output text with the position vectors corresponding to these Chinese characters in the output text, splitting it into Chinese words, Chinese phrases, and Chinese intonation phrases. Then, based on the association relationship between the second differential naming and the first differential naming, determine the prosody structures corresponding to the split Chinese words, Chinese phrases, and Chinese intonation phrases.

[0099] This application integrates Chinese BERT and a multi-head attention classifier into a preset Chinese prosody structure prediction model. The feature sequence of the input text is obtained through the pre-trained Chinese BERT, and the Chinese prosody structure prediction model is modeled using the pre-trained Chinese BERT, avoiding the complex and delicate text feature design process, and being convenient to be migrated to other scenarios, saving the high feature design cost and prosody data annotation cost. Then, the prosody pauses of each Chinese character in the input text are classified through the preset multi-head attention classifier. This classifier includes a multi-head attention layer and a one-dimensional convolutional layer. The dot product of vectors between different word vectors in the query value and the word vectors of each reference character in different prosody pause classifications is calculated respectively through the multi-head attention layer, and the dot products corresponding to the preset different prosody pause classifications are concatenated to obtain a concatenation result. Then, the concatenation result is linearly transformed a second time through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the dot product of vectors after concatenation. Based on the first differential naming and the dot products of vectors in the linear sequence, different prosody pause classifications corresponding to each Chinese character in the input text are identified. Using the multi-head attention classifier can make more effective use of context information and has higher classification accuracy. Finally, the Chinese prosody structure of the input text is predicted through the feature sequence, the prosody pause classification result, and the preset prosody structure category.

[0100] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0101] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning. This application can be applied to the field of Chinese prosodic structure prediction, so as to ensure the rhythm and authenticity of the synthesized speech when the standard Chinese front-end module performs speech synthesis.

[0102] This application belongs to the field of Chinese prosodic structure prediction. Through this solution, it is possible to ensure the rhythm and authenticity of the synthesized speech when the standard Chinese front-end module performs speech synthesis.

[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When this program is executed, it can include the processes of the embodiments of the above various methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0104] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments, and their execution order does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0105] Further reference Figure 5 to Figure 2 As an implementation of the method shown above, this application provides an embodiment of a device for predicting the Chinese prosodic structure. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.

[0106] Such as Figure 5As shown in the figure, the apparatus 500 for predicting the Chinese prosodic structure according to this embodiment includes: an input text acquisition module 501, a feature sequence acquisition module 502, a prosodic pause classification module 503, and a Chinese prosodic structure prediction module 504. Among them:

[0107] The input text acquisition module 501 is configured to acquire the input text;

[0108] The feature sequence acquisition module 502 is configured to obtain the feature sequence of the input text based on the pre-trained Chinese BERT;

[0109] The prosodic pause classification module 503 is configured to perform prosodic pause classification on each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier, and obtain the classified output text;

[0110] The Chinese prosodic structure prediction module 504 is configured to predict the Chinese prosodic structure of the input text based on the output text, the feature sequence, and a preset prosodic structure category.

[0111] In some optional implementation manners of this embodiment, the apparatus for predicting the Chinese prosodic structure further includes a text composition type recognition module, configured to recognize the composition type of the input text. If the input text only contains a single Chinese sentence, obtain the word vectors and position vectors of each Chinese character in the Chinese sentence, and use the word vectors and position vectors as the feature sequence. If the input text contains multiple Chinese sentences, obtain the sentence vectors of each Chinese sentence, the word vectors and position vectors of each Chinese character in each Chinese sentence, and use the sentence vectors, word vectors, and position vectors as the feature sequence.

[0112] In some optional implementation manners of this embodiment, the apparatus for predicting the Chinese prosodic structure further includes a parameter update judgment module, configured to judge the difference degree between the classification result of each Chinese character in the input text and the pre-classification result of the corresponding Chinese character in the data source based on the cross-entropy loss function; if the difference degree does not meet the preset difference degree threshold, update the configuration parameters of the classifier and the pre-trained Chinese BERT until the difference degree meets the preset difference degree threshold, and then the model update is completed.

[0113] Refer to Figure 6 , which is a schematic structural diagram of a specific implementation manner of the prosodic pause classification module. The prosodic pause classification module includes a first difference naming sub-module 5031, a first linear transformation sub-module 5032, a classifier parameter acquisition sub-module 5033, a feature vector acquisition sub-module 5034, a vector dot product splicing sub-module 5035, a second linear transformation sub-module 5036, and a prosodic pause classification and recognition sub-module 5037,

[0114] Among them, the first differential naming sub-module 5031 is used to perform first differential naming on different prosodic pause classifications in advance based on the prosodic pause classification categories, where the prosodic pause classification categories include: no prosodic pause category, prosodic word category, prosodic phrase category, and intonation phrase category. Use I to represent the no prosodic pause category, B1 to represent the prosodic word category, B2 to represent the prosodic phrase category, and B3 to represent the intonation phrase category;

[0115] The first linear transformation sub-module 5032 is used to use the word vectors in the feature sequence as query values, and use the word vectors of each Chinese character in the reference character sets corresponding to different prosodic pause classifications as key values and attribute values, and perform first linear transformation on the query values, key values, and attribute values respectively based on preset different parameters;

[0116] The classifier parameter acquisition sub-module 5033 is used to acquire the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value after the first linear transformation;

[0117] The feature vector acquisition sub-module 5034 is used to use the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value as parameters, and respectively acquire the feature vectors obtained by the word vectors of each Chinese character in the input text at different attention layers through the Attention(Q, K, V) function of the attention layer, where Q represents the linear sequence corresponding to the query value, K represents the linear sequence corresponding to the key value, and V represents the linear sequence corresponding to the attribute value;

[0118] The vector dot product splicing sub-module 5035 is used to calculate the vector dot product between the feature vectors of each Chinese character in the input text and the word vectors of each reference character in different prosodic pause classifications respectively through the multi-head attention layer, splice the vector dot products corresponding to the preset different prosodic pause classifications, and obtain a splicing result;

[0119] The second linear transformation sub-module 5036 is used to perform a second linear transformation on the splicing result through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the vector dot product after splicing as the output linear sequence;

[0120] The prosodic pause classification recognition sub-module 5037 is used to recognize the prosodic pause classification corresponding to each Chinese character in the input text based on the first differential naming and each vector dot product in the output linear sequence.

[0121] To solve the above technical problems, the embodiments of the present application also provide a computer device. For details, please refer to Figure 7 , Figure 7 This is the basic structural block diagram of the computer device in this embodiment.

[0122] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 7 with components 71-73 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0123] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0124] The memory 71 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 71 can be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 can also be an external storage device of the computer device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 7. Of course, the memory 71 can also include both the internal storage unit of the computer device 7 and its external storage device. In this embodiment, the memory 71 is generally used to store the operating system installed on the computer device 7 and various application software, such as computer-readable instructions for the method of predicting the Chinese prosody structure. In addition, the memory 71 can also be used to temporarily store various data that have been output or will be output.

[0125] In some embodiments, the processor 72 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to run the computer-readable instructions stored in the memory 71 or process data, such as running the computer-readable instructions of the method for predicting the Chinese prosody structure.

[0126] The network interface 73 may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device 7 and other electronic devices.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0128] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

Claims

1. A method for predicting the prosodic structure of Chinese, characterized in that, including the following steps: Obtain the input text; Obtain the feature sequence of the input text based on the pre-trained Chinese BERT; Classify the prosodic pauses of each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier to obtain the classified output text. Wherein, the preset multi-head attention classifier includes a multi-head attention layer and a one-dimensional convolutional layer. The step of classifying the prosodic pauses of each Chinese character in the input text based on the feature sequence and the preset multi-head attention classifier specifically includes: Previously, based on the prosodic pause classification categories, give different prosodic pause classifications a first distinct naming; Take the word vectors in the feature sequence as query values, take the word vectors of each Chinese character in the corresponding reference character sets of different prosodic pause classifications as key values and attribute values, and perform a first linear transformation on the query values, key values, and attribute values respectively based on preset different parameters; Obtain the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value after the first linear transformation; Taking the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value as parameters, through the function, respectively obtain the feature vectors obtained by the word vectors of each Chinese character in the input text at different attention layers, where Q represents the linear sequence corresponding to the query value, K represents the linear sequence corresponding to the key value, and V represents the linear sequence corresponding to the attribute value; Calculate the dot product of the vector between the feature vector of each Chinese character in the input text and the word vector of each reference character in different prosodic pause classifications through the multi-head attention layer, splice the dot products corresponding to the preset different prosodic pause classifications, and obtain the splicing result; Perform a second linear transformation on the splicing result through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the dot product of the vectors after splicing, as the output linear sequence; Based on the first distinct naming and each dot product in the output linear sequence, identify the prosodic pause classification corresponding to each Chinese character in the input text; Predict the Chinese prosodic structure of the input text based on the output text, the feature sequence, and a preset prosodic structure category. Wherein, the step of predicting the Chinese prosodic structure of the input text based on the output text, the feature sequence, and the preset prosodic structure category specifically includes: Previously, set a second distinct naming for different Chinese prosodic structures; According to the association relationship between the preset Chinese prosodic structure and the prosodic pause classification category, associate the second distinct naming with the first distinct naming; Based on the first distinct naming corresponding to each Chinese character in the output text, the position vector, and the association relationship between the second distinct naming and the first distinct naming, identify the Chinese prosodic structure of the target spliced text as the Chinese prosodic structure of the input text. Wherein, the position vector corresponding to each Chinese character in the output text is determined based on the feature sequence, and the target spliced text is generated by reverse splicing each Chinese character in the input text.

2. The method for predicting the Chinese prosodic structure according to claim 1, characterized in that Before the step of obtaining the feature sequence of the input text based on the pre-trained Chinese BERT, it further includes: Identify the composition type of the input text; If the input text only contains a single Chinese sentence, obtain the word vectors and position vectors of each Chinese character in this Chinese sentence, and take the word vectors and position vectors as the feature sequence; If the input text contains multiple Chinese sentences, obtain the sentence vectors of each Chinese sentence in this input text, the word vectors and position vectors of each Chinese character in each Chinese sentence, and take the sentence vectors, word vectors, and position vectors as the feature sequence.

3. The method for predicting the Chinese prosodic structure according to claim 1, wherein Before the first differential naming step for classifying different prosodic pauses, it further includes: Based on the Chinese BERT, obtaining word vectors of each Chinese character and the corresponding Chinese characters in a batch of Chinese sentences for which Chinese prosodic structure prediction has been completed; Using each Chinese character in the batch of Chinese sentences as a data source, pre-classifying the data source based on prosodic pause classification to obtain a reference character set corresponding to different prosodic pause classifications. Among them, when performing the pre-classification, each Chinese character is used as a label name at the same time, and its corresponding word vector is used as an attribute value to construct a key-value pair.

4. The method for predicting the Chinese prosodic structure according to claim 1, characterized in that After the one-dimensional convolutional layer performs the second linear transformation step on the splicing result, it further includes: Based on the sigmoid activation function, numerically compressing the dot product of each vector in the output linear sequence to compress it to a value range between [0, 1].

5. The method for predicting the Chinese prosodic structure according to claim 1, characterized in that, The step of identifying the prosodic pause classification corresponding to each Chinese character in the input text based on the first differential naming and the dot product of each vector in the output linear sequence specifically includes: Step A: Obtaining multiple dot products corresponding to a single Chinese character in the input text in the output linear sequence, and selecting the maximum dot product among the multiple dot products through comparison; Step B: Identifying the reference character set corresponding to the Chinese character based on the maximum dot product, and identifying the prosodic pause classification corresponding to the Chinese character through the reference character set; Step C: Repeatedly executing Step A and Step B to identify the prosodic pause classification corresponding to each Chinese character in the input text; Step D: Taking the first differential naming of each Chinese character in the input text and its corresponding prosodic pause classification as the output text.

6. The method for predicting the Chinese prosodic structure according to claim 3, wherein After the step of classifying the prosodic pauses of each Chinese character in the input text based on the feature sequence and the preset multi-head attention classifier, it further includes: Based on the cross-entropy loss function, judging the difference degree between the classification result of each Chinese character in the input text and the pre-classification result of the corresponding Chinese character in the data source; If the difference degree does not meet the preset difference degree threshold, then update the configuration parameters of the classifier and the pre-trained Chinese BERT until the difference degree meets the preset difference degree threshold, and the model update is completed.

7. The method for predicting the Chinese prosodic structure according to claim 1, characterized in that After the step of obtaining the classified output text, it further includes: If the input text is a single Chinese sentence, then based on the position vectors of each Chinese character in the feature sequence, anti-splicing each Chinese character in the output text to generate a spliced text; If the input text contains multiple Chinese sentences, then based on the sentence vectors of each Chinese sentence in the feature sequence and the position vectors of each Chinese character in each Chinese sentence, anti-splicing each Chinese character in the output text to generate a spliced text.

8. An apparatus for predicting the prosodic structure of Chinese, characterized in that The device for predicting the Chinese prosodic structure is used to implement the steps of the method for predicting the Chinese prosodic structure described in any one of claims 1 to 7. The device includes: An input text acquisition module for acquiring an input text; A feature sequence acquisition module for acquiring a feature sequence of the input text based on the pre-trained Chinese BERT; The prosodic pause classification module is used to classify the prosodic pauses of each Chinese character in the input text based on the feature sequence and a preset multi-head attention classifier, and obtain the classified output text. The prosodic pause classification module includes a first differential naming sub-module, a first linear transformation sub-module, a classifier parameter acquisition sub-module, a feature vector acquisition sub-module, a vector dot product splicing sub-module, a second linear transformation sub-module, and a prosodic pause classification recognition sub-module. Among them, the first differential naming sub-module is used to pre-perform a first differential naming for different prosodic pause classifications based on the prosodic pause classification categories. The first linear transformation sub-module is used to use the word vectors in the feature sequence as query values, and use the word vectors of each Chinese character in the reference character sets corresponding to different prosodic pause classifications as key values and attribute values, and perform a first linear transformation on the query values, key values, and attribute values respectively based on preset different parameters. The classifier parameter acquisition sub-module is used to obtain the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value after the first linear transformation. A feature vector acquisition sub-module, which is used to take the linear sequence corresponding to the query value, the linear sequence corresponding to the key value, and the linear sequence corresponding to the attribute value as parameters, and through the function, respectively obtain the feature vectors obtained by the word vectors of each Chinese character in the input text at different attention layers, where Q represents the linear sequence corresponding to the query value, K represents the linear sequence corresponding to the key value, and V represents the linear sequence corresponding to the attribute value; The vector dot product splicing sub-module is used to calculate the vector dot products between the feature vectors of each Chinese character in the input text and the word vectors of each reference character in different prosodic pause classifications through the multi-head attention layer, splice the vector dot products corresponding to the preset different prosodic pause classifications, and obtain the splicing result. The second linear transformation sub-module is used to perform a second linear transformation on the splicing result through the one-dimensional convolutional layer to obtain the linear sequence corresponding to the vector dot product after splicing as the output linear sequence. The prosodic pause classification recognition sub-module is used to identify the prosodic pause classification corresponding to each Chinese character in the input text based on the first differential naming and each vector dot product in the output linear sequence. The Chinese prosodic structure prediction module is used to predict the Chinese prosodic structure of the input text based on the output text, the feature sequence, and a preset prosodic structure category. The Chinese prosodic structure prediction module includes a second differential naming sub-module, a differential naming association sub-module, and a Chinese prosodic structure recognition sub-module. Among them, the second differential naming sub-module is used to pre-set a second differential naming for different Chinese prosodic structures. The differential naming association sub-module is used to associate the second differential naming with the first differential naming according to the association relationship between the preset Chinese prosodic structure and the prosodic pause classification category. The Chinese prosodic structure recognition sub-module is used to identify the Chinese prosodic structure of the target splicing text as the Chinese prosodic structure of the input text based on the first differential naming corresponding to each Chinese character in the output text, the position vector, and the association relationship between the second differential naming and the first differential naming. The position vector corresponding to each Chinese character in the output text is determined based on the feature sequence, and the target splicing text is generated by reverse splicing each Chinese character in the input text.

9. A computer device, characterized in that, It includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the method for predicting the Chinese prosodic structure as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the method for predicting the Chinese prosodic structure as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Chinese error correction method based on a pinyin encoding and decoding model

    CN109492202A

  • Chinese rhythm hierarchy prediction method and system based on self-attention

    CN111354333A