An information processing method, apparatus and electronic device
By generating a set of pronunciation texts of polyphonic characters and calculating their similarity to the text to be predicted, the problem of low prediction accuracy of polyphonic characters caused by data imbalance in the prior art is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202210268350.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-18
AI Technical Summary
In the prior art, the multiphonetic prediction model has low prediction accuracy for uncommon pronunciations due to data imbalance.
By obtaining the target words in the text to be predicted, a set of pronunciation text corresponding to at least two pronunciations is generated, and the pronunciation of the target words is determined based on the similarity between these sets of pronunciation texts and the text to be predicted.
It improves the prediction accuracy of uncommon pronunciations, reduces the impact of training corpus imbalance, and enhances the accuracy of pronunciation prediction of polyphonic characters.
Smart Images

Figure CN114742044B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and more specifically, to an information processing method, device, and electronic device. Background Art
[0002] The goal of Chinese polyphone pronunciation prediction is to correctly predict the pronunciation of polyphones in text. With the help of this information, text-to-speech (TTS) can synthesize more natural and human-like sounds. Polyphone pronunciation prediction is an important part of speech synthesis technology.
[0003] The existing technology uses BiLSTM (Bi-directional Long Short-Term Memory) networks and other methods to train context-dependent polyphone prediction models. Although this utilizes contextual information, the acquired polyphone data is often highly unbalanced, with more common pronunciation data and less uncommon pronunciation data. The BiLSTM network is sensitive to data imbalance, resulting in low accuracy in the prediction of uncommon pronunciations by the polyphone prediction model. Summary of the Invention
[0004] In view of this, the present application provides an information processing method as follows:
[0005] An information processing method, comprising:
[0006] Obtaining a target word in a text to be predicted, where the target word has at least two pronunciations;
[0007] Based on the target character and the at least two pronunciations, generating at least two sets of pronunciation text sets corresponding to the at least two pronunciations, wherein each set of pronunciation text sets corresponds to one pronunciation, and each set of pronunciation texts includes at least one pronunciation text;
[0008] The pronunciation of the target word in the text to be predicted is determined based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0009] Optionally, the above method, based at least on the target character and the at least two pronunciations, generates at least two sets of pronunciation texts corresponding to the at least two pronunciations, including:
[0010] Based on the text to be predicted, the target character and the at least two pronunciations, at least two sets of pronunciation texts corresponding to the at least two pronunciations are generated, and the contents of the pronunciation texts in the at least two sets of pronunciation texts are related to the contents of the text to be predicted.
[0011] Optionally, the above method generates at least two sets of pronunciation texts corresponding to at least two pronunciations based on the text to be predicted, the target character, and the at least two pronunciations, including:
[0012] generating a first vector based on the text to be predicted, wherein the first vector represents content contained in the text to be predicted;
[0013] Based on the first vector, the target word and the at least two pronunciations, at least two groups of pronunciation text sets corresponding to the at least two pronunciations are generated, each group of pronunciation text sets corresponds to one pronunciation, each group of pronunciation texts contains at least one pronunciation text, and the content of each pronunciation text is related to the content of the text to be predicted.
[0014] Optionally, the method described in any one of the above items, determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation texts and the text to be predicted, includes:
[0015] Generate a vector for each pronunciation text in each pronunciation text set, where the vector generated by each pronunciation text represents the content of the pronunciation text;
[0016] Processing vectors generated from pronunciation texts belonging to the same pronunciation text set to obtain a mean vector corresponding to the pronunciation text set;
[0017] Determining similarities between the pronunciation texts corresponding to the two pronunciation text sets and the text to be predicted based on mean vectors corresponding to at least two pronunciation text sets and a first vector generated by the text to be predicted;
[0018] Based on the similarity, a first pronunciation is determined as the pronunciation of the target word in the text to be predicted, and a pronunciation text in the pronunciation text set corresponding to the first pronunciation meets a similarity condition with the text to be predicted.
[0019] Optionally, in the above method, generating at least two sets of pronunciation texts corresponding to at least two pronunciations based on the first vector, the target character, and the at least two pronunciations includes:
[0020] The target generator processes the first vector, the target word, and the at least two pronunciations to generate the at least two groups of pronunciation text sets.
[0021] Optionally, the above method, before obtaining the target word in the text to be predicted, further includes:
[0022] The original learning model is trained based on a training text set, a target polyphone and at least two pronunciations of the target polyphone to obtain a target learning model, wherein the target learning model includes a target generator and a target discriminator, and the training text set includes the target polyphone and annotated pronunciations, wherein the annotated pronunciations are the annotated pronunciations of the target polyphone in the training text.
[0023] Optionally, in the above method, the training of the original learning model based on the training text, the target polyphonetic character, and at least two pronunciations of the target polyphonetic character to obtain the target learning model includes:
[0024] Obtaining a training text set, wherein the training text set includes at least two groups of training texts, each group of training texts corresponding to a marked pronunciation;
[0025] generating a training vector based on the training text, wherein the training vector represents content contained in the training text;
[0026] At least the training vector, the target polyphone, and the annotated pronunciation are input into an original generator as input conditions to obtain a target generator, so that the target generator generates a first text;
[0027] The first text and any training text are used as input content of the original discriminator, so that the original discriminator judges the true situation of the first text and the training text until the first agreed training stop condition is met, thereby obtaining a target learning model, which includes a target generator and a target discriminator.
[0028] Optionally, the method described above, wherein the training vector, the target polyphone character, and the annotated pronunciation are input as input conditions into an original generator to obtain a target generator, comprises:
[0029] Inputting the training vector, the target polyphonetic character and the annotated pronunciation as input conditions into an original generator to generate a second text;
[0030] Adding a false label to the second text, and returning the labeled second text as training text to execute the step of generating a training vector based on the training text until a second agreed training stop condition is met, thereby obtaining a target generator;
[0031] The training vector, the target polyphonetic character and the annotated pronunciation are input into a target generator as input conditions, so that the target generator generates a first text.
[0032] An information processing device, comprising:
[0033] An acquisition module, configured to acquire a target word in a text to be predicted, wherein the target word has at least two pronunciations;
[0034] A generating module, configured to generate at least two sets of pronunciation texts corresponding to the at least two pronunciations based on the target character and the at least two pronunciations, wherein each set of pronunciation texts corresponds to one pronunciation, and each set of pronunciation texts includes at least one pronunciation text;
[0035] The determination module is configured to determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0036] An electronic device comprising: a memory and a processor;
[0037] Wherein, the memory stores a processing program;
[0038] The processor is used to load and execute the processing program stored in the memory to implement each step of the information processing method as described in any one of the above items.
[0039] As can be seen from the above technical solution, this application provides an information processing method, in which the text to be predicted includes one or more polyphonetic characters, each of which has at least two pronunciations. The polyphonetic character is used as the target character in the text to be predicted. Based on the target character in the text to be predicted and its at least two pronunciations, at least two sets of pronunciation texts corresponding to the at least two pronunciations are generated, and the pronunciation of the target character in the text to be predicted is further determined based on the multiple pronunciation texts and the text to be predicted. In this solution, multiple pronunciation texts are generated based on the polyphone in the text to be predicted, so that the pronunciation corresponding to the polyphone in the predicted text can be further selected from the multiple pronunciation texts and the text to be predicted, without being affected by the imbalance of the training corpus, and the accuracy of predicting the pronunciation of the polyphone is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0041] Figure 1 This is a flowchart of Example 1 of an information processing method provided by this application;
[0042] Figure 2 This is a flowchart of Example 3 of an information processing method provided by this application;
[0043] Figure 3 This is a flowchart of Example 3 of an information processing method provided by this application;
[0044] Figure 4 It is a flowchart of Embodiment 4 of an information processing method provided by this application;
[0045] Figure 5 It is a flowchart of Embodiment 5 of an information processing method provided by this application;
[0046] Figure 6 It is a schematic diagram of a learning model in Embodiment 5 of an information processing method provided by this application;
[0047] Figure 7 It is a flowchart of Embodiment 6 of an information processing method provided by this application;
[0048] Figure 8 It is a flowchart of Embodiment 7 of an information processing method provided by this application;
[0049] Figure 9 It is a schematic structural diagram of an embodiment of an information processing device provided by this application. Specific embodiments
[0050] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0051] As Figure 1 shown, it is a flowchart of Embodiment 1 of an information processing method provided by this application. This method is applied to an electronic device, and this method includes the following steps:
[0052] Step S101: Obtain the target character in the text to be predicted, and the target character has at least two pronunciations;
[0053] Among them, the text to be predicted contains one or more characters, including at least one target character, and the target character has two or more pronunciations.
[0054] For example, the text to be predicted is "All the money is gone, but it will come back again", and the character "还" in it has two pronunciations, namely "huán" and "hái".
[0055] Among them, the multiple pronunciations of the target character in the text to be predicted are known, and in this solution, it is to determine which of the pronunciations of the target character in the text to be predicted is.
[0056] Step S102: Generate at least two sets of pronunciation text sets corresponding to the at least two pronunciations based on the target character and the at least two pronunciations;
[0057] Among them, each set of pronunciation text collections corresponds to one pronunciation, and each set of pronunciation text collections contains at least one pronunciation text.
[0058] Among them, based on the target character and multiple pronunciations of the target character, multiple sets of pronunciation text collections are generated, and each set of pronunciation text collections corresponds to one pronunciation.
[0059] For example, if the target character is "还" which has two pronunciations, two sets of text collections corresponding to these two pronunciations are generated respectively. For example, there are three pronunciation texts in the text collection corresponding to the pronunciation "huán", and each pronunciation text contains the character "还" with the pronunciation "huán". There are three pronunciation texts in the text collection corresponding to the pronunciation "hái", and each pronunciation text contains the character "还" with the pronunciation "hái".
[0060] Among them, the pronunciation text contains the target character, and the number of characters in the pronunciation text is random, which can be the same as or different from the number of characters in the text to be predicted; the position of the target character in the pronunciation text is random, which can be the same as or different from the position of the target character in the text to be predicted.
[0061] Step S103: Determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text collections and the text to be predicted.
[0062] Among them, determine the similarity between each set of pronunciation text collections and the text to be predicted respectively, and select the pronunciation corresponding to the text collection with the highest similarity as the pronunciation of the target character in the text to be predicted.
[0063] Specifically, each set of pronunciation text collections is regarded as a whole, calculate the similarity between this whole and the text to be predicted, select the pronunciation text collection corresponding to the maximum value among the calculated multiple similarities, and determine the pronunciation corresponding to this pronunciation text collection as the pronunciation of the target character in the text to be predicted.
[0064] It should be noted that the higher the similarity between the pronunciation text collection and the text to be predicted, the more similar the pronunciation texts in the pronunciation text collection are to the text to be predicted, and the greater the possibility that the pronunciations of the target characters in the two are the same. Therefore, in this application, the similarity between the pronunciation text collection and the text to be measured is used to determine the pronunciation of the target character, which is not affected by the imbalance of the training corpus, and the accuracy of predicting the pronunciation of polyphonic characters is higher.
[0065] In summary, this embodiment provides an information processing method, wherein the text to be predicted includes one or more polyphonetic characters, each of which has at least two pronunciations. The polyphonetic characters serve as target characters in the text to be predicted. Based on the target characters in the text to be predicted and their at least two pronunciations, at least two sets of pronunciation texts corresponding to the at least two pronunciations are generated, and the pronunciation of the target characters in the text to be predicted is further determined based on the multiple pronunciation texts and the text to be predicted. In this solution, multiple pronunciation texts are generated based on the polyphonetic characters in the text to be predicted, so that the pronunciation corresponding to the polyphonetic characters in the predicted text can be further selected from the multiple pronunciation texts and the text to be predicted. This method is not affected by the imbalance of the training corpus, and the accuracy of predicting the pronunciation of the polyphonetic characters is higher.
[0066] like Figure 2 The flowchart shown is a second embodiment of an information processing method provided by the present application, and the method includes the following steps:
[0067] Step S201: obtaining a target word in a text to be predicted, where the target word has at least two pronunciations;
[0068] Among them, step S201 is consistent with step S101 in embodiment 1 and is not described in detail in this embodiment.
[0069] Step S202: generating at least two sets of pronunciation texts corresponding to the at least two pronunciations based on the text to be predicted, the target character, and the at least two pronunciations;
[0070] The contents of the pronunciation texts in the at least two groups of pronunciation text sets are related to the contents of the text to be predicted.
[0071] Wherein, based on the text to be predicted, the target word and the multiple pronunciations of the target word, multiple groups of pronunciation text sets are generated, and the generated pronunciation text sets are respectively related to the text to be predicted, the target word and the multiple pronunciations of the target word.
[0072] Among them, the content of the text to be predicted represents the field to which it belongs, such as the medical field, the mechanical field, the automotive field, etc. Therefore, in this solution, the factors of the generated pronunciation text set include the text to be predicted, and the pronunciation texts in the generated pronunciation text set are related to the text to be predicted, specifically, the contents of the two are related, that is, they belong to the same or similar fields.
[0073] For example, if the text to be predicted is in the medical field, the pronunciation texts in the generated pronunciation text set are also in the medical field, so as to improve the similarity between the entire pronunciation text set and the text to be predicted and improve the accuracy of predicting the pronunciation of polyphonetic characters.
[0074] Step S203: determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0075] Among them, step S203 is consistent with step S103 in embodiment 1 and is not described in detail in this embodiment.
[0076] In summary, in an information processing method provided by this embodiment, at least two sets of pronunciation text sets corresponding to at least two pronunciations are generated based on at least the target character and the at least two pronunciations, including: based on the text to be predicted, the target character and the at least two pronunciations, at least two sets of pronunciation text sets corresponding to at least two pronunciations are generated, and the contents of the pronunciation texts in the at least two sets of pronunciation text sets are related to the contents of the text to be predicted. In this solution, if the factors of the generated pronunciation text set include the text to be predicted, then the pronunciation texts in the generated pronunciation text set are related to the text to be predicted, so as to improve the similarity between the pronunciation text set as a whole and the text to be predicted, and improve the accuracy of predicting the pronunciation of polyphonetic characters.
[0077] like Figure 3 The flowchart shown is a third embodiment of an information processing method provided by the present application, which includes the following steps:
[0078] Step S301: obtaining a target word in a text to be predicted, where the target word has at least two pronunciations;
[0079] Among them, step S301 is consistent with step S201 in embodiment 2 and is not described in detail in this embodiment.
[0080] Step S302: generating a first vector based on the text to be predicted, where the first vector represents content contained in the text to be predicted;
[0081] The text to be predicted is specifically a sentence, and the sentence may be composed of one word or multiple words.
[0082] The text to be predicted is processed using a set word2vec network to generate a first vector. The first vector is specifically an embedding representation of the target word in the text to be predicted, and the first vector includes the content included in the text to be predicted.
[0083] Specifically, the word2vec network has two inputs, one is the text to be predicted, and the other is the relative position of the current word relative to the target word. The output of the word2vec network is the embedding representation of the target word in the text to be predicted.
[0084] The first vector is a vector obtained by processing the entire sentence of the text to be predicted, and includes all the contents included in the text to be predicted.
[0085] The first vector is a vector for the text to be predicted containing the target word, and the vector implicitly contains information such as the position of the target word in the text to be predicted.
[0086] Specifically, the word2vec network is trained in an unsupervised manner in advance so that the trained word2vec network can process the input sentence to generate a sentence-level vector.
[0087] It should be noted that the sentence-level vector obtained by sentence processing and the similarity between the sentences corresponding to the two vectors can be determined based on two sentence-level vectors. The closer the vectors are, the more similar the corresponding sentences are.
[0088] Step S303: generating at least two sets of pronunciation texts corresponding to the at least two pronunciations based on the first vector, the target character, and the at least two pronunciations;
[0089] Each set of pronunciation texts corresponds to one pronunciation, each set of pronunciation texts includes at least one pronunciation text, and the content of each pronunciation text is related to the content of the text to be predicted.
[0090] In the first vector, based on the first vector, the target word and the multiple pronunciations of the target word, multiple groups of pronunciation text sets corresponding to the multiple voices are generated.
[0091] Among them, since the first vector contains all the content contained in the text to be predicted, the first vector is used as a factor to generate the pronunciation text. The pronunciation text in the generated pronunciation text set is related to the text to be predicted, specifically, the content of the two is related, that is, they belong to the same or similar fields.
[0092] Correspondingly, based on the number of pronunciations of the target word, a pronunciation text set with an array corresponding to the number of pronunciations is generated. Each pronunciation text set contains one or more pronunciation texts. In order to improve the accuracy of the prediction, multiple pronunciation texts are generally generated in each pronunciation text set.
[0093] It should be noted that the sentence-level vector is a representation of the entire sentence text. It is a vector obtained by processing the content contained in the entire sentence. This vector is combined with the context information of the target word. In this solution, the first vector generated based on the text to be predicted is used to combine the target word and its pronunciation with a vector. Compared with predicting by combining the sentence text with the target word and its pronunciation, the complexity of data processing in the prediction process is low.
[0094] Step S304: determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0095] Among them, step S304 is consistent with step S203 in embodiment 2 and is not described in detail in this embodiment.
[0096] In summary, in an information processing method provided by this embodiment, based on the text to be predicted, the target word and the at least two pronunciations, at least two groups of pronunciation text sets corresponding to at least two pronunciations are generated, including: generating a first vector based on the text to be predicted, the first vector representing the content contained in the text to be predicted; based on the first vector, the target word and the at least two pronunciations, generating at least two groups of pronunciation text sets corresponding to at least two pronunciations, each group of pronunciation text sets corresponding to one pronunciation, each group of pronunciation texts containing at least one pronunciation text, and the content of each pronunciation text being related to the content of the text to be predicted. In this solution, based on the first vector generated by the text to be predicted, a vector is combined with the target word and its pronunciation for prediction. Compared with prediction based on the sentence text combined with the target word and its pronunciation, the complexity of data processing in the prediction process is low.
[0097] like Figure 4 The flowchart shown is a fourth embodiment of an information processing method provided by the present application, which includes the following steps:
[0098] Step S401: obtaining a target word in a text to be predicted, where the target word has at least two pronunciations;
[0099] Step S402: generating at least two sets of pronunciation texts corresponding to the at least two pronunciations based on the target character and the at least two pronunciations;
[0100] Among them, steps S401-402 are consistent with steps S101-102 in Example 1 and are not described in detail in this embodiment.
[0101] Step S403: generating a vector corresponding to each pronunciation text in each pronunciation text set;
[0102] The vector generated by each pronunciation text represents the content of the pronunciation text.
[0103] In this solution, the difference between the vectors is used to determine the similarity between the generated pronunciation text and the text to be predicted, and then the pronunciation of the target word in the text to be predicted is determined.
[0104] A sentence-level vector is generated for each pronunciation text in each pronunciation text set, and each vector represents the content contained in the corresponding pronunciation text.
[0105] Among them, each pronunciation text contains the target character. The sentence-level vector generated by the pronunciation text combines the information of the context of the target character and contains the overall content of the pronunciation text.
[0106] Step S404: Process the vectors generated by the pronunciation texts belonging to the same group of pronunciation text sets to obtain the mean vector corresponding to the pronunciation text set.
[0107] Among them, the vectors generated by each pronunciation text within the same pronunciation text set are processed to calculate the mean value, and the mean vector of the pronunciation text set is obtained.
[0108] Among them, the mean vector represents the overall vector situation of the pronunciation text set, and the mean vectors of different pronunciation text sets are different.
[0109] Step S405: Based on the mean vectors corresponding to at least two pronunciation text sets and the first vector generated by the text to be predicted, determine the similarity between the pronunciation texts corresponding to the two pronunciation text sets and the text to be predicted.
[0110] Among them, the mean vector corresponding to each pronunciation text set represents the content of the corresponding pronunciation text.
[0111] Specifically, calculate the cosine distance between the first vector generated by the text to be predicted and the mean vector. This distance characterizes the gap between the overall pronunciation text set and the text to be predicted. The smaller the distance, the higher the similarity between the two.
[0112] For example, the target character "还" has two pronunciation characters "huán" and "hái". The distance between the mean vector of the "huán" pronunciation text set and the first vector is the first distance, and the distance between the mean vector of the "hái" pronunciation text set and the first vector is the second distance. Among them, the first distance is less than the second distance, indicating that the similarity between the "huán" pronunciation text set and the text to be predicted is higher.
[0113] Step S406: Based on the similarity, determine the first pronunciation as the pronunciation of the target character in the text to be predicted.
[0114] Among them, the pronunciation texts in the pronunciation text set corresponding to the first pronunciation satisfy the similarity condition with the text to be predicted.
[0115] Among them, after determining the similarities between the pronunciation text sets corresponding to multiple pronunciations and the text to be predicted, the similarities are sorted, and the pronunciation corresponding to the pronunciation text set with the largest similarity is selected as the pronunciation of the target character.
[0116] Among them, since the higher the similarity between the pronunciation text set and the text to be predicted, the more similar the pronunciation text in the pronunciation text set is to the text to be predicted, and the greater the possibility that the pronunciation of the target characters in the two are the same, the pronunciation text in the pronunciation text set corresponding to the first pronunciation has the highest similarity with the text to be predicted, and the first pronunciation is determined as the pronunciation of the target character.
[0117] In the specific implementation, after determining the cosine distance between the mean vector of the pronunciation text set and the first vector, the pronunciation corresponding to the pronunciation text set with the smallest cosine distance is directly selected as the pronunciation of the target word. There is no need to determine the similarity between the pronunciation text set and the text to be predicted based on the pre-distance, and then select the pronunciation of the target word based on the similarity, so as to reduce the amount of data processing.
[0118] In summary, in an information processing method provided by this embodiment, based on the similarity between the at least two groups of pronunciation text sets and the text to be predicted, the pronunciation of the target word in the text to be predicted is determined, including: generating a vector corresponding to each pronunciation text in each group of pronunciation text sets, the vector generated by each pronunciation text represents the content contained in the pronunciation text; processing the vectors generated by the pronunciation texts belonging to the same group of pronunciation text sets to obtain the mean vector corresponding to the pronunciation text set; based on the mean vectors corresponding to at least two pronunciation text sets and the first vector generated by the text to be predicted, determining the similarity between the pronunciation texts corresponding to the two pronunciation text sets and the text to be predicted; based on the similarity, determining a first pronunciation as the pronunciation of the target word in the text to be predicted, the pronunciation text in the pronunciation text set corresponding to the first pronunciation meets the similarity condition with the text to be predicted. In this scheme, based on the vectors generated by each pronunciation text in the pronunciation text set, the mean vector corresponding to the pronunciation text set is determined, and then the cosine distance is calculated based on the mean vector and the first vector of the text to be predicted. Based on the cosine distance, the similarity between the pronunciation text set and the text to be tested is determined, and then the similarity between the pronunciation text set and the text to be tested is used to determine the pronunciation of the target word. This scheme is not affected by the imbalance of the training corpus, and the accuracy of predicting the pronunciation of polyphonetic words is higher.
[0119] like Figure 5 The flowchart shown is a fifth embodiment of an information processing method provided by the present application, and the method includes the following steps:
[0120] Step S501: obtaining a target word in a text to be predicted, where the target word has at least two pronunciations;
[0121] Step S502: generating a first vector based on the text to be predicted, where the first vector represents content contained in the text to be predicted;
[0122] Among them, steps S501-502 are consistent with steps S301-302 in Example 3 and are not described in detail in this embodiment.
[0123] Step S503: Processing the first vector, the target word, and the at least two pronunciations based on a target generator to generate the at least two sets of pronunciation text sets;
[0124] In this solution, a pronunciation text set is generated based on a target generator in a learning model, wherein the learning model adopts a condition GAN model, which is a trained model.
[0125] The condition GAN model includes a generator and a discriminator. The generator generates output content based on the input conditions, and the discriminator scores the output content. For example, if it is close to reality and meets the conditions, it will be scored 1 point. If the output content is of low quality or does not meet the conditions, it will be scored 0 points.
[0126] like Figure 6 The figure shows a schematic diagram of a learning model, comprising a generator 601 and a discriminator 602. The generator generates output content based on input conditions, while the discriminator discriminates against the output content. The generator's goal is to generate as realistic content as possible based on the input conditions to deceive the discriminator. The discriminator's goal, on the other hand, is to distinguish the generator's output content from the real content. The generator and the discriminator form a dynamic "game" process.
[0127] Specifically, the target generator based on the condition GAN model after the training is completed takes the first vector generated based on the text to be predicted, the target word and at least two pronunciations of the target word as input conditions and inputs them into the target generator so that the target generator generates at least two sets of pronunciation text sets based on the input conditions.
[0128] In a specific implementation, due to the working rules of the generator, in addition to the input conditions, a random signal needs to be input into the generator so that the target generator generates content based on the input conditions.
[0129] Among them, due to the generator in the condition GAN model that has completed the training of the target generator, the target generator can generate output content that meets the input conditions, that is, at least two sets of pronunciation text sets generated based on the target generator are consistent with the text to be predicted as the input conditions, and the consistency can specifically belong to the same field.
[0130] Step S504: determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0131] Among them, step S504 is consistent with step S304 in embodiment 3 and is not described in detail in this embodiment.
[0132] In summary, in an information processing method provided by this embodiment, the method generates at least two sets of pronunciation text sets corresponding to at least two pronunciations based on the first vector, the target word and the at least two pronunciations, including: processing the first vector, the target word and the at least two pronunciations based on the target generator to generate the at least two sets of pronunciation text sets. In this solution, the first vector, the target word and the at least two pronunciations of the target word generated based on the text to be predicted are used as input conditions and input into the target generator of the condition GAN model, so that the target generator generates at least two sets of pronunciation text sets based on the input conditions, and generates at least two sets of pronunciation text sets that meet the input conditions.
[0133] like Figure 7 The figure is a flowchart of Example 6 of an information processing method provided by the present application, and the method includes the following steps:
[0134] Step S701: training an original learning model based on a training text set, a target polyphonetic character, and at least two pronunciations of the target polyphonetic character to obtain a target learning model;
[0135] The target learning model includes a target generator and a target discriminator, and the training text set includes the target polyphone and annotated pronunciation, and the annotated pronunciation is the annotated pronunciation of the target polyphone in the training text.
[0136] A training text set is preset, and the training text set includes a large number of training texts. Each training text has at least one polyphonic character, and the pronunciation of the polyphonic character is marked.
[0137] The original learning model is trained based on the training text set to obtain a target learning model.
[0138] The training text set includes common pronunciations and uncommon pronunciations.
[0139] Among them, the target learning model includes a target generator and a target discriminator. In this solution, the target learning model is obtained by training the original learning model, and the training process includes training the generator and discriminator therein.
[0140] Among them, the target generator in the target learning model can generate output content that is consistent with the input conditions to ensure that in the subsequent steps, a pronunciation text containing the target word can be generated based on the text to be predicted and the target word contained therein and the multiple pronunciations of the target word. The generated pronunciation text is consistent with the target word and multiple pronunciations in the input text to be predicted. Specifically, the pronunciation text contains the target word and the corresponding multiple pronunciations, and the pronunciation text belongs to the same field as the text to be predicted.
[0141] It should be noted that there are more common pronunciation corpora and less uncommon pronunciation corpora in the training text set. However, in this solution, a pronunciation text containing the target word is generated based on the text to be predicted, the target word contained therein, and multiple pronunciations of the target word. The pronunciation of the target word in the text to be predicted is determined based on the similarity between the pronunciation text and the text to be predicted. It is not limited by the amount of pronunciation corpora in the training text set. The predicted pronunciation of the target word is not affected by the imbalance of the training corpus, and the prediction accuracy is higher.
[0142] Step S702: Obtain a target word in the text to be predicted, where the target word has at least two pronunciations;
[0143] Step S703: generating a first vector based on the text to be predicted, where the first vector represents content contained in the text to be predicted;
[0144] Step S704: Processing the first vector, the target word, and the at least two pronunciations based on the target generator to generate the at least two sets of pronunciation text sets;
[0145] Step S705: Determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted.
[0146] Among them, steps S702-705 are consistent with steps S501-504 in Example 5 and are not described in detail in this embodiment.
[0147] In summary, the information processing method provided by this embodiment further includes: training the original learning model based on the training text set, the target polyphone and at least two pronunciations of the target polyphone to obtain a target learning model, wherein the target learning model includes a target generator and a target discriminator, and the training text set includes the target polyphone and the annotated pronunciation, and the annotated pronunciation is the annotated pronunciation of the target polyphone in the training text. In this scheme, the polyphone contained in the set training text set is annotated with pronunciation, and the original learning model is trained on the training text set after the annotated pronunciation to obtain the target learning model, wherein the target generator in the target learning model can generate a pronunciation text containing the target word based on the text to be predicted and the target word contained therein and the multiple pronunciations of the target word, and the generated pronunciation text is consistent with the target word and multiple pronunciations in the input text to be predicted, and the pronunciation text contains the target word and the corresponding multiple pronunciations, and the pronunciation text belongs to the same field as the text to be predicted, so as to provide a basis for subsequently determining the pronunciation of the target word in the text to be predicted based on multiple pronunciation text sets.
[0148] like Figure 8 The flowchart shown is a seventh embodiment of an information processing method provided by the present application, and the method includes the following steps:
[0149] Step S801: obtaining a training text set, wherein the training text set includes at least two groups of training texts, each group of training texts corresponding to a marked pronunciation;
[0150] The training text set includes multiple groups of training texts, each group of training texts includes polyphonetic characters with the same pronunciation, and the pronunciation of the polyphonetic characters in each training text is annotated.
[0151] In a specific implementation, in the training text set, each pronunciation of each polyphonetic character corresponds to a group of training texts, and each group of training texts may include multiple training texts.
[0152] Step S802: generating a training vector based on the training text;
[0153] The training vector represents the content contained in the training text.
[0154] The training text is processed using a set word2vec network to generate a training vector. The training vector is a vector obtained by processing the entire sentence of the training text, which contains all the content contained in the training text.
[0155] Specifically, all training texts in the training text set are processed in sequence to generate a corresponding number of training vectors.
[0156] Step S803: inputting at least the training vector, the target polyphonetic character, and the annotated pronunciation as input conditions into an original generator to obtain a target generator, so that the target generator generates a first text;
[0157] The training vector, the target polyphone, and the annotated pronunciation of the target polyphone are used as input conditions and input into an original generator to train the original generator to obtain a target generator. The target generator can further process the training text set to generate a first text.
[0158] The generated first text meets the input conditions of the target generator.
[0159] Specifically, since the learning model includes two parts: a generator and a discriminator, in the specific training process, the two parts are trained separately. The generator is trained first to obtain the target generator, and then the discriminator is trained based on the target generator.
[0160] The step of inputting the training vector, the target polyphonetic character, and the annotated pronunciation as input conditions into an original generator to obtain a target generator includes:
[0161] Step S01: inputting the training vector, the target polyphone character and the annotated pronunciation as input conditions into an original generator to generate a second text;
[0162] The training vector generated based on the training text, the target polyphone in the training text and its standard pronunciation are input as input conditions into the original generator, and the original generator generates the output content second text based on the input conditions.
[0163] Among them, since the original generator has not been trained, the output content generated based on the input conditions is random. In order to ensure that the output content generated by the generator based on the input conditions meets the input conditions, it is necessary to adjust the input conditions of the original generator based on the output content of the original generator.
[0164] Step S02: adding a false label to the second text, and returning the labeled second text as training text to execute the step of generating training vectors based on the training text until a second agreed training stop condition is met, thereby obtaining a target generator;
[0165] Among them, during the training process of the original generator, if the second text generated by the original generator does not meet the input conditions, the trainer will add a label that is judged to be false to the second text, add it to the training text set as a training text, continue to use the training text set to generate training vectors, and input the generator based on the generated training vectors, so that the generator continues to generate text.
[0166] A second agreed stopping condition is set for the training generator, for example, the condition is the number of cyclic training times.
[0167] For example, the number of loop training is 100 times. When the original generator is trained based on the training text set 100 times, the training of the generator is stopped to obtain the target generator.
[0168] Step S03: inputting the training vector, the target polyphonetic character and the annotated pronunciation as input conditions into a target generator, so that the target generator generates a first text.
[0169] Among them, the training vector generated by the original training text, the target polyphone contained therein and its annotated pronunciation are used as input conditions, and the input conditions are input into the target generator. The target generator generates a first text, and the first text is used for subsequent input into the original discriminator to train the original discriminator.
[0170] In the process of training the generator, the random signal, training vector, target polyphone and annotated pronunciation are also input into the original generator to train the original generator to obtain the target generator.
[0171] Due to the working rules of the generator, in addition to the input conditions, a random signal needs to be input into the original generator so that the original generator generates content based on the input conditions.
[0172] Step S804: using the first text and any training text as inputs of an original discriminator, so that the original discriminator judges the true situation of the first text and the training text until a first agreed stopping condition is met, thereby obtaining a target learning model;
[0173] The target learning model includes a target generator and a target discriminator.
[0174] The first text generated by the target generator based on the training text and any training text are used as input content and input into the original discriminator, so that the original discriminator can judge the true situation of the first text and the training text.
[0175] Since the original discriminator has not been trained, its judgment of the actual situation may be correct or wrong, such as judging that the training text is false.
[0176] Among them, the first agreed training stop condition can be set according to the situation, such as setting the number of training times, the value of the loss function (LOSS) of the output result no longer decreasing, etc.
[0177] Specifically, when training the original discriminator, each time training is performed, the target generator generates a first text, and the first text and the training text are input into the original discriminator as input content for training.
[0178] When the first agreed training stop condition is the number of training times, the number of times the original discriminator is trained using the training text and the first text is recorded until the number of training times meets the set number, and the training is terminated.
[0179] Among them, when the first agreed condition for stopping training is that the loss function no longer decreases, the judgment result of each output is recorded, and the loss of the judgment result of the output is calculated based on the loss function until the loss function of the judgment result of each training output no longer decreases, and the training is ended.
[0180] Among them, the original discriminator training is completed, indicating that the learning model training is completed, and the target generator and target discriminator are obtained.
[0181] Step S805: obtaining a target word in the text to be predicted, where the target word has at least two pronunciations;
[0182] Step S806: generating a first vector based on the text to be predicted, where the first vector represents content contained in the text to be predicted;
[0183] Step S807: Processing the first vector, the target word, and the at least two pronunciations based on the target generator to generate the at least two sets of pronunciation texts;
[0184] Step S808: Determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two groups of pronunciation text sets and the text to be predicted.
[0185] Among them, steps S805-808 are consistent with steps S702-705 in Example 6 and are not described in detail in this embodiment.
[0186] In summary, in an information processing method provided by this embodiment, an original learning model is trained based on training text, a target polyphone and at least two pronunciations of the target polyphone to obtain a target learning model, including: obtaining a training text set, wherein the training text set includes at least two groups of training texts, each group of training texts corresponds to a marked pronunciation; generating a training vector based on the training text, wherein the training vector represents the content contained in the training text; inputting the original generator at least based on the training vector, the target polyphone and the marked pronunciation as input conditions to obtain a target generator, so that the target generator generates a first text; using the first text and any training text as input content of the original discriminator, so that the original discriminator judges the true situation of the first text and the training text until the first agreed stop training condition is met, thereby obtaining a target learning model, wherein the target learning model includes a target generator and a target discriminator. In this scheme, the training process of the target learning model is explained, and the original learning model is trained based on the training process to obtain a trained target learning model, and the target generator in the target learning model is used to generate multiple pronunciation texts in the subsequent prediction of the pronunciation of the target word in the text to be predicted.
[0187] Corresponding to the above-mentioned information processing method embodiment provided by the present application, the present application also provides an apparatus embodiment applying the information processing method.
[0188] like Figure 9 The figure shows a schematic diagram of the structure of an embodiment of an information processing device provided by the present application, the device includes the following structures: an acquisition module 901, a generation module 902 and a determination module 903;
[0189] The acquisition module 901 is configured to acquire a target word in the text to be predicted, wherein the target word has at least two pronunciations.
[0190] wherein the generating module 902 is configured to generate at least two sets of pronunciation texts corresponding to the at least two pronunciations based on the target character and the at least two pronunciations, wherein each set of pronunciation texts corresponds to one pronunciation, and each set of pronunciation texts includes at least one pronunciation text;
[0191] The determination module 903 is configured to determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation texts and the text to be predicted.
[0192] Optionally, the generation module is used to:
[0193] Based on the text to be predicted, the target character and the at least two pronunciations, at least two sets of pronunciation texts corresponding to the at least two pronunciations are generated, and the contents of the pronunciation texts in the at least two sets of pronunciation texts are related to the contents of the text to be predicted.
[0194] Optionally, the generation module includes:
[0195] A vector unit, configured to generate a first vector based on the text to be predicted, wherein the first vector represents content contained in the text to be predicted;
[0196] A generation unit is used to generate at least two groups of pronunciation text sets corresponding to at least two pronunciations based on the first vector, the target word and the at least two pronunciations, each group of pronunciation text sets corresponding to one pronunciation, each group of pronunciation texts containing at least one pronunciation text, and the content of each pronunciation text is related to the content of the text to be predicted.
[0197] Optionally, the determination module includes:
[0198] Generate a vector for each pronunciation text in each pronunciation text set, where the vector generated by each pronunciation text represents the content of the pronunciation text;
[0199] Processing vectors generated from pronunciation texts belonging to the same pronunciation text set to obtain a mean vector corresponding to the pronunciation text set;
[0200] Determining similarities between the pronunciation texts corresponding to the two pronunciation text sets and the text to be predicted based on mean vectors corresponding to at least two pronunciation text sets and a first vector generated by the text to be predicted;
[0201] Based on the similarity, a first pronunciation is determined as the pronunciation of the target word in the text to be predicted, and a pronunciation text in the pronunciation text set corresponding to the first pronunciation meets a similarity condition with the text to be predicted.
[0202] Optionally, the generating unit is used to:
[0203] The target generator processes the first vector, the target word, and the at least two pronunciations to generate the at least two groups of pronunciation text sets.
[0204] Optionally, also include:
[0205] A training module is used to train an original learning model based on a training text set, a target polyphone and at least two pronunciations of the target polyphone to obtain a target learning model, wherein the target learning model includes a target generator and a target discriminator, and the training text set includes the target polyphone and annotated pronunciation, wherein the annotated pronunciation is the annotated pronunciation of the target polyphone in the training text.
[0206] Optional training module, specifically for:
[0207] Obtaining a training text set, wherein the training text set includes at least two groups of training texts, each group of training texts corresponding to a marked pronunciation;
[0208] generating a training vector based on the training text, wherein the training vector represents content contained in the training text;
[0209] At least the training vector, the target polyphone, and the annotated pronunciation are input into an original generator as input conditions to obtain a target generator, so that the target generator generates a first text;
[0210] The first text and any training text are used as input content of the original discriminator, so that the original discriminator judges the true situation of the first text and the training text until the first agreed training stop condition is met, thereby obtaining a target learning model, which includes a target generator and a target discriminator.
[0211] Optional training module, specifically for:
[0212] Inputting the training vector, the target polyphonetic character and the annotated pronunciation as input conditions into an original generator to generate a second text;
[0213] Adding a false label to the second text, and returning the labeled second text as training text to execute the step of generating a training vector based on the training text until a second agreed training stop condition is met, thereby obtaining a target generator;
[0214] The training vector, the target polyphonetic character and the annotated pronunciation are input into a target generator as input conditions, so that the target generator generates a first text.
[0215] The functions of the information processing device in this application are explained with reference to the method embodiment and will not be described in detail in this embodiment.
[0216] In summary, the present application provides an information processing device that generates multiple pronunciation texts based on the polyphones in the text to be predicted, so as to further select the pronunciation corresponding to the polyphones in the predicted text based on the multiple pronunciation texts and the text to be predicted, without being affected by the imbalance of the training corpus, and the accuracy of predicting the pronunciation of the polyphones is higher.
[0217] Corresponding to the above-mentioned embodiment of an information processing method provided by the present application, the present application also provides an electronic device and a readable storage medium corresponding to the information processing method.
[0218] The electronic device includes: a memory and a processor;
[0219] Wherein, the memory stores a processing program;
[0220] The processor is used to load and execute the processing program stored in the memory to implement each step of the information processing method as described in any one of the above items.
[0221] For details on how to implement the information processing method of the electronic device, please refer to the aforementioned information processing method embodiments.
[0222] The readable storage medium stores a computer program thereon, and the computer program is called and executed by the processor to implement the steps of the information processing method according to any one of the above claims.
[0223] Specifically, the computer program stored in the readable storage medium is executed to implement the information processing method, and reference may be made to the aforementioned information processing method embodiments.
[0224] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices provided in the embodiments, since they correspond to the methods provided in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0225] The above description of the provided embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features provided herein.
Claims
1. An information processing method, comprising: obtaining a target character in the text to be predicted, where the target character has at least two pronunciations; generating at least two sets of pronunciation text sets corresponding to the at least two pronunciations based on the text to be predicted, the target character, and the at least two pronunciations, where each set of pronunciation text sets corresponds to one pronunciation, and each set of pronunciation text sets contains at least one pronunciation text; determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted; wherein the content of the pronunciation texts in the at least two sets of pronunciation text sets is related to the content of the text to be predicted.
2. The method according to claim 1, generating at least two sets of pronunciation text sets corresponding to the at least two pronunciations based on the text to be predicted, the target character, and the at least two pronunciations, comprising: generating a first vector based on the text to be predicted, the first vector representing the content included in the text to be predicted; a target generator generates at least two sets of pronunciation text sets corresponding to the at least two pronunciations based on the first vector, the target character, and the at least two pronunciations, each set of pronunciation text sets corresponds to one pronunciation, and each set of pronunciation text sets contains at least one pronunciation text, and the content of each pronunciation text is related to the content of the text to be predicted.
3. The method according to any one of claims 1-2, determining the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted, comprising: generating a vector corresponding to each pronunciation text in each set of pronunciation text sets, and the vector generated by each pronunciation text represents the content included in the pronunciation text; processing the vectors generated by the pronunciation texts belonging to the same set of pronunciation text sets to obtain the mean vector corresponding to the pronunciation text set; determining the similarity between the pronunciation texts corresponding to the two pronunciation text sets and the text to be predicted based on the mean vectors corresponding to the at least two pronunciation text sets and the first vector generated by the text to be predicted; determining a first pronunciation as the pronunciation of the target character in the text to be predicted, and the pronunciation texts in the pronunciation text set corresponding to the first pronunciation satisfy the similarity condition with the text to be predicted.
4. The method according to claim 2, generating at least two sets of pronunciation text sets corresponding to the at least two pronunciations based on the first vector, the target character, and the at least two pronunciations, comprising: generating the at least two sets of pronunciation text sets by a target generator processing the first vector, the target character, and the at least two pronunciations.
5. The method according to claim 4, before obtaining the target character in the text to be predicted, further comprising: training an original learning model based on a training text set, a target polyphonic character, and at least two pronunciations of the target polyphonic character to obtain a target learning model, where the target learning model includes a target generator and a target discriminator, the training text set includes the target polyphonic character and a labeled pronunciation, and the labeled pronunciation is the pronunciation labeled for the target polyphonic character in the training text set.
6. The method according to claim 5, wherein the original learning model is trained based on a training text set, a target polyphonic character, and at least two pronunciations of the target polyphonic character to obtain a target learning model, including: obtaining a training text set, where the training text set includes at least two groups of training texts, and each group of training texts corresponds to a marked pronunciation; generating a training vector based on the training text, where the training vector represents the content included in the training text; inputting at least the training vector, the target polyphonic character, and the marked pronunciation as input conditions into an original generator to obtain a target generator, so that the target generator generates a first text; using the first text and any one of the training texts as the input content of an original discriminator, so that the original discriminator determines the true situation of the first text and the training text until a first agreed stop training condition is met to obtain a target learning model, where the target learning model includes a target generator and a target discriminator.
7. The method according to claim 6, wherein inputting the training vector, the target polyphonic character, and the marked pronunciation as input conditions into an original generator to obtain a target generator includes inputting the training vector, the target polyphonic character, and the marked pronunciation as input conditions into an original generator to generate a second text; adding a label determined to be false to the second text, and using the second text with the added label as a training text to return to execute the step of generating a training vector based on the training text until a second agreed stop training condition is met to obtain a target generator; inputting the training vector, the target polyphonic character, and the marked pronunciation as input conditions into the target generator, so that the target generator generates a first text.
8. An information processing device, including: an obtaining module, configured to obtain a target character in a text to be predicted, where the target character has at least two pronunciations; a generating module, configured to generate at least two sets of pronunciation text sets corresponding to at least two pronunciations based on the text to be predicted, the target character, and the at least two pronunciations, where each set of pronunciation text sets corresponds to one pronunciation, and each set of pronunciation text sets includes at least one pronunciation text; a determining module, configured to determine the pronunciation of the target character in the text to be predicted based on the similarity between the at least two sets of pronunciation text sets and the text to be predicted; where the content of the pronunciation texts in the at least two sets of pronunciation text sets is related to the content of the text to be predicted.
9. An electronic device, including: a memory and a processor; wherein the memory stores a processing program; the processor is configured to load and execute the processing program stored in the memory to implement the steps of the information processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Polyphone disambiguation method and device
CN112580335A
Polyphone disambiguation method, electronic equipment and computer readable storage medium
CN113486672A