Polyphone disambiguation method and device, equipment, storage medium and program product

By introducing soft-connection coding and classification processing methods in polyphonic disambiguation, the dependence problem on word segmentation and polyphonic dictionary in the prior art is solved, and a higher accuracy of polyphonic disambiguation is achieved.

CN120235142AActive Publication Date: 2025-07-01IFLYTEK CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510455056.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-01
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing polyphonic disambiguation scheme relies on word segmentation accuracy and the interpretation information of polyphonic dictionary, resulting in the accuracy of polyphonic pronunciation cannot be guaranteed in the case of word segmentation errors or interpretation information incorrectly, and the accuracy needs to be improved.

Method used

A method of disambiguation of polyphonic characters is proposed, and the pronunciation of polyphonic characters is determined through encoding and classification processing based on soft connections. The specific steps include: building multiple soft connections for the polyphonic characters in the target text, encoding these soft connections to obtain the target encoding features, classifying the polyphonic characters based on these features, obtaining the probability distribution of the polyphonic characters, and determining the final pronunciation through the voting mechanism.

Benefits of technology

This method does not require word segmentation tools and polyphonic dictionary, avoids the influence of word segmentation errors and errors in interpretation information, and improves the accuracy of polyphonic disambiguation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235142A_ABST
    Figure CN120235142A_ABST
Patent Text Reader

Abstract

The invention discloses a polyphone disambiguation method and device, equipment, a storage medium and a program product, and relates to the technical field of natural language processing, and the method comprises the steps: for a target polyphone in a target text, coding each flexible connection based on the association relationship between a plurality of flexible connections corresponding to the target polyphone and the target text, obtaining a target coding feature of each flexible connection; each flexible connection is at least composed of a target polyphone or at least two continuous characters containing the target polyphone in the target text; performing classification processing on the target polyphone based on the target coding feature of the flexible connection to obtain probability distribution of the target polyphone corresponding to the flexible connection; the probability distribution represents the pronunciation of the target polyphone in the flexible connection; the plurality of pronunciations are all pronunciations of all polyphones in a preset polyphone set; and voting based on each probability distribution of the target polyphone to determine the pronunciation of the target polyphone in the target text. According to the invention, the accuracy of polyphone disambiguation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method, apparatus, device, storage medium, and program product for disambiguating polyphonic characters. Background Art

[0002] Polyphonic characters are a widespread linguistic phenomenon in Chinese. If a Chinese character has multiple pronunciations, then this character is a polyphonic character. In scenarios such as speech synthesis and question answering, it is necessary to select the correct pronunciation for a Chinese character with multiple pronunciations, and this process is the process of disambiguating polyphonic characters.

[0003] The current mainstream solution for disambiguating polyphonic characters is implemented based on a classification model and a polyphonic word dictionary. This solution requires tokenizing the text to obtain tokens containing polyphonic characters, then looking up the explanatory information of the tokens containing polyphonic characters in the polyphonic word dictionary, and classifying the polyphonic characters based on this explanatory information to determine the pronunciation of the polyphonic characters. This solution strongly depends on the accuracy of tokenization and the accuracy of the explanatory information of the tokens in the polyphonic word dictionary. That is to say, in the case of tokenization errors or errors in the explanatory information of the tokens, the correctness of the pronunciation of polyphonic characters cannot be guaranteed. Therefore, the accuracy of the polyphonic character disambiguation task needs to be further improved. Summary of the Invention

[0004] In view of the above problems, this application provides a method, apparatus, device, storage medium, and program product for disambiguating polyphonic characters to improve the accuracy of polyphonic character disambiguation. The specific solutions are as follows:

[0005] The first aspect of this application provides a method for disambiguating polyphonic characters, including:

[0006] For a target polyphonic character in a target text, based on the association relationship between multiple soft connections corresponding to the target polyphonic character and the target text, encode each of the soft connections to obtain the target encoding features of each of the soft connections; each soft connection is composed of at least the target polyphonic character or at least two consecutive characters in the target text that contain the target polyphonic character, and the maximum number of characters included in each soft connection is a preset number;

[0007] Based on the target encoding features of each soft connection, perform classification processing on the target polyphonic character in each soft connection to obtain the probability distribution of the target polyphonic character corresponding to each soft connection; the probability distribution is the probability that the target polyphonic character belongs to each of several pronunciations, representing the pronunciation of the target polyphonic character in the soft connection; the several pronunciations are all the pronunciations of all the polyphonic characters in a preset polyphonic character set;

[0008] Voting is performed based on the probability distribution of the target polyphonic characters corresponding to each of the soft connections to determine the pronunciation of the target polyphonic character in the target text.

[0009] In a possible implementation, encoding each of the soft connections based on the association relationship between the multiple soft connections corresponding to the target polyphonic character and the target text includes:

[0010] Encoding each character in the target text to obtain the encoding features of each character;

[0011] For each soft connection corresponding to the target polyphonic character, fusing the encoding features of each character included in the soft connection to obtain the initial encoding feature of the soft connection;

[0012] Based on the cross-attention mechanism, fusing the initial encoding feature of the soft connection with the encoding features of each character in the target text to obtain the target encoding feature of the soft connection.

[0013] In a possible implementation, the step of fusing the initial encoding feature of the soft connection with the encoding features of each character in the target text based on the cross-attention mechanism to obtain the target encoding feature of the soft connection includes:

[0014] Using the initial encoding feature of the soft connection and the encoding features of each character in the target text to calculate the attention weights of the soft connection to each character in the target text;

[0015] Based on the attention weights of the soft connection to each character in the target text, performing a weighted sum of the encoding features of each character in the target text to obtain the target encoding feature of the soft connection.

[0016] In a possible implementation, the step of voting based on the probability distribution of the target polyphonic character corresponding to each of the soft connections includes:

[0017] Fusing the probabilities in the probability distribution of the target polyphonic character corresponding to each of the soft connections where the target polyphonic character belongs to the same pronunciation to obtain the target probabilities of the target polyphonic character belonging to each pronunciation;

[0018] Determining the pronunciation corresponding to the maximum target probability as the pronunciation of the target polyphonic character in the target text.

[0019] In a possible implementation, the step of voting based on the probability distribution of the target polyphonic character corresponding to each of the soft connections includes:

[0020] Based on the probability distribution of the target polyphonic character corresponding to each soft connection, determining the pronunciation of the target polyphonic character in each soft connection;

[0021] Count the pronunciations of the target polyphonic character in each of the soft connections, and determine the pronunciation with the largest number as the pronunciation of the target polyphonic character in the target text.

[0022] In a possible implementation, it further includes:

[0023] If the numbers of multiple pronunciations of the target polyphonic character are the same, determine the pronunciation corresponding to the highest probability as the pronunciation of the target polyphonic character in the target text;

[0024] Or, if the numbers of multiple pronunciations of the target polyphonic character are the same, fuse the probabilities that the target polyphonic character belongs to the same pronunciation in the probability distributions of the target polyphonic character corresponding to each soft connection to obtain the target probabilities that the target polyphonic character belongs to each pronunciation; determine the pronunciation corresponding to the largest target probability as the pronunciation of the target polyphonic character in the target text.

[0025] In a possible implementation, encode each character in the target text to obtain the encoding features of each character. For each soft connection corresponding to the target polyphonic character, based on the cross-attention mechanism, fuse the initial encoding features of this soft connection with the encoding features of each character in the target text to obtain the target encoding features of this soft connection, and the process of classifying the target encoding features of each soft connection includes:

[0026] Encode each character in the target text through the encoding module of the classification model to obtain the encoding features of each character;

[0027] Through the connection module of the classification model, for each soft connection corresponding to the target polyphonic character, fuse the encoding features of each character included in this soft connection to obtain the initial encoding features of this soft connection;

[0028] Through the attention module of the classification model, for each soft connection, based on the cross-attention mechanism, fuse the initial encoding features of this soft connection with the encoding features of each character in the target text to obtain the target encoding features of this soft connection;

[0029] Through the classification module of the classification model, classify the target polyphonic character in each soft connection based on the target encoding features of each soft connection to obtain the probability distribution of the target polyphonic character corresponding to each software connection.

[0030] In a possible implementation, the classification model is trained in the following manner:

[0031] Encode each character in the text sample through the encoding module to obtain the encoding features of each character;

[0032] For each polyphonic character in the text sample through the connection module, obtain the initial encoding features of multiple soft connections corresponding to the polyphonic character; wherein, each soft connection is composed of at least the polyphonic character or at least two consecutive characters in the text sample that contain the polyphonic character, and the maximum number of characters included in each soft connection is a preset number; the initial encoding feature of each soft connection is obtained by fusing the encoding features of each character included in the soft connection;

[0033] For each soft connection through the attention module, based on the cross-attention mechanism, fuse the initial encoding feature of the soft connection with the encoding features of each character in the text sample to obtain the target encoding feature of the soft connection;

[0034] Through the classification module, classify the polyphonic character in each soft connection based on the target encoding feature of each soft connection to obtain the probability distribution of the polyphonic character corresponding to the software connection;

[0035] Based on the probability distribution of each polyphonic character corresponding to each soft connection in the text sample, determine the first pronunciation of each polyphonic character corresponding to the soft connection;

[0036] For each polyphonic character in the text sample, use the multiple probability distributions of the polyphonic character determined based on the corresponding soft connections of the polyphonic character to vote to determine the second pronunciation of each polyphonic character in the text sample;

[0037] Taking the first pronunciation and the second pronunciation of each polyphonic character approaching the pronunciation label of the polyphonic character as the goal, update the parameters of the classification model.

[0038] The second aspect of the present application provides a polyphonic character disambiguation device, including:

[0039] A soft connection unit, configured to encode each soft connection based on the association relationship between the multiple soft connections corresponding to the target polyphonic character and the target text for the target polyphonic character in the target text to obtain the target encoding features of each soft connection; each soft connection is composed of at least the target polyphonic character or at least two consecutive characters in the target text that contain the target polyphonic character, and the maximum number of characters included in each soft connection is a preset number;

[0040] A classification unit, configured to classify the target polyphonic character in each soft connection based on the target encoding feature of each soft connection to obtain the probability distribution of the target polyphonic character corresponding to each soft connection; the probability distribution is the probability that the target polyphonic character belongs to each pronunciation among several pronunciations, representing the pronunciation of the target polyphonic character in the soft connection; the several pronunciations are all the pronunciations of all the polyphonic characters in a preset polyphonic character set;

[0041] A voting unit, configured to vote based on the probability distributions of the target polyphonic characters corresponding to each of the soft connections, so as to determine the pronunciations of the target polyphonic characters in the target text.

[0042] A third aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the polyphonic character disambiguation method according to the first aspect or any implementation manner of the first aspect.

[0043] A fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0044] The memory is configured to store a computer program;

[0045] The processor is configured to execute the computer program, so that the electronic device can implement the polyphonic character disambiguation method according to the first aspect or any implementation manner of the first aspect.

[0046] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, can enable the electronic device to implement the polyphonic character disambiguation method according to the first aspect or any implementation manner of the first aspect.

[0047] With the above technical solution, the method, apparatus, device, storage medium, and program product for disambiguating polyphonic characters provided by this application, for the target polyphonic character in the target text, based on the association relationship between multiple soft connections corresponding to the target polyphonic character and the target text, encode each soft connection to obtain the target encoding features of each soft connection; each soft connection is composed of at least the target polyphonic character or at least two consecutive characters including the target polyphonic character in the target text, and the maximum number of characters included in each soft connection is a preset number; classify the target polyphonic character in each soft connection based on the target encoding features of each soft connection to obtain the probability distribution of the target polyphonic character corresponding to each soft connection; this probability distribution is the probability that the target polyphonic character belongs to each of several pronunciations, representing the pronunciation of the target polyphonic character in the soft connection; the several pronunciations are all the pronunciations of all the polyphonic characters in the preset polyphonic character set; vote based on the probability distributions of the target polyphonic characters corresponding to each soft connection to determine the pronunciation of the target polyphonic character in the target text. It can be seen that the polyphonic character disambiguation solution of this application proposes the concept of soft connection, extracts the word combinations composed of the polyphonic character and its context before and after without relying on a word segmentation tool as soft connections, determines the target encoding features of each soft connection corresponding to the target polyphonic character based on the association relationship between each soft connection corresponding to the target polyphonic character and the target text, classify the target polyphonic character in each soft connection respectively based on the target encoding features of each soft connection to obtain multiple possible classification results of the target polyphonic character, vote based on each classification result to determine the pronunciation of the target polyphonic character in the target text. This process does not require word segmentation of the target text, will not cause polyphonic character disambiguation errors due to word segmentation errors, nor does it need to use the explanatory information of the word segmentation containing polyphonic characters, and will not cause polyphonic character disambiguation errors due to incorrect explanatory information, thus avoiding the dependence of polyphonic character disambiguation on word segmentation and explanatory information and improving the accuracy of polyphonic character disambiguation. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0049] Figure 1 It is a flowchart of an implementation of the method for disambiguating polyphonic characters provided by this application;

[0050] Figure 2 It is an example of a soft connection provided by this application;

[0051] Figure 3 It is a flowchart of an implementation of encoding each soft connection based on the association relationship between multiple soft connections corresponding to the target polyphonic character and the target text to obtain the target encoding features of each soft connection provided by this application;

[0052] Figure 4 A flowchart of an implementation for voting based on various probability distributions of target polyphonic characters provided for this application;

[0053] Figure 5 Another flowchart of an implementation for voting based on various probability distributions of target polyphonic characters provided for this application;

[0054] Figure 6 A schematic structural diagram of a classification model provided for this application;

[0055] Figure 7 A schematic structural diagram of a polyphonic character disambiguation device provided for this application;

[0056] Figure 8 A schematic structural diagram of an electronic device provided for this application. Detailed implementation manners

[0057] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application. The terms used in the implementation part of this application are only used to explain the specific embodiments of this application, rather than intending to limit this application.

[0058] The embodiments of this application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0059] The terms "first", "second", etc. in the specification, claims and above-mentioned accompanying drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinction adopted when describing objects with the same attributes in the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0060] In order to avoid strong dependence on the accuracy of word segmentation and the accuracy of word segmentation interpretation information in polyphonic character disambiguation, the solution of this application is proposed.

[0061] As Figure 1 shown, a flowchart of an implementation of the polyphonic character disambiguation method provided in the embodiments of this application may include:

[0062] Step S101: For the target polyphonic character in the target text, based on the association relationship between multiple soft connections corresponding to the target polyphonic character and the target text, encode each soft connection to obtain the target encoding features of each soft connection; each soft connection consists of at least the target polyphonic character or at least two consecutive characters containing the target polyphonic character in the target text, and the maximum number of characters contained in each soft connection is a preset number.

[0063] The target text contains at least one polyphonic character, and the target polyphonic character is any polyphonic character in the target text.

[0064] Optionally, in some scenarios, such as in a question-and-answer scenario, the target polyphonic character can be specified by the user. For example, the user gives a sentence and asks what the pronunciation of a certain polyphonic character in the sentence is.

[0065] Optionally, in some scenarios, such as in a speech synthesis scenario, the target polyphonic character can be automatically recognized. For example, for each character in the target text, it can be searched in a preset set of polyphonic characters (which can also be called a polyphonic character dictionary) to determine whether the character exists. If it exists, the character is determined to be a polyphonic character; otherwise, the character is determined not to be a polyphonic character.

[0066] Among the multiple soft connections corresponding to the target polyphonic character, the number of characters contained in different soft connections may be the same or different.

[0067] Among different soft connections containing the same number of characters, the position of the target polyphonic character in the soft connection is different.

[0068] The soft connection with the least number of characters is the soft connection containing 1 character (i.e., the target polyphonic character), and the soft connection with the most number of characters is the soft connection containing N (i.e., the preset number of characters).

[0069] The target polyphonic character corresponds to n soft connections containing n (n = 1, 2, 3,..., N) characters. Based on this, the total number M of soft connections corresponding to the target polyphonic character is:

[0070] M = 1 + 2 + 3 +... + N = N(N + 1) / 2

[0071] When the number of characters n included in the soft link is greater than 1, one of the n characters is the target polyphonic character. For the other n - 1 characters, the n - 1 characters can be n - 1 consecutive characters adjacent to the target polyphonic character in the target text; or, the first k (k is a positive integer less than n - 1) characters of the n - 1 characters are k characters adjacent to and before the target polyphonic character, and the last n - 1 - k characters are n - 1 - k characters adjacent to and after the target polyphonic character; or, the last k characters of the n - 1 characters are k consecutive characters in the target text that are adjacent to and after the target polyphonic character, and the first n - 1 - k characters are preset information representing blanks, that is to say, the target polyphonic character is the first character in the target text; or, the first k characters of the n - 1 characters are k consecutive characters in the target text that are adjacent to and before the target polyphonic character, and the last n - 1 - k characters are preset information representing blanks, that is to say, the target polyphonic character is the last character in the target text; or, the last k characters of the n - 1 characters are k consecutive characters in the target text that are adjacent to and after the target polyphonic character, and the first n - 1 - k characters are all the characters in the target text that are adjacent to and before the target polyphonic character and preset information representing blanks; or, the first k characters of the n - 1 characters are k consecutive characters in the target text that are adjacent to and before the target polyphonic character, and the last n - 1 - k characters are all the characters in the target text that are adjacent to and after the target polyphonic character and preset information representing blanks.

[0072] Among the n soft links each containing n characters, the positions of the target polyphonic characters are different in different soft links.

[0073] As an example, N = 4. Of course, N can also take other values, such as N = 3, or N = 5, N = 6, etc.

[0074] As Figure 2 shown, it is an example of the soft link provided by the embodiment of the present application. In this example, the target text consists of 9 characters from t1 to t9. Among them, t5 is a polyphonic character, and the soft link corresponding to this polyphonic character contains at most 4 characters. Based on this, each polyphonic character corresponds to 10 soft links.

[0075] When the present application encodes each soft link, it considers the association relationship between the soft link and the target text, so that the encoding features of each soft link carry the information of the soft link itself and the association relationship between the soft link and the target text (such as semantic relevance).

[0076] Step S102: Classify the target polyphonic characters in the soft link based on the target encoding features of each soft link to obtain the probability distribution of the above-mentioned target polyphonic characters corresponding to the soft link; the probability distribution is the probability that the target polyphonic character belongs to each pronunciation among several pronunciations, and the probability distribution represents the pronunciation of the target polyphonic character in the soft link; the several pronunciations are all the pronunciations of all the polyphonic characters in the preset polyphonic character set.

[0077] For the i-th (i = 1, 2, 3, ……, M) soft link corresponding to the target polyphonic character, classify the target polyphonic character in the i-th soft link based on the target encoding features of the i-th soft link to obtain the probability distribution of the target polyphonic character corresponding to the i-th soft link (denoted as the i-th probability distribution of the target polyphonic character), that is, the i-th probability distribution of the target polyphonic character. Obviously, this application obtains M probability distributions corresponding to the target polyphonic character.

[0078] Step S103: Vote based on the probability distributions of the target polyphonic characters corresponding to each soft link to determine the pronunciation of the target polyphonic character in the target text (for the convenience of narration and distinction, denoted as the target pronunciation).

[0079] Soft voting can be performed based on the probability distributions of the target polyphonic character to determine the target pronunciation of the target polyphonic character in the target text.

[0080] Alternatively, hard voting can be performed based on the probability distributions of the target polyphonic character to determine the target pronunciation of the target polyphonic character in the target text.

[0081] Alternatively, hard voting and soft voting can be performed based on the probability distributions of the target polyphonic character to determine the target pronunciation of the target polyphonic character in the target text.

[0082] The polyphonic character disambiguation method provided by this application proposes the concept of a soft link. Without relying on a word segmentation tool, it extracts the word combinations formed by the polyphonic character and its context before and after as soft links. Based on the association relationship between each soft link corresponding to the target polyphonic character and the target text, it determines the target encoding features of each soft link corresponding to the target polyphonic character. Based on the target encoding features of each soft link, it classifies the target polyphonic character in each soft link respectively to obtain multiple possible classification results of the target polyphonic character. Based on each classification result, it votes to determine the pronunciation of the target polyphonic character in the target text. This process does not require word segmentation of the target text, will not cause polyphonic character disambiguation errors due to word segmentation errors, nor does it need to use the explanatory information of the word segmentation containing polyphonic characters, and will not cause polyphonic character disambiguation errors due to incorrect explanatory information. Thus, it avoids the dependence of polyphonic character disambiguation on word segmentation and explanatory information, and improves the accuracy of polyphonic character disambiguation.

[0083] In addition, in the polyphonic word disambiguation solution based on the polyphonic word dictionary, as the number of words containing polyphonic characters in the polyphonic word dictionary increases, the search efficiency of the entries (i.e., the word segmentation containing polyphonic characters) will be reduced, and ultimately directly affect the efficiency of the polyphonic word disambiguation task. However, this application does not require a polyphonic word dictionary. Therefore, it will not be affected by the search efficiency of the entries, and can improve the accuracy of polyphonic word disambiguation while ensuring the execution efficiency of the polyphonic word disambiguation task.

[0084] In an optional embodiment, a flowchart of an implementation of encoding each soft connection based on the association relationship between multiple soft connections corresponding to the target polyphonic character and the target text to obtain the target encoding feature of each soft connection is as Figure 3 shown, and may include:

[0085] Step S301: Encode each character in the target text to obtain the encoding feature of each character.

[0086] As an example, an encoding network can be used to encode each character in the target text to obtain the encoding feature of each character. The encoding feature of each character is a word embedding representation that integrates context information. The structure of the encoding network here can adopt the network structure of the encoding module in a pre-trained language model, where the language model can include but is not limited to any of the following language models:

[0087] BERT (Bidirectional Encoder Representation from Transformers) model, ELMo (Embeddings from Language Models) model, GPT (Generative Pre-trained Transformer) model, T5 (Text-to-Text Transfer Transformer) model, etc.

[0088] Step S302: For each soft connection corresponding to the target polyphonic character, fuse the encoding features of each character included in the soft connection to obtain the initial encoding feature of the soft connection.

[0089] In the case where the soft connection contains preset information representing blanks, the encoding feature of the preset information is a preset encoding feature, such as a zero vector.

[0090] Optionally, for the i-th soft connection, the encoding features of each character included in the i-th soft connection can be weighted and summed to obtain the initial encoding feature of the i-th soft connection. The initial encoding feature of each soft connection is a vector.

[0091] As an example, the sum of the weights of each character in the i-th soft connection is 1, and the weights of each character are the same, that is, the encoding features of each character included in the i-th soft connection are averaged to obtain the initial encoding feature of the i-th soft connection.

[0092] As an example, the sum of the weights of each character in the i-th soft connection is 1, and the weights of different characters may be the same or different. For example, the weight of the target polyphonic character is the largest, and the farther away from the target polyphonic character, the smaller the weight.

[0093] Step S303: Based on the cross-attention mechanism, fuse the initial encoding feature of the soft connection with the encoding features of each character in the target text to obtain the target encoding feature of the soft connection.

[0094] That is to say, the target encoding feature of the i-th soft connection not only contains the information of each character in the i-th soft connection, but also includes the information of each character in the target text, as well as the association relationship between the i-th soft connection and each character in the target text.

[0095] The initial encoding feature of the i-th soft connection can be fused with the encoding features of each character in the target text based on the attention weights of the i-th soft connection to each character in the target text to obtain the target encoding feature of the soft connection.

[0096] Optionally, the initial encoding feature of the i-th soft connection can be fused with the encoding features of each character in the target text in the following way to obtain the target encoding feature of the i-th soft connection:

[0097] Use the initial encoding feature of the i-th soft connection and the encoding features of each character in the target text to calculate the attention weights of the i-th soft connection to each character in the target text.

[0098] For the j-th (j = 1, 2, 3, ……, J; J is the total number of characters in the target text) character in the target text, the attention weight of the i-th soft connection to the j-th character can be calculated based on the initial encoding feature of the i-th soft connection and the encoding feature of the j-th character. As an example, the similarity between the initial encoding feature of the i-th soft connection and the encoding feature of the j-th character can be calculated, and after shrinking and normalizing the similarity, the attention weight of the i-th soft connection to the j-th character in the target text can be obtained.

[0099] The target encoding feature of the soft connection and the encoding feature of the character are vectors of the same length. The similarity between the initial encoding feature of the i-th soft connection and the encoding feature of the j-th character can be multiplied by a reduction factor (a positive number less than 1) to reduce the similarity. The reduction factor can be determined based on the length of the target encoding feature of the soft connection. As an example, the reduction factor can be: the reciprocal of the square root of the length of the target encoding feature. The softmax function can be used to normalize the reduced similarity to obtain the attention weight of the i-th soft connection to the j-th character in the target text.

[0100] Based on the attention weights of the i-th soft connection to each character in the target text, the encoding features of each character in the target text are weighted and summed to obtain the target encoding feature of this soft connection.

[0101] The weight of the encoding feature of the j-th character is the attention weight of the i-th soft connection to the j-th character.

[0102] In an optional embodiment, when voting based on the probability distribution of the target polyphonic character corresponding to each soft connection, soft voting can be used for voting, hard voting can be used for voting, or the two voting methods can be used in combination.

[0103] Optionally, in the case of soft voting, a flowchart of an implementation of voting based on the probability distribution of the target polyphonic character corresponding to each soft connection is as Figure 4 shown, and may include:

[0104] Step S401: Integrate the probabilities that the target polyphonic character belongs to the same pronunciation in the probability distributions of the target polyphonic character corresponding to each soft connection to obtain the target probability that the target polyphonic character belongs to each pronunciation.

[0105] Suppose the target polyphonic character corresponds to M soft connections, and the number of all pronunciations of all polyphonic characters in the preset polyphonic character set is D. Then, for each of the M soft connections, a probability distribution of the target polyphonic character is obtained, so a total of M probability distributions are obtained for the target polyphonic character. The i-th probability distribution (i = 1, 2, 3,..., M) is the probability that the target polyphonic character corresponding to the i-th soft connection belongs to each of the D pronunciations. That is, for the target polyphonic character and the d-th (d = 1, 2, 3,..., D) pronunciation among the D pronunciations, this application obtains M probabilities that the target polyphonic character belongs to the d-th pronunciation.

[0106] This application integrates the M probabilities corresponding to the d-th pronunciation in the M probability distributions to obtain the target probability that the target polyphonic character belongs to the d-th pronunciation. Optionally, the M probabilities corresponding to the d-th pronunciation can be averaged to obtain the target probability that the target polyphonic character belongs to the d-th pronunciation.

[0107] Step S402: Determine the pronunciation of the target polyphone in the target text as the pronunciation corresponding to the maximum target probability.

[0108] Optionally, in the case of hard voting, another implementation flowchart of the above voting based on the probability distributions of the target polyphones corresponding to each soft connection is as Figure 5 shown, and it may include:

[0109] Step S501: Based on the probability distribution of the target polyphone corresponding to each soft connection, determine the pronunciation of the target polyphone in each soft connection respectively.

[0110] For the i-th probability distribution among the M probability distributions, the pronunciation corresponding to the maximum probability in the i-th probability distribution can be determined as a pronunciation of the target polyphone (i.e., the pronunciation of the target polyphone in the i-th soft connection, which can be denoted as the i-th pronunciation of the target polyphone), and then M pronunciations are determined for the target polyphone.

[0111] Step S502: Statistically analyze the pronunciations of the target polyphone in each soft connection, and determine the pronunciation with the largest number as the pronunciation of the target polyphone in the target text.

[0112] Statistically analyze the determined M pronunciations to determine the number of the same pronunciations (i.e., count the number of votes for each pronunciation of the determined target polyphone), and determine the pronunciation with the largest number as the pronunciation of the target polyphone in the target text.

[0113] Optionally, in some cases, it is possible that the number of votes for different pronunciations of the target polyphone is the same. For example, when the target polyphone has two pronunciations and the target polyphone corresponds to 10 soft connections, it is possible that the number of votes for both pronunciations is 5 or both are 4, etc. In this case, the maximum probabilities corresponding to the two pronunciations with the most votes can be compared, and the pronunciation corresponding to the larger probability among the two maximum probabilities can be determined as the pronunciation of the target polyphone in the target text.

[0114] Optionally, when the number of votes for different pronunciations of the target polyphone is the same, a soft voting method can be used to determine the pronunciation of the target polyphone in the target text. That is, fuse the probabilities of the target polyphone belonging to the same pronunciation in the probability distributions of the target polyphone corresponding to each soft connection to obtain the target probabilities of the target polyphone belonging to each pronunciation; determine the pronunciation corresponding to the maximum target probability as the pronunciation of the target polyphone in the target text.

[0115] In an optional embodiment, the above process of encoding each character in the target text to obtain the encoding features of each character, fusing the initial encoding features of each soft connection corresponding to the target polyphonic character with the encoding features of each character in the target text based on the cross-attention mechanism to obtain the target encoding features of the soft connection, and classifying the encoding features of each soft connection can be implemented by a classification model. For example, Figure 6 As shown in

[0116] Figure Figure 6 , which is a schematic structural diagram of a classification model provided by an embodiment of the present application, may include:

[0117] An encoding module 601, a connection module 602, an attention module 603, and a classification module 604;

[0118] Among them, the encoding module 601 is configured to: encode each character in the target text to obtain the encoding features of each character.

[0119] The connection module 602 is configured to: for each soft connection corresponding to the target polyphonic character, fuse the encoding features of each character included in the soft connection to obtain the initial encoding features of the soft connection.

[0120] The attention module 603 is configured to: for each soft connection, fuse the initial encoding features of the soft connection with the encoding features of each character in the target text based on the cross-attention mechanism to obtain the target encoding features of the soft connection.

[0121] Optionally, the classification module 604 may adopt a deep neural network. As an example, the classification module 604 includes a plurality of fully connected layers.

[0122] Optionally, the classification model can be trained in the following manner:

[0123] Encode each character in the text sample through the encoding module 601 to obtain the encoding features of each character.

[0124] For each polyphonic character in the text sample, obtain the initial encoding features of multiple soft connections corresponding to the polyphonic character through the connection module 602; wherein, each soft connection is composed of at least the polyphonic character or at least two consecutive characters including the polyphonic character in the text sample, and the maximum number of characters included in each soft connection is a preset number; the initial encoding features of each soft connection are obtained by fusing the encoding features of each character included in the soft connection;

[0125] For each soft connection through the attention module 603, based on the cross-attention mechanism, the initial encoded features of the soft connection are fused with the encoded features of each word in the text sample to obtain the target encoded features of the soft connection;

[0126] Through the classification module 604, based on the target encoded features of each soft connection, the polyphone in the soft connection is classified to obtain the probability distribution of the polyphone corresponding to the software connection;

[0127] Taking the pronunciation characterized by the probability distribution of the polyphone corresponding to each soft connection in the text sample approaching the pronunciation label of the polyphone corresponding to the soft connection, and the pronunciation of the same polyphone determined by voting based on the probability distributions of the same polyphone in the corresponding text sample approaching the pronunciation label of the same polyphone in the text sample as the goal, the parameters of the classification model are updated. That is to say, the pronunciation of each polyphone corresponding to each soft connection in the text sample can be determined (for the convenience of narration and distinction, denoted as the first pronunciation); for each polyphone in the text sample, multiple probability distributions of the polyphone determined based on the corresponding soft connections of the polyphone are used for voting to determine the pronunciation of each polyphone in the text sample (for the convenience of narration and distinction, denoted as the second pronunciation); taking the first pronunciation and the second pronunciation of each polyphone approaching the pronunciation label of the polyphone as the goal, the parameters of the classification model are updated.

[0128] As can be seen from the foregoing, each polyphone corresponds to M soft connections, so M probability distributions are obtained for the corresponding polyphone. In this application, taking the pronunciation characterized by each probability distribution approaching the pronunciation label of the polyphone, and the pronunciation of the polyphone determined by voting based on M probability distributions approaching the pronunciation label of the polyphone as the goal, the parameters of the classification model are updated.

[0129] As an example, for the qth (q = 1, 2, 3,..., Q; Q is the number of polyphones in the text sample) polyphone in the text sample, the cross-entropy loss between the ith probability distribution of the qth polyphone and the pronunciation label of the qth polyphone can be calculated, and the mean value of the obtained M cross-entropy losses is calculated to obtain the first loss of the qth polyphone; based on the M probability distributions of the qth polyphone, soft voting is performed (for example, the mean value of M probability distributions is calculated) to obtain the target probability distribution of the qth polyphone, and the cross-entropy loss between the target probability distribution of the qth polyphone and the pronunciation label of the qth polyphone is calculated as the second loss. The first loss and the second loss are weighted and summed to obtain the comprehensive loss corresponding to the qth polyphone. Taking the comprehensive loss as the goal of becoming smaller (that is, the comprehensive loss calculated after updating the parameters of the classification model is smaller than the comprehensive loss calculated before updating the parameters of the classification model), the parameters of the classification model are updated.

[0130] In the above example, for each polyphonic character, two classification losses based on a voting mechanism are adopted as the comprehensive loss of the polyphonic character. Among them, the first loss is the loss based on the hard voting mechanism, and the second loss is the loss based on the soft voting mechanism. Based on this, the robustness and classification accuracy of the classification model are improved.

[0131] Among them, the weights of the first loss and the second loss can be the same or different. The sum of the weight of the first loss and the weight of the second loss is 1.

[0132] In the case where the classification module 604 includes multiple fully connected layers, when training the classification model, some of the fully connected layers are dropout layers. That is, during the training process of the classification model, the outputs of some neurons in the dropout layer are randomly set to zero, thereby reducing the dependence relationship between neurons. During the training process, each neuron in the dropout layer has a certain probability of being "dropped out" (i.e., set to zero), that is, it does not participate in the forward propagation and backward propagation in the current training batch. This can prevent the classification model from overly relying on certain specific neurons and improve the generalization ability of the classification model.

[0133] It should be noted that the dropout layer usually only "drops out" the outputs of some neurons during the training process. After the training is completed, in the prediction stage, the outputs of all neurons will participate in the calculation.

[0134] Corresponding to the method embodiment, the present application also provides a polyphonic character disambiguation device. A schematic structural diagram of the polyphonic character disambiguation device provided by the present application is as Figure 7 shown, and may include:

[0135] A soft connection unit 701, a classification unit 702, and a voting unit 703;

[0136] Among them, the soft connection unit 701 is used to encode each of the soft connections for the target polyphonic character in the target text based on the association relationship between the multiple soft connections corresponding to the target polyphonic character and the target text, to obtain the target encoding features of each of the soft connections; each of the soft connections is at least composed of the target polyphonic character or at least composed of at least two consecutive characters in the target text that contain the target polyphonic character, and the maximum number of characters included in each soft connection is a preset number;

[0137] The classification unit 702 is used to classify the target polyphonic characters in each of the soft connections based on the target encoding features of each soft connection, so as to obtain the probability distribution of the target polyphonic characters corresponding to each soft connection; the probability distribution is the probability that the target polyphonic character belongs to each of several pronunciations, which represents the pronunciation of the target polyphonic character in the soft connection; the several pronunciations are all the pronunciations of all the polyphonic characters in the preset polyphonic character set.

[0138] The voting unit 703 is used to vote based on the probability distributions of the target polyphonic characters corresponding to each soft connection to determine the pronunciation of the target polyphonic character in the target text.

[0139] The polyphonic character disambiguation device provided by the embodiment of the present application proposes the concept of a soft connection, extracts the word combination composed of a polyphonic character and its context before and after without relying on a word segmentation tool as a soft connection, determines the target encoding features of each soft connection corresponding to the target polyphonic character based on the association relationship between each soft connection corresponding to the target polyphonic character and the target text, classifies the target polyphonic characters in each soft connection respectively based on the target encoding features of each soft connection to obtain multiple possible classification results of the target polyphonic character, and votes based on each classification result to determine the pronunciation of the target polyphonic character in the target text. This process does not require word segmentation of the target text, will not cause polyphonic character disambiguation errors due to word segmentation errors, nor does it need to use the explanatory information of word segmentation containing polyphonic characters, and will not cause polyphonic character disambiguation errors due to incorrect explanatory information, thus avoiding the dependence of polyphonic character disambiguation on word segmentation and explanatory information and improving the accuracy of polyphonic character disambiguation.

[0140] In addition, in the polyphonic character disambiguation scheme based on a polyphonic word dictionary, as the number of polyphonic words in the polyphonic word dictionary increases, the efficiency of looking up entries will be reduced, and ultimately directly affect the efficiency of the polyphonic character disambiguation task. And the present application does not require a polyphonic word dictionary. Therefore, while ensuring the execution efficiency of the polyphonic character disambiguation task, the accuracy of polyphonic character disambiguation can be improved.

[0141] In an optional embodiment, when the soft connection unit 701 encodes each soft connection based on the association relationship between the multiple soft connections corresponding to the target polyphonic character and the target text, it is used for:

[0142] Encode each character in the target text to obtain the encoding features of each character;

[0143] For each soft connection corresponding to the target polyphonic character, fuse the encoding features of each character included in the soft connection to obtain the initial encoding feature of the soft connection;

[0144] Based on the cross-attention mechanism, fuse the initial encoded features of the soft connection with the encoded features of each character in the target text to obtain the target encoded features of the soft connection.

[0145] In an optional embodiment, when the soft connection unit 701 fuses the encoded features of each character included in the soft connection to obtain the initial encoded features of the soft connection, it is used for:

[0146] Perform a weighted sum of the encoded features of each character included in the soft connection to obtain the initial encoded features of the soft connection.

[0147] In an optional embodiment, when the soft connection unit 701 fuses the initial encoded features of the soft connection with the encoded features of each character in the target text based on the cross-attention mechanism to obtain the target encoded features of the soft connection, it is used for:

[0148] Use the initial encoded features of the soft connection and the encoded features of each character in the target text to calculate the attention weights of the soft connection to each character in the target text;

[0149] Based on the attention weights of the soft connection to each character in the target text, perform a weighted sum of the encoded features of each character in the target text to obtain the target encoded features of the soft connection.

[0150] In an optional embodiment, when the voting unit 703 votes based on the probability distributions of the target polyphonic characters corresponding to each soft connection, it is used for:

[0151] Fuse the probabilities in the probability distributions of the target polyphonic characters corresponding to each soft connection where the target polyphonic characters belong to the same pronunciation to obtain the target probabilities of the target polyphonic characters belonging to each pronunciation;

[0152] Determine the pronunciation of the target polyphonic character in the target text as the pronunciation corresponding to the maximum target probability.

[0153] In an optional embodiment, when the voting unit 703 votes based on the probability distributions of the target polyphonic characters corresponding to each soft connection, it is used for:

[0154] Based on the probability distribution of the target polyphonic character corresponding to each soft connection, determine the pronunciation of the target polyphonic character in each soft connection;

[0155] Statistically analyze the pronunciations of the target polyphonic character in each soft connection, and determine the pronunciation with the largest quantity as the pronunciation of the target polyphonic character in the target text.

[0156] In an optional embodiment, the voting unit 703 is further used for:

[0157] If the number of pronunciations of the target polyphonic character is the same, determine the pronunciation corresponding to the highest probability as the pronunciation of the target polyphonic character in the target text;

[0158] Alternatively, if the number of pronunciations of the target polyphonic character is the same, fuse the probabilities of the target polyphonic character belonging to the same pronunciation in the probability distributions of the target polyphonic character corresponding to each soft connection to obtain the target probabilities of the target polyphonic character belonging to each pronunciation; determine the pronunciation corresponding to the highest target probability as the pronunciation of the target polyphonic character in the target text.

[0159] In an optional embodiment, the soft connection unit 701 encodes each character in the target text to obtain the encoding features of each character. For each soft connection corresponding to the target polyphonic character, based on the cross-attention mechanism, fuse the initial encoding features of the soft connection with the encoding features of each character in the target text to obtain the target encoding features of the soft connection. The process of the classification unit 702 classifying the target encoding features of each soft connection includes:

[0160] The soft connection unit 701 encodes each character in the target text through the encoding module of the classification model to obtain the encoding features of each character;

[0161] The soft connection unit 701 fuses the encoding features of each character included in each soft connection corresponding to the target polyphonic character through the connection module of the classification model to obtain the initial encoding features of the soft connection;

[0162] The soft connection unit 701 fuses the initial encoding features of each soft connection with the encoding features of each character in the target text based on the cross-attention mechanism through the attention module of the classification model to obtain the target encoding features of the soft connection;

[0163] The classification unit 702 classifies the target polyphonic character in each soft connection based on the target encoding features of each soft connection through the classification module of the classification model to obtain the probability distribution of the target polyphonic character corresponding to each software connection.

[0164] In an optional embodiment, the polyphonic character disambiguation device further includes a training module for training the classification model. When the training module trains the classification model, it is used for:

[0165] Encode each character in the text sample through the encoding module to obtain the encoding features of each character;

[0166] For each polyphonic character in the text sample through the connection module, obtain the initial encoding features of multiple soft connections corresponding to the polyphonic character; wherein, each soft connection is at least composed of the polyphonic character or at least composed of at least two consecutive characters in the text sample that contain the polyphonic character, and the maximum number of characters included in each soft connection is a preset number; the initial encoding feature of each soft connection is obtained by fusing the encoding features of each character included in the soft connection;

[0167] For each soft connection through the attention module, based on the cross-attention mechanism, fuse the initial encoding feature of the soft connection with the encoding features of each character in the text sample to obtain the target encoding feature of the soft connection;

[0168] Through the classification module, classify the high-frequency polyphonic characters in each soft connection based on the target encoding feature of each soft connection to obtain the probability distribution of the polyphonic character corresponding to the software connection;

[0169] Based on the probability distribution of each polyphonic character corresponding to each soft connection in the text sample, determine the first pronunciation of each polyphonic character corresponding to the soft connection;

[0170] For each polyphonic character in the text sample, use the multiple probability distributions of the polyphonic character determined based on the respective soft connections corresponding to the polyphonic character to vote to determine the second pronunciation of each polyphonic character in the text sample;

[0171] With the goal that the first pronunciation and the second pronunciation of each polyphonic character both approach the pronunciation label of the polyphonic character, update the parameters of the classification model.

[0172] An electronic device is also provided in an embodiment of the present application. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application can be a terminal device (such as, a car machine, a large screen device, a smart home, a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc.), or a server (it can be a single server, or a server cluster, or, it can be a cloud server, etc.). Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present application.

[0173] As Figure 8As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0174] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a memory card, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0175] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any of the polyphone disambiguation methods provided by the embodiments of the present application.

[0176] In an embodiment of the present application, there is also provided a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any of the polyphone disambiguation methods provided by the embodiments of the present application.

[0177] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which may be specifically implemented as one or more communication buses or signal lines.

[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0179] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. Professional technicians can use different methods to implement the described functions for each specific solution, but such implementation should not be considered to exceed the scope of this application.

[0180] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device or data center to another website, computer, training device or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0181] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0182] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for disambiguating polyphones, characterized in that: include: For a target polyphone in a target text, encoding each of the soft links based on the association between the multiple soft links corresponding to the target polyphone and the target text, and obtaining a target encoding feature of each of the soft links; each of the soft links is at least composed of the target polyphone or at least composed of at least two consecutive words in the target text that contain the target polyphone, and the maximum number of words contained in each soft link is a preset number; Classifying the target polyphone in each soft connection based on the target coding feature of each soft connection to obtain a probability distribution of the target polyphone corresponding to each soft connection; The probability distribution is the probability that the target polyphone belongs to each of the plurality of pronunciations, representing the pronunciation of the target polyphone in the soft connection; the plurality of pronunciations are all the pronunciations of all the polyphones in the preset polyphone set; Voting is performed based on the probability distribution of the target polyphonetic character corresponding to each of the soft links to determine the pronunciation of the target polyphonetic character in the target text.

2. The method according to claim 1, characterized in that The encoding of each of the soft links based on the association relationship between the multiple soft links corresponding to the target polyphone and the target text includes: Encoding each word in the target text to obtain encoding features of each word; For each soft connection corresponding to the target polyphone, the coding features of each character contained in the soft connection are merged to obtain an initial coding feature of the soft connection; Based on the cross attention mechanism, the initial encoding features of the soft connection are fused with the encoding features of each word in the target text to obtain the target encoding features of the soft connection.

3. The method according to claim 2, characterized in that The method of fusing the initial coding features of the soft connection with the coding features of each word in the target text based on the cross attention mechanism to obtain the target coding features of the soft connection includes: Using the initial coding features of the soft connection and the coding features of each word in the target text, calculating the attention weight of the soft connection to each word in the target text; Based on the attention weight of the soft connection to each word in the target text, the encoding features of each word in the target text are weighted and summed to obtain the target encoding features of the soft connection.

4. The method according to claim 1, characterized in that The voting based on the probability distribution of the target polyphone corresponding to each of the soft links includes: Merging the probabilities of the target polyphone belonging to the same pronunciation in the probability distribution of the target polyphone corresponding to each of the soft links to obtain the target probability of the target polyphone belonging to each pronunciation; The pronunciation corresponding to the maximum target probability is determined as the pronunciation of the target polyphone in the target text.

5. The method according to claim 1, characterized in that The voting based on the probability distribution of the target polyphone corresponding to each of the soft links includes: Determine the pronunciation of the target polyphonetic character in each of the soft links based on the probability distribution of the target polyphonetic character corresponding to each of the soft links; The pronunciations of the target polyphone in each of the soft links are counted, and the pronunciation with the largest number is determined as the pronunciation of the target polyphone in the target text.

6. The method according to claim 5, characterized in that Also includes: If the number of the multiple pronunciations of the target polyphone is the same, determining the pronunciation corresponding to the maximum probability as the pronunciation of the target polyphone in the target text; Alternatively, if the number of multiple pronunciations of the target polyphone is the same, the various probabilities of the target polyphone belonging to the same pronunciation in the probability distribution of the target polyphone corresponding to each soft link are merged to obtain the target probability of the target polyphone belonging to each pronunciation; and the pronunciation corresponding to the maximum target probability is determined as the pronunciation of the target polyphone in the target text.

7. The method according to claim 2, characterized in that The process of encoding each character in the target text to obtain encoding features of each character, fusing the initial encoding features of the soft connection with the encoding features of each character in the target text based on a cross attention mechanism for each soft connection corresponding to the target polyphone, and obtaining a target encoding feature of the soft connection, and classifying the target encoding feature of each soft connection includes: Encoding each word in the target text by using an encoding module of a classification model to obtain encoding features of each word; For each soft connection corresponding to the target polyphone, the connection module of the classification model fuses the coding features of each character contained in the soft connection to obtain an initial coding feature of the soft connection; For each soft link, the attention module of the classification model fuses the initial coding features of the soft link with the coding features of each word in the target text based on a cross attention mechanism to obtain a target coding feature of the soft link; The target polyphone in each soft connection is classified based on the target coding feature of each soft connection by the classification module of the classification model to obtain the probability distribution of the target polyphone corresponding to each soft connection.

8. The method according to claim 7, characterized in that The classification model is trained in the following way: Encoding each word in the text sample by the encoding module to obtain encoding features of each word; For each polyphone in the text sample, the connection module obtains initial coding features of multiple soft connections corresponding to the polyphone; wherein each soft connection is at least composed of the polyphone or at least composed of at least two consecutive characters in the text sample that contain the polyphone, and the maximum number of characters contained in each soft connection is a preset number; the initial coding feature of each soft connection is obtained by fusing the coding features of each character contained in the soft connection; For each soft connection, the attention module fuses the initial coding features of the soft connection with the coding features of each word in the text sample based on a cross attention mechanism to obtain a target coding feature of the soft connection; The classification module classifies the polyphone in each soft connection based on the target coding feature of each soft connection to obtain a probability distribution of the polyphone corresponding to the soft connection; Determine the first pronunciation of each polyphonic character corresponding to each soft link in the text sample based on the probability distribution of each polyphonic character corresponding to the soft link; For each polyphone in the text sample, voting is performed using multiple probability distributions of the polyphone determined based on each soft connection corresponding to the polyphone, to determine a second pronunciation of each polyphone in the text sample; With the goal of making the first pronunciation and the second pronunciation of each polyphonetic character close to the pronunciation label of the polyphonetic character, the parameters of the classification model are updated.

9. A polyphonetic character disambiguation device, characterized in that: include: A soft connection unit is used to encode each soft connection for a target polyphone in a target text based on the association relationship between a plurality of soft connections corresponding to the target polyphone and the target text, so as to obtain a target encoding feature of each soft connection; each soft connection is at least composed of the target polyphone or at least composed of at least two consecutive characters in the target text that contain the target polyphone, and a maximum number of characters contained in each soft connection is a preset number; A classification unit, configured to classify the target polyphone in each soft connection based on the target coding feature of each soft connection, and obtain a probability distribution of the target polyphone corresponding to each soft connection; The probability distribution is the probability that the target polyphone belongs to each of the plurality of pronunciations, representing the pronunciation of the target polyphone in the soft connection; the plurality of pronunciations are all the pronunciations of all the polyphones in the preset polyphone set; The voting unit is used to vote based on the probability distribution of the target polyphone corresponding to each of the soft links to determine the pronunciation of the target polyphone in the target text.

10. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the polyphone disambiguation method as claimed in any one of claims 1 to 8.

11. An electronic device, characterized in that: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the polyphone disambiguation method as described in any one of claims 1 to 8.

12. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the polyphone disambiguation method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method, device and system for determining pronunciation of polyphone

    CN107515850A

  • Polyphone recognition method and device, readable medium and electronic equipment

    CN111798834A

  • Text processing method and device

    CN112257420A

  • Polyphone pronunciation determination method, device, electronic equipment and storage medium

    CN112818657A

  • Polyphone pronunciation labeling method and device, equipment and storage medium

    CN113268974A