A method, device, equipment and storage medium for screening polyphonetic characters to be marked

By generating and traversing the string dictionary, filtering out the comprehensive context of polyphonic characters, solving the problems of low multiphonic character labeling efficiency and low model prediction accuracy in the prior art, and achieving efficient manual labeling and polyphonic character pinyin prediction.

CN112201221BActive Publication Date: 2025-05-16GUANGZHOU DUOYI NETWORK TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011037697.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-28
Publication Date
2025-05-16
Estimated Expiration
2040-09-28

AI Technical Summary

Technical Problem

In the prior art, the corpus labeling efficiency of polyphonic characters is low, and due to the similar context environment of polyphonic characters in the corpus, the accuracy of the model prediction of pinyin is not high.

Method used

By obtaining the original text corpus, generating a Chinese character string dictionary and a string text dictionary, looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary, generating a candidate text list, selecting Chinese characters to be marked, recording text information, and generating an output text list to filter out polyphonic Chinese characters to be marked with a comprehensive coverage and a small number of polyphonic Chinese characters to be marked.

Benefits of technology

It has realized the selection of a small number of polyphonic characters to be tagged from a large number of text corpuses, which has improved the efficiency and value of manual annotation, and enhanced the prediction accuracy of the polyphonic characters' pinyin disambiguation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112201221B_ABST
    Figure CN112201221B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and storage medium for screening polyphonic characters to be marked, including: obtaining original text corpus; generating a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record a list of Chinese characters mapped to all strings containing the Chinese characters, and the string text dictionary is used to record a list of strings mapped to all texts containing the strings; looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary so that the number of texts reaches a preset value, and generating a candidate text list; selecting Chinese characters to be marked, obtaining a list of texts to be marked through the candidate text list; and recording the information of each text in the list of texts to be marked in turn to obtain an output text list. The present invention can collect original text corpus with comprehensive subject matter types, and ensure that the text corpus covers the subject matter types and language styles comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and medium for screening polyphonetic characters to be marked. Background Art

[0002] Grapheme-to-phonetic conversion is an essential module of the text-to-speech system (TTS), and its accuracy directly affects the intelligibility of the text-to-speech system. In the Chinese text-to-speech system, the task of grapheme-to-phonetic conversion is to convert a text sequence into a corresponding pinyin sequence. Some Chinese characters correspond to multiple pinyins, and the key and difficulty of grapheme-to-phonetic conversion is how to solve this problem of multiple pinyins for one character. In most cases, grapheme-to-phonetic conversion is to search the current word in the dictionary and match it with the corresponding pinyin. However, many polyphonetic characters are single words, or the words where polyphonetic characters are located are not in the pinyin dictionary. It is necessary to train the model to learn the pronunciation rules of polyphonetic characters in order to make correct pinyin predictions. However, there are currently few corpora with pinyin annotated at the sentence level, and the contextual environments of polyphonetic characters in the corpus are often the same, which results in the inability to fully cover the language environment of polyphonetic characters, resulting in a low accuracy rate in predicting pinyin for the model trained with the above corpus.

[0003] With the development of the Internet industry, a large amount of text data is generated, but human resources are limited, so the amount of data that can be manually annotated is limited. If a large amount of data is annotated, the language phenomena covering polyphones are relatively comprehensive, but it takes a lot of manpower and time. If a part of the data is selected from a large amount of data, so that this part of the data covers the language phenomena of the original data, then only this small amount of data can achieve good results, but there is currently no method to select a small amount of data from a large amount of data to annotate the pronunciation of polyphones. Summary of the invention

[0004] The technical problem to be solved by the embodiments of the present invention is to provide a method, device, equipment and medium for screening polyphone corpus to be marked, which can collect original text corpus with comprehensive subject matter types and ensure that the text corpus covers comprehensive subject matter types and language styles.

[0005] In order to achieve the above object, an embodiment of the present invention provides a method for screening polyphonetic characters to be marked, comprising:

[0006] Acquire an original text corpus; wherein the original text corpus includes at least one polyphonetic character;

[0007] Generate a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record the mapping of Chinese characters to a list of all strings containing the Chinese characters, and the string text dictionary is used to record the mapping of strings to a list of all texts containing the strings;

[0008] Looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary so that the number of texts reaches a preset value, and generating a candidate text list;

[0009] Select the Chinese characters to be marked, and obtain a list of texts to be marked through the candidate text list;

[0010] The information of each text is recorded in turn from the list of texts to be marked. After examining a text, the number of texts currently containing preset polyphones, the number of polyphones and the Chinese characters of polyphones are recorded to obtain an output text list.

[0011] Furthermore, the obtaining of the original text corpus specifically includes:

[0012] Collect original text corpus, perform sentence processing on the original text corpus, add the sentence processed sentences as text to the input text list, perform deduplication operation on the text of the input text list, so as to obtain the deduplicated input text list.

[0013] Furthermore, after generating the Chinese character string dictionary and the string text dictionary, the following steps are further included:

[0014] Set a string window and perform the following operations on each text in the list consisting of all the texts: move the string window from the beginning to the end of the sentence, examine the string in the string window, and if there are polyphonetic Chinese characters, add the string to the string list of the Chinese characters in the Chinese character string dictionary, and add the text to the text list of the string in the string text dictionary.

[0015] Furthermore, the loop traverses the dictionary, extracts polyphonic Chinese characters from the Chinese character string dictionary, so that the number of texts reaches a preset value, and generates a candidate text list, specifically including:

[0016] Take a polyphonic Chinese character from the Chinese character string dictionary as the target Chinese character, search all strings mapped to the target Chinese character from the Chinese character string dictionary, take a string from the dictionary as the target string, and search all texts mapped to the target string from the string text dictionary as a mapping text list;

[0017] Intersecting the mapping text list and the candidate text list to obtain an intersection text list, and if the intersection text list is empty, selecting a text from the mapping text list as the target text, otherwise selecting a text from the intersection text list as the target text;

[0018] Add the target text to the candidate text list, examine the Chinese character string dictionary, delete the target string from the list of all strings mapped to the target Chinese character, examine the string text dictionary, and delete the target text from the list of all texts mapped to the target string.

[0019] Furthermore, the step of selecting the Chinese characters to be marked and obtaining a list of texts to be marked through the candidate text list specifically includes:

[0020] Operate all the texts in the candidate text list, examine each Chinese character in the text in turn, and if it is a polyphone, traverse all strings containing the polyphone and having a length of the string window as a window string list, and search the string list including the polyphone from the Chinese character string dictionary as a global string list, and take the intersection of the window string list and the global string list as an intersection string list;

[0021] If the intersection string list is not empty, the polyphone is set as the polyphone Chinese character to be marked, and the string in the intersection string list is deleted from the global string list;

[0022] A candidate text list with the polyphones to be marked annotated is obtained as a text list to be marked.

[0023] Further, the information of each text is recorded in turn from the list of texts to be marked, and after examining a text, the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters are recorded to obtain an output text list, which specifically includes:

[0024] The information of each text is recorded in turn from the list of texts to be marked. After examining a text, the number of texts and the number of polyphones currently containing preset polyphones are recorded. If the number of texts and the number of polyphones do not reach the preset value, the text is added to the empty output text list, and finally a confirmed output text list is obtained.

[0025] Further, the information of each text is recorded in turn from the list of texts to be marked, and after examining a text, the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters are recorded to obtain an output text list, which also includes:

[0026] If the number of the texts and the number of the polyphones reach a preset value, the examination of the text in the list of texts to be marked is stopped.

[0027] The embodiment of the present invention further provides a device for screening polyphonetic characters to be marked, comprising:

[0028] An acquisition module, used for acquiring an original text corpus; wherein the original text corpus includes at least one polyphonic character;

[0029] A generating module, used to generate a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record a list of Chinese characters mapped to all strings containing the Chinese characters, and the string text dictionary is used to record a list of strings mapped to all texts containing the strings;

[0030] A loop module, for looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary so that the number of texts reaches a preset value, and generating a candidate text list;

[0031] A selection module, used for selecting Chinese characters to be marked, and obtaining a list of texts to be marked through the candidate text list;

[0032] The output text list module is used to record the information of each text in the to-be-marked text list in turn, and after examining a text, record the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters to obtain an output text list.

[0033] An embodiment of the present invention further provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for screening polyphone corpus to be marked as described in any one of the above items is implemented.

[0034] An embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the above-mentioned methods for screening polyphonetic characters to be marked.

[0035] Compared with the prior art, the beneficial effects of the method, device, equipment and storage medium for screening polyphone corpus to be marked provided by the embodiment of the present invention are: the embodiment of the present invention is a method for screening a small amount of text from a large amount of text corpus, so as to construct a corpus for manual annotation. The method is automatic screening, and the screened text corpus can retain the information about polyphones in the original corpus to a large extent, with less information loss. The number of screened texts and polyphones is smaller, which effectively improves the efficiency and value of manual annotation work, and better results can be achieved with less annotated corpus, that is, less manual annotation workload. The screened text corpus helps train the polyphone pinyin disambiguation model and improves the accuracy of predicting the polyphone pinyin. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flow chart of a preferred embodiment of a method for screening polyphonetic characters to be marked provided by an embodiment of the present invention;

[0037] Figure 2It is an operational schematic diagram of some steps in a preferred embodiment of a method for screening polyphonetic characters to be marked provided by an embodiment of the present invention;

[0038] Figure 3 It is a structural schematic diagram of a preferred embodiment of a device for screening polyphonetic characters to be marked provided by an embodiment of the present invention;

[0039] Figure 4 It is a structural schematic diagram of a preferred embodiment of a terminal device provided by the present invention. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] See also Figure 1 , Figure 1 The present invention provides a method for screening polyphonetic characters to be marked in a preferred embodiment of the present invention. The method for screening polyphonetic characters to be marked in a preferred embodiment of the present invention comprises:

[0042] S1, obtaining an original text corpus; wherein the original text corpus includes at least one polyphonetic character;

[0043] S2, generating a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record a list of all strings including the Chinese characters mapped to the Chinese characters, and the string text dictionary is used to record a list of all texts including the strings mapped to the strings;

[0044] S3, looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary, so that the number of texts reaches a preset value, and generating a candidate text list;

[0045] S4, selecting Chinese characters to be marked, and obtaining a list of texts to be marked through the candidate text list;

[0046] S5, recording the information of each text in the list of texts to be marked in turn, and after examining a text, recording the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters to obtain an output text list.

[0047] It should be noted that in step S1, the input text list is constructed and the initial state is empty;

[0048] Collect original text corpora with comprehensive subject matter, divide the text corpora into sentences, treat each sentence as a text, and add it to the input text list. Perform deduplication operations on the text in the input text list to ensure that each text is unique;

[0049] After processing all the original text corpus, we get the input text list;

[0050] Through step S1, it can be ensured that the text corpus covers the subject matter types and language styles comprehensively, providing rich materials for a comprehensive language environment.

[0051] Please see again Figure 2 , Figure 2 The diagram is a schematic diagram of the operations of steps S2 and S3 in a method for screening polyphonetic characters to be marked provided by an embodiment of the present invention.

[0052] In step S2, first, a mapping dictionary is constructed to record the mapping of Chinese characters to a list of all strings containing the Chinese characters, called a Chinese character string dictionary. In the initial state, the string list mapped to each Chinese character is empty. In addition, a mapping dictionary is constructed to record the mapping of strings to a list of all texts containing the string, called a string text dictionary. In the initial state, the text list mapped to each string is empty;

[0053] Set up a string window, and perform the following operations on each text in the above text list. Move the string window from the beginning to the end of the sentence. Check the string in the string window. If there are polyphonic Chinese characters, add the string to the string list of the Chinese character in the Chinese character string dictionary, and add the text to the text list of the string in the string text dictionary. Otherwise, do not process it.

[0054] After processing all the texts, we first obtain a mapping dictionary that records the mapping of each polyphonic Chinese character to all strings containing the Chinese character, namely the Chinese character string dictionary. We also obtain a mapping dictionary that records the mapping of each string to all texts containing the string, namely the string text dictionary.

[0055] Through step S2, the context of each polyphone can be recorded in detail, thus retaining a comprehensive language environment.

[0056] In step S3, construct an empty candidate text list;

[0057] Taking out a polyphonetic Chinese character from the Chinese character string dictionary one by one, the above process is executed cyclically, that is, after taking out all polyphonetic characters one by one, starting from the beginning one by one until the number of constructed texts reaches a preset value;

[0058] The single operation method is as follows: Take a polyphonic Chinese character from the Chinese character string dictionary, called the target Chinese character. Find all the strings mapped to the target Chinese character from the Chinese character string dictionary, and take a string from them, called the target string. Find all the texts mapped to the target string from the string text dictionary, called the mapping text list;

[0059] Intersecting the above mapping text list with the above candidate text list to obtain an intersection text list. If the intersection text list is empty, then select a text from the mapping text list, which is called the target text. Otherwise, select a text from the intersection text list as the target text.

[0060] Add the target text to the candidate text list. Check the Chinese character string dictionary and delete the target string from the list of all strings mapped to the target Chinese character. Check the string text dictionary and delete the target text from the list of all texts mapped to the target string;

[0061] After executing the above operations in a loop, a confirmed candidate text list is obtained.

[0062] Through step S3, it is ensured that the screened text contains all polyphones and the context of the polyphones, and the text is selected in the intersection text list first, so that the text is less redundant, so that less text can cover the language environment more comprehensively, and the marking efficiency of polyphones is effectively improved.

[0063] In step S4, operations are performed on all texts in the candidate text list;

[0064] The operation method for each text is as follows: examine each Chinese character in the text in turn. If it is a polyphone, then traverse all strings containing the polyphone with a length equal to the above string window, which is called the window string list, and search the string list including the polyphone from the Chinese character string dictionary, which is called the global string list. Take the intersection of the above window string list and the above global string list, which is called the intersection string list;

[0065] If the intersection string list is not empty, the polyphone is set as the Chinese character to be marked as a polyphone, and the string in the intersection string list is deleted from the global string list. Otherwise, no processing is performed.

[0066] After the above process is completed, a candidate text list with polyphonetic characters to be marked is obtained, which is called a text list to be marked.

[0067] Through step S4, the process specifies the polyphones that need to be manually annotated, effectively reducing the number of polyphones to be annotated. Under the premise of retaining the comprehensive context of the polyphones, the strings that have appeared will be deleted in the global string list, that is, no repeated annotations will be made, so that the number of polyphones to be annotated is reduced, and the efficiency of manual annotation is improved.

[0068] In step S5, the information of each text is recorded in turn from the list of texts to be marked. After examining a text, the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters are recorded to obtain an output text list.

[0069] The information of each text is recorded in turn from the above-mentioned list of texts to be marked. After examining a text, the number of texts and the number of polyphones currently containing preset polyphones are recorded. If the number of texts and the number of polyphones do not reach the preset value, the text is added to the output text list, otherwise, the downward examination is terminated. After the above operation is completed, a confirmed output text list is obtained.

[0070] Through step S5, this operation is used to adapt to the actual manual annotation budget. Even if the budgeted annotation amount is small, the best corresponding corpus can be obtained based on this annotation amount, and the corpus can be dynamically retrieved according to actual needs. Because all polyphones are traversed in the loop through the dictionary step, even if a small amount of text is screened from the beginning, all polyphones and more context environments can be covered. The higher the number of selected texts, the more comprehensive the coverage of the polyphone context. And the corpus can be dynamically increased based on the original small amount of corpus according to demand, and the increased corpus is complementary to the original small amount of corpus, that is, the polyphone context information is not repeated.

[0071] The following is an example of a specific implementation plan:

[0072] In step S1, input text is collected

[0073] Construct an input text list, which is initially empty.

[0074] Collect original text corpus with comprehensive subject matter types.

[0075] For example, the collected original text is 2 sentences:

[0076] ```

[0077] What's the weather like in Chongqing? It's sunny today in Chongqing and it will rain tomorrow.

[0078] Happiness is important! Happiness is important! Happiness is important! Repeat this important word three times.

[0079] ```

[0080] Divide the above text corpus into sentences, treat each sentence as a text, and add it to the "input text list".

[0081] For example, by dividing sentences by punctuation marks, we can get:

[0082] ```

[0083] What's the weather like in Chongqing?

[0084] It is sunny in Chongqing today and it will rain tomorrow.

[0085] Happiness is important!

[0086] Happiness is important!

[0087] Happiness is important!

[0088] Repeat important words three times.

[0089] ```

[0090] Scan the "input text list" and delete the duplicate texts, leaving only one duplicate text.

[0091] For example, if we remove the repeated "Happiness is important!", we can get (the number in front of the sentence indicates the sequence number of the sentence):

[0092] ```

[0093] 1.What’s the weather like in Chongqing?

[0094] 2. It is sunny in Chongqing today and it will rain tomorrow.

[0095] 3. Happiness is important!

[0096] 4. Repeat important words three times.

[0097] ```

[0098] After processing all the original text corpus, we get the "input text list".

[0099] In step S2, a mapping dictionary is constructed.

[0100] First, a mapping dictionary is constructed to record the mapping of Chinese characters to a list of all strings containing the Chinese characters, called the "Chinese character string dictionary". In the initial state, the string list mapped to each Chinese character is empty.

[0101] For example, the constructed "Chinese character string dictionary" is called zi_chuan_dict, and at this time the dictionary is empty, that is, zi_chuan_dict={}.

[0102] In addition, a mapping dictionary is constructed that records the mapping of strings to lists consisting of all texts of the strings, called "string text dictionary". In the initial state, the text list mapped to each string is empty.

[0103] For example, the constructed "string text dictionary" is called chuan_wen_dict, and the dictionary is empty at this time, that is, chuan_wen_dict = {}.

[0104] Set a "string window" and perform the following operations on each text in the above-mentioned entire text list. The string window moves from the beginning of the sentence to the end, and examines the string in the string window.

[0105] For example, when processing the first sentence "What's the weather like in Chongqing?", the string window is set to 2. The strings obtained by moving the string window are: Chongqing, Qing's, 's sky, weather, sky how, how to, to what, what?

[0106] If the examined string contains polyphonic Chinese characters, add the string to the string list of that Chinese character in the "Chinese character string dictionary", and add the text to the text list of that string in the "string text dictionary", otherwise do not process.

[0107] For example, in this sentence, the polyphonic characters are: Chong (重), de (的), me (么).

[0108] The "Chinese character string dictionary" zi_chuan_dict obtained after processing is: {Chong: [Chongqing], de: [Qing's, 's sky], me: [how to, to what]}.

[0109] The "string text dictionary" chuan_wen_dict obtained after processing is: {Chongqing: [1], Qing's: [1], 's sky: [1], how to: [1], to what: [1]}.

[0110] After processing all the texts, one is to obtain a mapping dictionary that records each polyphonic Chinese character mapped to all the strings containing that Chinese character, that is, the "Chinese character string dictionary". Additionally, a mapping dictionary that records each string mapped to all the texts containing that string is obtained, that is, the "string text dictionary".

[0111] In step S3, loop through the dictionary.

[0112] Construct an empty "candidate text list".

[0113] For example, if the candidate text list is named hou_wen_list, then hou_wen_list = [].

[0114] Take out a polyphonic Chinese character from the above-mentioned "Chinese character string dictionary" in turn, and the above process is looped, that is, after taking out all the polyphonic characters in turn, start from the beginning and take them out in turn.

[0115] For example, in the case of the "Chinese character string dictionary" zi_chuan_dict = {Chong: [Chongqing], de: [Qing's, 's sky], me: [how to, to what]}, take out in turn: Chong, de, me, Chong, de, me...

[0116] The single operation method is as follows: Take out a polyphonic Chinese character from the "Chinese character string dictionary", which is called the target Chinese character.

[0117] For example, take the first character: "重" (zhong), then the target Chinese character is "重" (zhong).

[0118] Search for all the strings mapped by the target Chinese character in the "Chinese character string dictionary", and select one of them, which is called the target string.

[0119] For example, search for the character "重" (zhong) in the "Chinese character string dictionary" zi_chuan_dict = {"重": ["重庆"], "的": ["庆的", "的天"], "么": ["怎么", "么样"]}. All the strings are ["重庆"], and select one string: "重庆", then the target string is "重庆".

[0120] Search for all the texts mapped by the target string in the "string text dictionary", which is called the "mapped text list".

[0121] For example, the "string text dictionary" chuan_wen_dict = {"重庆": ["1"], "庆的": ["1"], "的天": ["1"], "怎么": ["1"], "么样": ["1"]}, the target string is "重庆", and the obtained mapped text list is ["1"].

[0122] Take the intersection of the above "mapped text list" and the above "candidate text list" to get the "intersection text list". If the "intersection text list" is empty, then select a text from the "mapped text list" as the "target text", otherwise select a text from the "intersection text list" as the "target text".

[0123] For example, at the initial stage, the "candidate text list" hou_wen_list = [], and the intersection with the "mapped text list" ["1"] is empty []. Then select a text from the "mapped text list" ["1"], here it is to select text 1.

[0124] For example, when executed to a certain point in time, the "candidate text list" is ["1", "3"], and the "mapped text list" is ["3", "6"]. The intersection is ["3"], which is not empty. Then select text 3 from the intersection ["3"].

[0125] Add the "target text" to the "candidate text list", delete the "target string" from the list of all strings mapped by the "target Chinese character", and delete the "target text" from the list of all texts mapped by the "target string".

[0126] For example, if the "target text" is text 1, after adding it to the "candidate text list", the result is [1]. Delete the "target string" which is "Chongqing", and after deletion, the "Chinese character string dictionary" zi_chuan_dict = {Chong: [], de: [qing de, de tian], me: [zen me, me yang]}. Delete the target text 1, and after deletion, the "string text dictionary" chuan_wen_dict = {Chongqing: [], qing de: [1], de tian: [1], zen me: [1], me yang: [1]}.

[0127] After repeatedly performing the above operations, the "candidate text list" is obtained.

[0128] For example, if all the texts [1, 2, 3, 4] are operated on, finally the candidate text list [1, 2, 3, 4] is obtained. Different "input text lists" may result in different candidate text lists, which may be [1, 4].

[0129] In step S4, select the Chinese character to be labeled.

[0130] Perform operations on all the texts in the above "candidate text list".

[0131] For example, perform operations on all the texts in the "candidate text list" which is [1, 2, 3, 4].

[0132] The operation method for each text is as follows. Examine each Chinese character of the text in turn. If it is a polyphonic character, traverse all the character strings of the length of the above character string window that contain the polyphonic character, which is called the "window character string list".

[0133] For example, examine text 1: How's the weather in Chongqing?

[0134] Examine in turn: Chong, qing, de, tian, qi, zen, me, yang.

[0135] For example, when examining the polyphonic character "me", the length of the character string window is 2, and the obtained "window character string list" is: [zen me, me yang].

[0136] Search in the "Chinese character string dictionary" for the list of character strings that include the polyphonic character, which is called the "global character string list".

[0137] For example, when the polyphonic character being examined is "me", search the "Chinese character string dictionary" zi_chuan_dict = {Chong: [Chongqing], de: [qing de, de tian], me: [zen me, me yang]}, and the obtained "global character string list" is: [zen me, me yang].

[0138] Take the intersection of the above "window character string list" and the above "global character string list", which is called the "intersection character string list".

[0139] For example, take the intersection of the "window string list" which is [zen me, me yang] and the "global string list" which is also [zen me, me yang]. The intersection is [zen me, me yang]. Since this example is the first sentence, the three sets are the same. As the process continues, the "intersection string list" may be different from the "window string list" and the "global string list".

[0140] If the "intersection string list" is not empty, set the polyphonic character as the polyphonic character to be marked, and set the default pinyin for the annotation to be reviewed.

[0141] For example, if the "intersection string list" is [zen me, me yang] and is not empty, set the character "me" in Text 1 as the polyphonic character to be marked. After setting the default pinyin, the text is: What's the weather like in Chongqing? (me)

[0142] If the "intersection string list" is not empty, delete the strings in the intersection string list from the above "global string list".

[0143] For example, if the "global string list" is: [zen me, me yang] and the "intersection string list" is: [zen me, me yang], after deletion, the "global string list" is empty, i.e., [].

[0144] If the "intersection string list" is empty, skip directly without any processing.

[0145] After completing the above process, obtain a list of candidate texts with polyphonic characters to be marked, called the "text list to be marked".

[0146] For example, after operating on Text 1, the text to be marked in Text 1 becomes: What's the weather like in Chongqing? (chong2) de (de) zen me (me) yang?

[0147] In step S5, customize the output text.

[0148] Preset the number of texts to be output, the number of polyphonic characters, and the polyphonic characters, and construct an empty "output text list".

[0149] For example, preset the number of texts to be output as 2, the number of polyphonic characters as 3, and the polyphonic characters as: chong, de, me.

[0150] Record the information of each text in turn from the above "text list to be marked". After examining a text, record the number of texts currently containing the preset polyphonic characters and the number of polyphonic characters. If neither the number of texts nor the number of polyphonic characters reaches the preset value, add the text to the "output text list", otherwise, stop examining further down.

[0151] For example, consider Text 1: What's the weather like in Chongqing? This is the first text, so the current text count is 1, the number of polyphonic characters is 3, and the current number of polyphonic characters has reached the preset number of polyphonic characters, which is 3. Therefore, stop examining downward, and the output text list only contains Text 1.

[0152] Suppose the preset text count is 2 and the preset number of polyphonic characters is 5. Then add Text 1 to the output text list and continue to examine Text 2. Text 2 is: The weather in Chongqing is sunny today and will rain tomorrow. When examining Text 2, the current text count is 2 and the number of polyphonic characters is 4. Since the preset polyphonic characters do not include the character "will", the character "will" is not recorded. Since the text count of 2 has reached the preset text count, add Text 2 to the "output text list" and stop examining downward.

[0153] After completing the above operations, an output text list is obtained.

[0154] For example, the output text list is: [1, 2], that is:

[0155] ```

[0156] What's the weather like in Chongqing?

[0157] The weather in Chongqing is sunny today and will rain tomorrow.

[0158] ```

[0159] Please refer to Figure 3 , Figure 3 which is a structural schematic diagram of a preferred embodiment of a screening device for polyphonic character candidate language materials provided by an embodiment of the present invention.

[0160] Based on the above screening method for polyphonic character candidate language materials, the present invention also provides a screening device for such polyphonic character candidate language materials, including:

[0161] An acquisition module 301, configured to acquire an original text corpus; wherein, the original text corpus includes at least one polyphonic character;

[0162] A generation module 302, configured to generate a Chinese character string dictionary and a string text dictionary, where the Chinese character string dictionary is used to record a list of all strings containing the Chinese character mapped by the Chinese character, and the string text dictionary is used to record a list of all texts containing the string mapped by the string;

[0163] A loop module 303, configured to loop through the dictionary, take out polyphonic Chinese characters from the Chinese character string dictionary, so that the text count reaches a preset value, and generate a candidate text list;

[0164] A selection module 304 is used to select Chinese characters to be marked, and obtain a list of texts to be marked through the candidate text list;

[0165] The output text list module 305 is used to record the information of each text in the to-be-marked text list in turn, and after examining a text, record the number of texts currently containing preset polyphones, the number of polyphones and polyphone Chinese characters to obtain an output text list.

[0166] In the specific implementation, the working principle, control process and technical effect of the screening device for polyphonetic characters to be marked provided in the embodiment of the present invention are the same as those of the screening method for polyphonetic characters to be marked in the above embodiment, and will not be repeated here.

[0167] See also Figure 4 , Figure 4 4 is a schematic diagram of a preferred embodiment of a terminal device provided by the present invention. The terminal device includes a processor 401, a memory 402, and a computer program stored in the memory 402 and configured to be executed by the processor 401. When the processor 401 executes the computer program, the method for screening polyphonetic characters to be marked in any of the above embodiments is implemented.

[0168] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2, ...), and the one or more modules / units are stored in the memory 402 and executed by the processor 401 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0169] The processor 401 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 401 can also be any conventional processor. The processor 401 is the control center of the terminal device, and various parts of the terminal device are connected using various interfaces and lines.

[0170] The memory 402 mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application program required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory 402 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card), etc., or the memory 402 can also be other volatile solid-state storage devices.

[0171] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 4 The structural diagram is only an example of the above-mentioned terminal device and does not constitute a limitation of the above-mentioned terminal device. It may include more or less components than shown in the figure, or combine certain components, or different components.

[0172] An embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for screening polyphonetic characters to be marked in any of the above embodiments.

[0173] In summary, the embodiments of the present invention provide a method, device, equipment and storage medium for screening polyphone corpus to be marked, and the present invention has the following advantages: collecting original text corpus with comprehensive subject matter types to ensure that the text corpus covers comprehensive subject matter types and language styles.

[0174] The process of building the mapping dictionary records the context of each polyphone, so the comprehensive language environment is retained. Currently, there is no screening solution that can fully cover the context of Chinese characters, so the texts screened by other existing solutions will lose a lot of subject matter.

[0175] The process of looping through the dictionary ensures that the filtered text contains all polyphones and the context of polyphones, and gives priority to selecting text in the intersection text list, so that the text is less redundant, so that even with less text, the language environment can be covered more comprehensively, effectively improving the annotation efficiency of polyphones. The current text filtering scheme cannot achieve a more comprehensive language environment with less text, so the filtered text has little improvement in the annotation efficiency of polyphones.

[0176] The process of selecting Chinese characters to be marked specifies the polyphones to be manually marked, which effectively reduces the number of polyphones to be marked. Under the premise of preserving the comprehensive context of the polyphones, the strings that have appeared will be deleted from the global string list, that is, they will not be marked repeatedly, which reduces the number of polyphones to be marked and improves the efficiency of manual marking. The existing scheme does not reduce the number of items to be marked with repeated information, so that manual marking will do some redundant marking work, which reduces the marking efficiency.

[0177] The process of outputting the text list is used to adapt to the actual manual annotation budget. Even if the budgeted annotation amount is small, the best corresponding corpus can be obtained based on this annotation amount. In addition, the corpus can be dynamically retrieved according to actual needs. Because all polyphones are traversed in the step of looping through the dictionary, even if a small amount of text is screened from the beginning, all polyphones and more contextual environments can be covered. The higher the number of selected texts, the more comprehensive the coverage of the polyphone context. And the corpus can be dynamically added to the original small amount of corpus according to needs. The added corpus and the original small amount of corpus are complementary, that is, the polyphone context information is not repeated. The existing corpus screening scheme cannot dynamically retrieve the corpus according to actual needs. If the target quantity changes, it needs to be set from the beginning and re-screened. Moreover, the results of each screening are different, and the corpus screened each time cannot be complementary.

[0178] The embodiment of the present invention is a method for screening a small amount of text from a large amount of text corpus, thereby constructing a corpus for manual annotation. The method is automatic screening, and the screened text corpus can retain the information about polyphones in the original corpus to a large extent, with less information loss. The number of screened texts and polyphones is smaller, which effectively improves the efficiency and value of manual annotation work, and better results can be achieved with less annotated corpus, that is, the workload of manual annotation is less. The screened text corpus helps train a polyphone pinyin disambiguation model and improves the accuracy of predicting the pinyin of polyphones.

[0179] It should be noted that the device embodiments described above are merely schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the system embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art may understand and implement it without paying creative labor.

[0180] The above are preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications are also considered to be within the protection scope of the present invention.

Claims

1. A method for screening polyphonetic characters to be marked, characterized in that: include: Acquire an original text corpus; wherein the original text corpus includes at least one polyphonetic character; Generate a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record the mapping of Chinese characters to a list of all strings containing the Chinese characters, and the string text dictionary is used to record the mapping of strings to a list of all texts containing the strings; Looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary so that the number of texts reaches a preset value, and generating a candidate text list; Select the Chinese characters to be marked, and obtain a list of texts to be marked through the candidate text list; Recording the information of each text in the list of texts to be marked in turn, after examining a text, recording the number of texts currently containing the preset polyphones, the number of polyphones and the Chinese characters of the polyphones, so as to obtain an output text list; The step of selecting the Chinese characters to be marked and obtaining a list of texts to be marked through the candidate text list specifically includes: Operate all the texts in the candidate text list, examine each Chinese character in the text in turn, if it is a polyphone, traverse all strings containing the polyphone and the length of the string window as a window string list, and search the string list including the polyphone from the Chinese character string dictionary as a global string list, and take the intersection of the window string list and the global string list as an intersection string list; If the intersection string list is not empty, the polyphone is set as the polyphone Chinese character to be marked, and the string in the intersection string list is deleted from the global string list; A candidate text list with the polyphones to be marked annotated is obtained as a text list to be marked.

2. The method for screening polyphonetic characters to be marked as claimed in claim 1, characterized in that: The obtaining of the original text corpus specifically includes: Collect original text corpus, perform sentence processing on the original text corpus, add the sentence processed sentences as text to the input text list, perform deduplication operation on the text of the input text list, so as to obtain the deduplicated input text list.

3. The method for screening polyphonetic characters to be marked as claimed in claim 1, characterized in that: The generating of the Chinese character string dictionary and the string text dictionary further includes: Set a string window and perform the following operations on each text in the list consisting of all the texts: move the string window from the beginning to the end of the sentence, examine the string in the string window, and if there are polyphonetic Chinese characters, add the string to the string list of the Chinese characters in the Chinese character string dictionary, and add the text to the text list of the string in the string text dictionary.

4. The method for screening polyphonetic characters to be marked as claimed in claim 1, characterized in that: The loop traverses the dictionary, extracts polyphonic Chinese characters from the Chinese character string dictionary, so that the number of texts reaches a preset value, and generates a candidate text list, specifically including: Take a polyphonic Chinese character from the Chinese character string dictionary as the target Chinese character, search all strings mapped to the target Chinese character from the Chinese character string dictionary, take a string from the dictionary as the target string, and search all texts mapped to the target string from the string text dictionary as a mapping text list; Intersecting the mapping text list and the candidate text list to obtain an intersection text list, and if the intersection text list is empty, selecting a text from the mapping text list as the target text, otherwise selecting a text from the intersection text list as the target text; Add the target text to the candidate text list, examine the Chinese character string dictionary, delete the target string from the list of all strings mapped to the target Chinese character, examine the string text dictionary, and delete the target text from the list of all texts mapped to the target string.

5. The method for screening polyphonetic characters to be marked as claimed in claim 1, characterized in that: The method records the information of each text in the list of texts to be marked in turn, and after examining a text, records the number of texts, the number of polyphonic characters, and the polyphonic Chinese characters that currently contain the preset polyphonic characters to obtain an output text list, specifically including: The information of each text is recorded in turn from the list of texts to be marked. After examining a text, the number of texts and the number of polyphones currently containing preset polyphones are recorded. If the number of texts and the number of polyphones do not reach the preset value, the text is added to the empty output text list, and finally a confirmed output text list is obtained.

6. The method for screening polyphonetic characters to be marked as claimed in claim 5, characterized in that: The method of sequentially recording the information of each text from the list of texts to be marked, and after examining a text, recording the number of texts currently containing preset polyphonic characters, the number of polyphonic characters, and polyphonic Chinese characters to obtain an output text list, further includes: If the number of the texts and the number of the polyphones reach a preset value, the examination of the text in the list of texts to be marked is stopped.

7. A device for screening polyphonetic characters to be marked, characterized in that: include: An acquisition module, used for acquiring an original text corpus; wherein the original text corpus includes at least one polyphonic character; A generating module, used to generate a Chinese character string dictionary and a string text dictionary, wherein the Chinese character string dictionary is used to record a list of Chinese characters mapped to all strings containing the Chinese characters, and the string text dictionary is used to record a list of strings mapped to all texts containing the strings; A loop module, for looping through the dictionary, taking out polyphonic Chinese characters from the Chinese character string dictionary so that the number of texts reaches a preset value, and generating a candidate text list; A selection module, used for selecting Chinese characters to be marked, and obtaining a list of texts to be marked through the candidate text list; An output text list module is used to record the information of each text in the list of texts to be marked in turn, and after examining a text, record the number of texts currently containing preset polyphones, the number of polyphones and the Chinese characters of polyphones to obtain an output text list; The step of selecting the Chinese characters to be marked and obtaining a list of texts to be marked through the candidate text list specifically includes: Operate all the texts in the candidate text list, examine each Chinese character in the text in turn, if it is a polyphone, traverse all strings containing the polyphone and the length of the string window as a window string list, and search the string list including the polyphone from the Chinese character string dictionary as a global string list, and take the intersection of the window string list and the global string list as an intersection string list; If the intersection string list is not empty, the polyphone is set as the polyphone Chinese character to be marked, and the string in the intersection string list is deleted from the global string list; A candidate text list with the polyphones to be marked annotated is obtained as a text list to be marked.

8. A terminal device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for screening polyphonetic characters to be marked as claimed in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for screening polyphonetic characters to be marked as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for generating polyphone annotating template

    CN105225657A

  • Method and apparatus for screening valid term of a pronunciation lexicon

    CN105893414A