Text error correction method, electronic device and computer-readable storage medium
Through word-grained slicing and word-frequency editing distance calculation, the high cost and low accuracy problems caused by the existing text error correction methods relying on word-grained slicing are solved, and efficient and accurate text error correction is achieved.
Patent Information
- Application Number
- CN202111012472.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-08-31
AI Technical Summary
Existing text error correction methods rely on word granularity slicing and editing distance, resulting in error correction results relying on slicing methods, which are costly and have low accuracy, and cannot effectively distinguish approximate errors.
The word-grained segmentation is used, and the editing distance is calculated using the preset index word element set and word frequency, and the word frequency of the candidate words is scored to determine the error correction result.
It significantly reduces the cost of text error correction, improves the accuracy and accuracy of error correction, and ensures that the candidate words are consistent with the user's intention.
Smart Images

Figure CN113743094B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a text error correction method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Natural language processing is an important direction in the fields of computer science and artificial intelligence. Natural language processing is a science that integrates linguistics, computer science, and mathematics, and is a bridge that enables effective communication between humans and computers using natural language. Natural language processing technologies are mainly applied in aspects such as machine translation, public opinion monitoring, automatic summarization, opinion extraction, text recognition, text semantic comparison, text error correction, Chinese optical character recognition (Optical Character Recognition, abbreviated as: OCR), etc. Among them, text error correction refers to the process of correcting the incorrect content in the text. Text error correction technology has broad application scenarios and benefits, such as improving the work efficiency of typists, enhancing the search accuracy of search engines, reducing low-level mistakes in production caused by text errors, and so on.
[0003] However, the inventor found that in general text error correction, such as the technical solution disclosed in the patent with the application number "CN202010164805.5", most of them first segment the text to be processed at the word granularity, and then correct the segmented result at the word granularity. The error correction result of this method depends on the segmentation result, and the segmentation result depends on the segmentation method. In the case of an inappropriate segmentation method, the final error correction result will be incorrect; moreover, when segmenting at the word granularity, it depends on the corresponding vocabulary dictionary, so the vocabulary dictionary needs to be maintained for a long time, which results in too high a cost for text error correction; at the same time, in the traditional scoring rules, the edit distance is usually used for calculation, and in the edit distance, the costs of substitution, addition, and deletion are generally set to 1, which cannot distinguish between approximately incorrect inputs, resulting in a relatively low accuracy of the final error correction result. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a text error correction method, an electronic device, and a computer-readable storage medium, which can significantly reduce the cost of text error correction, greatly improve the accuracy of text error correction, and at the same time improve the precision of text error correction.
[0005] To solve the above technical problems, an embodiment of the present application provides a text error correction method, including the following steps: segmenting the vocabulary to be error-corrected at the character granularity to obtain a number of retrieval segments; wherein the type of the retrieval segment is a single letter or Chinese pinyin; determining a target index term consistent with the retrieval segment in a preset index term set; wherein the types of the index terms in the index term set include single letters and Chinese pinyin; retrieving in a preset index according to the target index term to obtain a number of proper nouns in the same order as the target index term as candidate words; wherein the index is a set of mapping relationships between the preset index terms and the proper nouns; calculating an edit distance based on the character frequency of the vocabulary to be error-corrected and the character frequency of the candidate words, scoring the candidate words, and obtaining the scores corresponding to the candidate words; using the candidate word with the highest score as the error correction result to replace the vocabulary to be error-corrected.
[0006] An embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above text error correction method.
[0007] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and the computer program realizes the above text error correction method when executed by a processor.
[0008] The text error correction method, electronic device, and computer-readable storage medium provided by this application. After the server obtains the vocabulary to be corrected, it first performs character-level segmentation on the vocabulary to be corrected to obtain several retrieval segments of the vocabulary to be corrected. The type of the retrieval segment is a single letter or Chinese pinyin. After the server obtains the retrieval segments, it then determines the target index element that is the same as the retrieval segment in the preset index element set. The types of the index elements in the index element set include single letters and Chinese pinyin. The server performs a search in the preset index based on the target index element to obtain several proper nouns that are in the same order as the target index element as candidate words. The server calculates the edit distance based on the character frequency of the vocabulary to be corrected and the character frequency of the candidate words, scores each candidate word, and obtains the score corresponding to each candidate word. Finally, the candidate word with the highest score is used as the error correction result to replace the vocabulary to be corrected. In the embodiments of this application, when performing text error correction on the vocabulary to be corrected, the vocabulary to be corrected is segmented at the character level, and the obtained retrieval segments are single letters or Chinese pinyin in the vocabulary to be corrected. The index elements in the preset index element set are also single letters or Chinese pinyin. Whether it is a single letter or Chinese pinyin, it is a small-scale data set that can be enumerated, that is, this application does not have a dictionary that needs to be continuously maintained, which can significantly reduce the cost of text error correction; at the same time, although each of the retrieval segments after character-level segmentation represents a single character and does not have a specific semantics, this application retrieves candidate words based on the vocabulary to be corrected input by the user, that is, the determination of the candidate words completely depends on the input situation of the user. The obtained candidate words are in the same order as the target index element, so the obtained candidate words basically conform to the user's intention. At the same time, the calculation of the edit distance in this application is based on the character frequency, using statistical rules to reflect the situation in the actual scenario, refining the calculation of the edit distance, thereby improving the scoring accuracy and making the final error correction result more accurate, greatly improving the accuracy of text error correction.
[0009] In addition, the index element set and the index are obtained through the following steps: obtaining a preset proper noun set, where the proper noun set includes several of the proper nouns; traversing the proper nouns, performing character-level segmentation on the proper nouns to obtain several index segments; where the type of the index segment includes a single letter and Chinese pinyin; where the single letter includes the original letter and Chinese character letter, the original letter is the letter that exists in the proper noun itself, and the Chinese character letter is the first letter of the pinyin of each Chinese character in the proper noun. The Chinese pinyin includes the original pinyin and approximate pinyin. The original pinyin is the pinyin of each Chinese character in the proper noun itself, and the approximate pinyin is the approximate sound determined from the preset near-sound dictionary according to the original pinyin; using the index segment as the index element to obtain the index element set, and constructing a mapping relationship between the index element and the proper noun to obtain the index.
[0010] In addition, the types of the indexing tokens in the indexing token set further include single Chinese characters. Calculating the edit distance based on the word frequencies of the error-prone word to be corrected and the candidate word, and scoring the candidate word to obtain the score corresponding to the candidate word includes: segmenting the error-prone word to be corrected at the character granularity to obtain a plurality of scoring segments, where the types of the scoring segments are any one of the following: single letters, Chinese pinyin, or single Chinese characters; counting the word frequencies of each indexing token in the indexing token set, and determining the word frequency of a first target word and the word frequency of a second target word according to the word frequencies of each indexing token; where the first target word is the scoring segment, and the second target word is the indexing token corresponding to the scoring segment in the candidate word; calculating the edit cost between the first target word and the second target word according to the word frequency of the first target word, the word frequency of the second target word, and a preset cost function; calculating the edit distance between the error-prone word to be corrected and the candidate word according to the edit cost; calculating the edit similarity between the error-prone word to be corrected and the candidate word according to the edit distance; and scoring the candidate word according to the edit similarity to obtain the score corresponding to the candidate word.
[0011] In addition, counting the word frequencies of each indexing token in the indexing token set includes: obtaining the initial word frequencies of each proper noun in the proper noun set corresponding to the indexing token set; determining the number of times each proper noun appears in the historical error correction record as the error correction word frequency of the proper noun according to the historical error correction record; determining the cumulative word frequency of the proper noun according to the initial word frequency and the error correction word frequency; obtaining the cumulative word frequencies of a plurality of proper nouns corresponding to each indexing token, and statistically summing the plurality of cumulative word frequencies to obtain the word frequency of the indexing token.
[0012] In addition, calculating the edit cost between the first target word and the second target word according to the word frequency of the first target word, the word frequency of the second target word, and a preset cost function includes: if the first target word is the same as the second target word, determining that the edit cost is: cost = 0, where cost is the edit cost; if the first target word is different from the second target word, determining that the edit cost is: cost = norm[lg(A / B)], norm(x) = (m x -m -x ) / (m x +m -x ), where norm[lg(x)] is the preset cost function, m is a preset constant, A is the word frequency of the first target word, and B is the word frequency of the second target word.
[0013] In addition, if the first target word is different from the second target word, calculating the edit cost between the first target word and the second target word further includes: if the first target word is a single Chinese character and the pinyin of the first target word is the same as the pinyin of the second target word, using the word frequency of the pinyin of the first target word as the word frequency of the first target word; if the first target word is a single Chinese character and the approximate sound of the pinyin of the first target word is the same as the approximate sound of the pinyin of the second target word, using the sum of the word frequency of the pinyin of the first target word and the word frequency of the pinyin of the second target word as the word frequency of the first target word; if the first target word is a Chinese pinyin and the approximate sound of the first target word is the same as the approximate sound of the pinyin of the second target word, using the sum of the word frequency of the Chinese pinyin represented by the first target word and the word frequency of the pinyin of the second target word as the word frequency of the first target word.
[0014] In addition, through the following formula, calculate the edit distance between the to-be-corrected word and the candidate word according to the edit cost: where a i is the i-th scoring segment of the to-be-corrected word, a i-1 is the (i - 1)-th scoring segment of the to-be-corrected word, b j is the j-th index token of the candidate word, b j-1 is the (j - 1)-th index token of the candidate word, lev(a i , b j ) is the edit distance between the first i scoring segments of the to-be-corrected word and the first j index tokens of the candidate word, lev(a i-1 , b j ) is the edit distance between the first (i - 1) scoring segments of the to-be-corrected word and the first j index tokens of the candidate word, lev(a i , b j-1 ) is the edit distance between the first i scoring segments of the to-be-corrected word and the first (j - 1) index tokens of the candidate word, lev(a i-1 , b j-1 ) is the edit distance between the first (i - 1) scoring segments of the to-be-corrected word and the first (j - 1) index tokens of the candidate word, cost(a i-1 , b j-1 ) is the edit cost between the (i - 1)-th scoring segment of the to-be-corrected word and the (i - 1)-th index token of the candidate word, and min is the minimum value function.
[0015] In addition, scoring the candidate word according to the edit similarity includes: obtaining the cumulative word frequency of the proper noun corresponding to the candidate word and the total word frequency of the proper noun set; wherein, the total word frequency is the sum of the cumulative word frequencies of each proper noun in the proper noun set; scoring the candidate word according to the edit similarity, the total word frequency, and the cumulative word frequency of the proper noun corresponding to the candidate word.
[0016] In addition, segmenting the word to be corrected at the character granularity to obtain a plurality of retrieval segments includes: obtaining the type of the word to be corrected; wherein, the type includes a mixed type of letters and Chinese characters, an all-letter type, and an all-Chinese character type; segmenting the word to be corrected at the character granularity corresponding to the type to obtain a plurality of retrieval segments.
[0017] In addition, before scoring the candidate word according to the edit similarity to obtain the score corresponding to the candidate word, it includes: determining whether the edit similarity is less than a preset first threshold; if the edit similarity is less than the preset first threshold, terminating the correction of the word to be corrected. Description of the Drawings
[0018] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings, and these exemplary illustrations do not limit the embodiments.
[0019] Figure 1 is a flowchart of the text error correction method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of segmenting the word to be corrected at the character granularity to obtain a plurality of retrieval segments according to an embodiment of the present application;
[0021] Figure 3 is a flowchart of obtaining a preset index word element set and a preset index according to an embodiment of the present application;
[0022] Figure 4 is a flowchart of calculating the edit distance according to the character frequency of the word to be corrected and the character frequency of the candidate word, scoring the candidate word, and obtaining the score corresponding to the candidate word according to an embodiment of the present application;
[0023] Figure 5 is a flowchart of counting the character frequency of each index word element in the index word element set according to an embodiment of the present application;
[0024] Figure 6 is a function image of a norm(A / B) and a norm[lg(A / B)] provided according to an embodiment of the present application;
[0025] Figure 7It is a flowchart for scoring candidate words according to edit similarity in an embodiment of the present application;
[0026] Figure 8 It is a flowchart for obtaining words to be corrected in an embodiment of the present application;
[0027] Figure 9 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present invention. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.
[0029] In a technical solution for text error correction based on splitting the text to be processed at the word granularity, the server can perform word granularity word segmentation on the proper nouns to be corrected, obtain multiple segmented fragments of the proper nouns to be corrected, and output the pinyin of each segmented fragment. Then, taking the pinyin of each segmented fragment as a keyword, the server retrieves the corresponding candidate words for the segmented fragment from a preset homophone dictionary to obtain the retrieval result. If the retrieval result is empty, the server will traverse and remove the characters in the segmented fragment to obtain multiple word groups, respectively use these word groups as keywords, and call the preset inverted index to retrieve the multiple candidate words corresponding to the word groups to obtain the retrieval result, so as to perform text error correction.
[0030] For example: The segmented fragment is "shunuan pipe". The server takes the pinyin of "shunuan pipe" as a keyword and retrieves this keyword from the preset homophone dictionary. The server determines that the retrieval result is empty. At this time, the server will traverse and remove each character in "shunuan pipe". If the character "shun" is removed, the remaining two characters are "nuan" and "pipe". The server does not retrieve candidate words using the inverted index; if the character "nuan" is removed, the remaining two characters are "shun" and "pipe". The server retrieves candidate words such as "fallopian tube", "vas deferens", and "blood vessel" using the inverted index; if the character "pipe" is removed, the remaining two characters are "shun" and "nuan". The server does not retrieve candidate words using the inverted index.
[0031] The inventors of the present application found that when the server splits the vocabulary to be corrected by word granularity, it needs to use a Chinese word segmenter for word segmentation, that is, it needs to use a dictionary segmented by word granularity. The server pre-processes the original corpus by word segmentation and retrieves it from the homophone dictionary during retrieval. The segmented fragments are input in the format of pinyin, and the pinyin of multiple word fragments is counted to obtain the homophone dictionary. These two dictionaries cannot be exhaustive. There are several vocabulary items stored in the dictionary segmented by word granularity. When the server identifies that there is a vocabulary item stored in the dictionary segmented by word granularity in the text to be corrected, this vocabulary item can be segmented out. However, when new words appear, the vocabulary items stored in the dictionary segmented by word granularity need to be updated. Since new words keep emerging, this requires manual long-term continuous update and maintenance. The corresponding new words need to be supplemented and recorded in the word segmentation dictionary and the homophone dictionary respectively, and then the index is constructed. New words are diverse and endless, which requires manual continuous maintenance and update, consuming time and effort, and the cost of text correction is too high.
[0032] At the same time, retrieving based on the pinyin of the segmented fragments in the homophone dictionary can only directly correct the situation of homophonic but different characters. For the situation of near-homophonic but different characters, it cannot be directly corrected. For the situation of near-homophonic but different characters, the server can only first obtain the set of single characters of the proper noun, then construct an inverted index according to the set of single characters, and then retrieve candidate words by removing each character in the segmented fragment one by one. Since some characters in the segmented fragment are completely removed, this will result in candidate words in the retrieval results that are very different from the user's input intention. For example, in the above example, "fallopian tube" is the correction result of "uterine tube", but "blood vessel" and "vas deferens" are candidate words that are obviously inconsistent with the user's input intention. In some cases, the introduction of such candidate words that are obviously inconsistent with the user's input intention will lead to relatively low accuracy of text correction. Moreover, when there are multiple character errors in the vocabulary to be corrected input by the user, the server needs to remove all the wrong characters. The more characters are removed, the more candidate words that are obviously inconsistent with the user's input intention are introduced, and the lower the accuracy of text correction.
[0033] To solve the above problems of too high text correction cost and relatively low accuracy of text correction, an embodiment of the present application provides a text correction method, which is applied to an electronic device. Herein, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The following specifically describes the implementation details of the text correction method of this embodiment. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0034] The application scenarios of the embodiments of the present application may include but are not limited to: correcting errors in proper nouns in a certain field that are input by users into a search engine; correcting errors in texts containing proper nouns in a certain field that are edited by users such as papers, journals, and newsletters; correcting errors in texts containing proper nouns in a certain field that are input by users into a chat box, a social media short message posting box, or a text message editing interface, etc.
[0035] The specific process of the text error correction method of this embodiment can be as follows Figure 1 As shown, including:
[0036] Step 101, segment the vocabulary to be corrected according to word granularity to obtain a number of search segments.
[0037] Specifically, after obtaining the vocabulary to be corrected, the server may segment the vocabulary to be corrected according to the character granularity to obtain several search segments of the vocabulary to be corrected, wherein the type of the search segment may be a single letter or Chinese pinyin.
[0038] In a specific implementation, after the server obtains the vocabulary to be corrected, it first lowercases the vocabulary to be corrected, that is, converts the uppercase letters in the vocabulary to be corrected into lowercase letters, and divides the vocabulary to be corrected after the lowercase processing according to the word granularity. The server uses a single letter in the vocabulary to be corrected as a search segment, uses the Chinese pinyin in the vocabulary to be corrected as a search segment, converts a single Chinese character in the vocabulary to be corrected into Chinese pinyin, and uses the Chinese pinyin converted from a single Chinese character as a search segment.
[0039] In an example, the word to be corrected is "PFa Bank". The server first lowercases "PFa Bank" to obtain the lowercase word to be corrected, namely "pfa Bank". The server divides "pfa Bank" according to word granularity, and takes "p" as the first search segment, "fa" as the second search segment, converts "silver" into "yin", and uses "yin" as the third search segment, converts "line" into "hang", and uses "hang" as the fourth search segment.
[0040] In one example, the server splits the error-correction target word by character granularity to obtain a number of retrieval fragments, which can be implemented by a preset first analyzer. Taking the error-correction target word "pfa bank" as an example, the first analyzer first splits "pfa bank" into "pfa", "silver", and "bank" according to the content of the error-correction target word. Then, it determines that the Chinese pinyin "fa" is included in "pfa", so it splits "pfa" into "p" and "fa". For "silver" and "bank", the first analyzer converts "silver" into the Chinese pinyin "yin" and "bank" into the Chinese pinyin "hang". Among them, the preset first analyzer can be an existing analyzer or an analyzer developed by those skilled in the art. The embodiments of the present application do not make specific limitations on this.
[0041] Step 102: Determine a target index term that is consistent with the retrieval fragment in a preset index term set.
[0042] Step 103: Retrieve in a preset index according to the target index term to obtain a number of proper nouns that are in the same order as the target index term as candidate words.
[0043] In a specific implementation, after the server splits the error-correction target word by character granularity to obtain a number of retrieval fragments of the error-correction target word, it can determine a target index term that is consistent with the retrieval fragment in a preset index term set, and then retrieve in a preset index according to the determined target index term that is consistent with the retrieval fragment to obtain a number of proper nouns that are in the same order as the target index term as candidate words. Among them, the types of index terms in the index term set include single letters and Chinese pinyin, and the index is a set of mapping relationships between preset index terms and proper nouns.
[0044] In one example, the error-correction target word is "p fa y bank". The server splits "p fa y bank" by character granularity to obtain four retrieval fragments in the order of: "p", "fa", "y", and "hang". The server then finds the target index terms: "p", "fa", "y", and "hang" in the index term set, and retrieves in the index. The proper nouns that are in the same order as the target index terms are "Shanghai Pudong Development Bank", "Popular Law Bank", and "Shanghai Pudong Development Technology Bank". The server uses "Shanghai Pudong Development Bank", "Popular Law Bank", and "Shanghai Pudong Development Technology Bank" as candidate words.
[0045] Step 104: Calculate the edit distance according to the character frequency of the error-correction target word and the character frequency of the candidate word, score the candidate word, and obtain the score corresponding to the candidate word.
[0046] Step 105: Use the candidate word with the highest score as the error-correction result to replace the error-correction target word.
[0047] In a specific implementation, after determining the candidate words, the server can obtain the character frequencies of the word to be corrected and the candidate words. Based on the character frequencies of the word to be corrected and the candidate words, the server calculates the edit distance between the word to be corrected and the candidate words. The edit distance is a quantitative measurement of the difference degree between two words, and the measurement method is to determine at least how many times of processing are required to change one word into another word. The server can score each candidate word according to the edit distance to obtain the scores corresponding to each candidate word, and take the candidate word with the highest score among all candidate words as the correction result to replace the word to be corrected.
[0048] In an example, the word to be corrected is "p fa y xing". The candidate words determined by the server include "Shanghai Pudong Development Bank", "Popularize Law Bank", and "Shanghai Pudong Development Technology Bank". The score of "Shanghai Pudong Development Bank" is 88, the score of "Popularize Law Bank" is 71, and the score of "Shanghai Pudong Development Technology Bank" is 25. The score of the candidate word "Shanghai Pudong Development Bank" is greater than the scores of the candidate words "Popularize Law Bank" and "Shanghai Pudong Development Technology Bank". The server takes "Shanghai Pudong Development Bank" as the correction result to replace "p fa y xing".
[0049] In this embodiment, when the server performs text correction on the word to be corrected, it divides the word to be corrected at the character granularity. The retrieved fragments obtained by the division are individual letters or Chinese pinyin in the word to be corrected. The index elements in the preset index element set are also individual letters or Chinese pinyin. In this application, whether it is an individual letter or Chinese pinyin, it is a small-scale data set that can be enumerated. Therefore, there is no dictionary that needs to be continuously maintained during the process of dividing the word to be corrected at the character granularity in this application. Thus, in subsequent maintenance, this application only needs to build an index according to new words and does not need to maintain a dictionary, which can significantly reduce the cost of text correction.
[0050] In this embodiment, the server first obtains a retrieval fragment according to the vocabulary to be corrected, then determines the target index token, and finally obtains a proper noun with the same order as the target index token from the preset index as the candidate word. The types of the retrieval fragment and the target index token are both single letters or Chinese pinyin. In this entire retrieval stage, this application only utilizes the pinyin and letter information in the vocabulary to be corrected. The determination of the candidate word completely depends on the pinyin and letter information contained in the vocabulary to be corrected itself, so that the determined candidate word is consistent with the pinyin and letter information of the vocabulary to be corrected. In the subsequent scoring stage, the edit distance is calculated and scored according to the character frequency of the vocabulary to be corrected and the character frequency of the candidate word, and all the information of the vocabulary to be corrected itself and the corresponding character frequency information are utilized. This application quantifies the replacement probability between two characters through the character frequency of the content represented by the vocabulary to be corrected itself and the character frequency of the characters replaced in the candidate word, and calculates the edit distance between the vocabulary to be corrected and the candidate word based on this. Different from setting it to 1 in the prior art, the calculation of the edit distance in this application is based on the character frequency, and uses statistical rules to reflect the usage of Chinese characters, pinyin, and letters in the actual scenario, refining the calculation of the edit distance, thereby improving the scoring accuracy and making the final error correction result more accurate. That is, the entire solution of this application makes full use of the pinyin and letter information, all the information of itself, and the character frequency information in the vocabulary to be corrected, processes the results segmented by character granularity using multiple types of information, ensures the accuracy of text error correction on the basis of reducing the cost of text error correction, and significantly improves the accuracy of error correction.
[0051] In one embodiment, the vocabulary to be corrected is segmented by character granularity to obtain a number of retrieval fragments, which can be implemented through the steps as Figure 2 shown, specifically including:
[0052] Step 201, obtain the type of the vocabulary to be corrected.
[0053] Step 202, segment the vocabulary to be corrected by character granularity corresponding to the type to obtain a number of retrieval fragments.
[0054] In a specific implementation, before segmenting the vocabulary to be corrected by character granularity, the server can first determine the type of the vocabulary to be corrected. Among them, the types of the vocabulary to be corrected include the mixed type of letters and Chinese characters, the all-letter type, and the all-Chinese-character type. The server can call an analyzer corresponding to the type of the vocabulary to be corrected according to the type of the vocabulary to be corrected, and segment the vocabulary to be corrected by character granularity corresponding to the type of the vocabulary to be corrected to obtain a number of retrieval fragments.
[0055] In one example, the word to be corrected is "Pufa Bank". The server determines that the type of "Pufa Bank" is all-Chinese characters. The server calls the preset first analyzer to split "Pufa Bank" into four Chinese characters: "Pu", "Fa", "Yin", and "Hang", and obtains the pinyin of these four Chinese characters as: "pu", "fa", "yin", and "hang". The server uses "pu", "fa", "yin", and "hang" as the four retrieval segments of "Pufa Bank".
[0056] In one example, the word to be corrected is "pfa Bank". The server determines that the type of "pfa Bank" is a mixture of letters and Chinese characters. The server calls the preset first analyzer to split "pfa Bank" into "p", "fa", "Yin", and "Hang". The server uses "p" as the first retrieval segment, "fa" as the second retrieval segment, converts "Yin" to "yin", and uses "yin" as the third retrieval segment, converts "Hang" to "hang", and uses "hang" as the fourth retrieval segment.
[0057] In one example, the word to be corrected is "pfayh". The server determines that the type of "pfayh" is all-letters. The server calls the preset second analyzer to split "pfayh" into "p", "fa", "y", and "h". The server uses "p", "fa", "y", and "h" as the four retrieval segments of "pfayh".
[0058] In this embodiment, the server can perform different character-level segmentations on words to be corrected of different types according to the type of the word to be corrected. Whether it is a word to be corrected that is a mixture of letters and Chinese characters, all-letters, or all-Chinese characters, the server can determine the target index tokens and determine the candidate words according to the target index tokens and the index, realizing the direct correction of different types of texts.
[0059] In one embodiment, the preset index token set and the preset index can be obtained through the following steps as shown Figure 3 follows:
[0060] Step 301, obtain the preset proper noun set.
[0061] In a specific implementation, the server can collect the proper nouns in a certain target field to generate a proper noun set, and use the generated proper noun set as the preset proper noun set; the server can also directly obtain the proper noun set of a certain field from an open-source database to obtain the preset proper noun set.
[0062] Step 302, traverse the proper nouns, and perform character-level segmentation on the proper nouns to obtain a number of index segments.
[0063] Specifically, the server can traverse each proper noun in the set of proper nouns, segment each proper noun at the character granularity to obtain a number of index segments. Among them, the types of index segments include single letters and Chinese pinyin. The single letters include original letters and Chinese character letters. The original letters are the letters existing in the proper noun itself, and the Chinese character letters are the first letters of the pinyin of each Chinese character in the proper noun. The Chinese pinyin includes original pinyin and approximate pinyin. The original pinyin is the pinyin of each Chinese character in the proper noun itself, and the approximate pinyin is the approximate sound determined from the preset near-sound dictionary according to the original pinyin.
[0064] In one example, segmenting a proper noun at the character granularity to obtain a number of index segments can be implemented by a preset third analyzer. Taking the proper noun "Vanke A" as an example, the third analyzer splits "Vanke A" into "Wan", "Ke", and "A", obtains the pinyin "wan" of "Wan", the pinyin "ke" of "Ke", and obtains the first letter "w" of "wan", the first letter "k" of "ke", and the approximate sound "wang" of "wan". The server obtains the following index segments of "Vanke A": "Wan", "w", "wan", "wang", "Ke", "k", "ke", and "A".
[0065] Step 303: Use the index segments as index tokens to obtain an index token set, and construct a mapping relationship between the index tokens and the proper nouns to obtain an index.
[0066] Specifically, after the server segments a proper noun at the character granularity to obtain a number of index segments, it can use these index segments as index tokens to obtain an index token set, and construct a mapping relationship between the index tokens and the proper nouns to obtain an index.
[0067] In a specific implementation, to prevent duplicate index tokens, the server can use the hash value of each index segment as the id of the index segment, and directly overwrite it when a duplicate id appears.
[0068] In one example, the set of proper nouns includes keywords such as "Shanghai Pudong Development Bank", "Kweichow Moutai", "China Vanke A", and "PT Jintian". The server splits "Shanghai Pudong Development Bank" at the character granularity into four single Chinese characters: "Pu", "Fa", "Yin", and "Hang". It obtains that the first letter of the pinyin of "Pu" is "p", the pinyin of "Pu" is "pu", the first letter of the pinyin of "Fa" is "f", the pinyin of "Fa" is "fa", the first letter of the pinyin of "Yin" is "y", the pinyin of "Yin" is "yin". The server finds that the approximate sound of "yin" in the preset near-sound dictionary is "ying", the first letter of the pinyin of "Hang" is "h", the pinyin of "Hang" is "hang". The server finds that the approximate sound of "hang" in the preset near-sound dictionary is "han". The segmentation of the proper noun "Kweichow Moutai" is similar to that of the proper noun "Shanghai Pudong Development Bank", which will not be elaborated here. The server splits "China Vanke A" at the character granularity into two single Chinese characters: "Wan" and "Ke" and the original letter "a". It obtains that the first letter of the pinyin of "Wan" is "w", the pinyin of "Wan" is "wan". The server finds that the approximate sound of "wan" in the preset near-sound dictionary is "wang", the first letter of the pinyin of "Ke" is "k", the pinyin of "Ke" is "ke". The segmentation of the proper noun "PT Jintian" is similar to that of "China Vanke A", which will not be elaborated here. The server uses these index fragments as index tokens, obtains an index token set, and constructs a mapping relationship between the index tokens and the proper nouns to obtain an index. The index constructed by the server can be as shown in Table 1:
[0069] Table 1: Schematic Table of the Index Constructed by the Server (I)
[0070]
[0071]
[0072] In this embodiment, the index tokens are obtained by splitting proper names at the character granularity. The index tokens can be single letters or Chinese pinyins, all of which can be enumerated, and the construction difficulty is small. The index tokens also include approximate sounds determined from the preset near-sound dictionary according to the original pinyin. That is, the text error correction method of this application can directly correct the situation of the same sound but different characters and can also directly correct the situation of near-sound but different characters. At the same time, the near-sound dictionary can also be enumerated. Only when the proper nouns need to be added, deleted, or modified, the index token set and the index need to be updated and maintained, which greatly reduces the time required for manual maintenance, thereby further reducing the cost of text error correction.
[0073] In one embodiment, the type of the index tokens in the index token set also includes single Chinese characters. The schematic table of the index constructed by the server can be as shown in Table 2:
[0074] Table 2: Schematic Index Table (II) Constructed by the Server
[0075]
[0076]
[0077] In this embodiment, the server calculates the edit distance based on the character frequency of the word to be corrected and the character frequency of the candidate word, scores the candidate word, and obtains the score corresponding to the candidate word. This can be achieved through the following steps: Figure 4 as shown below, specifically including:
[0078] Step 401: Segment the word to be corrected at the character level to obtain several scoring segments.
[0079] Specifically, after determining the candidate word, the server can segment the word to be corrected at the character level to obtain several scoring segments. The types of the scoring segments are any one of the following: a single letter, the Chinese pinyin of a Chinese character, or a single Chinese character.
[0080] In an example, the word to be corrected is "pfa Bank". The server determines that a candidate word for "pfa Bank" is "Shanghai Pudong Development Bank". The server can call the preset fourth analyzer to segment the word to be corrected "pfa Bank" at the character level. The fourth analyzer splits "pfa Bank" into: "pfa", "yin", and "hang", and then splits "pfa" into "p" and "fa", obtaining a total of four scoring segments: "p", "fa", "yin", and "hang". The type of "p" is a single letter, the type of "fa" is the Chinese pinyin of a Chinese character, and the types of "yin" and "hang" are single Chinese characters.
[0081] Step 402: Count the character frequencies of each index term in the index term set, and determine the character frequency of the first target character and the character frequency of the second target character according to the character frequencies of each index term. The first target character is the scoring segment, and the second target character is the index term corresponding to the scoring segment.
[0082] Specifically, after segmenting the word to be corrected at the character level to obtain several scoring segments, the server can use these scoring segments as the first target characters, and use the index terms corresponding to the scoring segments as the second target characters, count the character frequencies of each index term in the index term set, and determine the character frequency of the first target character and the character frequency of the second target character according to the character frequencies of each index term.
[0083] In an example, the server can count the character frequencies of each index term in the index term set and determine the character frequency of the first target character and the character frequency of the second target character through the following steps: Figure 5 as shown below, specifically including:
[0084] Step 501, obtain the initial word frequency of each proper noun in the set of proper nouns corresponding to the index token set.
[0085] Step 502, according to the historical error correction records, determine the number of times each proper noun appears in the historical error correction records as the error correction word frequency of the proper noun.
[0086] Step 503, determine the cumulative word frequency of the proper noun according to the initial word frequency and the error correction word frequency.
[0087] Step 503, obtain the several cumulative word frequencies of several proper nouns corresponding to each index token, and sum up the several cumulative word frequencies to obtain the character frequency of the index token.
[0088] In a specific implementation, the character frequency of each index token is determined based on the word frequency of each proper noun. The word frequency of the proper noun consists of two parts: the first part is the initial word frequency of the proper noun, and the initial word frequency of the proper noun is generally set to 1, or can be set to other values according to network statistical data; the second part is the error correction word frequency, which comes from the historical error correction records. The server can determine the number of times each proper noun appears in the historical error correction records according to the historical error correction records, and use the number of times the proper noun appears in the historical error correction records as the error correction word frequency of the proper noun. Each time the proper noun appears in the historical error correction records, the error correction word frequency of the proper noun is incremented by one. After the server obtains the initial word frequency and the error correction word frequency of the proper noun, it can use the sum of the initial word frequency and the error correction word frequency as the cumulative word frequency of the proper noun, and obtain the several cumulative word frequencies of several proper nouns corresponding to each index token, and sum up the several cumulative word frequencies to obtain the character frequency of the index token.
[0089] In an example, the index schematic table constructed by the server can be as shown in Table 3. Among them, the word frequency of the proper noun "Shanghai Pudong Development Bank" is 10, the word frequency of the proper noun "Kweichow Moutai" is 13, the word frequency of the proper noun "Vanke A" is 15, and the word frequency of the proper noun "PT Jintian" is 11. According to the word frequency of the proper noun, the server can determine the character frequency of each index token as shown in Table 4:
[0090] Table 3: Index Schematic Table Constructed by the Server (III)
[0091]
[0092] Table 4: Character Frequency Statistical Table of Each Index Token (I)
[0093] Index Term Pu Fa Yin Hang Gui Zhou Word Frequency 10 10 10 10 13 13 Index Term Mao Tai Wan Ke Jin Tian Word Frequency 13 13 15 15 11 11 Index Term p f y h g Z Word Frequency 21 10 10 10 13 13 Index Term m t a pu fa ym Word Frequency 13 24 15 10 10 10 Index Term ymg hang han gm zhou zou Word Frequency 10 10 10 13 13 13 Index Term mao tai wan wang ke jin Word Frequency 13 13 15 15 15 11 Index Term jmg tian Word Frequency 11 11
[0094] Step 403, calculate the edit cost between the first target word and the second target word according to the character frequency of the first target word, the character frequency of the second target word, and a preset cost function.
[0095] Specifically, after the server determines the word frequency of the first target word and the word frequency of the second target word, it can calculate the edit cost between the first target word and the second target word according to the word frequency of the first target word, the word frequency of the second target word, and a preset cost function. The preset cost function can be set by those skilled in the art according to actual needs.
[0096] In one example, if the first target word is the same as the second target word, the edit cost is determined as: cost = 0, where cost is the edit cost. If the first target word is different from the second target word, the edit cost is determined as: cost = norm(A / B), norm(x) = (m x -m -x ) / (m x +m -x ), where norm(x) is the preset cost function, m is a preset constant, A is the word frequency of the first target word, and B is the word frequency of the second target word. Using the norm(x) function can achieve normalization processing, so that the range of the finally calculated edit cost is within [0, 1).
[0097] In one example, if the first target word is the same as the second target word, the edit cost is determined as: cost = 0, where cost is the edit cost. If the first target word is different from the second target word, the edit cost is determined as: cost = norm[lg(A / B)], norm(x) = (m x -m -x ) / (m x +m -x ), where norm[lg(x)] is the preset cost function, m is a preset constant, A is the word frequency of the first target word, and B is the word frequency of the second target word.
[0098] The function graph of norm(A / B) and the function graph of norm[lg(A / B)] can be as Figure 6As shown, for norm[lg(A / B)], if the value of A / B approaches 1, then norm[lg(A / B)] approaches 0 at this time, that is, the editing cost approaches 0. In this embodiment, the word frequencies of the first target word and the second target word are derived from the word frequencies of the indexed tokens, and the word frequencies of the indexed tokens are in turn derived from the initial word frequencies and corrected word frequencies of the proper nouns. Therefore, the values of A and B depend on the initial word frequencies and corrected word frequencies of the proper nouns. Different from the situation where the word frequencies of proper nouns remain unchanged, if the values of A and B only depend on the initial word frequencies of the proper nouns, the editing costs calculated at different times are the same, and ultimately the same corrected result will be returned; while in this application, the corrected word frequencies are combined. As the corrected word frequencies increase, when the corrected word frequency of a certain proper noun is much larger than that of other proper nouns, the word frequency of the indexed token corresponding to this proper noun is greatly affected by the corrected word frequency of this proper noun, and the word frequencies of other proper nouns have little influence on it. At this time, the values of A and B are close, that is, the value of A / B approaches 1, so that the editing cost tends to 0. In this embodiment, the value of A / B changes continuously with the user's correction behavior. By statistically quantifying the user's correction behavior for calculating the editing cost, the finally obtained corrected result is more in line with the user's intention, thereby improving the accuracy of the corrected result. Compared with norm(A / B), the change trend of norm[lg(A / B)] is flatter, and the value range of A / B in the region with an obvious change trend is wider, that is, A / B can be significantly distinguished in a larger range, which is beneficial to more refined, scientific, and reasonable determination of the editing cost between the first target word and the second target word.
[0099] As Figure 6 shown, in this application, since the value range associated with the value represented by A is very large, the value of A must be greater than or equal to the value of B, A / B is greater than or equal to 1. Therefore, for the norm(A / B) function, the value cannot be 0, and the value range is very small. When m takes 40, the value of norm(1) is already very close to 1. So for the range where A / B is greater than or equal to 1, the value of norm(A / B) approaches 1 and the change trend is not obvious, that is, different values of A / B cannot be significantly distinguished; while for the norm[lg(A / B)] function, when A / B is equal to 1, lg(A / B) is equal to 0, and the value of norm[lg(A / B)] is 0, that is, the value range of norm[lg(A / B)] is [0, 1). When A / B is equal to 10, lg(A / B) is equal to 1, and the editing cost, that is, the value of norm[lg(A / B)], tends to 1. That is, the norm[lg(A / B)] function has an obvious change trend in the interval [1, 10] and can be significantly distinguished. In order to more refinedly calculate the editing cost, if it is desired that when A is 15 times of B, the editing cost is approximately equal to 1, that is, the norm[lg(A / B)] function has an obvious change trend in the interval [1, 15], then norm[log15 (A / B)] function; if it is desired that A is 20 times B, it is edited to be approximately equal to 1, and the norm[lg(A / B)] function has an obvious changing trend in the interval [1, 20], then the norm[log 20 (A / B)] function can be used.
[0100] Step 404: Calculate the edit distance between the to-be-corrected word and the candidate word according to the editing cost.
[0101] In a specific implementation, after the server calculates the editing cost between the first target word and the second target word, it can calculate the edit distance between the to-be-corrected word and the candidate word according to these editing costs.
[0102] Step 405: Calculate the edit similarity between the to-be-corrected word and the candidate word according to the edit distance.
[0103] In an example, after the server calculates the edit distance between the to-be-corrected word and the candidate word, it can calculate the edit similarity between the to-be-corrected word and the candidate word based on the edit distance, the length of the to-be-corrected word, and the length of the candidate word.
[0104] In an example, the server can calculate the edit similarity between the to-be-corrected word and the candidate word according to the following formula based on the edit distance: Sim = 1 - lev / max(K, L), where Sim is the edit similarity, lev is the edit distance, K is the total number of scoring segments of the to-be-corrected word, and L is the total number of indexed word elements of the candidate word.
[0105] In an example, after the server calculates the edit similarity between the to-be-corrected word and the candidate word, it can determine whether the calculated edit similarity is less than a preset first threshold. If the edit similarity is less than the preset first threshold, the correction of the to-be-corrected word is terminated. If the edit similarity is too small, it indicates that the gap between the to-be-corrected word and the proper nouns in the target field is too large, and the to-be-corrected word is not a reasonable correction object. The server does not need to correct it, which can save correction resources.
[0106] Step 406: Score the candidate words according to the edit similarity to obtain the scores corresponding to the candidate words.
[0107] In a specific implementation, after the server calculates the edit similarity between the to-be-corrected word and the candidate word, it can score each candidate word according to the edit similarity and a preset scoring rule to obtain the scores corresponding to each candidate word. The higher the edit similarity of the candidate word, the higher the score, and the lower the edit similarity of the candidate word, the lower the score. The preset scoring rule can be set by those skilled in the art according to actual needs.
[0108] In this embodiment, considering that the related art simply considers the editing cost between two identical characters to be 0 and the editing cost between two different characters to be 1 when determining the editing cost between two characters, the editing cost of replacing the first target character with the second target character in various cases is regarded as the same, which cannot fully, detailed and accurately characterize the difference between the two characters. For example, replacing "d" with "silver", "y" with "silver", "yin" with "silver" and "yin" with "silver" are obviously not equivalent. Since the determination of the editing cost is a one-size-fits-all determination, the process of determining the editing distance is merely a simple superposition of the editing costs. Such an editing distance cannot fully, detailed and accurately characterize the difference between the two words. In this embodiment, the calculation of the editing cost can be refined in combination with the word frequency of the first target character and the word frequency of the second target character, and the replacements in different cases can be distinguished, so that the editing distance between the determined text and the candidate word is more refined, thereby further improving the accuracy of text error correction.
[0109] In one embodiment, if the first target word is different from the second target word, calculating the edit cost between the first target word and the second target word may be divided into the following cases.
[0110] It can be understood that the second target word is the index word corresponding to the first target word in the candidate word. The candidate word is selected from proper nouns, and proper nouns only include single Chinese characters and single letters. Since the index word in the candidate word is obtained by retrieving the search fragment from the index word set, if the second target word is a single letter, the first target word must also be a single letter. Therefore, when the first target word is different from the second target word, the second target word cannot be a single letter, that is, the second target word must be a single Chinese character. The word frequency of the second target word in this embodiment is the word frequency of the second target word itself.
[0111] If the first target character is a single Chinese character, and the server determines that the pinyin of the first target character is the same as the pinyin of the second target character, that is, the first target character and the second target character are homophones but different characters, then the server can use the frequency of the pinyin of the first target character as the frequency of the first target character; if the first target character is a single Chinese character, and the server determines that the pinyin of the first target character and the approximate pronunciation of the pinyin of the second target character are the same, that is, the first target character and the second target character are approximate pronunciations but different characters, then the server can use the sum of the frequency of the pinyin of the first target character and the frequency of the pinyin of the second target character as the frequency of the first target character.
[0112] If the first target word is a Chinese pinyin, it means that the content of the first target word entered by the user is the Chinese pinyin. If the server determines that the pinyin of the first target word is the same as that of the second target word, it means that the user entered the pinyin of the second target word. At this time, the server determines that the word frequency of the first target word is the word frequency of the Chinese pinyin represented by the first target word. If the server determines that the approximate sound of the pinyin of the first target word is the same as that of the second target word, it means that the user entered the approximate sound of the pinyin of the second target word. At this time, the server takes the sum of the word frequency of the Chinese pinyin represented by the first target word and the word frequency of the pinyin of the second target word as the word frequency of the first target word.
[0113] If the first target word is a single letter, it means that the user entered the first letter of the pinyin of the second target word, and the server takes the word frequency of this single letter itself as the word frequency of the first target word.
[0114] In this embodiment, for the cases of homophonic different characters, near-homophonic different characters, the user only enters Chinese pinyin, and the user only enters the first letter of Chinese pinyin, the word frequency of pinyin (or the word frequency of a single letter) is used as the word frequency of the first target word. At this time, the greater the ratio of the word frequency of the first target word to the word frequency of the second target word, the more Chinese characters the first target word can be associated with. Therefore, the lower the probability of replacing the first target word with the second target word and the higher the required editing cost. The smaller the ratio of the word frequency of the first target word to the word frequency of the second target word, that is, the closer the word frequency of the first target word is to the word frequency of the second target word, the fewer other Chinese characters the first target word can be associated with, and the greater the probability of replacing the first target word with the second target word and the lower the required editing cost.
[0115] In one embodiment, the server can calculate the edit distance between the error-prone vocabulary and the candidate word according to the following formula based on the editing cost:
[0116]
[0117] In the formula, a i is the i-th scoring segment of the error-prone vocabulary, a i-1 is the (i - 1)-th scoring segment of the error-prone vocabulary, b j is the j-th indexed token of the candidate word, b j-1 is the (j - 1)-th indexed token of the candidate word, lev(a i , b j ) is the edit distance between the first i scoring segments of the error-prone vocabulary and the first j indexed tokens of the candidate word, lev(a i-1 , b j ) is the edit distance between the first (i - 1) scoring segments of the error-prone vocabulary and the first j indexed tokens of the candidate word, lev(a i , b j-1) is the edit distance between the first i scoring segments of the word to be error-corrected and the first j - 1 indexed tokens of the candidate word. lev(a i-1 , b j-1 ) is the edit distance between the first i - 1 scoring segments of the word to be error-corrected and the first j - 1 indexed tokens of the candidate word. cost(a i - 1, b j-1 ) is the edit cost between the (i - 1)-th scoring segment of the word to be error-corrected and the (j - 1)-th indexed token of the candidate word. min is the minimum function.
[0118] For ease of understanding, here is a specific example. Suppose the four words to be error-corrected are "p fa y hang", "pu fa y hang", "pu fa yin hang", and "pu fa ying hang", and these four words to be error-corrected can all match the candidate word "Shanghai Pudong Development Bank". According to the method of determining the edit distance in the existing technology, that is, the edit cost between two identical characters is 0, and the edit cost between two different characters is 1, and the edit distance is the simple superposition of the edit costs. The edit distances between these four words to be error-corrected and the candidate word are all 2. However, using the method of determining the edit distance provided in this embodiment, the server first determines that the edit cost between "p" and "Pu" is 0.913, the edit cost between "pu" and "Pu" is 0.669, the edit cost between "y" and "yin" is 0.882, the edit cost between "yin" and "yin" is 0.140, and the edit cost between "ying" and "yin" is 0.409. The schematic tables of the edit distances from these four words to be error-corrected to the candidate word "Shanghai Pudong Development Bank" are shown in Tables 5 to 8.
[0119] Table 5: Schematic table of the edit distance between the word to be error-corrected "p fa y hang" and the candidate word "Shanghai Pudong Development Bank"
[0120] 0 Pu Fa Yin Hang 0 0 1 2 3 4 p 1 0.913 1.913 2.913 3.913 Fa 2 1.913 0.913 1.913 2.913 y 3 2.913 1.913 1.796 2.796 Hang 4 3.913 2.913 2.796 1.796
[0121] Table 6: Schematic table of the edit distance between the word to be error-corrected "pu fa y hang" and the candidate word "Shanghai Pudong Development Bank"
[0122] 0 Pu Fa Yin Hang 0 0 1 2 3 4 pu 1 0.669 1.669 2.669 3.669 Fa 2 1.669 0.669 1.669 2.669 y 3 2.669 1.669 1.552 2.552 Hang 4 3.669 2.669 2.552 1.552
[0123] Table 7: Schematic table of the edit distance between the word to be error-corrected "pu fa yin hang" and the candidate word "Shanghai Pudong Development Bank"
[0124] 0 Pu Fa Yin Hang 0 0 1 2 3 4 pu 1 0.669 1.669 2.669 3.669 Fa 2 1.669 0.669 1.669 2.669 yin 3 2.669 1.669 0.809 1.809 Hang 4 3.669 2.669 1.809 0.809
[0125] Table 8: Schematic table of the edit distance between the word to be error-corrected "pu fa ying hang" and the candidate word "Shanghai Pudong Development Bank"
[0126] 0 Pu Fa Yin Hang 0 0 1 2 3 4 pu 1 0.669 1.669 2.669 3.669 Fa 2 1.669 0.669 1.669 2.669 ying 3 2.669 1.669 1.078 2.078 Hang 4 3.669 2.669 2.078 1.078
[0127] As shown in Tables 5 to 8, "pu" is the pinyin of "Pu", "yin" is the pinyin of "yin", and the edit distance between "pufayinhang" and "Pufa Bank" should be very small. In this embodiment, the edit distance between "pufayinhang" and "Pufa Bank" is 0.809, which is the smallest edit distance between the four words to be corrected and the candidate words, while the difference between "pfayhang" and "Pufa Bank" is large. The edit distance between "pfayhang" and "Pufa Bank" is 1.796, which is the largest edit distance between the four words to be corrected and the candidate words. According to the edit distance scheme in the prior art, the edit distance between all the words to be corrected and the candidate word "Pufa Bank" is 2, so it is impossible to further distinguish. This application uses the word frequency of the words to be corrected and the word frequency of the candidate words for statistical quantitative calculation, thereby realizing the refinement and improvement of the edit distance, so that various situations can be distinguished, thereby improving the accuracy of the text correction process.
[0128] In one embodiment, for the error-correcting word "p发y行", the server also determines the candidate word "普发科技银行". The edit distance between the error-correcting word "p发y行" and the candidate word "普发科技银行" is shown in Table 9.
[0129] Table 9: Schematic diagram of the edit distance between the word to be corrected "p發y行" and the candidate word "浦发科技银行"
[0130] 0 Pu Fa Ke Ji Yin Hang 0 0 1 2 3 4 5 6 p 1 0.913 1.913 2.913 3.913 4.913 5.913 Fa 2 1.913 0.913 1.913 2.913 3.913 4.913 y 3 2.913 1.913 1.913 2.913 3.796 4.796 Hang 4 3.913 2.913 2.913 2.913 3.913 3.796
[0131] As shown in Tables 5 and 8, the "科" and "技" in the candidate word "普发科技银行" are the second target words that cannot be matched by the to-be-corrected vocabulary "p发y行". Therefore, when "p发y行" is replaced with "普发科技银行", "科" and "技" are addition operations from scratch, so the editing costs are both 1, and the editing distance between "p发y行" and "普发科技银行" is 2 greater than the editing distance between "p发y行" and "普发银行".
[0132] In one embodiment, the error correction word is "yinhan", and the candidate words matched by the server include "英汉", "银行", "导航", "硬汉" and "隐含". The word frequency of each index word element determined by the server is shown in Table 10:
[0133] Table 10: Word frequency statistics of each index word (II)
[0134] Index Term Yin Hang Ying Han Ying Yin Word Frequency 170 112 84 69 1 1 Index Term Han Yin Hang yin ying han Word Frequency 1 3 138 242 257 123 Index Term hang Word Frequency 161
[0135] The server determines that the edit distance between "yinhan" and "bank" is 0.496, the edit distance between "yinhan" and "English-Chinese" is 0.840, the edit distance between "yinhan" and "tough guy" is 1.202, the edit distance between "yinhan" and "implied" is 1.891, and the edit distance between "yinhan" and "pilotage" is "1.913". According to Table 9 and the edit distances between "yinhan" and each candidate word, although the pinyin of the word to be corrected "yinhan" is the same as that of the candidate word "implied", the character frequencies of "yin" and "han" in "implied" are too small, while the character frequencies of "yin" and "han" are very large, indicating that there are many Chinese characters that can be associated with "yin" and "han". Therefore, the edit distance between "yinhan" and "implied" is still very high.
[0136] Considering that in the prior art, the edit distance from "yinhan" to each candidate word is 2, when the word frequencies of the proper nouns corresponding to each candidate word are all 1, it is impossible to distinguish which candidate word is used to correct "yinhan". If the word frequency of the candidate word "pilotage" becomes 2 at this time, the final corrected result returned is "pilotage", that is, in the prior art, the word frequency of the candidate word has a decisive effect on the final corrected result. Even if the change range of the word frequency is very small, it still has a decisive effect on the corrected result, which is inevitable. In this application, the word frequency of the proper noun is only the basis for determining the character frequency of the first target character and the character frequency of the second target character. When the word frequency of "pilotage" is 1, the edit distance between "yinhan" and "pilotage" is lev = norm[lg(242 / 1)] + norm[lg(284 / 1)] = norm[lg(242)] + norm[lg(284)] = 1.913. Even if the word frequency of "pilotage" is 2, the server calculates that the edit distance between "yinhan" and "pilotage" is lev = norm[lg(242 / 2)] + norm[lg(286 / 2)] = norm[lg(121)] + norm[lg(142)] = 1.197. Since the original word frequency of "pilotage" is too small, the edit distance has changed significantly. Although the edit distance between "yinhan" and "pilotage" has changed, it is still greater than the edit distance between "yinhan" and "bank". The change in the edit distance does not directly determine the final corrected result due to the change in the word frequency, and the final corrected result is still "bank".
[0137] In one embodiment, the server scores the candidate words according to the edit similarity, and the following steps can be implemented: Figure 7 as shown below, specifically including:
[0138] Step 601, obtain the cumulative word frequency of the proper nouns corresponding to the candidate words and the total word frequency of the proper noun set.
[0139] In a specific implementation, when the server scores the candidate words, it can first obtain the cumulative word frequency of the proper nouns corresponding to the candidate words and the total word frequency of the proper noun set. The total word frequency is the sum of the cumulative word frequencies of each proper noun in the proper noun set.
[0140] Step 602, score the candidate words according to the edit similarity, the total word frequency, and the cumulative word frequency of the proper nouns corresponding to the candidate words.
[0141] In a specific implementation, the server can score the candidate words according to the following formula based on the edit similarity, the total word frequency, and the cumulative word frequency of the proper nouns corresponding to the candidate words: In the formula, Scout is the score corresponding to the candidate word, N is a preset adjustment coefficient, Sim is the edit similarity, f is the cumulative word frequency of the proper nouns corresponding to the candidate word, and F is the total word frequency. The larger the ratio of the cumulative word frequency of the proper nouns corresponding to the candidate word to the total word frequency, the more frequently the word to be corrected is used, and thus the higher the score.
[0142] In this embodiment, when the server performs text correction, it can comprehensively score the candidate words from two perspectives: character frequency and word frequency, that is, improve the edit distance. The character frequency and word frequency are two quantities that affect each other and are closely related. This not only avoids the decisive influence of word frequency on the correction result but also eliminates the problem of a large correction matching range caused by correcting at the character granularity, improving the scientificity and rationality of the text correction process and further enhancing the accuracy of text correction.
[0143] In one embodiment, the word to be corrected can be obtained through the steps as Figure 8 shown, specifically including:
[0144] Step 701, filter the obtained target words according to the preset filtering rules.
[0145] In an example, the preset filtering rules include discarding target words that contain non-Chinese characters or letters, discarding target words with less than 3 characters, discarding target words with more than 15 characters, discarding target words in all capital letters, and discarding target words that are all letters and have less than 4 characters, etc.
[0146] Step 702, check whether there is a proper noun in the preset proper noun set that is exactly the same as the filtered target word.
[0147] Step 703, if there is no proper noun in the preset proper noun set that is exactly the same as the filtered target word, then use the target word as the word to be corrected.
[0148] In this embodiment, if the server can find a proper noun in the preset proper noun set that is exactly the same as the filtered target word, it indicates that the target word is correct and no error correction is required. That is, this embodiment does not perform error correction on correct words and words that do not conform to the filtering rules, which can save error correction resources and further reduce the cost of text error correction.
[0149] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of the algorithm and process, are all within the protection scope of this patent.
[0150] Another aspect of the present application relates to an electronic device, such as Figure 9 shown, including: at least one processor 801; and a memory 802 communicatively connected to the at least one processor 801; wherein, the memory 802 stores instructions executable by the at least one processor 801, and the instructions are executed by the at least one processor 801 to enable the at least one processor 801 to execute the text error correction method in the above various embodiments.
[0151] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art. Therefore, they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0152] The processor is responsible for managing the bus and normal processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when performing operations.
[0153] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.
[0154] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0155] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A text error correction method, characterized in that, Including: Segment the error correction word at the character granularity to obtain a number of retrieval segments; wherein, the type of the retrieval segment is a single letter or Chinese pinyin; In a preset set of index tokens, determine the target index token that is the same as the retrieval segment; wherein, the types of the index tokens in the set of index tokens include single letters and Chinese pinyin; Retrieve in a preset index according to the target index token, and obtain a number of proper nouns in the same order as the target index token as candidate words; wherein, the index is a set of mapping relationships between the preset index tokens and the proper nouns; Calculate the edit distance according to the character frequency of the error correction word and the character frequency of the candidate word, score the candidate word, and obtain the score corresponding to the candidate word; Use the candidate word with the highest score as the error correction result to replace the error correction word; Among them, the calculating the edit distance according to the character frequency of the error correction word and the character frequency of the candidate word, scoring the candidate word, and obtaining the score corresponding to the candidate word includes: Segment the error correction word at the character granularity to obtain a number of scoring segments; Count the character frequencies of each index token in the set of index tokens, and determine the character frequency of the first target character and the character frequency of the second target character according to the character frequencies of each index token; wherein, the first target character is the scoring segment, and the second target character is the index token corresponding to the scoring segment in the candidate word; Calculate the edit cost between the first target character and the second target character according to the character frequency of the first target character, the character frequency of the second target character, and a preset cost function; Calculate the edit distance between the error correction word and the candidate word according to the edit cost; Calculate the edit similarity between the error correction word and the candidate word according to the edit distance; Score the candidate word according to the edit similarity to obtain the score corresponding to the candidate word.
2. The text error correction method according to claim 1, characterized in that, The set of index tokens and the index are obtained through the following steps: Obtain a preset set of proper nouns, and the set of proper nouns includes a number of the proper nouns; Traverse the proper nouns, segment the proper nouns at the character granularity to obtain a number of index segments; wherein, the types of the index segments include single letters and Chinese pinyin, the single letters include original letters and Chinese characters letters, the original letters are the letters existing in the proper noun itself, the Chinese characters letters are the first letters of the pinyin of each Chinese character in the proper noun, the Chinese pinyin includes original pinyin and approximate pinyin, the original pinyin is the pinyin of each Chinese character in the proper noun itself, and the approximate pinyin is the approximate sound determined from a preset near-sound dictionary according to the original pinyin; Use the index segments as the index tokens to obtain the set of index tokens, and construct the mapping relationship between the index tokens and the proper nouns to obtain the index.
3. The text error correction method according to claim 1 or 2, characterized in that, The types of the index tokens in the set of index tokens also include single Chinese characters, and the types of the scoring segments are any one of the following: single letters, Chinese pinyin, or single Chinese characters.
4. The text error correction method according to claim 3, wherein Statistically calculating the character frequency of each indexing token in the indexing token set, including: Obtaining the initial word frequency of each proper noun in the proper noun set corresponding to the indexing token set; Determining, according to the historical error correction records, the number of occurrences of each proper noun in the historical error correction records as the error correction word frequency of the proper noun; Determining the cumulative word frequency of the proper noun according to the initial word frequency and the error correction word frequency; Obtaining the several cumulative word frequencies of several proper nouns corresponding to each indexing token, statistically summing the several cumulative word frequencies, and obtaining the character frequency of the indexing token.
5. The text error correction method according to claim 3, wherein Calculating the edit cost between the first target word and the second target word according to the character frequency of the first target word, the character frequency of the second target word, and a preset cost function, including: If the first target word is the same as the second target word, determine that the editing cost is: , where cost is the editing cost; If the first target word is different from the second target word, determining that the edit cost is: , , where is the preset cost function, m is a preset constant, A is the word frequency of the first target word, and B is the word frequency of the second target word.
6. The text error correction method according to claim 5, characterized in that, If the first target word is different from the second target word, the calculating the edit cost between the first target word and the second target word further includes: If the first target word is a single Chinese character and the pinyin of the first target word is the same as the pinyin of the second target word, taking the character frequency of the pinyin of the first target word as the character frequency of the first target word; If the first target word is a single Chinese character and the approximate sound of the pinyin of the first target word is the same as the pinyin of the second target word, taking the sum of the character frequency of the pinyin of the first target word and the character frequency of the pinyin of the second target word as the character frequency of the first target word; If the first target word is a Chinese pinyin and the approximate sound of the first target word is the same as the pinyin of the second target word, taking the sum of the character frequency of the Chinese pinyin represented by the first target word and the character frequency of the pinyin of the second target word as the character frequency of the first target word.
7. The text error correction method according to claim 5 or 6, characterized in that, Calculating the edit distance between the to-be-error-corrected vocabulary and the candidate word according to the following formula based on the edit cost: Among them, is the i-th scoring segment of the vocabulary to be error-corrected, is the (i - 1)-th scoring segment of the vocabulary to be error-corrected, is the j-th indexed token of the candidate word, is the (j - 1)-th indexed token of the candidate word, is the edit distance between the first i scoring segments of the vocabulary to be error-corrected and the first j indexed tokens of the candidate word, is the edit distance between the first (i - 1) scoring segments of the vocabulary to be error-corrected and the first j indexed tokens of the candidate word, is the edit distance between the first i scoring segments of the vocabulary to be error-corrected and the first (j - 1) indexed tokens of the candidate word, is the edit distance between the first (i - 1) scoring segments of the vocabulary to be error-corrected and the first (j - 1) indexed tokens of the candidate word, is the edit cost between the (i - 1)-th scoring segment of the vocabulary to be error-corrected and the (j - 1)-th indexed token of the candidate word, and min is the minimum value function.
8. The text error correction method according to claim 4, wherein The scoring the candidate word according to the edit similarity includes: Obtaining the cumulative word frequency of the proper noun corresponding to the candidate word and the total word frequency of the proper noun set; wherein, the total word frequency is the sum of the cumulative word frequencies of each proper noun in the proper noun set; Scoring the candidate word according to the edit similarity, the total word frequency, and the cumulative word frequency of the proper noun corresponding to the candidate word.
9. The text error correction method according to claim 1, characterized in that, Segmenting the to-be-error-corrected vocabulary at the character granularity to obtain several retrieval segments, including: Obtaining the type of the to-be-error-corrected vocabulary; wherein, the type includes a mixed type of letters and Chinese characters, a full-letter type, and a full-Chinese-character type; Segmenting the to-be-error-corrected vocabulary at the character granularity corresponding to the type to obtain several retrieval segments.
10. The text error correction method according to claim 3, characterized in that Before scoring the candidate word according to the edit similarity to obtain the score corresponding to the candidate word, including: Judging whether the edit similarity is less than a preset first threshold; If the edit similarity is less than the preset first threshold, terminating the error correction of the to-be-error-corrected vocabulary.
11. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text error correction method according to any one of claims 1 to 10.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the text error correction method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Intelligent error correction method and device for proper nouns, equipment and storage medium
CN111428494A
Chinese character pattern cognition similarity computing method
CN102393850A
Text error correction method and device
CN110032722A