A Chinese text error correction method based on online search assistance

Through the online search-assisted method, search engine query and GPT-2 model are used to solve the problem of small data set size and insufficient professional vocabulary correction capabilities of Chinese text error correction technology, and realize high accuracy error correction that is ready-to-use.

CN115033773BActive Publication Date: 2025-07-25ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210742412.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-07-25
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing Chinese text error correction technology has small scale, low accuracy and low universality, especially for professional vocabulary error correction capabilities, and requires targeted training to be used, so it cannot be used immediately.

Method used

Through online search-assisted methods, search engines use search engines to query word frequency tables and context information, filter candidate words based on word frequency, pinyin editing distance, structural similarity and topsis algorithms, and use the GPT-2 model to calculate the confusion, and automatically correct errors in the statement.

Benefits of technology

It realizes immediate Chinese text error correction, improves error correction accuracy and universality, can handle professional vocabulary errors, and avoids the need for repeated training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115033773B_ABST
    Figure CN115033773B_ABST
Patent Text Reader

Abstract

The present invention discloses a Chinese text error correction method based on online search assistance. First, the sentence to be corrected is segmented, and an online query is performed through a search engine to crawl and count the word frequencies to construct a word frequency table. Then, the original sentence is tokenized, and suspicious words are detected according to the word frequencies and perplexities of the tokenized results in the word frequency table. The suspicious words are searched based on the context of the original sentence and the context information retrieved from the search engine. The topsis algorithm is used to score according to the word frequency, pinyin edit distance, and structural similarity to form candidate words, and some near-sound and near-shape words are also added as candidate words. The candidate words are used to replace the suspicious words in the original sentence, and the perplexity is calculated using the original GPT-2 model. The sentence with the minimum perplexity is selected as the final corrected result. This method can automatically correct Chinese sentences that may contain errors immediately without the need for additional training of an error correction model or a large-scale data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and particularly relates to a Chinese text error correction method based on online search assistance. Background Art

[0002] Text error correction is a technology for automatically checking and correcting sentences. A text error correction system can output a correct sentence for an input sentence that may contain errors. Through text error correction technology, the quality of the text can be improved, which is one of the cornerstones of the field of natural language processing. According to statistics, in new media fields such as the Internet, the text error rate is higher than 2%; in the field of speech recognition, the error rate can reach up to 8-10% at most. Although the text error rate seems not high, the appearance of a single misspelled word in a sentence may completely change the original meaning of the whole sentence, which may cause readers to misunderstand the author's meaning, and then have a bad impact. For example, in the medical field, the impact caused by a single misspelled character may be fatal. Moreover, most natural language processing technologies need to operate on correct texts and cannot obtain good results on texts with errors. Therefore, it is very meaningful to have a method that can automatically correct text.

[0003] Nowadays, globally, the research on English text error correction has been relatively mature. However, the Chinese text error correction technology is still not perfect. At present, the scale of the Chinese text error correction dataset is small, resulting in low accuracy and universality of Chinese text error correction. Secondly, Chinese text error correction depends on the data on which the model is trained. However, many professional terms are difficult to appear in the dataset, which leads to the fact that most current Chinese text error correction models cannot correctly correct sentences containing professional word errors. Finally, most existing Chinese text error correction models need to be trained specifically using the corresponding dataset first and then can perform text error correction. If the text type is changed, the model needs to be retrained and cannot be used immediately. Summary of the Invention

[0004] The object of the present invention is to provide a Chinese text error correction method based on online search assistance in view of the deficiencies of the prior art. The method first divides the sentence to be corrected into clauses, and then sequentially queries the divided sentences through a search engine online. The content retrieved is crawled and the word frequencies therein are counted to construct a word frequency table. Then, the original sentence is segmented, and suspicious words are detected based on the word frequencies and perplexities of the segmented results in the word frequency table. The suspicious words are searched according to the context of the original sentence and the context information retrieved in the search engine, and scored using the topsis algorithm based on word frequency, pinyin edit distance, and structural similarity. The words with higher scores are selected as candidate words, and some additional words with similar pronunciations and shapes are also used as candidate words. The candidate words are used to replace the suspicious words in the original sentence, and the perplexity of the sentence after replacement is calculated using the original GPT-2 model. The sentence with the minimum perplexity is selected as the final corrected result. Based on the query results of the sentence to be corrected in the search engine, this method can intelligently check the error part of the sentence to be corrected, then correct it, and finally return the correct sentence. In addition, because the present invention is based on online queries, a large amount of data similar in meaning to the sentence to be corrected can be obtained through the search engine. Error detection is performed using methods such as word frequency, the current occurrence probability of words, and n-gram. The topsis algorithm is used to screen out more appropriate replacement words for the error words considering three factors: word frequency, pinyin edit distance, and structural similarity. And using the GPT-2 model with perplexity as the evaluation index can obtain the most appropriate correct sentence output. After obtaining the required model, the present invention does not require additional training and can be used immediately without being interfered by other factors.

[0005] The object of the present invention is achieved by the following technical solutions:

[0006] A Chinese text error correction method based on online search assistance, comprising the following steps:

[0007] S1: Divide the original sentence to be corrected into clauses, and the basis for clause division is the number of words contained in the original sentence.

[0008] S2: Query the sentences obtained by dividing in step S1 through a search engine, and crawl and save the titles and abstract parts of the first thirty query results locally.

[0009] S3: Based on the thirty query results obtained in step S2, perform word segmentation and word frequency statistics, and then construct a word frequency table.

[0010] S4: Based on the thirty query results obtained in step S2, perform new word discovery on them, add the results of new word discovery to the jieba word list, and then segment the original sentence according to the changed jieba word list to obtain the segmented result of the original sentence.

[0011] S5: Based on the word frequency table constructed by querying through the search engine and the word segmentation result of the original statement obtained in steps S3 and S4, error detection is performed. Use the word segmentation result of the original statement to query in the word frequency table. If the word frequency value of a certain word in the word frequency table is less than the threshold, it is considered that the word may be incorrect and is used as a suspicious word.

[0012] S6: Based on the word segmentation result of the original statement obtained in step S4, out-of-vocabulary word error detection and supplementation are performed, and the out-of-vocabulary words not in the jieba dictionary are added to the suspicious words.

[0013] S7: Probability error detection and supplementation are performed. The original statement passes through the original GPT-2 model to obtain the probability value of each character. If the probability value of a certain character is significantly less than the probability values of other characters, the character is added to the suspicious words.

[0014] S8: Based on the suspicious words obtained in step S7, the context information text_ori of the suspicious words in the original statement is obtained in turn and stored in the form of a pair of a suspicious word and the corresponding context information text_ori. The way to obtain the context information is to obtain the words within a distance of x and within from the suspicious word in the original statement as the context information text_ori according to the distance.

[0015] S9: Based on the context information text_ori of the suspicious words in the original statement obtained in steps S8 and S2 and the results queried through the search engine, the context information text_search of text_ori in the results queried through the search engine is obtained in turn. At this time, the context information text_search is used as the candidate word and stored in the form of a pair of a suspicious word and the corresponding candidate word. Still using the method according to the distance, the words at distances of 2, 4, and 6 from text_ori in the results queried through the search engine are respectively obtained as the context information text_search.

[0016] S10: Based on the candidate words obtained in step S9, the pinyin edit distance and structural similarity between the candidate words and the corresponding suspicious words are calculated respectively. Among them, the structural similarity is calculated using a pre-constructed siamese network.

[0017] S11: Based on the pinyin edit distance, structural similarity, and word frequency table between the candidate words and the corresponding suspicious words obtained in steps S10 and S3, the topsis algorithm is used to calculate the scores based on the word frequency, pinyin edit distance, and structural similarity, and the top 8 candidate words with the highest scores are selected as the candidate words for the suspicious words.

[0018] S12: Based on the suspicious words obtained in step S7, select the words with a relatively small edit distance in the jieba vocabulary from the pinyin of the suspicious words and the words with a relatively high structural similarity, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10.

[0019] S13: Based on the search engine query results obtained in step S2, construct a 3-gram vocabulary, use the n-gram algorithm to select the words that may appear in the position of the suspicious words, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10.

[0020] S14: Based on the candidate words for the suspicious words obtained in step S13 and the suspicious words obtained in step S6, perform permutation and combination replacement of the suspicious words in the original sentence with the corresponding candidate words for the suspicious words to obtain a set of candidate sentences. Since the original sentence may be correct, the original sentence is also added to the set of candidate sentences.

[0021] S15: Based on the set of candidate sentences obtained in step S14, use the original GPT-2 model to calculate the perplexity of the entire candidate sentence, and select the sentence with the lowest perplexity as the final result.

[0022] Further, the original sentence to be corrected is segmented, and the segmentation condition is as follows: first, use jieba to segment the original sentence. If the number of words in the segmentation result of the original sentence is greater than or equal to 15, then the original sentence is segmented according to the full stop, question mark, exclamation mark, and semicolon. If the number of words in the segmented short sentence is still greater than or equal to 15, then continue to segment according to the comma.

[0023] Further, for the segmented sentence, the website prefix used for searching in the search engine specifically refers to: splicing the query content into the website prefix, using a crawler to access, crawling the search results in the search engine, and saving the titles and abstracts of the first thirty pieces of information in the search results locally.

[0024] Further, according to the query results, perform word segmentation and statistical word frequency, and then construct a word frequency table, specifically referring to: using jieba to segment the titles and abstracts of the thirty pieces of query results crawled, dividing them into words, then counting the number of times each word appears, and the number of appearances is the word frequency of this word. Save each word and its corresponding word frequency as the word frequency table.

[0025] Furthermore, for the new word discovery of the query results, adding the results of new word discovery to the Jieba dictionary, and then performing word segmentation on the original statement, specifically refers to: The new word discovery algorithm can mine from existing corpora to find those unregistered phrase that may form words. New words refer to newly emerged words or words with new meanings for old words. Since Jieba word segmentation is used and it depends on the Jieba dictionary, and the Jieba dictionary is relatively old and cannot correctly segment some relatively new words. Using the new word discovery algorithm to screen out relevant phrases that may form words and re-segment the original statement can improve the word segmentation accuracy. The new word discovery algorithm mainly consists of three steps: a) Generate an n-gram table from the corpus text and count the word frequency of each word. b) Use the cohesion degree to screen out candidate new words from the previous n-gram table. c) Then screen out the final new words from the candidate new words through the freedom degree. Given an original statement S, the word segmentation results are x1, x2, …, x n 。

[0026] The cohesion degree is represented by pointwise mutual information, and the formula is:

[0027]

[0028] Among them, PMI(x,y) is the pointwise mutual information, p(x,y) refers to the probability that two words appear together, and p(x), p(y) refer to the probabilities of each word appearing. The greater the cohesion degree, the greater the probability that these two words appear together, and the greater the possibility of being a word.

[0029] The freedom degree is represented by left and right entropy, and the left and right entropy formulas are respectively:

[0030]

[0031]

[0032] Among them, E Left (PreW) represents the left entropy, E Right (SufW) represents the right entropy, PreW is the set of prefixes of word W, and SufW is the set of suffixes of word W. The greater the freedom degree, the richer its surrounding words, and the greater the possibility of becoming an independent word.

[0033] Furthermore, for the query in the word frequency table using the word segmentation results of the original statement for error detection, specifically refers to: In the already constructed word frequency table, sequentially query the word frequency values c1, c2, …, c n in the word frequency table for the word segmentation results x1, x2, …, x n , select the maximum value c n from c1, c2, …, c maxAs a benchmark, if the word frequency value c of other words k(k≠max) <5% * c max , then it is considered that the word may be incorrect and is added to the suspicious words.

[0034] Furthermore, the detection and supplementation of out-of-vocabulary words specifically refer to: searching for the word segmentation results x1, x2,..., x of the original sentence in the jieba word list n . If the word x is not in the jieba word list, then it is considered that the word may be incorrect and is added to the suspicious words.

[0035] Furthermore, the probability-based error detection and supplementation specifically refer to: the sentence S is composed of words x1, x2,..., x n , and the GPT-2 model can input the true previous context x1, x2,..., x m-1 to obtain the possible words and their corresponding probabilities of the next word x m '. According to the true word x of each original sentence k to obtain the probability p k , calculate the median value p m . If the probability value of a word x m is less than 10% * p of the median value m , then it is considered that the word may be incorrect. If the word is not in the suspicious words, then it is added to the suspicious words.

[0036] Furthermore, obtaining the context information text_ori of the suspicious word in the original sentence specifically refers to: assuming there is an original sentence S, S is composed of words x1, x2, x3, x4, x5, x6, x7, x8, and the suspicious word is x3. According to the distance formula:

[0037] dis = min(3, word num / / 2)

[0038] where word num represents the number of words after word segmentation. By calculating the distance dis, obtain all the words within a distance of dis from the suspicious word x in the original sentence k as the context information text_ori of x k . text_ori = x k-dis , x k-dis+1 , …, x k-1 , x k+1 , …, x k+dis .

[0039] Further, the context information text_search of the obtained context information text_ori in the results retrieved by the search engine specifically refers to: among the thirty result contents retrieved by the local search engine, the context information text_search of the context information text_ori is obtained according to a preset distance. For example: one result of the search engine is S′, which consists of words x′1, x′2,..., x j ′, and there is an existing original statement S, which consists of words x1, x2,..., x n ; the suspicious words are x k , x k-dis , x k-dis+1 ,..., x k-1 , x k+1 ,..., x k+dis is the context information text_ori of x k . Now for each context information text_ori, the context information text_search is searched in S′ according to the distance dis search , and the distances dis search are respectively selected as 2, 4, 6. For x k-dis in the context information text_ori, if x k-dis also appears in S′, then x′ k-dis-2 , x′ k-dis-1 , x k-dis+1 ′, x k-dis+2 ′ is the context information text_search of x k-dis in S′ with a distance of 2, and x′ k-dis-4 ,..., x′ k-dis-1 , x′ k-dis+1 ,..., x′ k-dis+4 is the context information text_search of x1 in S′ with a distance of 4, and x′ k-dis-6 ,..., x′ k-dis-1 , x′ k-dis+1 ,..., x′ k-dis+6 is the context information text_search of x1 in S′ with a distance of 6. And so on for x k-dis+1 ,..., x k-1 , x k+1 ,..., x k+dis , the same operation is performed to obtain the context information text_search. These context information text_search are used as the candidate words for the corresponding suspicious words.

[0040] Here is based on an assumption. Suppose there is a sentence BAB, where the word A is incorrect. The context information of A is B, and in a correct sentence CBC, B appears in this correct sentence CBC. Then it can be considered that the context information C of B may be used to replace the incorrect A. That is, there is a word B next to the word A in the incorrect sentence, and there is a word C next to the word B in the correct sentence, then the word C may be used as a candidate for the word A. In the present invention, the incorrect sentence can be considered as the original statement, and the correct sentence is the top thirty pieces of information crawled from the search engine query results.

[0041] Furthermore, calculating the pinyin edit distance and structural similarity between the candidate word and the corresponding suspicious word specifically refers to: The edit distance is the minimum number of edits to change one string to another, where each edit can only insert a character, delete a character, or modify a character in the string. And the pinyin edit distance is the edit distance after converting two Chinese characters into pinyin without tone marks. The pinyin edit distance formula is as follows:

[0042] pydis = LS(py1, py2)

[0043] Where pydis represents the pinyin edit distance, LS represents the edit distance calculation, and py1 and py2 respectively represent the pinyin without tone marks of the two words.

[0044] The structural similarity is obtained by using a pre-trained siamese network to calculate the graphic similarity. The siamese network is a conjoined neural network, and the two neural networks in it share parameter weights. The siamese neural network has two inputs graph1 and graph2. After putting the two inputs into the two neural networks Network1 and Network2, after obtaining the backbone feature extraction network, we can obtain a multi-dimensional feature, flatten it into one dimension, and then obtain the one-dimensional vectors of the two inputs. Subtract these two one-dimensional vectors, and then sum the absolute values, which is equivalent to obtaining the distance between the two one-dimensional vectors. Then fully connect this distance and take the sigmoid of the result to make its value between 0 and 1, representing the similarity degree of the two input pictures.

[0045] Because there is no good Chinese character picture dataset, the present invention designs a set of Chinese character picture datasets for calculating the graphic similarity of Chinese characters. Using OpenCV, generate pictures of different font forms for each Chinese character c k of the same Chinese character c generate pictures k of different font pictures of the same Chinese character c The parts are regarded as the same type, and different Chinese characters are regarded as different types. During training, when two inputs point to pictures of the same type, the label is 1 at this time. When two inputs point to pictures of different types, the label is 0 at this time. Then, the cross-entropy operation is performed on the output result of the network and the true label, which can be used as the final loss. The structural similarity formula is as follows:

[0046] similarity = Graphsimi(graph1, graph2)

[0047] Among them, similarity is the structural similarity, Graphsimi is the siamese network model, and graph1 and graph2 are the word pictures generated using two words respectively.

[0048] Furthermore, the TOPSIS algorithm is used to calculate the score based on word frequency, pinyin edit distance, and structural similarity. Specifically, it means that the TOPSIS algorithm is a method for ranking according to the degree of closeness of a finite number of evaluation objects to the ideal target. First, the data is normalized. Normalization means processing each evaluation index to be better as it gets larger. In the present invention, the evaluation indexes are the word frequency of the candidate word, the pinyin edit distance between the candidate word and the suspicious word, and the structural similarity between the candidate word and the suspicious word. The word frequency and structural similarity are better as they get larger, so no processing is required. However, the pinyin edit distance is better as it gets smaller, so the reciprocal of the pinyin edit distance is taken as the evaluation index. That is, the evaluation indexes are the following three points: ① word frequency, ② ③similarity. Then, the data is standardized to eliminate the influence of different data index dimensions. The standardization formula is as follows:

[0049]

[0050] Among them, z ij represents the value of the j-th index of the i-th scheme after standardization, and x ij represents the j-th index of the i-th scheme in the original data.

[0051] Then, the optimal ideal value z + and the worst ideal value z - are determined. The attribute values of the optimal ideal value z + are the best values among the candidate schemes, that is, the largest values in each index. The worst ideal value z - is the smallest value in each index. Then, the Euclidean distances between each scheme and the optimal ideal value and the worst ideal value are calculated. For the i-th scheme z i , its distance formulas from the optimal solution and the worst solution are as follows:

[0052]

[0053]

[0054] Among them, represents the distance between the i-th solution z i and the optimal solution, represents the distance between the i-th solution z i and the worst solution, m represents the number of indicators, represents the maximum value of the j-th indicator, z ij represents the value of the j-th indicator of the i-th solution after standardization.

[0055] The scoring formula for the i-th solution is as follows:

[0056]

[0057] Among them, S i represents the score of the i-th solution. The score of the i-th solution is obtained according to each dissearch

[0058] From this, the closeness of each solution to the optimal solution is obtained. Multiply the search obtained with different distances dis respectively by the corresponding as the final score, that is:

[0059]

[0060] Among them, score is the final score, used as the criterion for evaluating the quality of the solution, and select the top eight candidate words with the largest score as the more reasonable candidate words. Generally speaking, the formula is as follows:

[0061]

[0062] Among them, is the candidate word corresponding to the suspicious word x k TOP-8 means taking the top 8.

[0063] Furthermore, the words and structures with relatively small phonetic edit distance and relatively high structural similarity to the suspicious word in the jieba word list are screened out, specifically: the suspicious word x k is segmented into n characters c1, c2,..., c n at the character granularity, and then for each character c m (1 ≤ m ≤ n), all characters with a phonetic edit distance less than or equal to 1 are found in the jieba word list Arrange and combine these characters in the order of subscript m as the position to generate Words, if these words are logged in the Jieba word list, they are also used as candidate words for suspicious words. For structural similarity, a pre-trained Siamese network is used to calculate the similarity with the suspicious word x k Words with a picture similarity greater than 80% are used as candidate words for suspicious words and added to the candidate words

[0064] Furthermore, the use of the n-gram algorithm to select words that may appear in the position of the suspicious word specifically means that: for the n-gram algorithm, we assume that the probability of the nth word appearing is only related to the previous n - 1 words. Therefore, the probability distribution of the sentence is as follows:

[0065]

[0066] Among them, P(s) is the probability of the entire sentence, w n is the word that makes up the sentence, represents the historical sequence from the w i th word to the w i-n+1 th word, represents the probability that the current word appears given the historical sequence of words. In the present invention, using the 3-gram algorithm, the probability distribution of the sentence is as follows:

[0067] P(s) = P(w1|w0, w -1 )P(w2|w1, w0)…P(w i |w i-1 , w i-2 )

[0068] Construct a 3-gram table for the thirty pieces of information obtained by querying through the search engine using the 3-gram algorithm, and store it in the form of (w i-1 , w i-2 , w i ). Traverse the original sentence S to be corrected. S is composed of words x1, x2,..., x n . If there exists x k-2 = w k-2 , x k-1 = w k-1 , then w k is considered a candidate word that can be used to replace x k and added to the candidate words

[0069] Furthermore, the replacement of the suspicious words in the original sentence with the candidate words corresponding to the suspicious words in a permutation and combination manner to obtain a candidate sentence set specifically means that: for the original sentence S, S is composed of words x1, x2,..., x n . There are existing suspicious words x i , x j ,..., xk , the candidate words corresponding to the suspicious words are The suspicious words x m (m∈(i,j,...,k)) with x m Candidate words for Replace, make permutations and combinations, and generate a set of candidate sentences.

[0070] Furthermore, the use of the original GPT-2 model to calculate the perplexity of the entire candidate sentence specifically means: for a given sentence, if its length is n, first shift it one position to the left as a label, remove the last position as input, input the input to GPT-2 to obtain the output and label to perform cross entropy loss, and then raise it to a power with a natural number as the base to obtain the desired perplexity. Perplexity is a method used to measure the quality of language probability models. The smaller the perplexity, the more reasonable the sentence. The formula for calculating perplexity is as follows:

[0071] out = GPT-2 (input)

[0072] loss=CrossEntropyLoss(out,label)

[0073] PPL=ln loss

[0074] In which, given a sentence S, S consists of words x1, x2, ..., x n Composition, input = x1, x2, ..., x n-1 , label = x2, ..., x n , PPL is the required perplexity. The perplexity of all candidate sentences and the original sentence is calculated, and the sentence with the smallest perplexity is selected as the final error correction result.

[0075] The beneficial effects of the present invention are as follows: the present invention provides a new Chinese text error correction method, which uses an online query method, detects errors in sentences to be corrected based on word frequency factors and probability, and uses the topsis algorithm to screen replaceable words for the wrong words based on word frequency, pinyin edit distance, and structural similarity for the wrong part of the sentence. The wrong words are replaced with replaceable words, and finally the optimal solution is obtained as the error correction result by using the GPT-2 model to calculate the perplexity. The problem that different data sets need to be trained and used in batches and professional vocabulary cannot be corrected is solved, and Chinese text error correction can be performed instantly. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 is a flow chart of the method proposed by the present invention; DETAILED DESCRIPTION

[0077] The present invention discloses a Chinese text error correction method based on online search assistance. Based on the query results of the sentence to be corrected in a search engine, it intelligently checks the error part of the sentence to be corrected, then corrects it, and finally returns the correct sentence. Based on the Chinese text error correction technology provided by the present invention, it can correct Chinese sentences that may contain errors, enabling readers to better understand the author's intention, and also increasing the accuracy of subsequent natural language processing technologies.

[0078] The present invention discloses a Chinese text error correction method based on online search assistance, which can automatically correct Chinese sentences that may contain errors immediately without the need for additional training of an error correction model or a large-scale data set. This method first divides the sentence to be corrected into clauses, queries the divided sentences through a search engine in turn, crawls the retrieved content and counts the word frequencies therein to construct a word frequency table. Then, the original sentence is segmented, and suspicious words are detected based on the word frequencies and perplexities of the segmented results in the word frequency table. The suspicious words are searched according to the context of the original sentence and the context information of the content retrieved in the search engine, and scored using the topsis algorithm based on word frequency, pinyin edit distance, and structural similarity. The words with higher scores are selected as candidate words, and some additional near-sound and near-shape words are also used as candidate words. The candidate words are used to replace the suspicious words in the original sentence, and the perplexity of the sentence after replacement is calculated using the original GPT-2 model. The sentence with the minimum perplexity is selected as the final corrected result. The present invention can detect errors by means of online query according to word frequency and perplexity, obtain candidate words by comprehensively considering word frequency, pinyin edit distance, and structural similarity, use the candidate words to replace the corresponding suspicious words to obtain candidate sentences, and finally obtain the optimal correct sentence after modification according to the overall perplexity of the candidate sentences.

[0079] The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. The purpose and effect of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0080] As Figure 1 shown, a Chinese text error correction method based on online search assistance includes the following steps:

[0081] S1: Divide the original sentence to be corrected into clauses, and the basis for clause division is the number of words contained in the original sentence.

[0082] S2: Query the sentences obtained by clause division in step S1 through a search engine, and crawl and save the titles and abstract parts of the first thirty query results locally.

[0083] S3: Based on the thirty query results obtained in step S2, perform word segmentation and statistical word frequency, and then construct a word frequency table.

[0084] S4: Based on the thirty query results obtained in step S2, perform new word discovery on them, add the results of new word discovery to the jieba word list, and then segment the original sentence according to the changed jieba word list to obtain the segmented result of the original sentence.

[0085] S5: Based on the word frequency table and the segmented result of the original sentence constructed through search engine queries obtained in step S3 and step S4, perform error detection. Use the segmented result of the original sentence to query in the word frequency table. If the word frequency value of a certain word in the word frequency table is less than the threshold, it is considered that the word may be incorrect and is regarded as a suspicious word.

[0086] S6: Based on the segmented result of the original sentence obtained in step S4, perform out-of-vocabulary word error detection and supplementation, and add the out-of-vocabulary words not in the jieba word library to the suspicious words.

[0087] S7: Perform probability error detection and supplementation. Pass the original sentence through the original GPT-2 model to obtain the probability value of each character. If the probability value of a certain character is significantly less than the probability values of other characters, add the character to the suspicious words.

[0088] S8: Based on the suspicious words obtained in step S7, sequentially obtain the context information text_ori of the suspicious words in the original sentence, and store them in the form of a pair of a suspicious word and the corresponding context information text_ori. The way to obtain the context information is to obtain the words within a distance of x and within the original sentence from the suspicious word as the context information text_ori according to the distance.

[0089] S9: Based on the context information text_ori of the suspicious words in the original sentence obtained in step S8 and the results obtained through search engine queries in step S2, sequentially obtain the context information text_search of text_ori in the results obtained through search engine queries. At this time, the context information text_search is used as a candidate word. Store them in the form of a pair of a suspicious word and the corresponding candidate word. Still use the distance-based method to respectively obtain the words at distances of 2, 4, and 6 from text_ori in the results obtained through search engine queries as the context information text_search.

[0090] S10: Based on the candidate words obtained in step S9, calculate the pinyin edit distance and structural similarity between the candidate words and the corresponding suspicious words respectively. Among them, the structural similarity is calculated using a pre-constructed siamese network.

[0091] S11: Based on the pinyin edit distance, structural similarity, and word frequency table of the candidate words and the corresponding suspicious words obtained in steps S10 and S3, use the topsis algorithm to calculate scores based on word frequency, pinyin edit distance, and structural similarity, and select the top 8 candidate words with the highest scores as the candidate words for the suspicious words.

[0092] S12: Based on the suspicious words obtained in step S7, filter out the words with a smaller pinyin edit distance and a higher structural similarity to the suspicious words in the jieba word list, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10.

[0093] S13: Based on the search engine query results obtained in step S2, construct a 3-gram word list, use the n-gram algorithm to select the words that may appear in the position of the suspicious words, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10.

[0094] S14: Based on the candidate words for the suspicious words obtained in step S13 and the suspicious words obtained in step S6, perform permutation and combination replacement of the suspicious words in the original sentence with the corresponding candidate words for the suspicious words to obtain a candidate sentence set. Since the original sentence may be correct, the original sentence is also added to the candidate sentence set.

[0095] S15: Based on the candidate sentence set obtained in step S14, use the original GPT-2 model to calculate the perplexity of the entire candidate sentence, and select the sentence with the lowest perplexity as the final result.

[0096] 1. Data preprocessing

[0097] First, the original sentence to be corrected needs to be segmented. The segmentation condition is: if the number of words in the original sentence is greater than or equal to 15, then the original sentence is segmented according to the full stop, question mark, exclamation mark, and semicolon. Suppose sentence S is the current sentence to be corrected, and S is composed of words x1, x2,..., x n If n≥15, then S needs to be segmented. For example, when x i = ′. ′, then S is divided into x1, x2,..., x i and x i+1 , x i+2 ,..., x n . If the number of words in the segmented short sentence is still greater than or equal to 15, then continue to segment according to the comma.

[0098] Then, the segmented sentences are queried in the search engine in turn, and the first thirty titles and abstracts of the query results are crawled and saved. Then, they are segmented and a word frequency table Table Word is constructed. The word frequency table is a table composed of words and their corresponding occurrence times.

[0099] Then, new word discovery is performed on the crawled query results. The results of new word discovery are added to the Jieba dictionary, and then the original sentence is segmented. The new word discovery algorithm mainly consists of three steps: a) Generate an n-gram table from the corpus text and count the word frequency of each word. b) Use the cohesion degree to screen out candidate new words from the previous n-gram table. c) Then screen out the final new words from the candidate new words through the freedom degree. Given an original sentence S, the segmented results are x1, x2,..., x n 。

[0100] The cohesion degree is represented by the pointwise mutual information, and the formula is:

[0101]

[0102] Among them, PMI(x, y) is the pointwise mutual information, p(x, y) is the probability that two words appear together, and p(x), p(y) are the probabilities that each word appears. The greater the cohesion degree, the greater the probability that these two words appear together, and the greater the possibility that it is a word.

[0103] The freedom degree is represented by the left and right entropy, and the left and right entropy formulas are respectively:

[0104]

[0105]

[0106] Among them, E Left (PreW) represents the left entropy, E Right (SufW) represents the right entropy. PreW is the set of prefixes of word W, and SufW is the set of suffixes of word W. The greater the freedom degree, the richer its surrounding words, and the greater the possibility that it becomes an independent word. Add the results of new word discovery to the Jieba dictionary and segment the sentence again

[0107] 2. Error Detection Module

[0108] Then use the previously constructed word frequency table Table Word to perform error detection. In the already constructed word frequency table Table Word , sequentially query the word frequency values c1, c2,..., c n of the segmented results x1, x2,..., x n of the original sentence in the word frequency table, and select the maximum value c n in c1, c2,..., c max as the benchmark. If the word frequency value c k(k≠max) of other words < 5% * c max , then it is considered that the word may be incorrect.

[0109] Then, perform out-of-vocabulary word error detection. If the words x1, x2, ..., x n are not in the jieba vocabulary, then the word is considered incorrect.

[0110] Finally, perform probability error detection. The GPT-2 model can input the previous text to obtain the probability of the next word x′ appearing. The probability formula is as follows:

[0111] P(x k ) = GPT-2(x1, x2, ..., x k-1 )

[0112] where P(x k ) represents the probability of the k-th word appearing, and x1, x2, ..., x k-1 represents the phrase composed of the previous k - 1 words. According to the probabilities p1, p2, ..., p n of each real word, calculate the median value p m . If the probability value of a word x m is less than 10% * p m , then the word is considered possibly incorrect.

[0113] 3. Error correction module

[0114] After obtaining the suspicious words, the context information text_ori of the suspicious words in the original sentence needs to be obtained. According to the distance formula:

[0115] dis = min(3, word num / / 2)

[0116] where word num represents the number of words after word segmentation. By calculating dis, obtain all the words within a distance of dis from the suspicious word x k in the original sentence as the context information text_ori of x k . text_ori = x k-dis , x k-dis+1 , …, x k-1 , x k+1 , …, x k+dis .

[0117] Then, in the thirty result contents retrieved by the search engine, obtain the context information text_search of the context information text_ori according to the preset distance. For example: One result of the search engine is S′, which is composed of the words x′1, x′2, ..., x j ′. There is an original sentence S, which is composed of the words x1, x2, ..., x n , and the suspicious word is xk , x k-dis , x k-dis+1 , ..., x k-1 , x k+1 , ..., x k+dis is x k 's context information text_ori. Now for each context information text_ori, in S′ according to the distance dis search search for context information text_search, the distance dis search select 2, 4, 6 respectively. For the x in the context information text_ori k-dis , if x k-dis also appears in S′, then x′ k-dis-2 , x′ k-dis-1 , x k-dis+1 ′, x k-dis+2 ′ is the context information text_search of x in S′ with a distance of 2, x′ k-dis , ..., x′ k-dis-4 , ..., x′ k-dis-1 , x′ k-dis+1 , ..., x′ k-dis+4 is the context information text_search of x1 in S′ with a distance of 4, x′ k-dis-6 , ..., x′ k-dis-1 , x′ k-dis+1 , ..., x′ k-dis+6 is the context information text_search of x1 in S′ with a distance of 6. And so on for x k-dis+1 , ..., x k-1 , x k+1 , ..., x k+dis , ..., x do the same operation to obtain the context information text_search. Take these context information text_search as the candidate words for the corresponding suspicious words.

[0118] After obtaining the candidate words, because there will be a large amount of data, so it is necessary to perform a certain screening to select the more suitable candidate words. The present invention uses the word frequency of the candidate words, as well as the pinyin edit distance and structural similarity with the corresponding suspicious words as reference criteria.

[0119] The edit distance refers to the minimum number of edits to change from one string to another, where each edit can only insert a character, delete a character, or modify a character in the string. And the pinyin edit distance is the edit distance after converting two Chinese characters into pinyin without tone marks. The pinyin edit distance formula is as follows:

[0120] pydis = LS(py1, py2)

[0121] Among them, pydis represents the pinyin edit distance, LS represents the edit distance calculation, and py1 and py2 respectively represent the phonetic - less pinyin of two words.

[0122] The structural similarity is obtained by using a pre - trained siamese network to calculate the graphic similarity. The siamese network is a conjoined neural network, in which the two neural networks share parameter weights. The siamese neural network has two inputs, graph1 and graph2. After putting the two inputs into two neural networks, Network1 and Network2, and obtaining the backbone feature extraction network, we can obtain a multi - dimensional feature. Flattening it into one dimension, we get the one - dimensional vectors of the two inputs. Subtracting these two one - dimensional vectors and then summing the absolute values is equivalent to calculating the distance between the two one - dimensional vectors. Then, fully connect this distance and take the sigmoid of the result to make its value between 0 and 1, representing the similarity degree of the two input pictures.

[0123] Because there is no good Chinese character picture dataset, the present invention designs a set of Chinese character picture datasets for calculating the graphic similarity of Chinese characters. Using OpenCV, generate pictures of different font forms of each Chinese character c k For the same Chinese character c k pictures of different fonts are all regarded as the same type, and different Chinese characters are regarded as different types. During training, when the two inputs point to pictures of the same type, the label is 1 at this time. When the two inputs point to pictures of different types, the label is 0 at this time. Then, perform cross - entropy operation on the output result of the network and the true label, which can be used as the final loss. The structural similarity formula is as follows:

[0124] similarity=Graphsimi(graph1, graph2)

[0125] Among them, similarity is the structural similarity, Graphsimi is the siamese network model, and graph1 and graph2 are the word pictures generated using two words respectively.

[0126] ​In order to reasonably use the three evaluation indicators of word frequency, pinyin edit distance, and structural similarity, the present invention uses the topsis algorithm for comprehensive consideration. The topsis algorithm is a method of ranking according to the degree of proximity of a finite number of evaluation objects to the ideal target. First, the data is normalized, and the normalization process means processing each evaluation indicator to be better the larger it is. In the present invention, the evaluation indicators are the word frequency of the candidate word, the pinyin edit distance between the candidate word and the suspicious word, and the structural similarity between the candidate word and the suspicious word. The word frequency and the structural similarity are both better the larger they are, so no processing is required, while the pinyin edit distance is better the smaller it is, so the reciprocal of the pinyin edit distance is taken as the evaluation indicator. That is, the evaluation indicators are the following three points: ① word frequency, ② ③ similarity.

[0127] Then, the data is standardized to eliminate the influence of different data index dimensions. The standardization formula is as follows:

[0128]

[0129] where z ij represents the value of the j-th indicator of the i-th scheme after standardization, and x ij represents the j-th indicator of the i-th scheme in the original data.

[0130] Then, the optimal ideal value z + and the worst ideal value z - of each indicator are determined. The attribute values of the optimal ideal value z + are the best values among the candidate schemes, that is, the largest value in each indicator. The worst ideal value z - is the smallest value in each indicator. Then, the Euclidean distances between each scheme and the optimal ideal value and the worst ideal value are calculated. For the i-th scheme z i , its distance formulas from the optimal solution and the worst solution are as follows:

[0131]

[0132]

[0133] where represents the distance between the i-th scheme z i and the optimal solution, represents the distance between the i-th scheme zi and the worst solution, m represents the number of indicators, represents the maximum value of the j-th indicator, and z ij represents the value of the j-th indicator of the i-th scheme after standardization.

[0134] The scoring formula for the i-th scheme is as follows:

[0135]

[0136] Among them, S i represents the score of the i-th solution. According to each dis search obtain the score of the i-th solution From this, the degree of closeness of each solution to the optimal solution is obtained. For different distances dis search obtained are respectively multiplied by the corresponding as the final score, that is:

[0137]

[0138] Among them, score is the final score, which is used as the criterion for evaluating the quality of the solution. Select the top eight candidate words with the largest score as the more reasonable candidate words. Generally speaking, the formula is as follows:

[0139]

[0140] Among them, is the candidate word corresponding to the suspicious word x k TOP-8 means taking the top 8.

[0141] Considering that the candidate words of misspelled words may not appear in the data obtained by querying through the search engine, words similar to the suspicious words in the jieba word list are further screened out. The pinyin edit distance and structural similarity are used as the evaluation criteria. For the pinyin edit distance, the suspicious word x k is segmented into n characters c1, c2,..., c n by character granularity, and then for each character c m (1≤m≤n), all characters with a pinyin edit distance less than or equal to 1 are found in the jieba word list These characters are arranged and combined according to the subscript m as the position to generate words. If these words are registered in the jieba word list, they are also used as candidate words for suspicious words. For structural similarity, a pre-trained siamese network is used to calculate words with a picture similarity greater than 80% with the suspicious word x k as candidate words for suspicious words and added to the candidate words

[0142] Additionally, the n-gram algorithm is used to select words that may appear in the position of the suspicious word. For the n-gram algorithm, we assume that the probability of the n-th word appearing is only related to the previous n-1 words. Therefore, the probability distribution of the sentence is as follows:

[0143]

[0144] Among them, P(s) is the probability of the entire sentence, and w n is a word that makes up the sentence, representing the historical sequence from the w i -th word to the w i_n+1 -th word, and representing the probability of the current word given the historical sequence of words. In the present invention, the 3-gram algorithm is used, and the probability distribution of the sentence is as follows:

[0145] P(s) = P(w1|w0, w -1 )P(w2|w1, w0)...P(w i |w i-1 , w i-2 )

[0146] Thirty pieces of information obtained by querying through a search engine are used to construct a 3-gram table with the 3-gram algorithm and stored in the form of (w i-1 , w i-2 , w i ). Traverse the original sentence S to be corrected. S is composed of words x1, x2,..., x n . If there exists x k-2 = w k-2 , x k-1 = w k-1 , then w k is considered a candidate word that can be used to replace x k and is added to the candidate word

[0147] After obtaining all the candidate words corresponding to the suspicious words, the original sentence S needs to be modified using the candidate words. The suspicious words in the original sentence are replaced by the candidate words corresponding to the suspicious words in a permutation and combination manner to obtain a candidate sentence set. For the original sentence S, S is composed of words x1, x2,..., x n . Currently, there are suspicious words x i , x j ,..., x k , and the candidate words corresponding to the suspicious words are Successively replace the suspicious words x m (m ∈ (i, j,..., k)) with the candidate words m of x for permutation and combination to generate a candidate sentence set.

[0148] Finally, the optimal sentence needs to be selected from the candidate sentence set as the final output result. The present invention uses the perplexity of a sentence to evaluate the overall smoothness of the sentence. Perplexity is a method for measuring the quality of a language probability model. The smaller the perplexity, the more reasonable the sentence is. The sentence with the lowest perplexity, that is, the smoothest sentence, is output as the correct sentence. The present invention uses the GPT-2 model to calculate the perplexity of the entire candidate sentence. For the GPT-2 model, given a sentence, if its length is n, first shift it one position to the left as the label, remove its last digit as the input, input the input into GPT-2, calculate the cross-entropy loss between the output obtained and the label, and then take the natural logarithm of the loss to obtain the required perplexity. The formula for calculating perplexity is as follows:

[0149] out = GPT-2(input)

[0150] loss = CrossEntropyLoss(out, label)

[0151] PPL = ln loss

[0152] where, given a sentence S, S is composed of words x1, x2,..., x n input = x1, x2,..., x n-1 , label = x2,..., x n , and PPL is the required perplexity. Calculate the perplexity of all candidate sentence sets and the original sentence, and select the sentence with the smallest perplexity as the final error correction result.

[0153] So far, the text error correction based on online query has been completed, and it is possible to input a text with errors and output the correct form of this text.

[0154] For those skilled in the art, the technical solutions described in the foregoing examples can be modified, or some of the technical features can be equivalently replaced. All modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.

Claims

1. A Chinese text error correction method based on online search assistance, characterized in that, The following steps are involved: S1: Divide the original sentence to be corrected into sentences, and the basis for sentence division is the number of words contained in the original sentence; S2: Search the sentence after the sentence division in step S1 through a search engine, crawl the titles and abstracts of the first thirty query results and save them locally; S3: Based on the thirty query results obtained in step S2, perform word segmentation and word frequency statistics, and then construct a word frequency table; S4: Based on the thirty query results obtained in step S2, new words are discovered, the results of the new word discovery are added to the jieba vocabulary, and then the original sentence is segmented according to the changed jieba vocabulary to obtain the result of the original sentence after segmentation; S5: Based on the word frequency table obtained by search engine query in step S3 and the original sentence segmentation result obtained in step S4, error detection is performed, and the original sentence segmentation result is used to query in the word frequency table. If the word frequency value of a word in the word frequency table is less than a threshold, it is considered that the word may be wrong and is regarded as a suspicious word; S6: Based on the original sentence segmentation result obtained in step S4, perform error detection and supplement of unregistered words, and add unregistered words that are not in the jieba word library to suspicious words; S7: Perform probability error detection and supplementation, pass the original sentence through the original GPT-2 model, and obtain the probability value of each word. If the probability value of a word is significantly smaller than the probability values of other words, add the word to the suspicious words; S8: Based on the suspicious words obtained in step S7, context information text_ori of the suspicious words in the original sentence is obtained in sequence, and stored in a pair of a suspicious word and the corresponding context information text_ori; the context information is obtained by obtaining words within a distance x from the suspicious word in the original sentence as the context information text_ori according to the distance; S9: Based on the context information text_ori of the suspicious word in the original sentence obtained in step S8 and the result obtained by the search engine query in step S2, the context information text_search of text_ori in the result obtained by the search engine query is obtained in sequence. At this time, the context information text_search is used as a candidate word, and is stored in a pair of a suspicious word and a corresponding candidate word. The distance method is still used to obtain the words in the result obtained by the search engine query that are 2, 4, and 6 away from text_ori as the context information text_search; S10: Based on the candidate words obtained in step S9, the phonetic edit distance and structural similarity between the candidate words and the corresponding suspicious words are calculated respectively; wherein the structural similarity is calculated using a pre-built twin network; S11: Based on the pinyin edit distance and structural similarity between the candidate word obtained in step S10 and the corresponding suspicious word and the word frequency table obtained in step S3, the topsis algorithm is used to calculate the score based on word frequency, pinyin edit distance and structural similarity, and the top 8 candidate words with the highest scores are selected as candidate words for the suspicious words; S12: Based on the suspicious words obtained in step S7, select the words with a small edit distance in the jieba vocabulary from the suspicious words' pinyin and the words with a high structural similarity, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10; S13: Based on the search engine query results obtained in step S2, construct a 3-gram vocabulary, use the n-gram algorithm, and select the words that appear in the position of the suspicious words, and also add them as candidate words for the suspicious words to the candidate words for the suspicious words obtained in step S10; S14: Based on the candidate words for the suspicious words obtained in step S13 and the suspicious words obtained in step S6, replace the suspicious words in the original sentence with the corresponding candidate words in a permutation and combination manner to obtain a set of candidate sentences. Since the original sentence is correct, the original sentence is also added to the set of candidate sentences; S15: Based on the set of candidate sentences obtained in step S14, use the original GPT-2 model to calculate the perplexity of the entire candidate sentence, and select the sentence with the lowest perplexity as the final result.

2. The Chinese text error correction method based on online search assistance according to claim 1, wherein In S1, the original sentence to be corrected is segmented. The segmentation condition is that first, the original sentence is tokenized using jieba. If the number of words in the tokenization result of the original sentence is greater than or equal to 15, then the original sentence is split according to the full stop, question mark, exclamation mark, and semicolon. If the number of words in the segmented short sentence is still greater than or equal to 15, then it is split according to the comma; in S3, the titles and abstracts of the thirty query results crawled are tokenized using jieba, divided into words, and then the number of times each word appears is counted. The number of appearances is the word frequency of this word. Each word and its corresponding word frequency are saved as the word frequency table.

3. The Chinese text error correction method based on online search assistance according to claim 1, wherein, In S4, new word discovery is performed on the query results. The new word discovery algorithm mainly consists of 3 steps: a) Generate an n-gram table from the corpus text and count the word frequency of each word. b) Use the cohesion degree to screen out candidate new words from the previous n-gram table. c) Then screen out the final new words from the candidate new words through the freedom degree; Given an original sentence S, the result after word segmentation is x1, x2, …, x n ; The cohesion degree is represented by pointwise mutual information, and the formula is: Among them, PMI(x,y) is the pointwise mutual information, p(x,y) refers to the probability that two words appear together, and p(x), p(y) refer to the probabilities of each word appearing; the greater the cohesion degree, the greater the probability that these two words appear together, and the greater the possibility of being a word; The freedom degree is represented by left and right entropy, and the left and right entropy formulas are respectively: Among them, E Left (PreW) represents the left entropy, and E Right (SufW) represents the right entropy. PreW is the set of prefixes of word W, and SufW is the set of suffixes of word W. The greater the degree of freedom, the richer its surrounding words are, and the greater the possibility that it becomes an independent word.

4. The Chinese text error correction method based on online search assistance according to claim 1, wherein S5 uses the original sentence segmentation results to query the word frequency table for error detection. Specifically, it refers to: in the constructed word frequency table, query the original sentence segmentation results x1, x2, ..., x n The word frequency values c1, c2, ..., c in the word frequency table n , select c1,c2,…,c n The maximum value c in max As a benchmark, if the frequency value of other words is c k(k≠max) Less than 5%*c max , then the word is considered to be possibly wrong. If the word is not in the suspicious words, it is added to the suspicious words. The unregistered word error detection and supplement described in S6 specifically refers to: searching the original sentence segmentation results x1, x2, …, x in the jieba word list. n If word x is not in the jieba vocabulary, it is considered that the word may be wrong and added to the suspicious words; the probability error detection supplement described in S7 specifically refers to: sentence S consists of words x1, x2, ..., x n The GPT-2 model inputs the real previous text x1, x2, …, x m-1 , get the next word x m ′ A probability table consisting of possible words and their corresponding probability values; according to the real word x of each original sentence m Get the probabilities p1, p2, ..., p in the probability value table n , calculate the median value p m , if there is a word x m The probability value is less than the median value 10%*p m , then the word is considered to be possibly wrong. If the word is not in the suspicious words, it is added to the suspicious words.

5. The Chinese text error correction method based on online search assistance according to claim 1, wherein Obtain the context information text_ori of the suspicious word in the original statement in S8, specifically referring to: Assume there is an existing original statement S, which is composed of words x1, x2, ……, x n and the suspicious word is x k . According to the distance formula: dis = min(3, work num / / 2) Among them, word num represents the number of words after word segmentation. By calculating the distance dis, all words within the distance dis from the suspicious word x in the original statement are obtained k as the context information text_ori of x, text_ori = x k , x k-dis , …, x k-dis+1 , …, x k-1 , x k+1 , …, x k+dis ; The context information text_search in the search results of the context information text_ori obtained in S9 from the search engine is specifically: Among the thirty search results saved to the local search engine, according to the preset distance, the context information text_search of the context information text_ori is obtained. One result of the search engine is S′, which consists of words x′1, x′2, …, x j ′. There is an original statement S, which consists of words x1, x2, …, x n , and the suspicious word is x k , x k-dis , x k-dis+1 , …, x k-1 , x k+1 , …, x k+dis is the context information text_ori of x k ; For each context information text_ori, the context information text_search needs to be found in S′ according to the distance, and the distances are respectively selected as 2, 4, and 6; For x k-dis in the context information text_ori, if x k-dis also appears in S′, then x′ k-dis-2 , x′ k-dis-1 , x k-dis+1 ′, x k-dis+2 ′ is the context information text_search of x k-dis with a distance of 2 in S′, x′ k-dis-4 , …, x′ k-dis-1 , x′ k-dis+1 , …, x′ k-dis+4 is the context information text_search of x1 with a distance of 4 in S′, x′ k-dis-6 , …, x′ k-dis-1 , x′ k-dis+1 , …, x′ k-dis+6 is the context information text_search of x1 with a distance of 6 in S′; And so on for x k-dis+1 , …, x k-1 , x k+1 , …, x k+dis Perform the same operation to obtain context information text_search; use this context information text_search as the candidate words for the corresponding suspicious words.

6. The Chinese text error correction method based on online search assistance according to claim 1, wherein In S10, calculate the edit distance in pinyin and the structural similarity between the candidate word and the corresponding suspicious word. Specifically, the edit distance is the minimum number of edits to change from one string to another, where each edit can only insert a character, delete a character, or modify a character in the string; and the edit distance in pinyin is the edit distance after converting two Chinese characters into pinyin without phonetic symbols; the edit distance in pinyin formula is as follows: pydis = LS(py1,py2) Among them, pydis represents the edit distance in pinyin, LS represents the edit distance calculation, and py1, py2 respectively represent the pinyin without phonetic symbols of the two words; The structural similarity is obtained by using a pre-trained siamese network to calculate the graph similarity; the siamese network is a conjoined neural network, and the two neural networks therein share parameter weights; the siamese neural network has two inputs, graph1 and graph2. After putting the two inputs into the two neural networks, Network1 and Network2, and obtaining the backbone feature extraction network, a multi-dimensional feature is obtained. Flattening it into one dimension, one-dimensional vectors of the two inputs are obtained. Subtracting these two one-dimensional vectors and then summing the absolute values is equivalent to calculating the distance between the two one-dimensional vectors. Then, fully connecting this distance and taking the sigmoid of the result to make its value between 0 and 1 represents the similarity degree of the two input pictures; Using OpenCV, generate images of different font forms for each Chinese character c k Same for different font forms of the same Chinese character c Generate images of different fonts k for the same Chinese character c are regarded as the same type, and different Chinese characters are regarded as different types; during training, when two inputs point to images of the same type, the label is 1 at this time, and when two inputs point to images of different types, the label is 0 at this time. Then, the cross-entropy operation is performed on the output result of the network and the true label, which is used as the final loss. The structural similarity formula is as follows: similarity = Graphsimi(graph1, graph2) where similarity is the structural similarity, Graphsimi is the siamese network model, and graph1 and graph2 are the word pictures generated using two words respectively, serving as the two inputs of the siamese neural network.

7. The Chinese text error correction method based on online search assistance according to claim 1, wherein Using the topsis algorithm described in S11, calculate the scores based on word frequency, pinyin edit distance, and structural similarity. Specifically, it means that the topsis algorithm is a method of ranking according to the degree of proximity of a finite number of evaluation objects to the ideal target. First, the data is normalized. Normalization means processing each evaluation index to be better as it gets larger. The evaluation indexes are the word frequency of the candidate word, the pinyin edit distance between the candidate word and the suspicious word, and the structural similarity between the candidate word and the suspicious word. The word frequency and structural similarity are better as they get larger, so no processing is needed. However, the pinyin edit distance is better as it gets smaller, so the reciprocal of the pinyin edit distance is taken as the evaluation index. That is, the evaluation indexes are the following three points: ① word frequency, ② pydis is the pinyin edit distance, ③ structural similarity similarity; Then, the data is normalized. This is to eliminate the influence of different data metric dimensions. The normalization formula is as follows: Among them, z ij represents the value of the j-th index of the i-th solution after standardization, and x ij represents the j-th index of the i-th solution in the original data; Then determine the optimal ideal value \(z\) of each index + and the worst ideal value \(z\) - , the attribute values of the optimal ideal value \(z\) + are the best values among the candidate solutions, that is, the maximum values in each index, while the worst ideal value \(z\) - is the minimum value in each index. Then calculate the Euclidean distances between each solution and the optimal ideal value and the worst ideal value. For the \(i\)-th solution \(z\) i , its distance formula from the optimal solution is as follows: Among them, represents the distance between the $i$-th solution $z$ i and the optimal solution, $m$ represents the number of indicators, represents the maximum value of the $j$-th indicator, $z$ ij represents the value of the $j$-th indicator of the $i$-th solution after standardization; For the $i$-th solution $z$ i , its distance formula from the worst solution is as follows: Among them, represents the i-th solution z i the distance from the worst solution, m represents the number of indicators, represents the maximum value of the j-th indicator, z ij represents the value of the j-th indicator of the i-th solution after standardization; The scoring formula for the i-th scheme is as follows: Among them, S i represents the score of the i-th solution; From this, the degree of closeness of each scheme to the optimal scheme is obtained, serving as the criterion for evaluating the quality of the scheme. Finally, the quality values of each scheme are obtained.

8. The Chinese text error correction method based on online search assistance according to claim 1, wherein In S12, words with a small phonetic edit distance from the suspicious word in the jieba vocabulary and words with a high structural similarity are screened out. Specifically, it means: splitting the suspicious word x into n characters c1, c2, …, c at the character granularity n , and then for each character c m , find all characters c1′, c2′, …, c l ′ in the jieba vocabulary with a phonetic edit distance less than or equal to 1. Arrange these characters c1′, c2′, …, c l ′ according to the positions of the characters c m in the suspicious word x. That is, the characters with a phonetic edit distance less than or equal to 1 from c1 are in the first position, the characters with a phonetic edit distance less than or equal to 1 from c2 are in the second position, and so on. If the generated words after combination are registered in the jieba vocabulary, these words are considered normal words and also serve as the corresponding candidate words for the suspicious word.

9. The Chinese text error correction method based on online search assistance according to claim 1, characterized in that The use of the n-gram algorithm described in S13 to select the words that appear at the suspicious word positions specifically means: for the n-gram algorithm, assuming that the probability of the n-th word appearing is only related to the previous n - 1 words, the probability distribution of the sentence is as follows: Among them, P(s) is the probability of the entire sentence, w n is the word that makes up the sentence, represents the historical sequence from the w i -th word to the w i-n+1 -th word, represents the probability of the current word given the historical sequence of words; using the 3-gram algorithm, the probability distribution of the sentence is as follows: P(s) = P(w1|w0, w -1 )P(w2|w1, w0)…P(w i |w i-1 , w i-2 ) Construct a 3-gram table for the thirty pieces of information obtained by querying the search engine using the 3-gram algorithm, and store them in the form of (w i-1 , w i-2 , w i ). Traverse the original sentence S to be corrected. S is composed of words x1, x2, …, x n . If there exists x i-2 = w i-2 , x i-1 = w i-1 , then consider w i as a candidate word to replace x i . In S14, replace the suspicious words in the original sentence with the candidate words corresponding to the suspicious words in a permutation and combination manner to obtain a set of candidate sentences. Specifically, for the original sentence S composed of words x1, x2, …, x n , given the existing suspicious words x i , x j , …, x k , successively replace x i with the candidate words of x i , replace x j with the candidate words of x j , ……, replace x k with the candidate words of x k to generate a set of candidate sentences.

10. The Chinese text error correction method based on online search assistance according to claim 1, wherein The use of the original GPT-2 model to calculate the perplexity of the entire candidate sentence described in S15 specifically means: for a given sentence with a length of n, first shift it one position to the left as the label, remove its last character as the input, input the input into GPT-2, and calculate the cross-entropy loss Cross Entropy Loss between the output obtained from GPT-2 and the label. Then, taking the natural logarithm of the result is the required perplexity; perplexity is a method for measuring the quality of a language probability model. The smaller the perplexity, the more reasonable the sentence. The formula for calculating perplexity is as follows: out = GPT-2(input) loss = CrossEntropyLoss(out, label) PPL = ln loss Among them, given a sentence S, which is composed of words x1, x2, …, x n forming, input = x1, x2, …, x n-1 , label = x2, …, x n , and PPL is the perplexity to be obtained; calculate the perplexity of all candidate sentence sets and the original sentence, and select the sentence with the smallest perplexity as the final error correction result.

Citation Information

Patent Citations

  • Automatic error correction method for Chinese search term of search engine

    CN106095778A

  • Artificial intelligence-based text error correction method and device, equipment, and storage medium

    CN113962215A