Improved Chinese text error correction method based on MACBERT and GEORR

By adopting the combination method of MACBERT and GECTOR models in Chinese text error correction technology, problems such as scarcity of training data and high expression flexibility in the prior art are solved, and more efficient spelling and grammatical error recognition is achieved, and the accuracy and recall of error correction are improved.

CN120046601APending Publication Date: 2025-05-27SHANGHAI HUIZHOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112080.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing Chinese text error correction technology faces problems such as scarce training data, flexible Chinese expression, confusion of similar characters, excessive error correction, insufficient vocabulary support and insufficient detection capabilities in professional fields.

Method used

The improved Chinese text error correction method based on MACBERT and GECTOR is adopted to identify spelling errors through the macbert model decoding mechanism, simulate macbert's mask mechanism to predict [mask] characters, combine the sequence labeling mechanism of geoctor to identify syntax errors, and reduce false positives through custom conflict processing rules and post-processing methods.

Benefits of technology

It effectively improves the recall and accuracy of Chinese text error correction, reduces false positives of new words, rare words and named entities, and enhances the detection capabilities of professional and general fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046601A_ABST
    Figure CN120046601A_ABST
Patent Text Reader

Abstract

The invention discloses an improved Chinese text error correction method based on MACBERT and GEOROR. The improved Chinese text error correction method comprises the steps that spelling error recognition is conducted on an input text through an improved macbert model, and grammar error recognition is conducted on the input text through a selector model; conflict processing; carrying out post-processing on a model identification result; other types of errors are detected, including custom common language detection, sensitive word detection, time format detection, important character name detection, sorting detection, entity matching affiliation relation detection and paragraph or word repetition error type detection; the results are fused; and based on intra-domain rule post-processing, false alarms are reduced, and finally an error detection report is output. According to the error correction model architecture, a macbert model decoding mechanism is used for recognizing spelling errors, some new words are filtered through a post-processing means, an error correction result is finally output through a special use method, and the accuracy rate is high.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An improved Chinese text error correction method based on MACBERT and GECTOR, characterized in that: Includes steps: S1: The input text is identified by spelling errors through the improved MacBert model and grammatical errors through the Gector model; S2: Conflict handling, merging model recognition results through custom conflict handling rules; S3: Post-process the model recognition results to reduce the model's false positives for new words, rare words, and other sparse named entities; S4: Other types of error detection, including detection of colloquialisms, sensitive words, time format, names of important people, sorting, entity collocation and affiliation, paragraph or word duplication errors; S5: Fusion the results of steps S3 and S4; S6: Reduce false positives based on domain rule post-processing and finally output an error detection report.

2. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The step S1 specifically includes the following steps: S101: Build a network structure, the network structure is MacBert+Softmax, where the Macabert network structure is consistent with the Bert network structure; S102: Prepare training data, collect network data such as WeChat, Weibo, websites, and some People's Daily data as training corpus, then clean the data and use preset strategies to generate training data; S103: Model training. The purpose of model training is to make the training loss converge as quickly as possible. The loss functions of the above MacBert and Softmax models are both cross entropy loss functions. S104: Train N-gram language model, collect massive corpus, cover multiple industries and fields, and train 3-gram, 4-gram, and 5-gram models based on single characters.

3. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 2, characterized in that: In the process of model training, step S103 first uses the grid parameterization verification method to determine an optimal parameter group in a small batch size, and uses the optimal parameter group to train the first epoch. After the loss value converges to a smaller value, a smaller batch_size and a smaller lr are used in the second epoch to further reduce the loss. This method can quickly train the model to fit the training data.

4. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 2, characterized in that: If the model cannot converge well in step S103, it is necessary to check whether the training data has a stable distribution.

5. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The conflict handling in step S2 includes: inputting the text to be corrected into two models, and merging the correction results according to the post-processing rules and the conflict handling rules; Specifically: First, the input text is cut into single complete sentences and input into the AI ​​model. For the Csc model, if the prob of the output optimal candidate character is lower than 0.7, the top 10 are taken. The N-gram language model trained by the above steps evaluates the scores of the 10 sentences respectively, and the candidate character with the maximum score is taken; the average of the three N-gram models can further improve the overall recall and precision.

6. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: In step S3, a custom new word discovery algorithm based on solidification, degree of freedom and word frequency is used to discover new words. Calculation of 2-gram solidification degree: akin, The specific logic is to use the ratio of the joint probability of two words and the product of their respective marginal probabilities. The higher the ratio, the more "solidified" it is. If the two words a and b just happen to be together, with enough data, it should be statistically possible that p(a)p(b)≈p(ab), there is no correlation between them, and the solidification degree is ≈11; If the two characters a and b are extremely related, they must appear at the same time. It should be statistically p(a)≈p(ab), and the solidification degree ≈1 / p(b) is generally much greater than 1; The logic of the degree of freedom is: by counting the distribution of other words on the left and right sides of the candidate word (calculating information entropy), we can examine whether its context is rich enough (the entropy is large enough and the collocation is uncertain enough). A sufficiently independent word should be used in different contexts, that is, the entropy of the distribution of the left and right sides of the candidate word is calculated respectively, and the smaller value is selected as the final degree of freedom.

7. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The step S4 detects colloquialisms and idioms by using confusion sets, i.e., by constructing error templates with more wrong characters and fewer characters, and building a try tree to improve the matching speed. The post-processing method combines Chinese word segmentation, positive trigger word mechanism, reverse trigger word mechanism, core word mechanism, and whether to use a language model mechanism for error detection.