Improved Chinese text error correction method based on MACBERT and GEORR
By adopting the combination method of MACBERT and GECTOR models in Chinese text error correction technology, problems such as scarcity of training data and high expression flexibility in the prior art are solved, and more efficient spelling and grammatical error recognition is achieved, and the accuracy and recall of error correction are improved.
Patent Information
- Application Number
- CN202510112080.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing Chinese text error correction technology faces problems such as scarce training data, flexible Chinese expression, confusion of similar characters, excessive error correction, insufficient vocabulary support and insufficient detection capabilities in professional fields.
The improved Chinese text error correction method based on MACBERT and GECTOR is adopted to identify spelling errors through the macbert model decoding mechanism, simulate macbert's mask mechanism to predict [mask] characters, combine the sequence labeling mechanism of geoctor to identify syntax errors, and reduce false positives through custom conflict processing rules and post-processing methods.
It effectively improves the recall and accuracy of Chinese text error correction, reduces false positives of new words, rare words and named entities, and enhances the detection capabilities of professional and general fields.
Smart Images

Figure CN120046601A_ABST
Abstract
Claims
1. An improved Chinese text error correction method based on MACBERT and GECTOR, characterized in that: Includes steps: S1: The input text is identified by spelling errors through the improved MacBert model and grammatical errors through the Gector model; S2: Conflict handling, merging model recognition results through custom conflict handling rules; S3: Post-process the model recognition results to reduce the model's false positives for new words, rare words, and other sparse named entities; S4: Other types of error detection, including detection of colloquialisms, sensitive words, time format, names of important people, sorting, entity collocation and affiliation, paragraph or word duplication errors; S5: Fusion the results of steps S3 and S4; S6: Reduce false positives based on domain rule post-processing and finally output an error detection report.
2. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The step S1 specifically includes the following steps: S101: Build a network structure, the network structure is MacBert+Softmax, where the Macabert network structure is consistent with the Bert network structure; S102: Prepare training data, collect network data such as WeChat, Weibo, websites, and some People's Daily data as training corpus, then clean the data and use preset strategies to generate training data; S103: Model training. The purpose of model training is to make the training loss converge as quickly as possible. The loss functions of the above MacBert and Softmax models are both cross entropy loss functions. S104: Train N-gram language model, collect massive corpus, cover multiple industries and fields, and train 3-gram, 4-gram, and 5-gram models based on single characters.
3. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 2, characterized in that: In the process of model training, step S103 first uses the grid parameterization verification method to determine an optimal parameter group in a small batch size, and uses the optimal parameter group to train the first epoch. After the loss value converges to a smaller value, a smaller batch_size and a smaller lr are used in the second epoch to further reduce the loss. This method can quickly train the model to fit the training data.
4. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 2, characterized in that: If the model cannot converge well in step S103, it is necessary to check whether the training data has a stable distribution.
5. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The conflict handling in step S2 includes: inputting the text to be corrected into two models, and merging the correction results according to the post-processing rules and the conflict handling rules; Specifically: First, the input text is cut into single complete sentences and input into the AI model. For the Csc model, if the prob of the output optimal candidate character is lower than 0.7, the top 10 are taken. The N-gram language model trained by the above steps evaluates the scores of the 10 sentences respectively, and the candidate character with the maximum score is taken; the average of the three N-gram models can further improve the overall recall and precision.
6. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: In step S3, a custom new word discovery algorithm based on solidification, degree of freedom and word frequency is used to discover new words. Calculation of 2-gram solidification degree: akin, The specific logic is to use the ratio of the joint probability of two words and the product of their respective marginal probabilities. The higher the ratio, the more "solidified" it is. If the two words a and b just happen to be together, with enough data, it should be statistically possible that p(a)p(b)≈p(ab), there is no correlation between them, and the solidification degree is ≈11; If the two characters a and b are extremely related, they must appear at the same time. It should be statistically p(a)≈p(ab), and the solidification degree ≈1 / p(b) is generally much greater than 1; The logic of the degree of freedom is: by counting the distribution of other words on the left and right sides of the candidate word (calculating information entropy), we can examine whether its context is rich enough (the entropy is large enough and the collocation is uncertain enough). A sufficiently independent word should be used in different contexts, that is, the entropy of the distribution of the left and right sides of the candidate word is calculated respectively, and the smaller value is selected as the final degree of freedom.
7. The improved Chinese text error correction method based on MACBERT and GECTOR according to claim 1, characterized in that: The step S4 detects colloquialisms and idioms by using confusion sets, i.e., by constructing error templates with more wrong characters and fewer characters, and building a try tree to improve the matching speed. The post-processing method combines Chinese word segmentation, positive trigger word mechanism, reverse trigger word mechanism, core word mechanism, and whether to use a language model mechanism for error detection.