Chinese spelling error correction method and device based on detection-correction balance framework
By training and setting up Chinese spelling detectors in the detection-correction framework, and performing feature fusion and model training, the problems of detector effect limitations and inappropriate application of error detection results are solved, and more accurate and efficient Chinese spelling error correction is achieved.
Patent Information
- Application Number
- CN202411749299.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-06
AI Technical Summary
The detector effect of the Chinese spelling error correction method under the existing detection-error correction framework has certain limitations, and it is difficult to reasonably apply the error detection results to error correction.
The pre-built language detector is trained through preset Chinese spelling error correction data text, and a Chinese spelling detector is generated, and the accuracy threshold and recall threshold are set to generate accuracy error detection results and recall error detection results. Mask text data is obtained based on the recall error detection results, splicing and embedding features are extracted, feature fusion is performed, and the Chinese spelling error corrector is trained based on the fusion feature and the target loss function.
It effectively improves error correction capabilities, solves the problems of detector effect limitations and improper application of error detection results, and achieves more accurate and efficient Chinese spelling error correction.
Smart Images

Figure CN119940349A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text error correction, and in particular to a Chinese spelling error correction method and device based on a detection-correction balanced framework. Background Art
[0002] Chinese Spelling Correction (CSC) aims to detect and correct incorrect characters in a given Chinese sentence. It is a basic natural language processing task and plays an indispensable role in many downstream natural language processing tasks.
[0003] In recent years, pre-trained language models (PLMs) have developed rapidly. The existing technology can regard Chinese spelling correction as a sequence labeling task and fine-tune the BERT-based model on sentence pairs. The source sentence is directly input into the model and the target sentence is used as the output. In addition, the existing technology can also use error detectors as a prerequisite for error correction, converting Chinese spelling correction into a two-stage process. The results produced by the detector indicate which characters are incorrect.
[0004] However, detectors do not always lead to significant performance improvements due to two potential reasons:
[0005] 1. The detector may not be powerful enough. Error detectors often do not work as expected, and a poorly performing detector may even have a negative impact on the subsequent error correction process; therefore, such error detection results have limited improvement in error correction capabilities and may even have negative effects.
[0006] 2. Existing Chinese spelling correction methods under the detection-correction framework all use character-level detectors, that is, they directly point out the location where the error occurs in the original sentence in some way, so as to attach the error information to the corresponding position to assist in error correction; however, in the correction of Chinese sentences, there is a strong correlation between the incorrect characters and the context. Simply applying the indication to the corresponding position cannot instruct the model to pay attention to this correlation.
[0007] Existing text error correction technology chooses to use additional error detectors to determine the error location to assist in error correction; however, due to the inherent performance limitations of the error detector, good results cannot be obtained for both the accuracy and recall of error detection; in addition, the existing technology's simple application strategy for error location information is not perfect and needs to be solved urgently. Summary of the invention
[0008] The present application provides a Chinese spelling correction method and device based on a detection-correction balanced framework to solve the problems that the detector effect of the Chinese spelling correction method under the existing detection-correction framework has certain limitations and it is difficult to reasonably apply the error detection results to error correction.
[0009] The first aspect of the present application provides a Chinese spelling correction method based on a detection-correction balance framework, comprising the following steps: training a pre-constructed language detector with preset Chinese spelling correction data text to generate a Chinese spelling detector; setting a precision threshold and a recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and the recall rate threshold, generating a precision error detection result and a recall error detection result corresponding to a preset initial text data to be corrected, and obtaining mask text data corresponding to the initial text data to be corrected according to the recall error detection result; splicing the mask text data and the initial text data to be corrected to obtain spliced data, and extracting embedded features of the spliced data through a pre-constructed Chinese spelling corrector, and performing feature fusion on the embedded features and the precision error detection results to obtain fused features, and training the Chinese spelling corrector based on the fused features and a preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected.
[0010] Optionally, in one embodiment of the present application, the pre-constructed language detector is trained by using a preset Chinese spelling correction data text, comprising: inputting the Chinese spelling correction data text into the language detector to determine the error probability of each data position in the Chinese spelling correction data text; judging whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, marking the data position corresponding to the error probability as an error prediction position, and outputting the error prediction position to train the language detector.
[0011] Optionally, in one embodiment of the present application, obtaining the masked text data corresponding to the initial text data to be corrected based on the recall error detection result includes: obtaining the error prediction position corresponding to the initial text data to be corrected through the Chinese spelling detector, and calculating the fuzzy indication strength corresponding to each data position in the initial text data to be corrected based on the error prediction position and a preset probability density function; based on the fuzzy indication strength and a preset strength threshold, determining at least one target data position in the initial text data to be corrected that meets a preset importance requirement, and performing a masking operation on each target data position in the at least one target data position and the data within the target range of each target data position to obtain the masked text data corresponding to the initial text data to be corrected.
[0012] Optionally, in one embodiment of the present application, the embedded features of the spliced data are extracted by a pre-built Chinese spelling corrector, and the embedded features and the precision error detection results are subjected to feature fusion to obtain fused features, including: extracting the embedded features of the spliced data by the Chinese spelling corrector, and obtaining hidden layer dimensional features in the embedded features; copying the hidden layer dimensional features to obtain hidden layer dimensional copy features, and adding the hidden layer dimensional copy features to the precision error detection results to obtain the precision error detection results to be fused, and performing feature addition operations on the precision error detection results to be fused and the embedded features to generate the fused features.
[0013] The second aspect of the present application provides a Chinese spelling correction device based on a detection-correction balance framework, including: a training module, used to train a pre-constructed language detector through preset Chinese spelling correction data text to generate a Chinese spelling detector; a generation module, used to set a precision threshold and a recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and the recall rate threshold, generate a precision error detection result and a recall error detection result corresponding to a preset initial text data to be corrected, and obtain mask text data corresponding to the initial text data to be corrected according to the recall error detection result; a Chinese spelling correction module, used to splice the mask text data and the initial text data to be corrected to obtain spliced data, and extract embedded features of the spliced data through a pre-constructed Chinese spelling corrector, and perform feature fusion on the embedded features and the precision error detection results to obtain fused features, and train the Chinese spelling corrector based on the fused features and a preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected.
[0014] Optionally, in one embodiment of the present application, the training module includes: a determination unit, used to input the Chinese spelling correction data text into the language detector to determine the error probability of each data position in the Chinese spelling correction data text; a judgment unit, used to judge whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, the data position corresponding to the error probability is marked as an error prediction position, and the error prediction position is output to train the language detector.
[0015] Optionally, in one embodiment of the present application, the generation module includes: a calculation unit, used to obtain the error prediction position corresponding to the initial text data to be corrected through the Chinese spelling detector, and calculate the fuzzy indication strength corresponding to each data position in the initial text data to be corrected according to the error prediction position and a preset probability density function; a mask unit, used to determine at least one target data position in the initial text data to be corrected that meets a preset importance requirement based on the fuzzy indication strength and a preset strength threshold, and perform a mask operation on each target data position in the at least one target data position and the data within the target range of each target data position to obtain masked text data corresponding to the initial text data to be corrected.
[0016] Optionally, in one embodiment of the present application, the Chinese spelling correction module includes: an acquisition unit, used to extract the embedded features of the spliced data through the Chinese spelling corrector, and obtain the hidden layer dimensional features in the embedded features; a copying unit, used to copy the hidden layer dimensional features to obtain hidden layer dimensional copy features, and add the hidden layer dimensional copy features to the precision error detection results to obtain the precision error detection results to be fused, and perform feature addition operations on the precision error detection results to be fused and the embedded features to generate the fused features.
[0017] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the Chinese spelling error correction method based on the detection-correction balance framework as described in the above embodiment.
[0018] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned Chinese spelling error correction method based on the detection-correction balance framework.
[0019] The fifth aspect of the present application provides a computer program product, including a computer program, which is executed to implement the above-mentioned Chinese spelling error correction method based on the detection-correction balance framework.
[0020] Therefore, the embodiments of the present application have the following beneficial effects:
[0021] The embodiment of the present application can train a pre-built language detector through a preset Chinese spelling correction data text to generate a Chinese spelling detector; set the precision threshold and recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and recall rate threshold, generate the preset precision error detection result and recall error detection result corresponding to the initial text data to be corrected, and obtain the mask text data corresponding to the initial text data to be corrected according to the recall error detection result; splice the mask text data and the initial text data to be corrected to obtain the spliced data, and extract the embedded features of the spliced data through the pre-built Chinese spelling corrector, and perform feature fusion on the embedded features and the precision error detection results to obtain the fusion features, and train the Chinese spelling corrector based on the fusion features and the preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected. The present application constrains the output results of the detector, so as to fully explore the performance of the detector without adding additional training and detection processes, and can effectively improve the error correction capability. This solves the problem that the detector effect of the Chinese spelling error correction method under the existing detection-correction framework has certain limitations and it is difficult to reasonably apply the error detection results to error correction.
[0022] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0024] Figure 1 A flowchart of a Chinese spelling error correction method based on a detection-correction balance framework provided according to an embodiment of the present application;
[0025] Figure 2 A schematic diagram of the execution logic of a Chinese spelling error correction method based on a detection-correction balance framework provided for one embodiment of the present application;
[0026] Figure 3 This is an example diagram of a Chinese spelling error correction device based on a detection-correction balance framework according to an embodiment of the present application;
[0027] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0028] Among them, 10-a Chinese spelling error correction device based on a detection-correction balance framework; 100-a training module, 200-a generation module, 300-a Chinese spelling error correction module; 401-a memory, 402-a processor, 403-a communication interface. DETAILED DESCRIPTION
[0029] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0030] The following describes the Chinese spelling error correction method and device based on the detection-correction balance framework of the embodiment of the present application with reference to the accompanying drawings. In view of the problems mentioned in the above background technology, the present application provides a Chinese spelling error correction method based on the detection-correction balance framework, in which a pre-constructed language detector is trained by a preset Chinese spelling error correction data text to generate a Chinese spelling detector; the precision threshold and recall rate threshold corresponding to the Chinese spelling detector are set, and based on the precision threshold and the recall rate threshold, the precision error detection result and the recall error detection result corresponding to the preset initial text data to be corrected are generated, and the mask text data corresponding to the initial text data to be corrected is obtained according to the recall error detection result; the mask text data and the initial text data to be corrected are spliced to obtain spliced data, and the embedded features of the spliced data are extracted by a pre-constructed Chinese spelling corrector, and the embedded features and the precision error detection results are feature fused to obtain fused features, and based on the fused features and the preset target loss function, the Chinese spelling corrector is trained to use the trained Chinese spelling corrector to perform Chinese spelling error correction operations on the initial text data to be corrected. This application constrains the output results of the detector, thereby fully exploring the performance of the detector without adding additional training and detection processes, and can effectively improve the error correction ability. This solves the problem that the detector effect of the Chinese spelling error correction method under the existing detection-error correction framework has certain limitations, and it is difficult to reasonably apply the error detection results to error correction.
[0031] Specifically, Figure 1 A flowchart of a Chinese spelling error correction method based on a detection-correction balance framework provided in an embodiment of the present application.
[0032] like Figure 1 As shown, the Chinese spelling error correction method based on the detection-correction balance framework includes the following steps:
[0033] In step S101, a pre-built language detector is trained with preset Chinese spelling error correction data text to generate a Chinese spelling detector.
[0034] In the embodiment of the present application, the pre-trained language model can be used as a language detector and a language corrector respectively, and the language detector can be trained by using Chinese spelling correction data text to adapt to the task of Chinese text spelling error detection.
[0035] Optionally, in one embodiment of the present application, a pre-constructed language detector is trained using a preset Chinese spelling correction data text, including: inputting the Chinese spelling correction data text into the language detector to determine the error probability of each data position in the Chinese spelling correction data text; judging whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, marking the data position corresponding to the error probability as an error prediction position, and outputting the error prediction position to train the language detector.
[0036] It should be noted that the embodiment of the present application may first define the input Chinese spelling correction data text as shown in the following formula:
[0037] X={x1,x2,…,x m}
[0038] The Chinese spelling correction data text is input into the language detector, so as to use the language detector to predict the possible error positions from the Chinese spelling correction data text X, and when the error probability of each position is higher than the error threshold λ, it is marked as an error prediction position in the output, as shown in the following formula:
[0039]
[0040] Therefore, the embodiment of the present application provides reliable technical support for subsequent Chinese spelling error correction by training the language detector.
[0041] In step S102, the precision threshold and recall rate threshold corresponding to the Chinese spelling detector are set, and based on the precision threshold and the recall rate threshold, the precision error detection result and the recall error detection result corresponding to the preset initial text data to be corrected are generated, and the mask text data corresponding to the initial text data to be corrected is obtained according to the recall error detection result.
[0042] Furthermore, the embodiments of the present application also need to set a precision threshold and a recall threshold, such as Figure 2 As shown, the output of the language detector is constrained to obtain high-precision error detection results and high-recall error detection results corresponding to the initial text data to be corrected; thereafter, the embodiment of the present application can represent the high-recall error detection results on the original sentence (i.e., the initial text data to be corrected) through a selective masking strategy (Selective MaskingStrategy, SM), thereby obtaining a sentence containing a mask (i.e., masked text data).
[0043] Optionally, in one embodiment of the present application, masked text data corresponding to the initial text data to be corrected is obtained based on the recall error detection result, including: obtaining the error prediction position corresponding to the initial text data to be corrected through a Chinese spelling detector, and calculating the fuzzy indication strength corresponding to each data position in the initial text data to be corrected based on the error prediction position and a preset probability density function; based on the fuzzy indication strength and a preset strength threshold, determining at least one target data position in the initial text data to be corrected that meets a preset importance requirement, and performing a masking operation on each target data position in the at least one target data position and the data within a target range of each target data position to obtain the masked text data corresponding to the initial text data to be corrected.
[0044] Specifically, the embodiment of the present application sets a precision threshold p and a recall threshold q in the trained language detector, wherein p is used to limit the detector to output a detection result with sufficiently high precision, and q is used to limit the detector to output a detection result with sufficiently high recall, that is, when the predicted probability is higher than p, it is retained in the high-precision detection result, and when the predicted probability is higher than q, it is retained in the high-recall detection result;
[0045] For each input sentence (i.e., the initial text data to be corrected), two sets of error detection results are generated, namely, high-precision detection results and high-recall detection results, so as to take into account both precision and recall and reduce the loss of erroneous information caused by the trade-off between precision and recall.
[0046] Secondly, the embodiment of the present application can design a fuzzy indication method to map a vector containing only 0 and 1 to a discrete Gaussian distribution according to its position information, so that an indication related to strength and distance is obtained near the position with the feature 1. Specifically, for each position x i , and its corresponding fuzzy indication strength is:
[0047] ∈=Gauss(μ+(ig)s)
[0048] Wherein, g is the wrong predicted position, Gauss indicates that the embodiment of the present application uses Gaussian distribution as the probability density function, and its mean is μ.
[0049] Afterwards, for each fuzzy indication strength, the embodiment of the present application sets an intensity threshold θ to eliminate information from unimportant positions, as shown in the following formula:
[0050]
[0051] Finally, the embodiments of the present application can set the indication position corresponding to the detection result of the language detector to [MASK] by selecting a mask strategy. In addition, the characters within a certain distance before and after the indication position (i.e., within the target range) are also masked using [MASK] to obtain masked text data corresponding to the initial text data to be corrected.
[0052] Therefore, the embodiments of the present application make special arrangements to address the performance limitations of the detector under the existing detection-correction framework, and constrain the output results of the detector to fully explore the performance of the detector without adding additional training and detection processes.
[0053] In step S103, the masked text data and the initial text data to be corrected are spliced to obtain spliced data, and the embedded features of the spliced data are extracted by a pre-built Chinese spelling corrector, and the embedded features and the precision error detection results are feature fused to obtain fused features, and based on the fused features and a preset target loss function, the Chinese spelling corrector is trained to perform Chinese spelling correction operations on the initial text data to be corrected using the trained Chinese spelling corrector.
[0054] Furthermore, the embodiments of the present application can concatenate the masked sentences to the original sentences to obtain concatenated data, and input the concatenated data into the feature embedding layer of the error correction network (i.e., the Chinese spelling corrector); secondly, the high-precision error detection results are processed using a fuzzy indication method, and then fused with the embedded features of the original sentences to obtain fused features, which are input into the subsequent network, and the error correction network is trained by fusing the features and using a predefined loss function.
[0055] Therefore, the embodiments of the present application enable the language detector to produce high-precision detection results and high-recall detection results, and take into account the contextual relevance of the error occurrence and the existence of errors, thereby using innovative feature fusion strategies and selective masking strategies to integrate the error detection results into the Chinese spelling correction task.
[0056] Optionally, in one embodiment of the present application, embedded features of the spliced data are extracted by a pre-built Chinese spelling corrector, and feature fusion is performed on the embedded features and the precision error detection results to obtain fused features, including: extracting embedded features of the spliced data by the Chinese spelling corrector, and obtaining hidden layer dimensional features in the embedded features; copying the hidden layer dimensional features to obtain hidden layer dimensional copy features, and adding the hidden layer dimensional copy features to the precision error detection results to obtain the precision error detection results to be fused, and performing feature addition operations on the precision error detection results to be fused and the embedded features to generate fused features.
[0057] Furthermore, the embodiment of the present application can concatenate the masked text data to the original sentence to obtain concatenated data, and use it as the input of the error correction network, as shown in the following formula:
[0058] X SM =Concat(X,X m )
[0059] Afterwards, the embodiment of the present application may design an Error Position Information Fusion Strategy (EP) to fuse the detection result of the language detector with the embedded features of the original sentence after the fuzzy indication method (FI). However, since the embedded features of the original sentence have one more hidden layer dimension, the embodiment of the present application copies the detection result in this dimension to obtain the hidden layer dimension copy feature, and adds the hidden layer dimension copy feature to the precision error detection result to obtain the precision error detection result to be fused, and adds the precision error detection result to be fused and the embedded feature to obtain the fusion feature; further, the embodiment of the present application can realize error correction by training the error correction network according to the fusion feature and the preset loss function, and the loss function can use the cross entropy loss, as shown in the following formula:
[0060]
[0061] Therefore, the embodiments of the present application indicate the error correction process from two directions through the error position information fusion strategy and the selection mask strategy, and can consider the context relevance of the error position in the Chinese spelling correction task through the fuzzy indication method, thereby effectively improving the error correction capability.
[0062] In summary, the embodiment of the present application improves the effect of the Chinese spelling correction method under the detection-correction framework by constraining the use of detection results and using two different detection result application strategies. Specifically, the present application constrains the language detector by using a specific threshold so that the language detector can generate two different sets of outputs to adapt to the subsequent error information indication scheme, that is, to obtain detection results with high accuracy and high recall rate; for the obtained high-precision detection results, the present application uses an error position information fusion strategy, and considers the contextual relevance of the error generation, uses fuzzy indication technology to process it, and performs feature fusion with the original sentence; for the obtained high-recall detection results, the present application uses a selective masking strategy, that is, masks the corresponding position in the sentence and its context, and then connects the sentence after the original sentence as the input of the error correction module, and the final error correction result is obtained from the corresponding position of the output.
[0063] The Chinese spelling error correction method based on the detection-correction balance framework of the present application is described below through a specific embodiment.
[0064] As a feasible method, the specific embodiment of the present application can adopt the Pytorch deep learning framework, whose version is 1.11.0 and CUDA version is 11.6; the experimental hardware environment is NVIDIA 3090 graphics card, and the processor is Intel(R) Xeon(R) CPU E5-2680 v4@2.40GHz.
[0065] Specifically, the following can explain the execution process and execution effect of the Chinese spelling correction method based on the detection-correction balance framework of this application from four parts: experimental setting, training details, experimental results and experimental analysis according to the deep learning framework and hardware environment.
[0066] 1. Experimental setup
[0067] The specific embodiment of the present application experiments on four datasets: ECSpell's legal dataset LAW, medical dataset MED, official document writing dataset ODW, and SIGHAN15 dataset.
[0068] The specific embodiment of the present application evaluates the performance of the model of the present application in the Chinese spelling error correction task by using sentence-level error detection indicators, that is, only when the entire sentence is completely error corrected can the sentence be considered to be successfully corrected. The specific embodiment of the present application quantitatively analyzes the error correction results by calculating precision, recall rate and F1 value.
[0069] 2. Training details
[0070] The specific embodiment of the present application trains two neural networks, an error detection network and an error correction network; wherein the error detection network is based on the ELECTRA-base model architecture and the pre-training parameters of the chinese-electra-180g-base-discriminator to perform downstream Chinese spelling error detection task training, and the error correction network is based on the BERT-base model architecture and the pre-training parameters of the bert-base-chinese to perform downstream Chinese spelling error correction task training.
[0071] Specifically, the error detection network and the error correction network both have 12 layers of 12-head Transformer layers, and the hidden layer size is 768. In the specific embodiment of the present application, the threshold can be adjusted on the validation set so that the accuracy of the high-precision result of the language detector is close to 0.95, and the recall rate of the high-recall result is close to 0.95, thereby obtaining two thresholds p and q for constraining the output of the detector; during the training of the language detector, the batch size is 128, the training steps are 10,000, and the learning rate is set to 5e-5; during the training of the error correction network, the batch size is 32, the training steps are 5,000, and the learning rate is set to 3e-5.
[0072] 3. Experimental results
[0073] The results of the LAW dataset are shown in Table 1:
[0074] Table 1
[0075] Methods Precision Recall F1 BERT 73.2 79.2 76.1 REALISE 63.1 61.6 62.3 MDCSpell 77.5 83.9 80.6 ECSpell 78.3 74.9 76.6 RSpell 85.3 81.6 83.4 ReLM 89.9 94.5 92.2 COIN 93.5 96.1 94.8
[0076] Comparing the results of COIN of the present application with other methods on the LAW dataset, it can be seen that the present application has achieved the best results in terms of precision, recall rate and F1 value. Compared with the current best method, the F1 value of the present application has been significantly improved.
[0077] The results of the MED dataset are shown in Table 2:
[0078] Table 2
[0079] Methods Precision Recall F1 BERT 57.9 58.1 58.0 REALISE 55.0 46.0 50.1 MDCSpell 69.9 69.3 69.6 ECSpell 75.9 71.2 73.5 RSpell 86.1 77.0 81.3 ReLM 85.5 85.3 85.4 COIN 92.2 94.0 93.1
[0080] Comparing the results of COIN of the present application with other methods on the MED dataset, it can be seen that the present application has achieved the best results in terms of precision, recall and F1 value. Compared with the current best method, the present application has improved the F1 value by 7.7.
[0081] The results of the ODW dataset are shown in Table 3:
[0082] Table 3
[0083] Methods Precision Recall F1 BERT 59.7 58.8 59.2 REALISE 55.0 50.6 52.7 MDCSpell 65.7 68.2 66.9 ECSpell 82.3 74.5 78.2 RSpell 89.0 79.9 84.2 ReLM 85.7 87.8 86.7 COIN 91.8 91.5 91.7
[0084] Comparing the results of COIN of the present application with other methods on the ODW dataset, it can be seen that the present application has achieved the best results in terms of precision, recall and F1 value. Compared with the current best method, the present application has improved the F1 value by 5.0.
[0085] For the data sets in the above three special fields, the specific embodiments of the present application have achieved very good results. These data sets are small in scale and provide less training data. The error patterns within them are difficult to be captured by the networks trained on other large data sets. It can be seen that the specific embodiments of the present application have strong generalization capabilities on small data sets.
[0086] The results of the SIGHAN15 dataset are shown in Table 4:
[0087] Table 4
[0088]
[0089] By comparing the results of COIN, a specific embodiment of the present application, with other methods on the SIGHAN15 dataset, it can be seen that the specific embodiment of the present application has achieved better results in terms of precision, recall rate and F1 value.
[0090] It is worth noting that the specific embodiment of the present application is a novel framework, and its internal backbone network can be replaced by other methods; the basic method of the specific embodiment of the present application is based on the BERT model, and in the Chinese spelling error correction task, the correlation between the pronunciation and glyphs of Chinese characters is a very important cause of errors; most existing methods have been improved compared to the baseline after considering the pronunciation and glyph information. Therefore, in order to prove the scalability of the specific embodiment of the present application, the specific embodiment of the present application replaces the backbone network with a method SCOPE that incorporates Chinese pronunciation and glyphs, which is different from the basic method COIN (BERT) of the specific embodiment of the present application, and the resulting network structure is recorded as COIN (SCOPE) for experiment.
[0091] It can be seen that even when using BERT as the backbone, the specific embodiment of the present application still achieves competitive results; compared to BERT, the specific embodiment of the present application improves the F1 index by 3.6; when using SCOPE as the backbone, compared to the current best method, the specific embodiment of the present application improves the F1 value by 1.1; considering that the calculation of F1 is related to precision and recall, and precision and recall are used alone to measure the quality of error correction results, it is one-sided, so the F1 value is often used to evaluate the effect of a method. The specific embodiment of the present application is 0.2 lower than MDCSpell in precision, but is much higher than MDCSpell in recall, so the F1 value is also higher than MDCSpell.
[0092] The results of the ablation experiment on the model usage strategy on the LAW dataset are shown in Table 5:
[0093] Table 5
[0094] Methods Precision Recall F1 COIN 93.5 96.1 94.8 w / o FI 93.0 93.7 93.4 w / o EP 91.6 94.9 93.2 w / o SM 92.0 94.1 93.0 w / o EP&SM 86.2 97.7 91.5
[0095] 4. Experimental analysis
[0096] In view of the above situation, the specific embodiments of this application make the following analysis:
[0097] The error position information fusion strategy and the selective masking strategy proposed in the specific embodiments of the present application both significantly improve the Chinese spelling error correction capability. The fuzzy indication method used in the error position information fusion strategy also has a significant impact on the error correction effect.
[0098] When the fuzzy indication method is not used, the contextual relevance of the error is not explicitly provided to the error correction model, which causes the model to focus only on the precise location of the error; however, this precise location comes from the error detection module and cannot be guaranteed to be accurate. Therefore, when the detector indicates an error or deviation, the error location information is misjudged, resulting in the neglect of the erroneous character, interfering with the normal operation of the error corrector, and thus reducing the error correction results.
[0099] When the error position information fusion strategy is not used, only the high-recall error detection results are passed to the error correction network in the form of masks for error indication, hoping that the network can predict the correct characters from the masks, which brings a certain indication ability, but due to the lack of high-precision error detection results, a large number of simple erroneous characters need to be judged and corrected by the network, which makes the network unable to optimize in the direction of focusing on those difficult errors; therefore, the error correction network's ability to correct errors decreases.
[0100] When the selective masking strategy is not used, only high-precision error detection results are passed to the error correction network after fuzzy indication, and feature fused with the sentence to be corrected to indicate the location of the error; therefore, the network can easily judge those errors that are easy to detect; however, due to the lack of high-recall error detection results, slightly difficult errors will be directly input into the error correction network; if there is a certain rationality in the connection between these erroneous characters and the context, they are easily retained in the final output, which leads to the failure of error correction, and thus the error correction results are reduced.
[0101] When neither the error position information fusion strategy nor the selective masking strategy is used, the network does not receive any instructions from the detector and only uses the error correction network to process the masked sentences to be corrected, so the effect is significantly reduced. From the results of the ablation experiment, the precision is significantly reduced in this case, while the recall rate is improved to a certain extent. This is because when each original character is masked and no indication of the error position is given, the model may make additional corrections, which means that the model will correct more characters even if these characters may be correct. Under this error correction method, most errors will be corrected, but new errors are also introduced, so the recall rate is slightly improved and the precision is significantly reduced.
[0102] According to the Chinese spelling error correction method based on the detection-correction balance framework proposed in the embodiment of the present application, a pre-built language detector is trained by a preset Chinese spelling error correction data text to generate a Chinese spelling detector; the precision threshold and recall rate threshold corresponding to the Chinese spelling detector are set, and based on the precision threshold and recall rate threshold, the precision error detection result and the recall error detection result corresponding to the preset initial text data to be corrected are generated, and the mask text data corresponding to the initial text data to be corrected is obtained according to the recall error detection result; the mask text data and the initial text data to be corrected are spliced to obtain spliced data, and the embedded features of the spliced data are extracted by a pre-built Chinese spelling error corrector, and the embedded features and the precision error detection results are feature fused to obtain fused features, and based on the fused features and the preset target loss function, the Chinese spelling error corrector is trained to use the trained Chinese spelling error corrector to perform Chinese spelling error correction operations on the initial text data to be corrected. The present application constrains the output results of the detector, thereby fully exploring the performance of the detector without adding additional training and detection processes, and can effectively improve the error correction capability.
[0103] Secondly, a Chinese spelling error correction device based on a detection-correction balance framework proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0104] Figure 3 It is a block diagram of a Chinese spelling error correction device based on a detection-correction balance framework according to an embodiment of the present application.
[0105] like Figure 3 As shown, the Chinese spelling error correction device 10 based on the detection-correction balance framework includes: a training module 100 , a generation module 200 and a Chinese spelling error correction module 300 .
[0106] The training module 100 is used to train a pre-built language detector using preset Chinese spelling correction data text to generate a Chinese spelling detector.
[0107] The generation module 200 is used to set the precision threshold and recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and recall rate threshold, generate the precision error detection result and recall error detection result corresponding to the preset initial text data to be corrected, and obtain the mask text data corresponding to the initial text data to be corrected according to the recall error detection result.
[0108] The Chinese spelling correction module 300 is used to splice the masked text data and the initial text data to be corrected to obtain the spliced data, and extract the embedded features of the spliced data through a pre-built Chinese spelling corrector, and perform feature fusion on the embedded features and the precision error detection results to obtain fused features, and train the Chinese spelling corrector based on the fused features and a preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected.
[0109] Optionally, in one embodiment of the present application, the training module 100 includes: a determination unit and a judgment unit.
[0110] The determination unit is used to input the Chinese spelling error correction data text into the language detector to determine the error probability of each data position in the Chinese spelling error correction data text.
[0111] A judging unit is used to judge whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, the data position corresponding to the error probability is marked as an error prediction position, and the error prediction position is output to train the language detector.
[0112] Optionally, in one embodiment of the present application, the generation module 200 includes: a calculation unit and a mask unit.
[0113] Among them, the calculation unit is used to obtain the error prediction position corresponding to the initial text data to be corrected through the Chinese spelling detector, and calculate the fuzzy indication strength corresponding to each data position in the initial text data to be corrected according to the error prediction position and a preset probability density function.
[0114] The mask unit is used to determine at least one target data position in the initial text data to be corrected that meets the preset importance requirement based on the fuzzy indication strength and the preset strength threshold, and to perform a mask operation on each target data position in the at least one target data position and the data within the target range of each target data position to obtain masked text data corresponding to the initial text data to be corrected.
[0115] Optionally, in one embodiment of the present application, the Chinese spelling error correction module 300 includes: an acquisition unit and a copy unit.
[0116] The acquisition unit is used to extract the embedded features of the spliced data through the Chinese spelling corrector, and obtain the hidden layer dimensional features in the embedded features.
[0117] The copying unit is used to copy the hidden layer dimension features to obtain the hidden layer dimension copy features, and add the hidden layer dimension copy features to the precision error detection results to obtain the precision error detection results to be fused, and perform feature addition operation on the precision error detection results to be fused and the embedded features to generate fused features.
[0118] It should be noted that the aforementioned explanation of the embodiment of the Chinese spelling error correction method based on the detection-correction balance framework is also applicable to the Chinese spelling error correction device based on the detection-correction balance framework of this embodiment, and will not be repeated here.
[0119] According to the Chinese spelling correction device based on the detection-correction balance framework proposed in the embodiment of the present application, it includes a training module, which is used to train a pre-built language detector through a preset Chinese spelling correction data text to generate a Chinese spelling detector; a generation module, which is used to set the precision threshold and the recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and the recall rate threshold, generate a precision error detection result and a recall error detection result corresponding to the preset initial text data to be corrected, and obtain mask text data corresponding to the initial text data to be corrected according to the recall error detection result; a Chinese spelling correction module, which is used to splice the mask text data and the initial text data to be corrected to obtain spliced data, and extract the embedded features of the spliced data through a pre-built Chinese spelling corrector, and perform feature fusion on the embedded features and the precision error detection results to obtain fused features, and train the Chinese spelling corrector based on the fused features and the preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected. This application constrains the output results of the detector, so as to fully explore the performance of the detector without adding additional training and detection processes, and can effectively improve the error correction capability.
[0120] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0121] Memory 401 , processor 402 , and a computer program stored in the memory 401 and executable on the processor 402 .
[0122] When the processor 402 executes the program, the Chinese spelling error correction method based on the detection-correction balance framework provided in the above embodiment is implemented.
[0123] Furthermore, the electronic device further comprises:
[0124] The communication interface 403 is used for communication between the memory 401 and the processor 402 .
[0125] The memory 401 is used to store computer programs that can be executed on the processor 402 .
[0126] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0127] If the memory 401, the processor 402 and the communication interface 403 are implemented independently, the communication interface 403, the memory 401 and the processor 402 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0128] Optionally, in a specific implementation, if the memory 401, the processor 402 and the communication interface 403 are integrated on a chip, the memory 401, the processor 402 and the communication interface 403 can communicate with each other through an internal interface.
[0129] The processor 402 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0130] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the Chinese spelling error correction method based on the detection-correction balance framework as described above.
[0131] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned Chinese spelling error correction method based on the detection-correction balance framework.
[0132] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0133] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0134] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0135] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0136] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0137] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0138] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0139] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A Chinese spelling error correction method based on a detection-correction balanced framework, characterized in that: The following steps are involved: Train a pre-built language detector with preset Chinese spelling correction data text to generate a Chinese spelling detector; Setting a precision threshold and a recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and the recall rate threshold, generating a precision error detection result and a recall error detection result corresponding to a preset initial text data to be corrected, and obtaining mask text data corresponding to the initial text data to be corrected according to the recall error detection result; The masked text data and the initial text data to be corrected are spliced to obtain spliced data, and the embedded features of the spliced data are extracted by a pre-built Chinese spelling corrector, and the embedded features and the precision error detection results are feature fused to obtain fused features, and based on the fused features and a preset target loss function, the Chinese spelling corrector is trained to perform Chinese spelling correction operations on the initial text data to be corrected using the trained Chinese spelling corrector.
2. The method according to claim 1, characterized in that The method of training a pre-built language detector using preset Chinese spelling correction data text includes: Inputting the Chinese spelling error correction data text into the language detector to determine the error probability of each data position in the Chinese spelling error correction data text; Determine whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, mark the data position corresponding to the error probability as an error prediction position, and output the error prediction position to train the language detector.
3. The method according to claim 2, characterized in that The step of obtaining mask text data corresponding to the initial text data to be corrected according to the recalled error detection result includes: The error prediction position corresponding to the initial text data to be corrected is obtained by the Chinese spelling detector, and the fuzzy indication strength corresponding to each data position in the initial text data to be corrected is calculated according to the error prediction position and a preset probability density function; Based on the fuzzy indication strength and a preset strength threshold, at least one target data position in the initial text data to be corrected that meets the preset importance requirement is determined, and a mask operation is performed on each target data position in the at least one target data position and the data within the target range of each target data position to obtain masked text data corresponding to the initial text data to be corrected.
4. The method according to claim 3, characterized in that The method of extracting the embedded features of the spliced data by using a pre-built Chinese spelling error corrector and fusing the embedded features with the precision error detection results to obtain fused features includes: Extracting the embedded features of the spliced data through the Chinese spelling corrector, and obtaining the hidden layer dimensional features in the embedded features; The hidden layer dimensional features are copied to obtain hidden layer dimensional copy features, and the hidden layer dimensional copy features are added to the precision error detection results to obtain the precision error detection results to be fused, and feature addition operation is performed on the precision error detection results to be fused and the embedded features to generate the fused features.
5. A Chinese spelling error correction device based on a detection-correction balance framework, characterized in that: include: A training module, used for training a pre-built language detector through a preset Chinese spelling correction data text to generate a Chinese spelling detector; A generation module, used to set a precision threshold and a recall rate threshold corresponding to the Chinese spelling detector, and based on the precision threshold and the recall rate threshold, generate a precision error detection result and a recall error detection result corresponding to a preset initial text data to be corrected, and obtain mask text data corresponding to the initial text data to be corrected according to the recall error detection result; A Chinese spelling correction module is used to splice the masked text data and the initial text data to be corrected to obtain spliced data, and extract the embedded features of the spliced data through a pre-built Chinese spelling corrector, and perform feature fusion on the embedded features and the precision error detection results to obtain fused features, and train the Chinese spelling corrector based on the fused features and a preset target loss function, so as to use the trained Chinese spelling corrector to perform Chinese spelling correction operations on the initial text data to be corrected.
6. The device according to claim 5, characterized in that The training module includes: A determination unit, used for inputting the Chinese spelling error correction data text into the language detector to determine the error probability of each data position in the Chinese spelling error correction data text; A judging unit is used to judge whether the error probability is greater than a preset error threshold, wherein if the error probability is greater than the error threshold, the data position corresponding to the error probability is marked as an error prediction position, and the error prediction position is output to train the language detector.
7. The device according to claim 6, characterized in that The generation module comprises: A calculation unit, used for obtaining the error prediction position corresponding to the initial text data to be corrected through the Chinese spelling detector, and calculating the fuzzy indication strength corresponding to each data position in the initial text data to be corrected according to the error prediction position and a preset probability density function; A masking unit is used to determine at least one target data position in the initial text data to be corrected that meets a preset importance requirement based on the fuzzy indication strength and a preset strength threshold, and to perform a masking operation on each target data position in the at least one target data position and the data within a target range of each target data position to obtain masked text data corresponding to the initial text data to be corrected.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the Chinese spelling error correction method based on the detection-correction balance framework as described in any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the Chinese spelling error correction method based on a detection-correction balance framework as described in any one of claims 1 to 4.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed to implement the Chinese spelling error correction method based on the detection-correction balanced framework as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Chinese spelling error correction method and device based on comparative learning and medium
CN116127953A
Text error correction method, system and equipment and storage medium
CN116822464A