Grammar error correction method based on large model syntax preference optimization

By integrating structured syntax features and optimizing the parameter space of the syntax error correction model, the existing model's insufficient handling of complex syntax error errors is solved, and a more efficient and accurate syntax error correction effect is achieved.

CN120430302APending Publication Date: 2025-08-05KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510370031.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing grammatical error correction model lacks global understanding of the overall grammatical structure of the sentence when facing complex syntax errors, resulting in the error correction result that it may destroy the grammatical integrity of the sentence and is difficult to balance the calculation efficiency and precision.

Method used

By integrating the structured syntactic features in the search examples, using Monte Carlo tree search and syntactic preference alignment mechanisms, optimizing the parameter space of the syntactic error correction model, building a syntactic-aware error correction corpus, and improving the model's error correction ability for complex errors.

Benefits of technology

It significantly alleviates the overcorrection phenomenon of the syntax error correction model, improves the performance and robustness of the model in Chinese and English grammar error correction tasks, especially in complex syntax error scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430302A_ABST
    Figure CN120430302A_ABST
Patent Text Reader

Abstract

The invention discloses a grammar error correction method based on large model syntactic preference optimization. According to the grammar error correction method, syntactic relevance between retrieval example sentence pairs and syntactic relevance between the retrieval example sentence pairs are cooperatively utilized. On one hand, a Monte Carlo tree search algorithm is used for exploring the influence of syntactic structure differences among retrieval examples on error correction performance, search path selection is adjusted based on dynamic weights, and syntactic perception corpora containing hidden syntactic association are constructed. On the other hand, by constructing a syntactic preference alignment mechanism, syntactic knowledge contained in syntactic perception corpora is effectively utilized, dynamic adjustment of grammar error correction model parameters is achieved, and then the performance of the grammar error correction model in a complex grammar error correction task is improved. According to the method, a comprehensive contrast experiment is carried out on two Chinese error correction data sets and two English error correction data sets. Experimental results show that the method can effectively integrate the structured syntactic features in the retrieval examples, and effectively relieve the over-correction phenomenon of the grammar error correction model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Language

[0002] The present invention relates to a grammatical error correction method based on large-model syntactic preference optimization, belonging to the technical field of natural language processing. Background Art

[0003] In the field of natural language processing, grammatical error correction technology is one of the key tasks to improve the standardization and readability of texts, and is widely used in scenarios such as educational assistance, content review, and intelligent writing. From real-time correction of writing errors for non-native learners in educational assistance scenarios, to the standardization review of news releases and advertising copy in the content production field, to improving the semantic accuracy of intelligent customer service in human-computer interaction scenarios, grammatical error correction technology has become a core link in ensuring the quality of language expression. Especially in the post-processing of machine translation, this technology can effectively repair problems such as word order confusion or missing function words caused by differences in grammatical structure, significantly improving the readability of cross-language conversion. However, current mainstream error correction models still face severe challenges when faced with unique complex syntactic errors, such as logical breaks in complex sentences, long-distance modification ambiguity, or interrupted topic chains. The fundamental reason is that traditional methods over-rely on local word-level features and lack a global understanding of the overall grammatical structure of the sentence.

[0004] Existing technologies usually rely on local contextual information or word-level features, and lack explicit modeling of the overall syntactic structure of the sentence, resulting in error correction results that may undermine the grammatical integrity of the sentence. For example, probability corrections that rely solely on language models may ignore the dependencies or component boundaries of the sentence, resulting in grammatical correctness but semantic deviation after error correction. In addition, most methods do not fully utilize the structured information provided by syntactic analysis tools, making it difficult for the model to distinguish between seemingly similar error types. On the other hand, traditional end-to-end error correction systems often face the trade-off between computational efficiency and accuracy in real-time application scenarios, especially when combined with syntactic parsing modules, which may introduce additional computational overhead.

[0005] Although existing grammatical error correction research has attempted to introduce syntactic features, there is still a significant gap between its application methods and the needs of real scenarios. Most methods are limited to single-level modeling of dependencies or component boundaries, and fail to systematically integrate the dynamic relationship between syntactic levels, long-distance constraints and semantic coherence, resulting in insufficient coverage of complex error types. For example, in the tasks of complex sentence logic correction or topic chain reconstruction, traditional syntactic rules and neural networks often have decision conflicts, which aggravates the randomness of the correction results. In response to the above problems, the present invention proposes a grammatical error correction method based on large-model syntactic preference optimization, which effectively alleviates the over-correction phenomenon of the grammatical error correction model by integrating the structured syntactic features in the retrieval examples. Summary of the Invention

[0006] This paper proposes a grammatical correction method based on large-scale model syntactic preference optimization. This method effectively integrates structured syntactic features from retrieval examples, significantly alleviating the overcorrection problem commonly seen in grammatical correction models. This method was comprehensively validated using two publicly available Chinese and English grammatical correction datasets, and the results fully demonstrate its effectiveness.

[0007] The technical solution of the present invention is: a grammatical error correction method based on large-model syntactic preference optimization, the specific steps of the method are as follows:

[0008] Step 1: Based on the syntactic feature analysis of grammatically incorrect sentences, we select several typical example sentence pairs with similar structural features from the manually annotated high-quality grammatically corrected corpus using a syntactic structure matching algorithm to construct a context example set.

[0009] Step 2: Input only grammatical correction examples and grammatical correction examples with example sentence pairs into the grammatical correction model respectively. Through comparative experiments, evaluate the impact of example guidance on grammatical correction results and screen candidate corpora for syntactic preference correction.

[0010] Step 3: Based on affinity and diversity evaluation indicators, data enhancement is performed on the selected candidate syntactic bias correction corpus to enrich its grammatical error distribution characteristics and construct the syntactic bias correction corpus;

[0011] Step 4: Based on the syntactic-biased correction corpus, an example selection algorithm based on Monte Carlo tree search is constructed. By iteratively optimizing the effect of different example sentence pairs on grammatical correction, a syntactic-aware correction corpus is finally formed.

[0012] Step 5: Design a syntactic preference alignment mechanism based on the direct preference optimization algorithm, and use the positive and negative sample pairs in the syntax-aware error correction corpus to dynamically optimize the parameter space of the grammatical error correction model, thereby improving its performance in the grammatical error correction task.

[0013] As a further solution of the present invention, step 1 includes the following specific steps:

[0014] Step 1.1: Download high-quality, manually annotated grammar correction corpora. To avoid overly large sample selection space, select the HSK Chinese grammar correction corpus and the W&I+LOCNESS English grammar correction corpus as the sample sentence pairs.

[0015] Step 1.2: Use a syntactic parser based on grammatically incorrect sentences to extract and generate grammatically incorrect sentences and example sentence pairs. Select the syntactic features of grammatically incorrect and grammatically correct sentence pairs in the corpus.

[0016] Step 1.3: Use the polynomial distance algorithm and the tree kernel algorithm to calculate the syntactic feature similarity between the grammatically incorrect sentences and the sample sentence pairs in the selected corpus, and select the top 5 with the highest similarity as the context example set.

[0017] As a further solution of the present invention, step 2 includes the following specific steps:

[0018] Step 2.1: Input only grammatically corrected sentences and grammatically corrected sentences with example sentence pairs into the grammatical correction model to obtain two sets of model correction results.

[0019] Step 2.2: Compare and analyze the error correction results of the two groups of models with the corresponding grammatically correct examples. Four situations can be summarized: "all correct", "syntactically valid", "syntactically invalid" and "all wrong".

[0020] Step 2.3: Select instances belonging to the "syntactically valid" and "syntactically invalid" cases as candidate corpora for syntactic preference correction.

[0021] As a further solution of the present invention, step 3 includes the following specific steps:

[0022] Step 3.1: Calculate the affinity and diversity scores of example sentence pairs from grammatically incorrect sentences to grammatically correct sentences in the candidate syntactic preference correction corpus.

[0023] Step 3.2: Based on the calculated affinity and diversity scores, the candidate syntactic preference correction corpus is augmented with random replacement to further enrich its grammatical error distribution characteristics, thereby constructing a more diverse and complete syntactic preference correction corpus.

[0024] As a further solution of the present invention, step 4 includes the following specific steps:

[0025] Step 4.1: For grammatically incorrect sentences in the syntactic bias correction corpus, select the sentence pair with the highest syntactic feature similarity from the five sample sentence pairs selected for it as the initial node. Then, randomly select one of the remaining four sample sentence pairs and add it as a new child node to the search tree, completing the construction of the initial structure.

[0026] Step 4.2: Starting from the initial node of the current search tree, explore downward along the branches of the tree until you reach an unexplored node or a leaf node. During this process, use the UCT formula to balance exploration and exploitation to select the optimal path.

[0027] Step 4.3: If the node reached in step 4.2 is an undeveloped node, randomly pick one from the unused example sentence pairs, generate a new child node and add it to the search tree, further expanding the tree structure.

[0028] Step 4.4: Randomly sample the newly expanded nodes to generate several candidate example combinations, and test the performance of these combinations in the grammatical correction task using the grammatical correction accuracy indicator.

[0029] Step 4.5: The evaluation results obtained in the simulation phase are transmitted back layer by layer along the path to update the number of visits and average performance of each node, providing a basis for subsequent path selection.

[0030] Step 4.6: Repeat steps 4.2 to 4.5 until the predetermined number of iterations or time limit is reached. Finally, positive examples (effective example combinations) and negative examples (ineffective example combinations) that have a significant impact on the grammatical error correction task are screened from the search tree to form high-quality syntax-aware error correction corpus.

[0031] As a further solution of the present invention, step 5 includes the following specific steps:

[0032] Step 5.1: Design a syntactic preference alignment mechanism based on a direct preference optimization algorithm. Using positive samples (valid example combinations) and negative samples (invalid example combinations) from the syntax-aware error correction corpus, define the preference relationship and clarify the criteria for positive samples to outperform negative samples in the grammatical error correction task.

[0033] Step 5.2: Based on the constructed syntactic preference alignment mechanism, dynamically optimize the parameter space of the large grammatical error correction model. During training, by comparing the output generated by the model with the preference relationship between positive and negative samples, the model parameters are adjusted to make it more inclined to generate output that conforms to the syntactic rules and task requirements.

[0034] Step 5.3: After each round of optimization, use the validation set to verify the model's performance, focusing on its grammatical error correction capabilities. Based on the verification results, adjust the optimization strategy to further strengthen the model's grammatical error correction capabilities and significantly enhance the model's performance in grammatical error correction tasks.

[0035] The present invention also provides a grammatical error correction system based on large-model syntactic preference optimization, the system comprising: a module for executing the above-mentioned grammatical error correction method based on large-model syntactic preference optimization.

[0036] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned grammatical error correction method based on large-model syntactic preference optimization is implemented.

[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the above-mentioned grammatical error correction method based on large model syntactic preference optimization.

[0038] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned grammatical error correction method based on large model syntactic preference optimization.

[0039] The beneficial effects of the present invention are:

[0040] 1. The present invention constructs a grammatical error correction and device based on syntax perception, which adjusts the syntactic knowledge preferences within the large model by utilizing the syntactic features of grammatically incorrect sentences, thereby improving its performance in Chinese and English grammatical error correction tasks.

[0041] 2. This paper analyzes the syntactic deviations of the large grammar error correction model in detail and constructs syntactic preference data based on this analysis. By building a syntactic preference alignment mechanism, the generalization ability and robustness of the large model are enhanced.

[0042] 3. The present invention has conducted comprehensive experimental verification on two public Chinese grammar correction datasets and two public English grammar correction datasets, verifying the effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the present invention;

[0044] Figure 2 A diagram of a grammatical error correction method for optimizing syntax preference of a large model in the present invention; DETAILED DESCRIPTION

[0045] Example 1: Figure 1 and Figure 2 As shown, a grammatical error correction method based on large-model syntactic preference optimization is provided, and the specific steps of the method are as follows:

[0046] a1. Based on the syntactic feature analysis of grammatically incorrect sentences, we selected five typical example sentence pairs with similar structural features from the manually annotated high-quality grammatically corrected corpus using a syntactic structure matching algorithm to construct a context example set. This includes:

[0047] 1) Download high-quality, manually annotated Chinese and English grammar correction corpora, and select high-quality and small-scale HSK Chinese grammar correction corpora and W&I+LOCNESS English grammar correction corpora as sample sentence pair selection corpora to avoid overly large sample selection space.

[0048] 2) Using a syntactic parser based on grammatically incorrect sentences, extract and generate grammatically incorrect sentences and example sentences to select syntactic features such as phrase structure, dependency, component boundary, central word, syntactic distance, and syntactic rules in the corpus.

[0049] 3) Using the polynomial distance algorithm and the tree kernel algorithm, the similarity of syntactic features between grammatically incorrect sentences and sample sentence pairs is calculated, and the top 5 with the highest similarity are selected as the context example set.

[0050] a2. Input only grammatical correction examples and grammatical correction examples with example sentence pairs into the grammatical correction model. Through comparative experiments, evaluate the impact of example guidance on grammatical correction results and screen candidate syntactic preference corpora, including:

[0051] 1) Input only grammatical correction sentences and grammatical correction sentences with example sentence pairs into the grammatical correction model respectively to obtain two sets of model correction results.

[0052] 2) The error correction results of the two groups of models were compared and analyzed with the corresponding grammatically correct examples, and four situations were summarized: "Both correct": The model can directly perform correct grammatical correction without referring to the example, and can still perform correct correction after referring to the example. "Syntactically valid": The model cannot perform correct correction without referring to the example, but can perform correct grammatical correction after referring to the example. "Syntactically invalid": The model can perform correct correction without referring to the example, but cannot perform correct grammatical correction after referring to the example. "Both incorrect": Regardless of whether the example is referred to, the model cannot complete correct grammatical correction.

[0053] 3) Select instances belonging to the two cases of "syntactically valid" and "syntactically invalid" as candidate corpora for syntactic preference correction.

[0054] a3. Based on affinity and diversity evaluation indicators, data enhancement is performed on the selected candidate syntactic preference correction corpus to enrich its grammatical error distribution characteristics and construct the syntactic preference correction corpus, including:

[0055] 1) Calculate the affinity and diversity scores of example sentence pairs from grammatically incorrect sentences to grammatically correct sentences in the candidate syntactic preference correction corpus. The calculation formulas for affinity and diversity are as follows:

[0056]

[0057] in, and represents pseudo-corpus and real corpus; x and y represent grammatically incorrect sentences and grammatically correct sentences; P r and P prepresents the probability of grammatical errors in the pseudo-corpus and the real corpus; KL represents the use of KL divergence to calculate the distance between the two distribution probabilities, and its calculation formula is as follows:

[0058]

[0059] Where x and y represent grammatically incorrect and grammatically correct sentences; P(·) represents P p (·) and P r (·)The affinity value between them.

[0060] 2) Based on the calculated affinity and diversity scores, the candidate syntactic preference correction corpus is enhanced by random replacement to further enrich its grammatical error distribution characteristics, thereby constructing a more diverse and complete syntactic preference correction corpus.

[0061] a4. Based on the syntactic-biased correction corpus, we construct an example selection algorithm based on Monte Carlo tree search. By iteratively optimizing the effect of different example sentence pairs on grammatical correction, we ultimately form a syntax-aware correction corpus, including:

[0062] 1) For grammatically incorrect sentences in the syntactic bias correction corpus, select the sentence pair with the highest syntactic feature similarity from the five sample sentence pairs with high syntactic feature similarity as the initial node. Then, randomly select one of the remaining four sample sentence pairs and add it as a new child node to the search tree, completing the construction of the initial structure.

[0063] 2) Starting from the initial node of the current search tree, explore downward along the branches of the tree until you reach an unexplored node or a leaf node. During this process, the UCT formula is used to balance exploration and exploitation to select the optimal path. The UCT formula is calculated as follows:

[0064]

[0065] Where Q(s,c) is defined as the reward of node s under the correction c, N(s) is defined as the number of visits to node s, p is the parent node, and w is the exploration weight.

[0066] 3) If the node reached in step 4.2 is an undeveloped node, randomly pick one from the unused example sentence pairs, generate a new child node and add it to the search tree to further expand the tree structure.

[0067] 4) Randomly sample the newly expanded nodes to generate several candidate example combinations, and test the performance of these combinations in the grammatical error correction task using the grammatical error correction accuracy indicator.

[0068] 5) The evaluation results obtained in the simulation phase are transmitted back layer by layer along the path, and the number of visits and average performance of each node are updated to provide a basis for subsequent path selection.

[0069] 6) Repeat steps 2) to 5) until the predetermined number of iterations or time limit is reached. Finally, positive examples (effective example combinations) and negative examples (ineffective example combinations) that have a significant impact on the grammatical error correction task are screened from the search tree to form the syntax-aware error correction corpus.

[0070] a5. Design a syntactic preference alignment mechanism based on the direct preference optimization algorithm, and use the positive and negative sample pairs in the syntax-aware error correction corpus to dynamically optimize the parameter space of the grammatical error correction model, thereby improving its performance in grammatical error correction tasks.

[0071] include:

[0072] 1) Design a syntactic preference alignment mechanism based on a direct preference optimization algorithm. Using positive samples (valid example combinations) and negative samples (invalid example combinations) from the syntax-aware error correction corpus, define preference relationships and clarify the criteria for positive samples to outperform negative samples in grammatical error correction tasks.

[0073] 2) Based on the constructed syntactic preference alignment mechanism, the parameter space of the grammatical error correction model is dynamically optimized. During the training process, by comparing the output generated by the model with the preference relationship between positive and negative samples, the model parameters are adjusted to make it more inclined to generate output that conforms to the syntactic rules and task requirements. The following formula is used to update the model parameters:

[0074]

[0075] Among them, (x,y w ,y l ) is a triple sampled from the syntax-aware error correction corpus, x represents a grammatically incorrect sentence, y w Represents samples with positive feedback for error correction, y l represents samples with negative feedback on error correction; represents the syntactic-aware error correction corpus; π θ (y|x) represents the probability distribution of the current model outputting y given a grammatically incorrect sentence x; π ref (y|x) represents the probability distribution of the reference model outputting y given a grammatically incorrect sentence x; β represents the temperature parameter that controls the optimization intensity; σ represents the Sigmoid function; Represents the expected value of a triple.

[0076] 3) After each round of optimization, the model's performance was verified using the validation set, with a focus on its grammatical error correction capabilities. Based on the verification results, the optimization strategy was adjusted to further strengthen the model's grammatical error correction capabilities, significantly enhancing the model's performance in Chinese and English grammatical error correction tasks.

[0077] The present invention also provides a grammatical error correction system based on large-model syntactic preference optimization, the system comprising:

[0078] The context example set construction module is used to analyze the syntactic features of grammatically incorrect sentences. It uses a syntactic structure matching algorithm to select several typical example sentence pairs with similar structural features from high-quality, manually annotated grammatically corrected corpus to construct a context example set.

[0079] The candidate syntactic bias correction corpus screening module is used to input only grammatical correction examples and grammatical correction examples with example sentence pairs into the grammatical correction model, evaluate the impact of example guidance on grammatical correction results through comparative experiments, and screen candidate syntactic bias correction corpora;

[0080] The syntactic bias error correction corpus construction module is used to perform data enhancement on the selected candidate syntactic bias error correction corpus based on affinity and diversity evaluation indicators, enrich its grammatical error distribution characteristics, and construct the syntactic bias error correction corpus;

[0081] The syntax-aware error correction corpus acquisition module is used to construct an example selection algorithm based on Monte Carlo tree search based on the syntax-biased error correction corpus. It iteratively optimizes the effect of combining different example sentence pairs on grammatical error correction, and ultimately forms the syntax-aware error correction corpus.

[0082] The optimization module is used to design a syntactic preference alignment mechanism based on the direct preference optimization algorithm. It uses the positive and negative sample pairs in the syntax-aware error correction corpus to dynamically optimize the parameter space of the grammatical error correction model, thereby improving its performance in grammatical error correction tasks.

[0083] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned grammatical error correction method based on large-model syntactic preference optimization is implemented.

[0084] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the above-mentioned grammatical error correction method based on large model syntactic preference optimization.

[0085] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned grammatical error correction method based on large model syntactic preference optimization.

[0086] Experiments in this paper were conducted on multiple authoritative grammatical error correction datasets, including the MuCGEC and FCGEC error correction datasets for Chinese, and the CoNLL-2014 and BEA-2019 error correction datasets for English. These datasets cover a rich variety of linguistic phenomena and error types, providing a solid foundation for a comprehensive evaluation of model performance and further verifying the robustness and generalization capabilities of the large error correction model in cross-lingual scenarios.

[0087] In order to measure the effectiveness of the grammatical error correction method based on large model syntactic preference optimization in Chinese and English grammatical error correction tasks, this paper uses precision, recall and F 0.5 Evaluation indexes are used for evaluation.

[0088] Precision: The percentage of sentences that were correctly identified as incorrect among all sentences that were incorrectly identified by the model. The calculation formula is as follows:

[0089]

[0090] Recall: The percentage of sentences correctly identified as correct among all sentences correctly identified by the model. Its calculation formula is as follows:

[0091]

[0092] F 0.5 Value: The weighted harmonic mean of precision and recall, used to measure the performance of the model. Its calculation formula is as follows:

[0093]

[0094] The present invention verifies the performance of the grammatical correction method based on large-model syntactic preference optimization on the MuCGEC and FCGEC error correction datasets in the Chinese field, and the CoNLL-2014 and BEA-2019 error correction datasets in the English field. The present invention combines this method with GECToR (using a customized word-level conversion method to map the input word to its target correction result), SynGEC (integrating dependency syntactic information into the grammatical correction model by using a parser specifically for grammatical correction tasks), and DeCoGLM (detecting the correction structure by combining autoregressive mask filling). At the same time, the present invention has been fully experimentally verified on multiple representative Chinese and English grammatical correction models, including a general Chinese and English correction model (BART), two large-scale models focusing on Chinese grammatical correction (Qwen2-7B and GLM4-9B), and two large models with outstanding performance in the field of English grammatical correction (Llama3-8B and Mistral-7B). These models cover different language scenarios and technical architectures, providing a strong guarantee for the comprehensiveness and reliability of the experimental results.

[0095] Table 1 shows the comparative experimental results on the MuCGEC and FCGEC datasets. It can be clearly seen that the improvement effect of the present invention on the FCGEC dataset is more significant. This is mainly because the errors in the FCGEC dataset come from native speakers, and their error types have strong regularity and typical syntactic patterns. In contrast, the errors in the MuCGEC dataset mainly come from second language learners. These errors not only occur less frequently, but also have more diverse and complex forms, resulting in a relatively limited improvement effect of the model. This comparison result further highlights the importance of syntactic preference learning. By introducing the syntactic preference alignment mechanism, the model can more efficiently utilize syntactic structure information, thereby improving the ability to capture and correct grammatical errors.

[0096] Table 1 Results of different methods on MuCGEC and FCGEC datasets

[0097]

[0098] Table 2 shows the experimental results of the present invention on the CoNLL-2014 and BEA-2019 datasets. The results show that the present invention can also significantly improve the performance in the English field, especially on BEA-2019. This is because the BEA-2019 data source is wider, the error types are more complex and the annotation quality is higher, covering context-related semantic problems and high-order language phenomena, and putting higher demands on the model's context understanding ability. In contrast, the error types of CoNLL-2014 are more concentrated on basic grammatical problems with strong regularity and are less difficult. The present invention fully demonstrates its versatility and robustness in multilingual tasks.

[0099] Table 2 Results of different methods on CoNLL-2014 and BEA-2019 datasets

[0100]

[0101] Furthermore, experimental results demonstrate the robustness of this method, showing consistent performance improvements across error correction datasets and large error correction models in different languages. This further validates the critical role of syntax-aware enhancement technology, particularly its effectiveness in addressing complex grammatical errors. By integrating syntactic information, this invention not only improves the accuracy of the model but also demonstrates its potential for application in a variety of scenarios.

[0102] The above content has been described in detail with reference to the accompanying drawings for the specific embodiments of the present invention. However, the present invention is not limited to the above embodiments. Various modifications can be made within the scope of knowledge generally known to those skilled in the art without violating the purpose of the present invention.

Claims

1. A grammatical error correction method based on large-scale model syntactic preference optimization, characterized by: The method comprises the following steps: Step 1: Based on the syntactic feature analysis of grammatically incorrect sentences, we select several typical example sentence pairs with similar structural features from the manually annotated high-quality grammatically corrected corpus using a syntactic structure matching algorithm to construct a context example set. Step 2: Input only grammatical correction examples and grammatical correction examples with example sentence pairs into the grammatical correction model respectively. Through comparative experiments, evaluate the impact of example guidance on grammatical correction results and screen candidate corpora for syntactic preference correction. Step 3: Based on affinity and diversity evaluation indicators, data enhancement is performed on the selected candidate syntactic bias correction corpus to enrich its grammatical error distribution characteristics and construct the syntactic bias correction corpus; Step 4: Based on the syntactic-biased correction corpus, an example selection algorithm based on Monte Carlo tree search is constructed. By iteratively optimizing the effect of different example sentence pairs on grammatical correction, a syntactic-aware correction corpus is finally formed. Step 5: Design a syntactic preference alignment mechanism based on the direct preference optimization algorithm, and use the positive and negative sample pairs in the syntax-aware error correction corpus to dynamically optimize the parameter space of the grammatical error correction model, thereby improving its performance in the grammatical error correction task.

2. The grammatical error correction method based on large-model syntactic preference optimization according to claim 1, characterized in that: The step 1 includes the following specific steps: Step 1.1: Download high-quality, manually annotated Chinese and English grammar correction corpora. To avoid overly large sample selection space, select the HSK Chinese grammar correction corpus and the W&I+LOCNESS English grammar correction corpus as the sample sentence pairs. Step 1.2: Use a syntactic parser based on grammatically incorrect sentences to extract and generate grammatically incorrect sentences and sample sentence pairs. Select the syntactic features of grammatically incorrect and grammatically correct sentence pairs in the corpus. Step 1.3: Use the polynomial distance algorithm and the tree kernel algorithm to calculate the syntactic feature similarity between the grammatically incorrect sentences and the sample sentence pairs. Select the top several with the highest similarity as the context example set.

3. The grammatical error correction method based on large-model syntactic preference optimization according to claim 1 is characterized in that: The step 2 includes the following specific steps: Step 2.1: Input only grammatically corrected sentences and grammatically corrected sentences with example sentence pairs into the grammatical correction model to obtain two sets of model correction results; Step 2.2: Compare and analyze the error correction results of the two models with the corresponding grammatically correct examples, and summarize the four situations: "all correct", "syntactically valid", "syntactically invalid", and "all incorrect"; Step 2.3: Select instances belonging to the "syntactically valid" and "syntactically invalid" cases as candidate corpora for syntactic preference correction.

4. The grammatical error correction method based on large-model syntactic preference optimization according to claim 1, characterized in that: The step 3 includes the following specific steps: Step 3.1: Calculate the affinity and diversity scores of example sentence pairs from grammatically incorrect sentences to grammatically correct sentences in the candidate syntactic preference correction corpus; Step 3.2: Based on the calculated affinity and diversity scores, the candidate syntactic preference correction corpus is augmented with random replacement to further enrich its grammatical error distribution characteristics, thereby constructing a more diverse and complete syntactic preference correction corpus.

5. The grammatical error correction method based on large-model syntactic preference optimization according to claim 1 is characterized in that: The step 4 includes the following specific steps: Step 4.1: For grammatically incorrect sentences in the syntactic bias correction corpus, select the sentence pair with the highest syntactic feature similarity from among several sample sentence pairs with high syntactic feature similarity screened out for the sentence as the initial node; Then, one of the remaining example sentence pairs is randomly selected and added to the search tree as a new child node to complete the construction of the initial structure. Step 4.2: Starting from the initial node of the current search tree, explore downward along the branches of the tree until you reach an unexplored node or a leaf node. During this process, use the UCT formula to balance exploration and exploitation to select the optimal path. Step 4.3: If the node reached in step 4.2 is an undeveloped node, randomly pick one from the unused sample sentence pairs, generate a new child node and add it to the search tree to further expand the tree structure; Step 4.4: Randomly sample the newly expanded nodes to generate several candidate example combinations, and test the performance of these combinations in the grammatical error correction task using the grammatical error correction accuracy metric; Step 4.5: The evaluation results obtained in the simulation phase are transmitted back layer by layer along the path, and the number of visits and average performance of each node are updated to provide a basis for subsequent path selection; Step 4.6: Repeat steps 4.2 to 4.5 until the predetermined number of iterations or time limit is reached; finally, positive and negative examples that have a significant impact on the grammatical error correction task are screened out from the search tree to form high-quality syntax-aware error correction corpus, where positive examples are valid example combinations and negative examples are invalid example combinations.

6. The grammatical error correction method based on large-model syntactic preference optimization according to claim 1, characterized in that: The step 5 includes the following specific steps: Step 5.1: Design a syntactic preference alignment mechanism based on the direct preference optimization algorithm. Utilize the positive and negative samples in the syntax-aware error correction corpus to define the preference relationship and clarify the criteria for positive samples to outperform negative samples in the grammatical error correction task. Step 5.2: Based on the constructed syntactic preference alignment mechanism, dynamically optimize the parameter space of the grammatical error correction model; During the training process, by comparing the output generated by the model with the preference relationship between positive and negative samples, the model parameters are adjusted to make it more inclined to generate output that conforms to the syntactic rules and task requirements; Step 5.3: After each round of optimization, use the validation set to verify the model's performance, focusing on its grammatical error correction capabilities. Based on the verification results, adjust the optimization strategy to further strengthen the model's grammatical error correction capabilities and significantly enhance the model's performance in Chinese and English grammatical error correction tasks.

7. A grammar error correction system based on large-scale model syntactic preference optimization, characterized in that: The system includes: a module for executing a grammatical error correction method based on large model syntactic preference optimization as described in any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a grammatical error correction method based on large model syntactic preference optimization as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the grammatical error correction method based on large-model syntactic preference optimization as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the grammatical error correction method based on large-model syntactic preference optimization as described in any one of claims 1 to 6 is implemented.