A vietnamese grammar correction method based on syntax error interpretation enhancement
Patent Information
- Application Number
- CN202610549501.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-23
- Publication Date
- 2026-09-11
AI Technical Summary
[0007]本发明要解决的技术问题是:针对越南语语法纠错任务中存在的语料稀缺、大模型在低资源语言上泛化性差以及传统检索方法语义匹配度低的问题,本发明提供一种基于语法错误解释增强的越南语语法纠错方法
[0032] 1. This invention introduces semantic matching retrieval, ensuring a high degree of semantic and structural consistency between the context examples and the current input sentence, thus avoiding noise interference from traditional random retrieval. Experiments show that this method significantly improves Vietnamese grammar correction on multiple large model bases. value.
Smart Images

Figure CN122735698A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a Vietnamese grammar correction method based on grammatical error explanation enhancement, belonging to the fields of natural language processing and low-resource language intelligent processing technology. Specifically, it relates to a Vietnamese grammar correction method based on grammatical error explanation enhancement that utilizes a large language model combined with semantic retrieval technology and improves model performance by introducing high-quality grammatical error explanations. Background Technology
[0002] As a fundamental and core task in the field of natural language processing, grammatical error correction aims to automatically detect and correct grammatical errors in text. It has extremely high application value for improving the output quality of downstream tasks such as automatic speech recognition and machine translation, as well as assisting non-native speakers in language learning.
[0003] However, existing research and applications in grammar correction mainly focus on high-resource languages such as English and Chinese, while research on low-resource languages such as Vietnamese lags behind and faces many challenges:
[0004] The complexity of Vietnamese language characteristics: Vietnamese is a typical analytic language. Unlike inflectional languages that express grammar through word inflection (such as verb conjugation and plural suffixes), Vietnamese lacks morphological inflection. Its grammatical relations are mainly reflected by strict word order and a rich vocabulary of function words (such as classifiers and modal particles). This leads to grammatical errors in Vietnamese often manifesting as subtle word order confusion, misuse of function words, or inappropriate semantic collocation (such as the difference between "drink rice" and "eat rice"), rather than explicit spelling or morphological errors, making it difficult for traditional error correction models to accurately detect them.
[0005] Limitations of Large Language Models: While large language models such as ChatGPT, LLaMA, and Qwen possess powerful generative capabilities, they are prone to "illusions" or misjudgments when dealing with specific low-resource languages due to a lack of domain-specific guidance. Existing retrieval augmentation (RAG) methods are mostly based on keyword matching (such as BM25), easily overlooking deep semantic structural similarities. Furthermore, most existing systems only output correction results, lacking interpretability of the reasons for errors, thus limiting the model's logical reasoning ability in complex grammatical scenarios.
[0006] It is particularly important to develop a Vietnamese grammar correction method that fully utilizes limited Vietnamese corpus resources, combining semantic matching retrieval mechanisms with high-quality grammatical error explanations to enhance the capabilities of large-scale models. By retrieving examples that are highly similar to the current input in semantics and structure, and combining this with detailed error attribution explanations, the language reasoning ability of large-scale models can be effectively activated, thereby strengthening their ability to correct grammatical errors in low-resource languages. Summary of the Invention
[0007] The technical problem this invention aims to solve is the scarcity of corpora, poor generalization of large models on low-resource languages, and low semantic matching accuracy of traditional retrieval methods in Vietnamese grammar correction tasks. This invention provides a Vietnamese grammar correction method based on enhanced grammar error explanations. This method effectively improves the ability of large models to identify and correct Vietnamese grammar errors by constructing a dataset containing detailed explanations of grammar errors and using semantic retrieval techniques to retrieve similar examples as contextual cues.
[0008] The technical solution of this invention is: a method for constructing Lao grammar correction training data based on iterative optimization, the method comprising:
[0009] Step 1: Construction of the Vietnamese Explanation Dataset: The existing Vietnamese grammar error correction dataset is annotated using annotation tools to obtain error types and correction information; then, a large language model is used to generate corresponding natural language explanations based on prompt templates, constructing an explanation dataset containing source sentences, target sentences, and detailed error explanations;
[0010] Step 2, Contextual Example Retrieval: A pre-trained vector encoding model is used to vectorize the source sentences in the explanation dataset, constructing a semantic index. For the input sentence to be corrected, the cosine similarity between the sentence and the vectors in the index is calculated to retrieve several examples with the closest semantic and syntactic structure as contextual references.
[0011] Step 3: Integrate into the large model for grammatical error correction: Combine the retrieved examples containing erroneous explanations with the current sentence to be corrected to form a prompt input, guide the large language model to learn the context, analyze the input sentence by analogy with the error correction logic and explanation information in the examples, and generate the final correction result.
[0012] Further, Step 1 includes:
[0013] Step 1.1: Select the manually annotated Vietnamese grammar correction dataset ViSGEC as the basic corpus, denoted as dataset.
[0014] Step 1.2: Use the modified ChERRANT annotation tool to annotate the dataset. The sentence pairs in the text are labeled at the word level, and the error type set and correction vocabulary for each sentence pair are extracted. ,in Error type For the correction word, For the first The number of errors in each sentence pair;
[0015] Step 1.3: Construct an explanation generation prompt template tailored to the characteristics of the Vietnamese language. The template includes task role settings, content requirements, and output format;
[0016] Step 1.4: Fill the annotated data into the prompt template, and use the DeepSeek-V3 large language model to generate grammatical error explanations for each sentence pair. The formula is as follows:
[0017]
[0018] The final form includes Explanation of dataset .
[0019] Furthermore, Step 2 includes:
[0020] Step 2.1: Constructing a Semantic Vector Index Library. The Sentence Transformers model, which supports multiple languages, is selected as the vector encoder. All source sentences (i.e., sentences containing errors) in the explanation dataset are mapped to fixed-dimensional semantic vectors to construct a semantic vector database.
[0021] Step 2.2: Vectorization of the sentence to be corrected. For any input sentence to be corrected, its embedding vector is calculated using the same vector encoder.
[0022] Step 2.3: Semantic Similarity Retrieval. Calculate the cosine similarity between the semantic vector of the input sentence to be corrected and all vectors in the vector database, and sort them from highest to lowest similarity score. Retrieve the vector with the highest similarity. These examples form a retrieval example set.
[0023] Furthermore, Step 3 includes:
[0024] Step 3.1: Construct an error correction prompt template that combines contextual retrieval enhancement. The template structure includes: a task instruction area, a retrieval example display area, and an input area to be corrected.
[0025] Step 3.2: In the search example display area, format and enter the search results. Each example contains a "src" (source sentence), a "tgt" (target sentence), and an "Explanation" (a brief description of the syntax error and the reason for the correction), explicitly showing the model the reason for the error and the correction logic;
[0026] Step 3.2: Large-Scale Model Inference and Output. The assembled complete prompts are input into the multilingual large-scale model (e.g., LLaMA3.1-8B, Qwen3-8B, Qwen3-14B). Based on the examples in the prompts, the model learns by analogy how errors in the examples are identified, explained, and corrected, and then analyzes the sentence to be corrected to generate the corrected sentence.
[0027] The present invention also provides a Vietnamese grammar correction system based on enhanced grammar error interpretation, the system comprising: a module for executing the Vietnamese grammar correction method based on enhanced grammar error interpretation.
[0028] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the Vietnamese grammar correction method based on grammar error interpretation enhancement.
[0029] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the Vietnamese grammar correction method based on grammar error interpretation enhancement.
[0030] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the Vietnamese grammar correction method based on grammar error interpretation enhancement.
[0031] The beneficial effects of this invention are:
[0032] 1. This invention introduces semantic matching retrieval, ensuring a high degree of semantic and structural consistency between the context examples and the current input sentence, thus avoiding noise interference from traditional random retrieval. Experiments show that this method significantly improves Vietnamese grammar correction on multiple large model bases. value.
[0033] 2. This invention adds detailed explanations of grammatical errors to the prompts, which is different from the traditional black-box mode that only provides error-correction pairs. This not only tells the model how to correct it, but also explains why it is wrong through natural language, thereby effectively activating the logical reasoning ability of the large model and making it perform better when faced with complex Vietnamese grammatical phenomena.
[0034] 3. This invention does not rely on full-scale fine-tuning of a large model, but rather utilizes a retrieval-enhanced generation technique to fully exploit the potential of existing small amounts of high-quality labeled data. By constructing an interpretation library and performing retrieval matching, it achieves efficient syntax correction at a relatively low computational cost, demonstrating high practical value and promising prospects for wider application.
[0035] 4. Compared with methods that rely solely on literal matching, the vector-based semantic retrieval of this invention can capture deep semantic features. Even if the input sentence and the example are not exactly the same in terms of word usage, as long as the grammatical structure or semantic logic is similar, the model can still perform effective transfer, thereby improving the robustness of the system when dealing with diverse user inputs. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the process in this invention. Detailed Implementation
[0037] Example 1: As Figure 1 As shown, a Vietnamese grammar correction method based on grammatical error interpretation enhancement is described, the method comprising:
[0038] Step 1: Construct a Vietnamese grammar error explanation dataset. Using annotation tools, fine-grained annotations are applied to the existing Vietnamese grammar error correction dataset to extract error types and correction words. Then, using prompt templates that incorporate Vietnamese language characteristics, a large language model is guided to generate detailed natural language explanations for each sentence pair, including error type, error cause, and correction logic. The original "source sentence-target sentence" binary data is reconstructed into a high-quality explanation dataset containing "source sentence-target sentence-error explanation" triples, providing a rich semantic knowledge base for subsequent retrieval enhancement.
[0039] Further, Step 1 includes:
[0040] Step 1.1: Data Preparation and Preprocessing. The training set of the ViSGE dataset, a high-quality Vietnamese grammar correction dataset with manual annotations, was selected as the basic corpus. Let this dataset be denoted as... ,in The source sentence contains grammatical errors. This corresponds to the correct target sentence. The ViSGEC dataset covers various types of common spelling, word order, and word usage errors in Vietnamese;
[0041] Step 1.2: Fine-grained Error Labeling To generate more accurate explanations for the large model, it's essential to first identify the specific errors in each sentence. This invention uses a modified ChERRANT scorer as the labeling tool. The original ChERRANT is based on character-level spans, which is unsuitable for Vietnamese, which is word-based. This invention modifies it to support word-level span editing and alignment. The tool automatically compares the source sentences. and target sentence Extract the first Error set in each sentence pair ,in:
[0042] This indicates the total number of errors in the sentence;
[0043] Indicates the first Error types (e.g., M - missing, R - redundant, S - substitution, W - word order error).
[0044] Indicates the corresponding correction term or operation;
[0045] Obtain the labeled dataset ;
[0046] Step 1.3: Construct an explanation generation prompt template tailored to the characteristics of the Vietnamese language. This template guides the generation of explanations from large models and includes:
[0047] Character setting: "Imagine you are a Vietnamese grammar expert";
[0048] Task Description: The model is required to analyze the errors present in the provided source sentence, target sentence, and labeled error types.
[0049] Content requirements: The explanation must cover the type of error (e.g., clearly indicating whether it is a collocation error or a missing element), the reason for the error (which Vietnamese grammar rule was violated), and the logic of the correction (why the modification of the target sentence is correct);
[0050] Step 1.4: Generate Explanation Dataset. Fill the annotated data into the prompt template, and use the DeepSeek-V3 large language model with powerful reasoning capabilities to generate natural language explanations for each sentence pair. The formula is expressed as:
[0051]
[0052] Finally, the number of explained numbers dataset was constructed. = ;
[0053] Step 2: Establish a context example retrieval mechanism based on semantic matching. A pre-trained vector encoding model is used to vectorize the source sentences in the explanation dataset, constructing a semantic index. For the input sentence to be corrected, the cosine similarity between the sentence and the vectors in the index is calculated to retrieve several examples whose semantics and syntactic structure are closest as contextual references.
[0054] Furthermore, Step 2 includes:
[0055] Step 2.1: Construct a semantic vector index library. The Sentence Transformers model, which supports multiple languages, is selected as the vector encoder. The dataset will be explained. All source sentences (i.e., sentences containing errors) Mapped to a fixed-dimensional semantic vector Building a semantic vector database i=1, ;
[0056] Step 2.2: Vectorization of the Sentence to be Corrected. For any input sentence (Query) to be corrected... They compute their embedding vectors using the same vector encoder. ;
[0057] Step 2.3, Semantic Similarity Retrieval Calculation With vector library Calculate the cosine similarity of all vectors in the dataset and sort them from highest to lowest similarity score. Retrieve the vector with the highest similarity score. Examples (in this embodiment) ), forming a search example set ;
[0058] Step 3: Integrate with the large model for grammatical error correction. Similar examples containing detailed error explanations are retrieved and combined with the sentence to be corrected to form a specific formatted prompt, which is then input into the large language model. The model utilizes contextual learning capabilities to analyze the input sentence and generate the final correction result by comparing the error correction logic and explanation information in the examples.
[0059] Furthermore, Step 3 includes:
[0060] Step 3.1: Construct an error correction prompt template that combines contextual retrieval enhancement. The template structure includes: a task instruction area, a retrieval example display area, and an input area to be corrected.
[0061] Step 3.2: In the search example display area, format and enter the search results. Each example contains a "src" (source sentence), a "tgt" (target sentence), and an "Explanation" (a brief description of the syntax error and the reason for the correction), explicitly showing the model the reason for the error and the correction logic;
[0062] Step 3.3: Large-Scale Model Inference and Output. The assembled complete prompts are input into the multilingual large-scale model (e.g., LLaMA3.1-8B, Qwen3-8B, Qwen3-14B). Based on the examples in the prompts, the model learns through analogy how errors in the examples are identified, explained, and corrected, and then uses this knowledge to understand the sentence to be corrected. Analyze the data and generate the corrected sentence. The formula is expressed as:
[0063]
[0064] To verify the feasibility of this solution, experiments were conducted on the publicly available ViSGEC test set, and precision (P), recall (R), and... The value is used as an evaluation indicator.
[0065] The specific formula for calculating precision (P) is shown below:
[0066]
[0067] Among them, TP (True Positive): The modifications made by the model are consistent with the correct modifications made by the human annotation (i.e., the model made the correct modifications).
[0068] FP (False Positive): The model made a modification, but the modification was incorrect (e.g., changing a correct word to a wrong word, or changing it to an inappropriate word).
[0069] The specific formula for calculating recall (R) is shown below:
[0070]
[0071] Among them, FN (False Negative): There is an error in the original text, but the model did not identify it or did not correct it (i.e. the model missed the correction).
[0072] The specific formula for calculating the value is as follows:
[0073]
[0074] 2) In this invention, the training set was used to generate the explanatory dataset, the test set was selected for evaluating model performance, and the validation set was used to analyze the model's generalization ability. As shown in Table 1:
[0075] Table 1. Statistics of Vietnamese grammar correction dataset
[0076] Table 2 shows the performance of three multilingual large language models in Vietnamese grammar correction tasks under different retrieval methods (random retrieval, text similarity retrieval, and BM25 retrieval), based on dataset testing. The large models effectively enhance the model's ability to identify and correct Vietnamese grammatical errors by providing high-quality grammatical reference information using the semantic retrieval-based contextual example method.
[0077] Table 2. Experimental results of various major models using different retrieval methods.
[0078] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A Vietnamese grammar correction method based on syntax error interpretation enhancement, characterized in that: The method includes the following steps: Step 1: Construction of the Vietnamese Explanation Dataset: The existing Vietnamese grammar error correction dataset is annotated using annotation tools to obtain error types and correction information; then, a large language model is used to generate corresponding natural language explanations based on prompt templates, constructing an explanation dataset containing source sentences, target sentences, and detailed error explanations; Step 2, Contextual Example Retrieval: A pre-trained vector encoding model is used to vectorize the source sentences in the explanation dataset, constructing a semantic index. For the input sentence to be corrected, the cosine similarity between the sentence and the vectors in the index is calculated to retrieve several examples with the closest semantic and syntactic structure as contextual references. Step 3: Integrate into the large model for grammatical error correction: Combine the retrieved examples containing erroneous explanations with the current sentence to be corrected to form a prompt input, guide the large language model to learn the context, analyze the input sentence by analogy with the error correction logic and explanation information in the examples, and generate the final correction result.
2. A Vietnamese grammar correction method based on syntax error interpretation enhancement according to claim 1, characterized in that: Step 1 includes: Step 1.1, select the Vietnamese grammar correction data set ViSGEC annotated by artificial as the basic corpus, denoted as data set ; Step 1.
2. Perform word-level annotation on the sentence pairs in the dataset using the modified ChERRANT annotation tool, and extract the error type set and correction words for each sentence pair , where is the error type, is the correction word, is the number of errors in the th sentence pair. Step 1.3, Constructing the interpretation generation prompt template for the language characteristics of Vietnamese The template contains task role setting, content requirements and output format; Step 1.4, fill in the prompt template with the annotated data, and use the DeepSeek-V3 large language model to generate the syntax error explanation for each sentence pair The formula is as follows: ; The final formation comprises An explanatory dataset .
3. The Vietnamese grammar correction method based on syntax error interpretation enhancement of claim 1, characterized in that: Step 2 includes: Step 2.1: Select the Sentence Transformers model, which supports multiple languages, as the vector encoder; map all source sentences in the explanation dataset into fixed-dimensional semantic vectors to build a semantic vector database; Step 2.2: For any input sentence to be corrected, calculate its embedding vector using the same vector encoder; Step 2.3, calculate the cosine similarity of the semantic vector of the input sentence to be corrected and all vectors in the vector library, and sort them from high to low according to the similarity score; retrieve the highest similarity In one example, a set of retrieval examples is formed.
4. The Vietnamese grammar correction method based on syntax error interpretation enhancement of claim 1, characterized in that: Step 3 includes: Step 3.1: Construct an error correction prompt template that combines contextual retrieval enhancement. The template structure includes: a task instruction area, a retrieval example display area, and an input area to be corrected. Step 3.2: In the search example display area, format and enter the search results. Each example contains "src" representing the source sentence, "tgt" representing the target sentence, and "Explanation" representing a brief description of the syntax error and the reason for correction, explicitly showing the model the reason for the error and the correction logic; Step 3.2, Large Model Inference and Output: The assembled complete prompts are input into the multilingual large model. Based on the examples in the prompts, the model learns by analogy to understand how errors in the examples are identified, explained, and corrected. Then, it analyzes the sentence to be corrected and generates the corrected sentence.
5. A Vietnamese grammar correction system based on grammatical error interpretation enhancement, characterized in that, The system includes a module for performing the Vietnamese grammar correction method based on grammar error interpretation enhancement as described in any one of claims 1 to 4.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the Vietnamese grammar correction method based on grammar error interpretation enhancement as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the Vietnamese grammar correction method based on grammar error interpretation enhancement as described in any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the Vietnamese grammar correction method based on grammar error interpretation enhancement as described in any one of claims 1 to 4.