Lao grammar error correction training data construction method and system based on iterative optimization

Through iterative optimization methods, combined with pseudo-data generation and feedback loop mechanisms, the problem of lack of corpus in Lao grammar correction was solved, and the model performance was significantly improved, providing new ideas for the research on natural language processing of low-resource languages.

CN120181075APending Publication Date: 2025-06-20KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347039.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The lack of high-quality labeling data and adaptation models in the Lao language grammar error correction research has led to limited application of deep learning models in low-resource language grammar error correction.

Method used

Using an iterative optimization method, through the pseudo-data generation and feedback loop mechanism, the data construction strategy is dynamically adjusted to generate training data that is closer to the actual error distribution.

Benefits of technology

It significantly improves the performance and practical application effect of the Lao grammar error correction model, solves the problem of corpus scarcity, and provides new solutions for the research on natural language processing in other low-resource languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181075A_ABST
    Figure CN120181075A_ABST
Patent Text Reader

Abstract

The invention relates to a Lao grammar error correction training data construction method and system based on iterative optimization, and belongs to the field of natural language processing. Comprising the following steps: performing initial grammar error correction prediction on an initial corpus by utilizing a pre-training language model, performing statistical analysis on residual errors in an initial prediction result, automatically generating a new sentence covering a specific error type by utilizing a rule method or a large model based on common error types and distributed statistical data, and performing initial grammar error correction prediction on the new sentence. The expansion module is used for expanding a grammar error correction training data set; fusing the expanded data with the original corpus and carrying out quality evaluation; and retraining the pre-training language model by using the training data set after quality evaluation, further optimizing the error correction pre-training language model, and screening out a high-quality Lao grammar error correction training data set covering various error distributions. Lao grammar error correction data covering wide error distribution is dynamically generated, the performance of a grammar error correction model is effectively improved, and the problem that Lao grammar error correction training data is scarce is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for constructing Lao grammar error correction training data based on iterative optimization, belonging to the technical field of natural language processing. Background Art

[0002] In recent years, with the rapid development of natural language processing technology, grammar error correction has gradually become an important research direction. The grammar error correction task aims to identify and correct grammar errors in text, which is of great significance for improving the accuracy of language learning, machine translation, and human-computer interaction systems. However, most current research mainly focuses on high-resource languages such as English and Chinese, while the research on grammar error correction for low-resource languages such as Lao lags far behind. This lag is mainly reflected in the lack of high-quality labeled data and the ability to adapt models, thus restricting the application of deep learning models in grammar error correction for low-resource languages.

[0003] As an isolating language, Lao has a flexible grammar structure, lacks strict inflection rules, and uses a unique Lao writing system. These characteristics pose unique challenges in the grammar error correction task. For example, the incorrect or missing tone marks can significantly affect the meaning of a sentence, and the subtle differences between characters may also lead to completely different interpretations. In addition, since grammar errors in Lao are usually caused by multiple factors such as spelling mistakes, word order confusion, or character confusion, existing general grammar error correction techniques are difficult to directly transfer and apply to Lao.

[0004] To make up for the shortage of labeled data, pseudo-data generation has become a common strategy in grammar error correction research for low-resource languages. By using existing correct sentences and combining rules or language models to generate sentences with grammar errors, the scale of training data can be quickly expanded. However, traditional pseudo-data generation methods have two main problems. First, the generated data is often limited to specific error types and is difficult to comprehensively cover the diverse actual error distributions in Lao; second, the generated pseudo-data is inconsistent with the error distributions in real scenarios, resulting in limited generalization performance after model training and difficulty in adapting to complex actual application environments.

[0005] The mechanism based on the feedback loop has gradually become an important method in low-resource language processing. Its core lies in using the errors that the model fails to correct during the prediction process to optimize the training data. Through continuous iteration, this method dynamically adjusts the data construction strategy to make the generated data closer to the actual error distribution. In the grammar error correction task, the feedback loop can not only uncover the weaknesses of the model but also effectively expand the coverage of the training data, thereby improving the error correction ability and actual application effect of the model.

[0006] In response to the specific need for Lao grammar error correction, the present invention proposes a method for constructing training data that combines pseudo-data generation and a feedback loop mechanism. By using a large amount of unannotated Lao data to generate initial pseudo-data, and analyzing the types of errors that cannot be corrected during the model prediction stage, this error information is used as a reference to further optimize the pseudo-data generation process. Such a feedback loop can dynamically capture common error types in actual scenarios and gradually improve the quality and diversity of the training data.

[0007] The application of this method not only solves the problem of the lack of Lao grammar error correction corpus, but also significantly improves the model's ability to handle complex grammar errors. At the same time, with the development of pre-trained language model technology, the application potential of this method in other low-resource languages is also very broad, providing a new solution idea for natural language processing research on low-resource languages. Summary of the Invention

[0008] The technical problem solved by the present invention is: The present invention provides a method and system for constructing Lao grammar error correction training data based on iterative optimization, which is used to solve the problem of the lack of Lao grammar error correction corpus. The present invention adopts an iterative optimization method to construct a high-quality Lao grammar error correction corpus.

[0009] The technical solution of the present invention is: A method for constructing Lao grammar error correction training data based on iterative optimization, the method includes:

[0010] Step1, Data collection and initial prediction: Collect Lao sentences containing grammar errors as the initial corpus, and construct a basic dataset for Lao grammar error correction; Use a pre-trained language model to perform initial grammar error correction prediction on the initial corpus, and generate the initial prediction results of the pre-trained language model;

[0011] Step2, Error distribution statistics and analysis, targeted data generation and expansion: Statistically analyze the residual errors in the initial prediction results, and generate statistical data including common error types and distributions; Based on the statistical data of common error types and distributions, use rule methods or large models to automatically generate new sentences covering specific error types for expanding the grammar error correction training dataset; The generated new sentences are processed by cleaning and deduplication;

[0012] Step3, Integrate the data expanded in Step2 with the original corpus, evaluate the quality of the integrated dataset, and obtain the final training dataset;

[0013] Step4, Model iterative training and optimization: Use the final training dataset to retrain the pre-trained language model, evaluate the model performance through a feedback loop, and further optimize the error correction pre-trained language model;

[0014] Step 5. After iterative optimization through the feedback loop, a high-quality Lao grammar error correction training dataset covering various error distributions is screened out.

[0015] Furthermore, the said Step 1 includes:

[0016] Step 1.1. Utilize the existing Lao sentence database containing grammar errors to collect a large-scale initial corpus.

[0017] Step 1.2. Divide the text in the initial corpus into independent sentences to ensure that each line contains only one sentence.

[0018] Step 1.3. Screen the sentence lengths.

[0019] Step 1.4. Preprocess the screened sentences and clean out the sentences containing garbled characters or unrecognizable special symbols.

[0020] Step 1.5. Screen out high-quality sentences from the preprocessed corpus to form a Lao grammar error correction basic dataset for subsequent feedback loop training.

[0021] Step 1.6. For the Lao grammar error correction basic dataset, use the pre-trained language model to make initial grammar error correction predictions.

[0022] Furthermore, the said Step 2 includes:

[0023] Step 2.1. Statistically analyze the residual errors in the initial prediction results to identify common error types and error distributions.

[0024] Step 2.2. According to the identified error distributions, use large models or manually set rules to automatically generate pseudo data containing various grammar error types.

[0025] Step 2.3. De-duplicate the generated pseudo data to remove duplicate and redundant sentences; at the same time, clean out the sentences that do not meet the quality requirements to ensure that the quality of the pseudo data is suitable for training.

[0026] Furthermore, the said Step 3 includes:

[0027] Step 3.1. Integrate the initial corpus in Step 1 and the pseudo data generated in Step 2 according to the characteristics of the error distributions to form a comprehensive dataset containing various error types.

[0028] Step 3.2. Evaluate the quality of the integrated dataset, screen out the sentences with the best effect, and use them as the final training dataset.

[0029] Further, Step 4 includes:

[0030] Step 4.1: Use the final training dataset to preliminarily train the Lao grammar error correction model to form a preliminary version of the error correction model.

[0031] Step 4.2: Conduct predictive evaluation on the error correction model obtained from the preliminary training, collect the error types and difficulties that cannot be corrected in the model output as feedback information.

[0032] Step 4.3: Continuously optimize the Lao grammar error correction model through the feedback information, perform iterative training, and continuously improve the error correction performance and accuracy of the Lao grammar error correction model.

[0033] Further, Step 5 includes:

[0034] Continuously optimize the Lao grammar error correction model through the feedback information, perform iterative training, feedback the optimization results, screen out the optimal Lao grammar error correction dataset, and finally form a Lao grammar error correction training dataset with a wide coverage of error distributions and high quality.

[0035] The present invention also provides a Lao grammar error correction training data construction system based on iterative optimization. The system includes a module for executing the above-mentioned Lao grammar error correction training data construction method based on iterative optimization.

[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the above-mentioned Lao grammar error correction training data construction method based on iterative optimization.

[0037] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned Lao grammar error correction training data construction method based on iterative optimization.

[0038] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-mentioned Lao grammar error correction training data construction method based on iterative optimization.

[0039] The beneficial effects of the present invention are:

[0040] 1. By introducing an iterative optimization mechanism with a feedback loop, the present invention incorporates the uncorrected errors in the model prediction into the optimization process of pseudo-data generation, accurately captures the complex error patterns in Lao grammar, and thus significantly improves the performance of the error correction model.

[0041] 2. The present invention utilizes pseudo - data generation technology to massively expand the corpus resources, effectively reducing the dependence on manual annotation and significantly reducing the time and economic costs in the data construction process;

[0042] 3. The present invention ensures that the generated training data has high quality and high accuracy through the cleaning and optimization of the corpus, providing a solid data foundation for the Lao grammar error - correction model;

[0043] 4. The present invention adopts a dynamic data optimization strategy, which can continuously adjust and improve the pseudo - data generation method, making the training corpus gradually meet the actual application requirements and enhancing the long - term usability and practical value of the grammar error - correction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Embodiment 1: As Figure 1 shown, a method for constructing Lao grammar error - correction training data based on iterative optimization, the method includes:

[0046] Step1. Data collection and initial prediction: Collect Lao sentences containing grammar errors as the initial corpus, and construct a basic data set for Lao grammar error - correction; Use a pre - trained language model (such as mBART) to perform initial grammar error - correction prediction on the initial corpus, and generate the initial prediction results of the pre - trained language model.

[0047] Further, the Step1 includes:

[0048] Step1.1. Utilize an existing Lao sentence database containing grammar errors to collect a large - scale initial corpus, ensuring coverage of common grammar error types;

[0049] Step1.2. Divide the text in the initial corpus into independent sentences, ensuring that each line contains only one sentence;

[0050] Step1.3. Screen the sentence lengths, removing sentences that are too short (e.g., less than 3 words) or too long (e.g., more than 50 words);

[0051] Step1.4. Pre - process the screened sentences, cleaning the sentences containing garbled characters or unrecognizable special symbols;

[0052] Step1.5. Screen out high - quality sentences from the pre - processed corpus to form a basic data set for Lao grammar error - correction for subsequent feedback loop training;

[0053] Step1.6. Perform initial grammar error correction prediction on the Lao grammar error correction basic dataset using a pre-trained language model.

[0054] Step2. Error distribution statistics and analysis, targeted data generation and augmentation: Statistically analyze the residual errors in the initial prediction results to generate statistical data including common error types and distributions; Determine the weaknesses of the model's error correction and the deficiencies in data coverage by analyzing the uncorrected error types; Based on the statistical data of common error types and distributions, use rule-based methods or large models to automatically generate new sentences covering specific error types for augmenting the grammar error correction training dataset; The generated new sentences are subjected to cleaning and deduplication processes.

[0055] Furthermore, Step2 includes:

[0056] Step2.1. Statistically analyze the residual errors in the initial prediction results to identify common error types and error distributions.

[0057] Step2.2. According to the identified error distributions, use large models or manually set rules to automatically generate pseudo-data containing multiple grammar error types to expand the diversity and coverage of the training dataset.

[0058] Step2.3. Deduplicate the generated pseudo-data to remove duplicate and redundant sentences; At the same time, clean up sentences that do not meet the quality requirements to ensure that the quality of the pseudo-data is suitable for training.

[0059] Step3. Integrate the augmented data in Step2 with the original corpus, evaluate the quality of the integrated dataset to obtain the final training dataset.

[0060] Furthermore, Step3 includes:

[0061] Step3.1. Integrate the initial corpus in Step1 and the pseudo-data generated in Step2 according to the characteristics of the error distributions to form a comprehensive dataset containing multiple error types, ensuring that the integrated dataset can cover various common grammar error types and has a relatively balanced error distribution, providing sufficient training data support for subsequent training and error correction tasks.

[0062] Step3.2. Evaluate the quality of the integrated dataset, select the sentences with the best effects, and use them as the final training dataset.

[0063] Step4. Model iterative training and optimization: Retrain the pre-trained language model or fine-tune the small language model using the final training dataset, evaluate the model performance through a feedback loop, and further optimize the error correction pre-trained language model.

[0064] Further, Step 4 includes:

[0065] Step 4.1: Use the final training dataset to preliminarily train the Lao grammar error correction model to form a preliminary version of the error correction model;

[0066] Step 4.2: Conduct prediction evaluation on the error correction model obtained from the preliminary training, collect the error types and difficulties that cannot be corrected in the model output as feedback information;

[0067] Step 4.3: Continuously optimize the Lao grammar error correction model through the feedback information, perform iterative training, and continuously improve the error correction performance and accuracy of the Lao grammar error correction model.

[0068] Step 5: After iterative optimization through the feedback loop, screen out a high-quality Lao grammar error correction training dataset covering various error distributions.

[0069] Further, Step 5 includes:

[0070] Continuously optimize the Lao grammar error correction model through the feedback information, perform iterative training, and feedback the optimization results. Screen out the optimal Lao grammar error correction dataset to ensure that the dataset covers various error distributions and has high quality. Finally, form a Lao grammar error correction training dataset covering a wide range of error distributions and with high quality as the basic dataset for model training and verification.

[0071] The present invention also provides a Lao grammar error correction training data construction system based on iterative optimization. The system includes:

[0072] A data collection and initial prediction module, which is used to collect Lao sentences containing grammar errors as the initial corpus, construct a basic dataset for Lao grammar error correction; use a pre-trained language model to perform initial grammar error correction prediction on the initial corpus and generate the initial prediction results of the pre-trained language model;

[0073] An error distribution statistics and data generation and expansion module, which is used to statistically analyze the remaining errors in the initial prediction results, generate statistical data including common error types and distributions; based on the statistical data of common error types and distributions, use rule methods or large models to automatically generate new sentences covering specific error types for expanding the grammar error correction training dataset; the generated new sentences are subjected to cleaning and deduplication processing;

[0074] A fusion module, which is used to fuse the expanded data with the original corpus, evaluate the quality of the fused dataset, and obtain the final training dataset;

[0075] The model iterative training and optimization module is used to retrain the pre-trained language model using the final training dataset, evaluate the model performance through a feedback loop, and further optimize the error correction pre-trained language model;

[0076] The Lao grammar error correction training dataset screening module is used to screen out high-quality Lao grammar error correction training datasets covering various error distributions after iterative optimization through a feedback loop.

[0077] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the method for constructing Lao grammar error correction training data based on iterative optimization.

[0078] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for constructing Lao grammar error correction training data based on iterative optimization is implemented.

[0079] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for constructing Lao grammar error correction training data based on iterative optimization is implemented.

[0080] The present invention uses a pre-trained language model (such as mBART) to correct and predict Lao sentences containing grammar errors, statistically analyzes the remaining errors in the prediction results to generate error distribution data; then, constructs a new training dataset targeted according to the error distribution data; fuses the correct sentences and the simulated error sentences to cover more typical error types; then, uses the newly constructed data to fine-tune the error correction model; makes predictions again on the fine-tuned model, and evaluates and statistically analyzes the new prediction results. The present invention dynamically generates Lao grammar error correction data covering a wide range of error distributions through an iterative optimization method with a feedback loop, effectively improving the performance of the grammar error correction model and solving the problem of scarce Lao grammar error correction training data at the same time.

[0081] To verify the feasibility of the present solution, the present invention conducts experimental verification:

[0082] 1) The evaluation metrics selected in the experiments of the present invention are CER (Character Error Rate), WER (Word Error Rate), and BLEU (Bilingual Evaluation Understudy).

[0083] The specific calculation method of CER is as follows:

[0084]

[0085] Among them, S represents the number of characters with substitution errors, D represents the number of characters with deletion errors, I represents the number of characters with insertion errors, and N represents the total number of characters in the reference text.

[0086] The specific calculation method of WER is as follows:

[0087]

[0088] Among them, S (Substitutions) represents the number of misrecognized words, D (Deletions) represents the number of missing words; I (Insertions) represents the number of extra words, and N (Number) represents the total number of words in the reference text.

[0089] The specific calculation method of BLEU is as follows:

[0090]

[0091] P n represents the precision of the n-gram, and w n represents the weight of each n-gram (usually uniformly distributed), and BP represents the short sentence penalty factor.

[0092] 2) In the experiment of the present invention, the pre-trained language model mbart is selected as the benchmark model, and the sources and scales of the corpora involved are shown in Table 1:

[0093] Table 1 shows the sources and scales of the corpora involved in the verification part

[0094]

[0095] The final experimental results on the model fusion corpus are shown in Table 2 below. It can be seen that through the iteratively optimized Lao corpus, the generalization effect of the model can be further enhanced and the error correction performance of the model can be improved.

[0096] Table 2 shows the final experimental results on the model fusion corpus

[0097]

[0098] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above implementation manners, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A method for constructing Lao grammar error correction training data based on iterative optimization, characterized in that: The method comprises: Step 1, data collection and initial prediction: collect Lao sentences containing grammatical errors as initial corpus, build a basic data set for Lao grammatical error correction; use the pre-trained language model to perform initial grammatical error correction prediction on the initial corpus, and generate the initial prediction results of the pre-trained language model; Step 2. Error distribution statistics and analysis, targeted data generation and expansion: Statistical analysis is performed on the residual errors in the initial prediction results to generate statistical data containing common error types and distributions. Based on the statistical data of common error types and distributions, new sentences covering specific error types are automatically generated using rule methods or large models to expand the grammar correction training data set. The generated new sentences are cleaned and deduplicated. Step 3: Fuse the expanded data in Step 2 with the original corpus, conduct a quality assessment on the fused data set, and obtain the final training data set; Step 4, iterative model training and optimization: Use the final training data set to retrain the pre-trained language model, evaluate the model performance through a feedback loop, and further optimize the error-correcting pre-trained language model; Step 5. After iterative optimization through the feedback loop, a high-quality Lao grammar correction training dataset covering a variety of error distributions is screened out.

2. The method for constructing Lao grammar error correction training data based on iterative optimization according to claim 1, characterized in that: The Step 1 includes: Step 1.

1. Collect a large-scale initial corpus using an existing Lao sentence database containing grammatical errors; Step 1.2, divide the text in the initial corpus into independent sentences, making sure each line contains only one sentence; Step1.3, filter the sentence length; Step 1.4, pre-process the filtered sentences to remove sentences containing garbled characters or unrecognizable special symbols; Step 1.5, select high-quality sentences from the preprocessed corpus to form a basic dataset for Lao grammar error correction for subsequent feedback loop training; Step 1.6: Use the pre-trained language model to perform initial grammatical correction prediction on the Lao grammatical correction basic dataset.

3. The method for constructing Lao grammar error correction training data based on iterative optimization according to claim 1, characterized in that: The Step 2 includes: Step 2.

1. Count the residual errors in the initial prediction results and identify common error types and error distributions. Step 2.2, based on the identified error distribution, use a large model or manually set rules to automatically generate pseudo data containing multiple types of grammatical errors; Step 2.3: De-duplicate the generated pseudo data to remove repeated and redundant sentences; at the same time, clean up the sentences that do not meet the quality requirements to ensure that the quality of the pseudo data is suitable for training.

4. The method for constructing Lao grammar error correction training data based on iterative optimization according to claim 1, characterized in that: The Step 3 includes: Step 3.1, merge the initial corpus in Step 1 and the pseudo data generated in Step 2 according to the characteristics of error distribution to form a comprehensive data set containing multiple error types; Step 3.2: Perform a quality assessment on the fused dataset, select the sentences with the best effect, and use them as the final training dataset.

5. The method for constructing Lao grammar error correction training data based on iterative optimization according to claim 1, characterized in that: The Step 4 includes: Step 4.

1. Use the final training data set to conduct preliminary training on the Lao grammar error correction model to form a preliminary version of the error correction model. Step 4.2, perform prediction evaluation on the error correction model obtained through preliminary training, and collect the types and difficulties of errors that cannot be corrected in the model output as feedback information; Step 4.

3. Continuously optimize the Lao grammar correction model through feedback information, iterate training, and continuously improve the error correction performance and accuracy of the Lao grammar correction model.

6. The method for constructing Lao grammar error correction training data based on iterative optimization according to claim 1, characterized in that: The Step 5 includes: The Lao grammar correction model is continuously optimized through feedback information, training is iterated, and optimization results are fed back to screen out the optimal Lao grammar correction dataset, ultimately forming a high-quality Lao grammar correction training dataset that covers a wide range of error distributions.

7. A Laotian grammar error correction training data construction system based on iterative optimization, characterized in that: The system comprises: a module for executing the method for constructing Lao grammar error correction training data based on iterative optimization as claimed in any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, a method for constructing Lao grammar error correction training data based on iterative optimization as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing Lao grammar error correction training data based on iterative optimization as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing Lao grammar error correction training data based on iterative optimization as described in any one of claims 1 to 6 is implemented.