A method and system for constructing a laos language grammar correction corpus based on error distribution to guide a large model, and an electronic device
By combining speech recognition and data generated by large models, we constructed Lao grammatical correction corpus, which solved the problem of lack of data for Lao grammatical correction models, achieved efficient and accurate grammatical correction effects, and improved the stability and accuracy of the model.
Patent Information
- Application Number
- CN202411848270.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-16
AI Technical Summary
The existing Lao grammatical correction models suffer from data scarcity and insufficient training, resulting in unsatisfactory correction accuracy and efficiency. In addition, existing methods find it difficult to effectively utilize the unlabeled data generated by speech recognition.
By combining the unlabeled data generated by the speech recognition model with the data generated by the large model, data cleaning, screening and fusion are performed to identify common error types and construct high-quality Lao grammatical correction corpus. The Meta-Llama large model is used to generate pseudo data and train it to build a Lao grammatical correction model.
The accuracy and efficiency of the Lao grammatical error correction model have been significantly improved, the dependence on manual annotation has been reduced, and the generalization ability and error correction effect of the model have been improved.
Smart Images

Figure CN119783661B_ABST
Abstract
Description
[0001] Technical Language
[0002] The present invention relates to a method, system and electronic device for constructing Laotian grammatical error correction corpus based on an error distribution guidance large model, and belongs to the technical field of natural language processing. Background Art
[0003] With the increasing globalization of the world, learning and using Lao has become increasingly important. However, Lao learners and users often face grammatical errors, which not only hinder the fluency of communication but also hinder the improvement of learners' language proficiency. Common grammatical errors include inconsistent verb tenses, inconsistent subject and predicate, and improper use of prepositions. These errors are prevalent in learning and using the language. Currently, traditional methods for correcting grammatical errors rely mainly on rule-based grammar checking tools. These tools use preset grammatical rules to detect errors, which can address the problem to a certain extent, but they are difficult to adapt to the complexities of language use. In addition, with the development of artificial intelligence technology, grammar correction methods based on machine learning are gradually gaining popularity. However, existing models still suffer from a lack of data and insufficient training when processing Lao, resulting in unsatisfactory grammar correction accuracy and efficiency.
[0004] Speech recognition technology has also advanced rapidly in recent years. By converting speech into text, it is possible to efficiently generate large-scale corpora. However, the text generated by speech recognition often contains various types of errors and is not manually annotated. While this unannotated data presents significant training challenges, it also provides a rich source of information, offering potential improvements to subsequent error correction models. Large models such as MT5 and MBART have demonstrated outstanding performance in various natural language processing tasks. These models can be trained on large corpora and demonstrate strong capabilities in tasks such as grammatical correction and text generation. However, existing research has not yet explored how to effectively combine unannotated data generated by speech recognition with data generated by large models to construct high-quality grammatical correction corpora for the specific needs of Lao. Therefore, the construction of a Lao grammatical correction corpus guided by error distribution is particularly necessary. By thoroughly analyzing the generated corpus, identifying common error types, and integrating them, we can not only improve the quality of the corpus but also enhance model training effectiveness. Such research will not only promote the education and popularization of the Lao language but also provide new ideas and methods for grammatical correction research in other minority languages. Therefore, exploring how to effectively combine speech recognition models and corpora generated by large models will be of great significance to improving the grammatical error correction effect of Lao. Summary of the Invention
[0005] The present invention provides a method, system, and electronic device for constructing Lao grammar correction corpus based on an error distribution-guided large model to solve the problem of insufficient Lao grammar correction corpus. The present invention obtains relatively high-quality Lao grammar correction corpus by fusing data obtained from speech recognition with data generated by the large model.
[0006] The technical solution of the present invention is: a method for constructing Laotian grammatical error correction corpus based on an error distribution guidance large model, the method comprising:
[0007] Step 1: Collection of unlabeled data. First, we use the existing Lao speech recognition model to automatically generate a large-scale Lao corpus. Then, we clean, filter, and preprocess the corpus. Finally, we generate Lao unlabeled data with different error distributions.
[0008] Step 2: First, relatively standard texts are selected from the basic corpus as the original correct sentences. Then, based on certain rules and constraints, a large model is used to automatically generate sentences with various error distributions to construct pseudo data. Finally, the generated pseudo data is cleaned and deduplicated to form the initial grammatically corrected pseudo data.
[0009] Step 3: Further improve Lao grammatical error correction performance by fusing data from all error distributions. First, analyze the error distribution of Lao unlabeled data and grammatical error correction pseudo data to identify common error types. Then fuse the unlabeled data and pseudo data containing different error types to form a corpus with different error distributions.
[0010] Step 4: Use the fused corpus to train the language model and build a Lao grammatical correction model for Lao grammatical correction; evaluate the correction prediction results to obtain the final Lao grammatical correction corpus.
[0011] As a further solution of the present invention, step 1 includes the following:
[0012] Step 1.1: Generate unlabeled data with various error distributions using the existing Lao speech recognition model.
[0013] Step 1.2: Divide the unlabeled data into independent sentences per line;
[0014] Step 1.3: Screen the sentences and remove those that are too short or too long.
[0015] Step 1.4: Preprocess the filtered sentences to remove sentences containing unrecognizable special symbols; the processed sentences are used as Lao unlabeled data with different error distributions.
[0016] As a further solution of the present invention, step 2 includes the following:
[0017] Step 2.1: Select sentences from the Lao unlabeled data base corpus as the original correct sentences;
[0018] Step 2.2: Use the Meta-Llama model to generate corresponding incorrect sentences with different error distributions based on correct sentences based on certain rules and constraints.
[0019] Step 2.3: Clean and preprocess the generated erroneous sentences, remove sentences containing impurities, and form initial grammatical error correction pseudo data.
[0020] As a further solution of the present invention, step 3 includes the following:
[0021] Step 3.1: Analyze the error distribution of Lao unlabeled data and pseudo-grammatical correction data to identify common error types, including missing subjects, improper word order, and repeated words.
[0022] Step 3.2: Fuse the Lao unlabeled data containing different error types with the grammatical correction pseudo data to form a fused corpus containing a variety of error distributions.
[0023] As a further solution of the present invention, step 4 includes the following:
[0024] Step 4.1: Use the fused corpus containing various error distributions to train the mbart language model and build a Lao grammar error correction model;
[0025] Step 4.2: Use word error rate to evaluate the prediction effect of the Lao grammatical correction model to obtain high-quality Lao grammatical correction corpus.
[0026] The present invention also provides a Laotian grammar correction corpus construction system based on the error distribution guidance large model, including: a module for executing the above-mentioned Laotian grammar correction corpus construction method based on the error distribution guidance large model.
[0027] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for constructing Laotian grammatical correction corpus based on the error distribution guidance large model is implemented.
[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for constructing a Laotian grammatical error correction corpus based on an error distribution guidance large model is implemented.
[0029] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for constructing Laotian grammatical error correction corpus based on the error distribution guidance large model.
[0030] The beneficial effects of the present invention are:
[0031] 1. By combining the unlabeled data generated by the speech recognition model with the data generated by the large model, the present invention can effectively capture and analyze common grammatical errors in Lao and significantly improve the accuracy of the error correction model.
[0032] 2. By effectively utilizing unlabeled data, the present invention reduces dependence on manual labeling, reduces the time and cost of data processing, and improves research efficiency.
[0033] 3. The present invention improves the generalization ability of the model by integrating multiple data sources, so that the error correction model can show better stability and accuracy when facing different users and texts.
[0034] 4. The present invention can construct a high-quality Lao grammatical error correction corpus by conducting in-depth analysis and cleaning of the generated corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flow chart of the present invention; DETAILED DESCRIPTION
[0036] Example 1: Figure 1 As shown, a method for constructing Laotian grammatical correction corpus based on an error distribution guidance large model, the method comprising:
[0037] Step 1: Obtain unlabeled Laotian data with different error distributions. The specific steps are as follows:
[0038] Step 1.1: Generate unlabeled data with various error distributions using the existing Lao speech recognition model.
[0039] Step 1.2: Divide the unlabeled data into independent sentences per line;
[0040] Step 1.3: Screen the sentences and remove those that are too short or too long.
[0041] Step 1.4: Preprocess the filtered sentences to remove sentences containing unrecognizable special symbols; the processed sentences are used as Lao unlabeled data with different error distributions.
[0042] Step 2: Select relatively standardized text from the unlabeled data as the original correct sentences. Based on certain rules and constraints, use the large model to automatically generate sentences with various error distributions to construct pseudo data. The generated pseudo data is cleaned and deduplicated to obtain the initial grammatically corrected pseudo data. Step 2 includes the following:
[0043] Step 2.1: Select sentences from the Lao unlabeled data base corpus as the original correct sentences;
[0044] Step 2.2: Use the Meta-Llama model to generate corresponding incorrect sentences with different error distributions based on correct sentences based on certain rules and constraints.
[0045] Step 2.3: Clean and preprocess the generated erroneous sentences, remove sentences containing impurities, and form initial grammatical error correction pseudo data.
[0046] Step 3: Fusing Lao unlabeled data with pseudo-grammatical correction data to further improve Lao grammatical correction performance; Step 3 includes the following:
[0047] Step 3.1: Analyze the error distribution of Lao unlabeled data and pseudo-grammatical correction data to identify common error types, such as missing subjects, improper word order, and repeated words.
[0048] Step 3.2: Fuse the Lao unlabeled data containing different error types with the grammatical correction pseudo data to form a fused corpus containing a variety of error distributions.
[0049] Step 4: Use the fused corpus to train the language model and construct a Lao grammar correction model for Lao grammar correction; evaluate the correction prediction results to obtain the final Lao grammar correction corpus. Step 4 includes the following:
[0050] Step 4.1: Use the fused corpus containing various error distributions to train the mbart language model and build a Lao grammar error correction model;
[0051] Step 4.2: Use word error rate to evaluate the prediction effect of the Lao grammatical correction model to obtain high-quality Lao grammatical correction corpus.
[0052] The present invention also provides a Laotian grammar correction corpus construction system based on the error distribution guidance large model, comprising:
[0053] The first acquisition module is used to obtain Lao unlabeled data containing different error distributions;
[0054] The second acquisition module is used to obtain grammatical error correction pseudo data;
[0055] The fusion module is used to fuse Lao unlabeled data with grammatical correction pseudo data;
[0056] The training and evaluation module is used to train the language model using the fused corpus, build a Lao grammatical correction model for Lao grammatical correction, and evaluate the correction prediction results to obtain the final Lao grammatical correction corpus.
[0057] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for constructing Laotian grammatical correction corpus based on the error distribution guidance large model is implemented.
[0058] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for constructing a Laotian grammatical error correction corpus based on an error distribution guidance large model is implemented.
[0059] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for constructing Laotian grammatical error correction corpus based on the error distribution guidance large model.
[0060] 1) This paper uses a fused corpus containing a variety of error distributions as experimental data, and uses the word error rate (WER) as an evaluation metric for the prediction performance of the Laotian grammar correction model. WER is a commonly used indicator of text error correction system performance; lower WER indicates better error correction performance. The specific calculation method is as follows:
[0061]
[0062] in:
[0063] S(Substitutions): The number of incorrectly recognized words.
[0064] D (Deletions): The number of omitted words.
[0065] I (Insertions): The number of extra words.
[0066] N (Number): The total number of words in the reference text.
[0067] 2) In order to verify the effect of the proposed method on improving the performance of the Lao grammatical error correction model, the present invention uses the official mbart language model as the benchmark model. The corpus sources and corpus sizes involved are shown in Table 1:
[0068] Table 1: Corpus sources and corpus size involved in the verification part
[0069]
[0070] The final experimental results of the proposed model on the fusion corpus are shown in Table 2 below. First, it can be seen that the increase in the original corpus can enhance the error correction performance of the baseline model and reduce the word error rate. Second, it is found that the fusion of different error corpora can effectively improve the generalization performance of the error correction model. This proves that high-quality error correction corpora help the model capture more common knowledge between languages and can fully utilize the language model knowledge, further verifying the effectiveness of this method.
[0071]
[0072]
[0073] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A method for constructing Laotian grammatical correction corpus based on an error distribution guidance model, characterized by: The method comprises: Step 1: Obtain unlabeled Laotian data with different error distributions; Step 2: Obtain grammatical error correction pseudo data; Step 3: Fuse the Lao unlabeled data and the grammatical correction pseudo data; Step 4: Use the fused corpus to train the language model and build a Lao grammatical correction model for Lao grammatical correction. Evaluate the correction prediction results to obtain the final Lao grammatical correction corpus. The step 1 includes the following: Step 1.1: Generate unlabeled data with various error distributions using the existing Lao speech recognition model. Step 1.2: Divide the unlabeled data into independent sentences per line; Step 1.3: Screen the sentences and remove those that are too short or too long. Step 1.4: Preprocess the filtered sentences to remove sentences containing unrecognizable special symbols; the processed sentences are used as Lao unlabeled data with different error distributions; The step 2 includes the following: Step 2.1: Select sentences from the Lao unlabeled data base corpus as the original correct sentences; Step 2.2: Use the Meta-Llama model to generate corresponding incorrect sentences with different error distributions based on correct sentences based on certain rules and constraints. Step 2.3: Clean and preprocess the generated erroneous sentences to remove sentences containing impurities and form initial grammatical error correction pseudo data; The step 3 includes the following: Step 3.1: Analyze the error distribution of Lao unlabeled data and pseudo-grammatical correction data to identify common error types, including missing subjects, improper word order, and repeated words. Step 3.2: Fuse the Lao unlabeled data containing different error types with the grammatical correction pseudo data to form a fused corpus containing a variety of error distributions; The step 4 includes the following: Step 4.1: Use the fused corpus containing various error distributions to train the mbart language model and build a Lao grammar error correction model; Step 4.2: Use word error rate to evaluate the prediction effect of the Lao grammatical correction model to obtain high-quality Lao grammatical correction corpus.
2. A Laotian grammar correction corpus construction system based on an error distribution guidance model, characterized in that: include: A module for executing the method for constructing Laotian grammatical correction corpus based on an error distribution guidance large model as described in claim 1.
3. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for constructing Laotian grammatical correction corpus based on the error distribution guidance large model as described in claim 1 is implemented.
4. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing Laotian grammatical error correction corpus based on an error distribution guidance large model as described in claim 1 is implemented.
5. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing Laotian grammatical error correction corpus based on an error distribution guidance large model as described in claim 1 is implemented.
Citation Information
Patent Citations
Text error correction data generation method and related device
CN111048065A
Text error correction method, device and equipment based on large model and storage medium
CN118520869A