The invention relates to a Vietnamese grammar error correction corpus construction method based on error type
perception of a
large model, and belongs to the field of
natural language processing. According to the method, firstly, a voice recognition model is used for simulating Vietnamese grammar errors in a real scene, a preliminary error correction
data set is generated, then through deep analysis of distribution rules and grammar structure features of typical errors in the
data set, a chain thinking cue (CoT) mechanism fusing error type features is designed in a targeted mode, and the error type features of the Vietnamese grammar errors are extracted. Guiding a large
language model (LLM) to generate synthetic statements containing predetermined grammar errors in batches; thirdly, in order to enhance corpus quality, synchronously implementing a
web crawler to collect a native Vietnamese text, and constructing a pure monolingual corpus through multi-layer filtering and cleaning; and finally, the generated
synthetic data needs to be strictly verified and processed to ensure that the error type is consistent with a preset target, and a pre-training model normal form and a
large model normal form are strengthened in a two-stage
fine tuning manner, so that the generalization ability of a grammar
error correction model is effectively improved, and the problem that Vietnamese grammar error correction corpus is deficient is solved.