The application provides a grammar error correction
large model training method and device based on
reinforcement learning, equipment and medium, relating to the technical field of
data processing. The method comprises the following steps: obtaining an initial grammar correction corpus containing error sentences and corresponding correct
sentence pairs,
processing the initial grammar correction corpus to generate a reasoning correction
training set, adjusting a first preset
language model according to the reasoning correction
training set to obtain an initial strategy model, performing
reinforcement learning training on the initial strategy model based on a preset
composite function by using a
reinforcement learning algorithm, and finally obtaining a target grammar correction model. The target grammar correction model can fully utilize the reasoning ability, effectively improve the performance of the model in terms of
precision and recall, better meet the actual needs of grammar correction, and generate correction results containing reasoning processes, which provides more transparent and interpretable basis for the correction process.