The invention relates to the technical field of power dispatching, and discloses a knowledge
distillation-based low-resource power big language big model training method,
system and device and a medium, and the method comprises the steps: obtaining accident case data of the power industry, carrying out the problem construction and task setting, introducing a quality evaluation mechanism, carrying out the refusal sampling through a
language model, and carrying out the training of a low-resource power big language big model. Generating a
distillation data set for model
distillation; introducing a LoRA module into the student model for
fine tuning, constructing multi-source heterogeneous
fine tuning data, and setting a training strategy to optimize the performance of the model; and performing training by adopting
reinforcement learning, introducing language consistency rewards until the
reinforcement learning achieves convergence on the reasoning task, and generating a final language
large model. According to the method, through knowledge distillation and
reinforcement learning, the
deep knowledge and the reasoning ability of the super-
large model are successfully migrated to the small model, so that the parameter quantity of the finally deployed model is greatly reduced, and the computing power resource required by reasoning is sharply reduced.