The invention discloses an
interpretability-based ability fusion and
recovery method after large
language model training, and belongs to the technical field of
artificial intelligence, and the method comprises the steps: carrying out the
interpretability analysis of a plurality of candidate large language models, and forming an ability missing sample
library; counting the missing degree of each model on each capability dimension, and generating a capability
feature vector; selecting at least one complementary model which is complementary with the basic model in capability and a
negative sample model according to a target scene requirement; performing parameter fusion on the basic model, the complementary model and the
negative sample model by adopting a geometric interpolation
algorithm, wherein the weight of the
negative sample model is lower than that of the complementary model; and carrying out standardized evaluation on the fused model, when an
evaluation result shows that a certain capability dimension is degraded, positioning a degraded sample according to
interpretability analysis, updating a capability
feature vector, and carrying out iterative optimization until the comprehensive capability meets the requirement. According to the method, multi-capability directional enhancement and overall performance balance are realized, and capability collapse caused by post-training of a large
language model is avoided.