The invention discloses a
database question and answer model training method and device, a storage medium and
computer equipment, and the method comprises the steps: associating standard structured
query language statements, standard execution result answers and standard
natural language questions, and generating training
annotation data; and collecting
simulation derivation problems possibly proposed for the
database to obtain non-
labeled data for training. Based on a GRPO
reinforcement learning framework and a scoring reward function provided by a double-
tower model, training is carried out on the scoring reward function by utilizing training labeling data, supervised
fine tuning training is carried out on a
database question and answer model, and non-labeling data for training, format rewards,
executable rewards and scoring rewards of the scoring reward function are combined, so that the scoring reward function of the database question and answer model is obtained. And continuing to
train the database question and answer model after supervised
fine tuning training. Preliminary training is carried out through a small amount of
annotation data, then subsequent training is carried out through non-
annotation data, the reasoning ability of the model can be stimulated, the annotation cost is reduced, and the training efficiency is improved.