The invention provides a grouping relative strategy optimization method with stable dynamic reference, which comprises the following steps of: performing multiple independent sampling on the same input prompt by a current strategy model, generating a candidate answer set, calculating an absolute
quality score of each candidate answer through a preset determinacy
evaluation function, and calculating a
quality score of each candidate answer; calculating an average
score of all candidate answers in the current group; dynamically updating a global historical reference value according to the
score of the current batch of candidate answers through a dynamic global historical
reference line maintenance module; through a composite
advantage calculation module, in-group relative comparison and global historical comparison are fused, and a composite
advantage value of each candidate answer is calculated; and through a strategy gradient updating module, the logarithmic probability gradient of the strategy model is calculated and the parameters of the strategy model are updated by taking the composite
advantage value as a weighted weight, so that the generation probability of answers with high advantage values is enhanced, and the generation probability of answers with low advantage values is inhibited. The method has the beneficial effect that the answer which is absolutely progressive compared with the historical performance of the model can be accurately identified.