Adaptive advantage learning method and application thereof

By introducing an adaptive dynamic scaling factor to adjust the weights of the advantage term in reinforcement learning, the contradiction between robustness and convergence speed in existing technologies is resolved, resulting in faster value function convergence and improved stability, thus enhancing the efficiency of reinforcement learning.

CN122114047APending Publication Date: 2026-05-29JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN UNIVERSITY
Filing Date
2026-01-09
Publication Date
2026-05-29

Smart Images

  • Figure CN122114047A_ABST
    Figure CN122114047A_ABST
Patent Text Reader

Abstract

The application discloses a self-adaptive advantage learning method, comprising the following steps: S1, introducing a dynamic scaling factor in a target update expression of a value function, and self-adapting weighting of an advantage term; S2, when an action gap is less than or equal to a preset threshold, the dynamic scaling factor tends to 1, the advantage term is reserved to enhance the action gap; S3, when the action gap is greater than the preset threshold, the dynamic scaling factor tends to 0, the advantage term is inhibited or removed, and the target update tends to a Bellman optimal operator. Compared with the prior art, the application has the advantages that a self-adaptive advantage learning method and application thereof are provided, a theoretical optimal balance between action gap enhancement required by robustness and faster convergence speed of a value function is achieved by introducing a self-adaptive dynamic scaling factor, and a technical defect of slow convergence speed of an existing AL operator is overcome.
Need to check novelty before this filing date? Find Prior Art