Adaptive advantage learning method and application thereof
By introducing an adaptive dynamic scaling factor to adjust the weights of the advantage term in reinforcement learning, the contradiction between robustness and convergence speed in existing technologies is resolved, resulting in faster value function convergence and improved stability, thus enhancing the efficiency of reinforcement learning.
CN122114047APending Publication Date: 2026-05-29JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINAN UNIVERSITY
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-29
Smart Images

Figure CN122114047A_ABST
Abstract
The application discloses a self-adaptive advantage learning method, comprising the following steps: S1, introducing a dynamic scaling factor in a target update expression of a value function, and self-adapting weighting of an advantage term; S2, when an action gap is less than or equal to a preset threshold, the dynamic scaling factor tends to 1, the advantage term is reserved to enhance the action gap; S3, when the action gap is greater than the preset threshold, the dynamic scaling factor tends to 0, the advantage term is inhibited or removed, and the target update tends to a Bellman optimal operator. Compared with the prior art, the application has the advantages that a self-adaptive advantage learning method and application thereof are provided, a theoretical optimal balance between action gap enhancement required by robustness and faster convergence speed of a value function is achieved by introducing a self-adaptive dynamic scaling factor, and a technical defect of slow convergence speed of an existing AL operator is overcome.
Need to check novelty before this filing date? Find Prior Art