This invention discloses an immediate reward learning method based on self-supervised
reinforcement learning, belonging to the field of
artificial intelligence reinforcement learning technology. First, the invention calculates a first prediction error between the agent's predicted state and the actual state. Second, it reconstructs action data using an
inverse dynamics model and calculates a second prediction error. Then, it constructs a causal
confidence factor based on the second prediction error to generate an effective prediction error. Next, it generates a fast-channel reward and calculates the volatility of the reward
signal using an oscillation index. Finally, it generates the final immediate reward. By integrating a multi-layered mechanism of prediction error correction, causal relationship modeling, reward
stability assessment, and policy stochastic adjustment, the stability and reliability of the reward
signal can be effectively enhanced,
noise interference reduced, and training oscillations avoided. Simultaneously, it significantly improves the quality of immediate reward generation in sparse reward environments, accelerates policy convergence, and enhances the efficiency and stability of self-supervised
reinforcement learning in complex tasks.