Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2results about How to "Efficient feedback" patented technology

An instant reward learning method based on self-supervised reinforcement learning

This invention discloses an immediate reward learning method based on self-supervised reinforcement learning, belonging to the field of artificial intelligence reinforcement learning technology. First, the invention calculates a first prediction error between the agent's predicted state and the actual state. Second, it reconstructs action data using an inverse dynamics model and calculates a second prediction error. Then, it constructs a causal confidence factor based on the second prediction error to generate an effective prediction error. Next, it generates a fast-channel reward and calculates the volatility of the reward signal using an oscillation index. Finally, it generates the final immediate reward. By integrating a multi-layered mechanism of prediction error correction, causal relationship modeling, reward stability assessment, and policy stochastic adjustment, the stability and reliability of the reward signal can be effectively enhanced, noise interference reduced, and training oscillations avoided. Simultaneously, it significantly improves the quality of immediate reward generation in sparse reward environments, accelerates policy convergence, and enhances the efficiency and stability of self-supervised reinforcement learning in complex tasks.
Owner:MINZU UNIVERSITY OF CHINA

Steel enterprise product adjustment effect evaluation system, method, device and medium

PendingCN122198777AComprehensive quantitative assessmentobjective quantitative assessmentData processing applicationsBiological models
The application provides a steel enterprise product adjustment effect evaluation system, method, equipment and medium, and relates to the technical field of data processing, and comprises the following steps: a data preprocessing module is used for preprocessing original transaction data corresponding to an evaluated object to obtain target transaction data, wherein the original transaction data comprises steel order data and steel gross profit data corresponding to a plurality of steel products sold by the evaluated object; a product adjustment effect evaluation module is used for constructing a product adjustment evaluation matrix corresponding to the evaluated object based on initial index values corresponding to multidimensional index items, so as to determine a product adjustment evaluation index corresponding to the evaluated object, and elements of the product adjustment evaluation matrix are used to describe target index values of the evaluated object relative to the multidimensional index items; and an evaluation visualization module is used for visualizing the product adjustment evaluation index. The application can more comprehensively and objectively quantify the product adjustment effect, thereby providing effective feedback and guidance for product adjustment decision-making.
Owner:ANGANG DIGITAL TECHNOLOGY (LIAONING) CO LTD