Reinforcement Learning Value Function Coefficient Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional reinforcement learning techniques face challenges in accurately determining input values for controlled objects when the state is not directly observable and the immediate cost or reward is unknown, making it difficult to efficiently learn and apply control laws.
Innovation Solution
A reinforcement learning method that estimates coefficients of a value function represented in a quadratic form using past inputs and outputs, allowing for the determination of input values based on these estimates, even when the state equation, output equation, and immediate cost equation coefficients are unknown.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional reinforcement learning techniques are used to determine input values for controlled objects, then the control law can be learned based on immediate cost or reward, but accurate determination becomes difficult when the state is not directly observable and the immediate cost or reward is unknown
Solution Approach 1:
The patent applies preliminary action by estimating the value function coefficients before determining the input values. The system pre-processes the relationship between inputs and outputs by estimating the quadratic value function coefficients from historical data, which then enables accurate input determination even when states are not directly observable. This preliminary estimation phase separates the complex measurement problem from the control decision phase.
Solution Approach 2:
The patent introduces the value function as an intermediary between the controlled object and the control input. By estimating the value function coefficients that represent the relationship between inputs and outputs, the system creates a mediator model that bridges the gap when direct state observation is unavailable. This intermediary value function enables indirect inference of optimal inputs without direct access to states or immediate costs.
2Productivity
If the value function is represented in quadratic form with estimated coefficients, then efficient learning and accurate input determination are enabled, but the complexity of estimating coefficients from past data increases
Solution Approach 1:
The patent applies parameter changes by representing the value function in quadratic form with specific parameters (coefficients) that can be estimated from data. By changing the functional form to a quadratic representation with estimable coefficients, the system transforms a complex control learning problem into a parameter estimation problem. This parameterization enables efficient learning from historical input-output data while maintaining mathematical tractability for determining optimal inputs.
Data Source
AI summary
A non-transitory, computer-readable recording medium stores therein a reinforcement learning program that uses a value function and causes a computer to execute a process comprising: estimating first coefficients of the value function represented in a quadratic form of inputs at times in the past than a present time and outputs at the present time and the times in the past, the first coefficients being estimated based on inputs at the times in the past, the outputs at the present time and the times in the past, and costs or rewards that corresponds to the inputs at the times in the past; and determining second coefficients that defines a control law, based on the value function that uses the estimated first coefficients and determining input values at times after estimation of the first coefficients.


