State Model Generation With Variance-Reduced Action Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating state models of controllable systems, such as Monte Carlo simulations, suffer from uncertainty due to reliance on random conditions, making it difficult to select actions that achieve desired states efficiently.

Innovation Solution

A method using variance reduction techniques, particularly Monte Carlo tree search, to enhance the accuracy of state model generation by incorporating control variates, reducing the influence of chance-based rewards and optimizing the state model based on determined rewards.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If Monte Carlo simulation methods are used to generate state models, then the system can learn behavior without rule specification, but the results have high uncertainty due to random conditions

Engineering Contradiction:
Improveability to learn without rulesVSAvoiduncertainty of results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the parameters of the simulation by introducing control variates that modify the reward calculation. Instead of using raw random rewards, the system uses adjusted rewards that account for the influence of control variates, thereby reducing variance while maintaining the exploratory nature of Monte Carlo methods

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback by using the determined rewards to optimize the state model iteratively. The results from simulations feed back into refining the state model, creating a closed-loop system that reduces uncertainty over time through continuous improvement based on accumulated experience

Inventive Principle:
Principle #23Feedback

2Ease of operation

If conventional simulation methods are used, then action selection can be performed, but the variance in reward determination is high

Engineering Contradiction:
Improveaction selection capabilityVSAvoidreward determination accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces control variates as intermediary elements that mediate between the random simulation outcomes and the final reward determination. These control variates act as intermediaries that filter out random noise while preserving the essential learning signal, thereby improving measurement precision without eliminating the exploratory nature of simulations

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the state model is optimized based on determined rewards, then accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvestate model accuracyVSAvoidoptimization complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-selecting control variates and pre-structuring the optimization process. By preparing the control variates and optimization framework in advance, the system reduces the complexity of real-time optimization while maintaining high accuracy in state model determination

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12429836B2Method for generating a state model describing a controllable system
Publication Date: 2025.09.30 ROBERT BOSCH GMBH
  • US12429836B2 patent drawing
  • US12429836B2 patent drawing

AI summary

A method for generating a state model describing a controllable system. The method includes: providing at least one part of the state model; selecting an action from a set of actions starting from the second state of the components; simulating further states of the components by a successive application of an action from the set of actions to the components in each case, an individual reward being determined for each of the applications of an action to the components; optimizing the at least one part of the state model based on the determined rewards, the optimizing of the at least one part of the state model taking place based on a variance reduction method and a maximum of the determined rewards; and adding the selected action and the second state to the at least one part of the state model.