Robot Action Synchronization for Sim-to-Real Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning struggles to ensure stationarity when transitioning from a simulated world to a real world due to differences in state changes, such as the time taken for a robot's joint to rotate, necessitating control over these differences to maintain learning effectiveness.

Innovation Solution

A method and apparatus for synchronizing actions between simulated and real-world learning devices by determining and correcting for delay times between reaching a target state, involving the addition of dummy time to align the movement trajectories of robots in both environments, ensuring synchronization through proportional, differential, or integral control based on physical states.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If reinforcement learning is performed in a simulated world with controlled variables, then learning stability is improved, but adaptability to the real world deteriorates

Engineering Contradiction:
Improvelearning stabilityVSAvoidadaptability to real world
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent creates a simulated world that copies the real world environment, robot dynamics, and task conditions. By training reinforcement learning agents in this simulated copy, the system achieves stable learning while the simulation accurately replicates real-world physics and constraints, enabling transfer to the actual robot without significant performance degradation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent systematically varies simulation parameters such as robot mass, friction coefficients, and environmental conditions to match real-world values. By carefully adjusting these parameters in the simulated world, the patent ensures that the learned policies remain effective when deployed to the real robot, bridging the simulation-reality gap.

Inventive Principle:
Principle #35Parameter changes

2Stability of the object's composition

If the difference between simulated world and real world is completely controlled, then stationarity of reinforcement learning is improved, but device complexity deteriorates

Engineering Contradiction:
Improvestationarity of reinforcement learningVSAvoidcomplexity of controlling differences
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent introduces a simulated world as an intermediary environment between theory and real-world deployment. This simulated intermediary captures the essential dynamics and constraints of the real system, allowing reinforcement learning to achieve stationarity in simulation while the same model transfers to the real world, avoiding the need to control every细微 difference between simulation and reality.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If delay times between simulated world and real world are not synchronized, then ease of operation is improved, but measurement precision deteriorates

Engineering Contradiction:
Improveease of independent operationVSAvoidprecision of state alignment
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms that measure the delay times between simulated and real-world state transitions. By continuously monitoring these temporal discrepancies and using them to adjust the synchronization of state updates, the system maintains precise alignment between simulated and real environments while allowing both to operate independently with their own timing characteristics.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230342665A1Method and apparatus for synchronizing actions of learning devices between simulated world and real world
Publication Date: 2023.10.26 ELECTRONICS & TELECOMM RES INST
  • US20230342665A1 patent drawing
  • US20230342665A1 patent drawing
  • US20230342665A1 patent drawing

AI summary

Provided is a method and apparatus for synchronizing actions of robots between a simulated world and a real world. The method may include determining whether the learning device of the simulated world and the learning device of the real world reach the target state after one unit time, when the learning device of the simulated world and the learning device of the real world reach the target state, determining a first delay time, which is a time until the learning device of the simulated world reaches the target state, and a second delay time, which is a time until the learning device of the real world reaches the target state, and performing a correction between a state of the learning device of the simulated world and a state of the learning device of the real world based on the first delay time and the second delay time.