An optimization method for an agent motion control strategy

By training expert policies in a simulation environment and transferring them to the real environment using progressive networks, combined with observational learning and reward function generation, the problem of insufficient autonomous decision-making and learning capabilities of robots in complex environments is solved, achieving efficient policy transfer and a safe learning process.

CN117706918BActive Publication Date: 2026-06-02OCEAN UNIV OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OCEAN UNIV OF CHINA
Filing Date
2023-11-15
Publication Date
2026-06-02

Smart Images

  • Figure CN117706918B_ABST
    Figure CN117706918B_ABST
Patent Text Reader

Abstract

The application discloses an optimization training method of an intelligent body motion control strategy, and comprises the following steps: a control strategy simulation training step; a control strategy migration step; and a control strategy interactive learning step in a real environment, which comprises the following steps: controlling the intelligent body to move in the real environment according to a target strategy, generating a plurality of sample data by a target generator, and generating a current track. The method combines the method of learning from observation with a progressive neural network, thereby avoiding the task failure caused by learning from only action pairs. The expert track based on the state pairs collected in the simulation environment task is used to guide the training of the intelligent body in the target task, and the source generator, the source strategy and the source discriminator are migrated by using the progressive network, so that the training speed of the target strategy is improved.
Need to check novelty before this filing date? Find Prior Art