Reinforcement learning coach-driven context learning general motion control method and system

By randomly sampling multiple robot structures in the parameter space and introducing Gaussian noise perturbations, and using a general contextual learning framework for cross-task meta-training, the flexibility and resource requirements problems of existing motion controllers in complex environments are solved, and stable control is achieved in a variety of environments and structures.

CN120755878APending Publication Date: 2025-10-10THE CHINESE UNIV OF HONG KONG (SHENZHEN) +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511013580.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing motion controllers lack flexibility when facing complex or dynamically changing environments. Methods that rely on precise models have high computational costs, while methods based on deep reinforcement learning have low sample efficiency, unstable training, and high resource requirements, making them difficult to apply in embedded systems.

Method used

A reinforcement learning coach-driven contextual learning method is adopted. By randomly sampling multiple robot structures in the preset parameter space, introducing Gaussian noise perturbations, and using a general contextual learning framework for cross-task meta-training, the robot control model is optimized to achieve rapid adaptation to new tasks.

Benefits of technology

It improves the stability and robustness of the robot control model in various environments and structures, reduces the actual training cost, and is suitable for deployment in resource-constrained embedded control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120755878A_ABST
    Figure CN120755878A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning coach-driven context learning general motion control method and system, and the method comprises the steps: receiving the structure parameters of a plurality of predefined robots, carrying out the random sampling in a preset parameter space, and generating a plurality of robots of different structures; target tasks are defined, and reinforcement learning coaches corresponding to the target tasks are trained based on robots of different structures; in the simulation environment, guiding the robot to execute tasks by using a reinforcement learning coach, and recording a task execution track corresponding to each target task; extracting a context sequence with a fixed length from the task execution trajectory by using a general context learning framework, and performing cross-task meta-training by taking maximization of expected accumulated rewards under each target task as a training target so as to optimize context learning model parameters and obtain a pre-trained robot control model; a pre-trained robot control model; and obtaining a target control action which is output by the pre-trained robot control model and corresponds to the current state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of intelligent robotics technology, and in particular to a context-based learning general motion control method and system driven by a reinforcement learning coach. Background Art

[0002] With the rapid development of robotics, their applications in industrial manufacturing, services, healthcare, and exploration continue to expand. The importance of motion controllers, their core control components, has become increasingly prominent. The motion controller's task is to convert high-level task instructions (such as trajectory planning and behavioral decisions) into low-level motor control signals, ensuring the robot can stably and accurately execute motion control tasks, such as precise grasping by a robotic arm and maintaining the dynamic gait of a quadruped robot.

[0003] Existing motion controllers are mainly divided into three categories: the first is traditional control methods, such as PID (Proportional Integral Derivative) control and fuzzy control, which rely on manual parameter adjustment and system models and lack flexibility when dealing with complex or dynamically changing environments. The second is model-based control methods, such as Model Predictive Control (MPC), which rely on precise dynamic modeling, but are difficult to model in nonlinear systems or high-uncertainty scenarios and have high computational costs. The third is methods based on deep reinforcement learning, such as Deep Q-Network (DQN) or policy gradient method, which have strong representation capabilities, but in practical applications still suffer from low sample efficiency, unstable training process, weak generalization ability, and high dependence on computing resources. Summary of the Invention

[0004] Based on the above problems, the embodiments of the present application provide a context-based learning general motion control method and system driven by a reinforcement learning coach, with the aim of quickly and stably acquiring accurate and feasible target control actions under low-sample conditions, and enhancing the generalization capability of the controller in actual deployment.

[0005] In a first aspect, embodiments of the present application provide a context-based general motion control method driven by a reinforcement learning coach, including:

[0006] Receive predefined structural parameters of multiple robots and randomly sample them within a preset parameter space to generate multiple robots with different structures; the structural parameters include: structural length, degree of freedom distribution, number of degrees of freedom, joint working angles, and dynamic parameters;

[0007] Define the target tasks and establish a simulation environment, allowing the robot to interact with the environment within the simulator and train a reinforcement learning coach for each target task based on the reward function defined by the target task;

[0008] A reinforcement learning coach is used to guide the robot to perform preset target tasks and record the task execution trajectory corresponding to each target task; the preset target task includes a preset target task definition, state space, action space, and reward function; the task execution trajectory includes the current state, control action, and reward obtained at each time step;

[0009] Adding different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory to obtain the task execution trajectory after Gaussian noise disturbance;

[0010] A general context learning framework is used to extract a fixed-length context sequence from the task execution trajectory. Cross-task meta-training is performed with the training objective of maximizing the expected cumulative reward under each target task to optimize the context learning model parameters and obtain a pre-trained robot control model. The context learning model is used to output the target control action in the current state based on the context sequence.

[0011] Inputting the obtained current state of the robot into the pre-trained robot control model;

[0012] Obtaining a target control action corresponding to the current state output by the pre-trained robot control model;

[0013] Based on the target control action, a target action execution instruction is determined, and the target action execution instruction is used to control the robot to execute the target control action.

[0014] In one embodiment, the context sequence is represented as:

[0015]

[0016] in, Indicates the current time context sequence; Indicates the step index; Indicates the The status of the step; Indicates the The action of the step; Indicates the Step reward; Indicates the The status of the step; Indicates the context window length.

[0017] In one embodiment, adding Gaussian noise to the current state and control action at each time step in the task execution trajectory includes:

[0018]

[0019] wherein, represents the state of the current time after adding Gaussian noise; represents the original state of the current time ; represents the state noise, that is, the state noise is subject to a normal distribution with a mean of 0 and a variance of ; represents the action of the current time after adding Gaussian noise; represents the original action of the current time ; represents the action noise, that is, the action noise is subject to a normal distribution with a mean of 0 and a variance of .

[0020] In an embodiment, the general context learning framework processes the context sequence through a self-attention mechanism, and the specific calculation formula is:

[0021]

[0022] wherein, represents the output attention; represents a query matrix, and represents a query vector generated according to the current state; represents a key matrix composed of the context sequence vector embedding; represents a similarity score matrix, which is used to calculate the similarity between the current input and the historical context; represents the dimension of the key vector, which is used to prevent gradient explosion caused by too large vector dimension, and scale the inner product; represents a value matrix composed of the context sequence vector embedding.

[0023] In an embodiment, the training target is to maximize the expected cumulative reward under each target task, and the specific expression is:

[0024]

[0025] wherein, represents the cumulative reward expected to be obtained when completing the target task under the policy with parameters ; is the probability-weighted sum / integral of all possible trajectories under , which represents the expected value when the trajectory is subject to the probability distribution of the policy . Represents the learnable parameters of the control model, and the learnable parameters include at least: convolution weight value and convolution bias value; Indicates that the current parameter The policy function is used to output a control action based on the current state and the context sequence. The probability distribution of Indicates the number of time steps of the target task; Indicates the current moment; represents the discount factor, , used to control the degree of attention paid to rewards in the future; Indicates time A combination of state, action, and reward at a given moment.

[0026] In one embodiment, obtaining a target control action corresponding to the current state output by the robot control model includes:

[0027] Read historical task execution trajectories and preset prior knowledge; the historical task execution trajectories include the current state, control actions, and rewards obtained at multiple time steps;

[0028] The current state of the robot at the current moment, the historical task execution trajectory and the prior knowledge are input into the strategy function to generate the target control action at the current moment.

[0029] In a third aspect, the present application also provides a context-based learning general motion control method driven by a reinforcement learning coach, including:

[0030] A structural parameter sampling module is used to receive predefined structural parameters of multiple robots and randomly sample them within a preset parameter space to generate multiple robots with different structures; the structural parameters include: structural length, degree of freedom distribution, number of degrees of freedom, joint working angles, and dynamic parameters;

[0031] The reinforcement learning coach training module is used to define the target task and establish a simulation environment, allowing the robot to interact with the environment within the simulator and train the reinforcement learning coach corresponding to each target task based on the reward function defined by the target task;

[0032] The task simulation and trajectory acquisition module is used to use the reinforcement learning coach to guide the robot to perform preset target tasks and record the task execution trajectory corresponding to each target task; the preset target task includes the preset target task definition, state space, action space and reward function; the task execution trajectory includes the current state, control action and reward obtained at each time step;

[0033] a noise perturbation module, configured to add different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory, thereby obtaining the task execution trajectory after being perturbed by Gaussian noise;

[0034] A context modeling and policy training module is used to extract a fixed-length context sequence from the task execution trajectory using a general context learning framework. The module performs cross-task meta-training with the goal of maximizing the expected cumulative reward under each target task to optimize the context learning model parameters and obtain a pre-trained robot control model. The context learning model is used to output the target control action in the current state based on the context sequence.

[0035] A robot current state acquisition module inputs the acquired current state of the robot into the pre-trained robot control model;

[0036] an action reasoning module, configured to obtain a target control action corresponding to the current state output by the pre-trained robot control model;

[0037] The action execution module is used to determine a target action execution instruction based on the target control action, and the target action execution instruction is used to control the robot to execute the target control action.

[0038] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0039] CPU, memory, input and output interfaces;

[0040] The memory is a transient storage memory or a persistent storage memory;

[0041] The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the method described in the first aspect of the embodiment of the present application and any specific implementation of the first aspect.

[0042] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the embodiment of the present application and any specific implementation method of the first aspect is executed.

[0043] In a fifth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect of the embodiment of the present application and any specific implementation method of the first aspect.

[0044] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0045] The embodiment of the application constructs a large number of robots with various structures by randomly sampling a plurality of robots (such as different bar lengths, masses, inertias, etc.) in a structural parameter space, so that the trained control strategy is no longer limited to a single robot form and has the adaptability to a plurality of robot physical structures. The introduction of Gaussian noise in the task execution trajectory data simulates the perception error and control deviation in the real environment, so that the robot control model can learn to handle various uncertainties in the training stage, thereby improving the stability and robustness in actual deployment. Moreover, the use of a general context learning framework combined with an in-context reinforcement learning (ICRL, In-Context Reinforcement Learning) mechanism can quickly understand the current task context based on historical task execution trajectory data, thereby outputting more targeted control actions. The meta-training of the robot control model in a plurality of tasks and a plurality of structural robots can realize the rapid transfer and strategy sharing between tasks, reduce the actual training cost, and make it more suitable for resource-limited embedded control system deployment. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only a part of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on the provided drawings.

[0047] Figure 1 A reinforcement learning coach driven context learning general motion control method flowchart provided by the embodiment of the present application;

[0048] Figure 2 A reinforcement learning coach driven context learning general motion control system structure schematic diagram provided by the embodiment of the present application;

[0049] Figure 3 An electronic device structure schematic diagram provided by the embodiment of the present application. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0051] With the rapid development of robotics, the design and optimization of motion controllers are becoming increasingly critical in robotic systems. Motion controllers are responsible for converting high-level task instructions into low-level motor control signals, ensuring that robots can accurately and stably execute actions, such as tracking the trajectory of a robotic arm or controlling the gait of a quadruped robot. In the prior art, motion controllers relevant to the present invention primarily include the following categories:

[0052] 1. Traditional control methods: such as PID control and fuzzy control, which are widely used in industrial robots and automation equipment;

[0053] 2. Model-based methods: such as MPC, which achieves precise control by building a dynamic model of the robot;

[0054] 3. Deep learning-based reinforcement learning (RL) methods, such as DQN and policy gradient methods, have shown potential in the field of robotic control in recent years.

[0055] However, these existing technologies have significant shortcomings when facing the needs of modern robotics applications, mainly manifested in strong model dependence and lack of flexibility:

[0056] 1) PID control and fuzzy control require precise parameter tuning and system models, making them less adaptable to nonlinear or time-varying systems, such as the dynamic control of a quadruped robot on uneven terrain. Control accuracy and flexibility are also significantly limited in complex tasks, such as the coordinated control of multi-degree-of-freedom manipulators or the gait adjustment of a quadruped robot in a changing environment. Model-based approaches have significant drawbacks, primarily including modeling difficulties and heavy computational overhead.

[0057] 2) Methods like MPC rely heavily on accurate dynamic models, but model building is extremely challenging in highly nonlinear robotic systems or environments with high uncertainty. Furthermore, these methods suffer from high computational complexity, making them difficult to meet real-time control requirements, and their application is particularly limited in embedded systems.

[0058] 3) RL methods based on deep learning have several drawbacks. First, they suffer from low sample efficiency. Training requires a large amount of environmental interaction data. For example, learning a new gait for a quadruped robot may require tens of thousands of trial-and-error iterations, which is costly and time-consuming for a physical robot. Second, the training process is unstable and exhibits large variance, which can lead to convergence difficulties or performance fluctuations. For example, training a quadruped robot's gait may fail due to parameter sensitivity.

[0059] 4) In addition, the generalization ability is weak, and the trained model performs poorly when the new task or environment changes (such as from flat ground gait to rugged terrain), and often needs to be retrained. Moreover, deep RL requires a large amount of computing resources, making it difficult to apply to resource-constrained embedded systems.

[0060] In view of the deficiencies of the prior art in sample efficiency, training stability, generalization ability and resource requirement, the application proposes a motion controller based on a general context learning framework. The controller can adapt to the control requirements of various forms of robots such as manipulators and quadruped robots through in-context reinforcement learning (ICRL). Relying on the general context learning framework, the controller performs meta-training in a large-scale randomized world to obtain in-context learning (ICL) ability, which can quickly adapt to new tasks without relying on gradient updates according to the learning tasks provided by the reinforcement learning coach (RL Coach, Reinforcement Learning Coach).

[0061] Firstly, the difference between the general context learning framework and the traditional artificial design reinforcement learning method will be briefly introduced. The biggest difference is that it no longer relies on the memory of specific tasks and the manually designed strategy activation method. The model trained by the traditional method has seen many tasks during training, and it is easy to learn in the form of conditioned reflex, that is, the model recognizes the current task and calls the previously remembered skills to respond. In this way, the model does not learn "how to learn", but learns to use skill A if it recognizes that it is task A, and uses skill B if it recognizes that it is task B. This way is commonly known as conditioned reflex or rote memorization, because the model trained in this way stores the specific solution (skill) of each task in memory, and when reasoning, it will activate the previously remembered skill according to the context. But actually, this is not a generalization learning, but simply activating the corresponding skill after identifying the task. The model trained by the traditional method has weak generalization ability and is difficult to adapt to new tasks that have not been seen before, and has not learned the "learning ability" (such as how to adjust the output strategy according to the feedback).

[0062] The general context learning framework, on the other hand, constructs a large-scale and diverse task set and uses long sequence context input to make the model unable to memorize each task and the corresponding specific solution or skill, and uses long sequence context input (such as multiple time step observations + actions + feedback) to make the model not only focus on how to complete the current task, but also master the meta-rules of strategy learning in different tasks, so that the model has stronger task generalization ability and autonomous learning ability, and can quickly adapt to new environments without gradient updates, embodying the essential transition from "memorizing answers" to "learning to cope with tasks and find solutions".

[0063] The following is a further detailed description of the various embodiments of the present application in conjunction with the accompanying drawings.

[0064] The present application embodiment provides a context-based learning general motion control method driven by a reinforcement learning coach, such as Figure 1 As shown, the method includes steps S101-S108.

[0065] S101: receiving predefined structural parameters of multiple robots, and performing random sampling within a preset parameter space to generate multiple robots with different structures.

[0066] To meet the control requirements of various robot forms, this application uses the Unified Robot Description Format (URDF) to describe the physical structure of multiple robot types. This aims to build a diverse population of robot models for training, thereby improving the generalization and task adaptability of the control model. URDF is an XML-formatted standard for describing robot structures, used in robotic platforms such as the Robot Operating System (ROS). It represents the structural parameters of a robot, including structural dimensions, degree-of-freedom distribution, number of degrees of freedom, joint working angles, and dynamic parameters. Structural dimensions refer to the physical geometric dimensions of each robot component (links, legs, arms, base, etc.), such as length, width, and height. The number of degrees of freedom refers to the total number of directions / axes in which a robot system can independently move, with each independently controllable rotation or translation counting as a degree of freedom. The degree-of-freedom distribution describes which robot components have which motion capabilities. The joint working angle represents the allowable angular range within which each joint can rotate or move. Dynamic parameters describe the robot's dynamic behavior, such as mass, inertia, friction, damping, and torque constants. The URDF file of a quadruped robot may include information such as the structural dimensions and length of each component, the connection relationship between the components, the type / joint working angle range / axis of each joint, the mass / inertia of each component, and dynamic parameters.

[0067] On this basis, in order to achieve structural diversity, this application defines a high-dimensional structural parameter space , you can refer to the following expression:

[0068]

[0069] in, is the length, mass, and inertia tensor of the first member; Indicates the torque limit of the first joint.

[0070] It is length, mass, inertia tensor of the i-th link; is the torque limit of the j-th joint; is the torque limit of the j-th joint; is the number of links.

[0071] Based on the above structure parameter space, a robot with structural diversity is generated by random sampling in the following preset range. The preset range can refer to the following expression:

[0072]

[0073] wherein, is the torque limit of the j-th joint; length, mass, inertia tensor of the i-th link, are minimum values of length, mass, inertia tensor, and torque limit, respectively; are maximum values of length, mass, inertia tensor, and torque limit, respectively.

[0074] By random combination sampling of the above structure parameters, a large number of robot models with structural diversity can be generated, such as robotic arms with different length and mass distribution, quadruped robots with different limb structure, etc. The diversified modeling robot helps the strategy model in the general context learning framework to have good transferability and adaptability when facing different physical structures of robots.

[0075] S102: define a target task and establish a simulation environment, make the robot interact with the environment in the simulator, and train the reinforcement learning coach corresponding to each target task based on the reward function defined by the target task.

[0076] The reinforcement learning coach refers to the RL Coach framework. For each target task, the RL Coach framework is used to train a corresponding reinforcement learning coach in a simulation environment. Through interaction between the robot and the simulation environment and learning based on a reward function, the coach acquires the necessary capabilities to complete the target task. Compared to the common prior art approaches to improving generalization capabilities, which employ different initial states, random perturbations, or environmental adjustments to mitigate the uncertainty of the real-world environment on a single-task or single-morphology robot, ensuring that the policy network maintains relatively stable output under diverse conditions, the proposed method offers significant advantages. Prior art approaches typically rely on a unified neural network to simultaneously learn policies for multiple tasks. This not only requires significant time and effort to design the network structure and loss function, but also presents challenges with modeling difficulty and computational complexity. It also easily causes interference between different tasks, making it difficult for the model training process to converge. In contrast, the present application trains a dedicated reinforcement learning coach for each target task, effectively avoiding interference between different tasks. This results in higher training efficiency, more tailored policy network output to the robot's current environment, and improved scalability and maintainability.

[0077] The following details the process of training a reinforcement learning coach for each target task in this application. First, the specific requirements of each target task need to be defined. For example, the target task could be to make a quadruped robot walk from a starting point to an end point, maintain balance, or climb stairs. Each target task requires defining the state space (sensor data or environmental information observable by the robot), the action space (the output of the robot's controllable joints), the termination conditions of the target task, and the criteria for determining whether the task succeeded or failed.

[0078] Secondly, build the corresponding physical simulation environment in the simulator. You can use common simulators such as Gazebo and MuJoCo to build the environment, and load the robot's structural model (for example, using the URDF file mentioned in S101) and the environment model (such as the ground, obstacles, target location, etc.). The simulator is used to receive control actions issued by the reinforcement learning coach and return feedback information such as the current state, reward value, and whether the target task is terminated.

[0079] Then, a reward function is designed based on the target task to clarify the direction of robot behavior optimization. For example, in a walking task, the reward function can be designed so that the farther the robot travels, the higher the reward value, while falling will result in a penalty. In a balancing task, the reward value can increase as the robot's body tilt angle decreases, while losing balance will result in a penalty.

[0080] Finally, after completing the above configuration, training begins for each target task's corresponding reinforcement learning coach. The reinforcement learning coach establishes communication with the simulation environment, configures the reinforcement learning algorithm and hyperparameters (such as the learning rate, discount factor, and batch size), and loads the reward function for the corresponding target task. During training, the reinforcement learning coach controls the robot to repeatedly try different actions in the simulation environment. Based on the reward signals obtained, the policy network model parameters are continuously updated, enabling the robot to gradually master the target task. Each target task is independently trained with a reinforcement learning coach and its corresponding policy network model, resulting in a specifically optimized control strategy for each target task.

[0081] S103: Using a reinforcement learning coach to guide the robot to perform preset target tasks, and recording the task execution trajectory corresponding to each target task.

[0082] The preset target task includes the preset target task, state space, action space and reward function; the task execution trajectory includes the current state, control action and reward obtained at each time step;

[0083] For the robot individuals with different structural parameters randomly generated in step S101, a target task is assigned to each robot individual. The reinforcement learning coach is then used to simulate and execute the target task and collect task execution trajectory data. The target task set in step S103 can be the same as the target task used to train the reinforcement learning coach in step S102, or it can be adjusted or expanded upon, such as by adjusting the termination conditions or reward function parameters of the target task or introducing new perturbation conditions, in order to verify the adaptability and generalization performance of the reinforcement learning coach under different task variations.

[0084] Specifically, the reinforcement learning coach receives a preset configuration, including the task objective, state space, action space, and reward function, and automatically applies it to individual robots with different structural parameters generated in step S101. During the simulation, it drives the robots to perform the task, collects state, action, and reward data at each step, and constructs a standardized trajectory sequence. By managing the allocation of target tasks, simulation execution, and data collection using the reinforcement learning coach, it is possible to efficiently obtain task execution trajectory data for a multi-task, multi-robot architecture.

[0085] Furthermore, in a feasible embodiment, if the general contextual learning framework is based on the architecture of the Transformer fusion gated slot attention (GSA) module, the task execution trajectory data collected by the above reinforcement learning coach can be written into the memory slot of the GSA module after vector encoding processing.

[0086] In S103, the reinforcement learning coach assigns a pre-defined target task to each structural robot, each of which pre-defines the following elements:

[0087] (1) Task execution goal: e.g. "grasp the target object", "walk to the designated position", or "walk forward while keeping balance";

[0088] (2) State space: define the perception variables for decision input, such as joint position, velocity, end position, robot posture, etc.;

[0089] (3) Action space: define the types and dimensions of controllable variables, such as joint torque, velocity command, position increment, etc.;

[0090] (4) Reward function: used to feedback the degree of completion of the target task, such as proximity to the target, energy consumption, stability score, etc.

[0091] Exemplarily, assume that one of the multiple pre-defined target task templates has the following content:

[0092] { "Task name": ["walk forward while keeping balance"],

[0093] "State space": ["joint angle", "body posture", "velocity"],

[0094] "Action space": ["torque of 12 joints"],

[0095] "Reward function": ["forward distance + stability - energy consumption"]}.

[0096] The reinforcement learning coach randomly assigns this task to multiple structural robots sampled at random, and drives the robots assigned to this task to simulate running according to the task, allowing the robots to attempt to complete the task under the state-action definition, and the reinforcement learning coach automatically records the task execution trajectory, including: the current state at each time step the control action taken under the current state , defined by the RL Coach according to the task goal (such as stability or target proximity).

[0097] The above task execution trajectory data is organized from time step 0 to t−1 as a state-action-reward trajectory sequence in the form of .

[0098] Each quadruple constitutes an interaction data at a time step, for example, the quadruple represents taking action under state , and obtaining reward , and enter a new state This sequence is used for subsequent context extraction and strategy training, and is the data basis for the context sequence of the general context learning framework to achieve task recognition and action prediction.

[0099] S104: Adding different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory to obtain the task execution trajectory after Gaussian noise disturbance.

[0100] For the task execution trajectory generated by the previous simulation, in order to more realistically simulate the uncertainty of the actual robot in the perception and execution process, Gaussian noise perturbation is introduced. Independently sampled Gaussian noise terms are added to the trajectory data at each time step to simulate sensor errors, measurement jitter, actuator errors, and motor deviations. In this way, the subsequent strategy learning process can adapt to realistic factors such as inaccurate perception and action execution errors, thereby improving the generalization ability of the control strategy.

[0101] S105: A fixed-length context sequence is extracted from the task execution trajectory using a general context learning framework, and cross-task meta-training is performed with the goal of maximizing the expected cumulative reward under each target task to optimize the context learning model parameters and obtain a pre-trained robot control model; the context learning model is used to output the target control action in the current state according to the context sequence.

[0102] For example, the general context learning framework can be based on the Transformer architecture and integrated with the GSA module to efficiently process context information and implement memory management. In this case, the general context learning framework consists of two main submodules:

[0103] (1) The Transformer model is used to model contextual dependencies in trajectory sequences and infer target control actions based on the current state and historical trajectory sequences. Furthermore, the Transformer's self-attention mechanism enables it to automatically focus on historical task execution trajectories that contributed to the current decision.

[0104] (2) GSA module: It is used to gate the time steps in the input trajectory and retain the "slots" (memory slots) that are highly relevant to the current task, thereby enhancing the model's ability to extract long sequence information and improving reasoning accuracy and generalization.

[0105] In other possible implementations, the general context learning framework can also be implemented based on other sequence modeling architectures, such as structures based on recurrent neural networks (RNNs), long short-term memory networks (LSTMs), to achieve the same context modeling, memory management, and control inference functions as described above, but for ease of understanding, the architecture based on the Transformer model and the GSA module will be used as an example for subsequent description.

[0106] In the training phase, the general context learning framework learns the ICL capability by meta-training on a large number of randomized Markov decision process (MDP) tasks. The MDP is defined as a five-tuple , where represents the state space, covering all states that the robot can be in; represents the action space, representing the control actions that the robot can perform; represents the state transition probability, representing the probability of transitioning from state to state after performing action ; represents the reward function, representing the immediate reward obtained when transitioning from state to after performing action ; is a discount factor, ranging from , used to balance immediate rewards and long-term benefits. In the meta-training process, the general context learning framework can extract general patterns from the historical interaction trajectories of different tasks, thereby achieving the ability to quickly adapt to new tasks.

[0107] Specifically, using the multiple robot structures and task trajectories (including Gaussian perturbations) obtained in step S103, a fixed-length context sequence is extracted from each task execution trajectory, for example, the last 10 steps of trajectory data are extracted as the context input of the policy model.

[0108] The general context learning framework uses the constructed context sequence as input, and meta-trains in a large number of different tasks to maximize the expected cumulative discounted reward in the target task, and optimizes the learnable parameters of the context learning model to optimize the overall performance of the context learning model in different tasks and different robot structures, so that the context learning model has the ability to quickly adapt and generalize to tasks.

[0109] It should be noted that the general context learning framework proposed in the present application is a whole architecture of training and inference method. The framework takes a well-constructed context sequence as input, adopts meta-training in a large number of diversified random tasks, and the goal is to maximize the expected cumulative discounted reward in each task. Through continuous iteration training and optimization, the finally trained model has context reasoning ability and cross-task generalization ability.

[0110] The context learning model is a specific strategy type trained based on the general context learning framework, and the learnable parameters thereof are continuously optimized in the training process based on the general context learning framework. The finally trained context learning model can be deployed in the control unit of the robot as a real-time control decision maker, and can quickly generate control actions according to the current state, historical trajectory and prior prompt. Therefore, the context learning model trained by the general context learning framework has high adaptability to different tasks and robots of different structures. More specifically, in a feasible implementation, the OmniRL framework can be used as the general context learning framework, and the OmniRL model trained in the framework is the real-time decision maker deployed on the robot.

[0111] S106: input the acquired current state of the robot into the robot control model. The robot control model is obtained based on the aforementioned disclosed reinforcement learning coach driven context learning general motion control method.

[0112] S107: acquire the target control action corresponding to the current state output by the robot control model.

[0113] S108: determine the target action execution instruction based on the target control action, and the target action execution instruction is used to control the robot to execute the target control action.

[0114] The current state of the robot includes but is not limited to sensor readings such as joint angle, IMU sensor data, end position, etc. In actual use, by collecting the current state of the robot and inputting it into the trained robot control model (which has context learning ability after being trained by the aforementioned disclosed steps), the target control action (such as joint torque, target position or speed, etc.) output by the robot control model is obtained. Then, the output target control action (usually a floating point value) is converted into an execution signal such as a control voltage, a PWM signal or a communication instruction, so that the target control action can be recognized and executed by the robot bottom control system, realizing the action control of the robot in the real environment.

[0115] In order to make the control model have context reasoning ability, in an embodiment, the fixed-length context sequence is extracted from the task execution trajectory by using the general context learning framework, which can be referred to as the following expression:

[0116]

[0117] in, Indicates the current time context sequence; Indicates the step index; Indicates the The status of the step; Indicates the The action of the step; Indicates the Step reward; Indicates the The status of the step; Indicates the context window length.

[0118] During robot training simulations, task execution trajectory data (states, actions) is typically obtained from an idealized simulation environment, lacking the perceptual errors and execution biases common in the real world. To enhance the model's adaptability to real-world interference, in this embodiment, the system actively introduces Gaussian noise into the task execution trajectory data at each time step to simulate the uncertainty inherent in real-world perception and control. In one feasible embodiment, the addition of Gaussian noise to the current state and control action at each time step in the task execution trajectory in step S103 of this application includes:

[0119]

[0120] in, Indicates the current moment after adding Gaussian noise Status; Represents the original current moment Status; represents the state noise, , that is, the state noise has a mean of 0 and a variance of Normal distribution; Indicates the current moment after adding Gaussian noise Actions; Represents the original current moment Action; represents motion noise, , that is, the action noise has a mean of 0 and a variance of Normal distribution.

[0121] For example, suppose the robot's actual joint sensors have limited accuracy and their readings are The deviation corresponds to , or motor torque control exists The output deviation corresponds to These biases can be modeled as Gaussian distributions and added to the simulation data, so that the robot control model learns how to deal with these uncertainties during training.

[0122] In an embodiment, the general context learning framework processes the context sequence through a self-attention mechanism, and the specific formula is as follows:

[0123]

[0124] wherein, represents the output attention; represents the query matrix, and represents the query vector generated according to the current state; represents the key matrix composed of the context sequence vector embedding; represents the similarity score matrix, which is used to calculate the similarity between the current input and the historical context; represents the dimension of the key vector, which is used to prevent gradient explosion caused by too large vector dimension, and scale the inner product; represents the value matrix composed of the context sequence vector embedding.

[0125] Exemplarily, in a feasible implementation, the general context learning framework adopts the architecture of the Transformer model and the GSA module, and the context learning model processes the context sequence information in the task trajectory based on the Transformer architecture. The core of the Transformer is its self-attention mechanism, which enables the context learning model to determine which part of the historical sequence is most helpful for the current decision at each time step, thereby generating a more accurate and more context-related control policy.

[0126] Specifically, a query vector is generated according to the current state, which represents how to perform the target control action in the next step according to the current state. The key matrix stores the state / action / reward data of the historical time steps in the context sequence, and the similarity between the query vector and the task execution trajectory data at each time step in the key matrix is calculated to determine which historical trajectory is most relevant to the current state, so as to determine which historical task execution trajectory to focus on to assist the current decision. After normalization by the softmax function, the similarity between the query vector and the task execution trajectory is used as the attention weight, and the attention weight is used to weight each value vector (i.e., the historical time step task execution trajectory), so as to weight and fuse the historical time step task execution trajectory into a context summary for use in the current policy decision.

[0127] In an embodiment, the training target is to maximize the expected cumulative reward under each target task, and the specific expression is as follows:

[0128]

[0129] in, Indicates that the parameter is Strategy The cumulative reward expected to be obtained when the target task is completed; In all possible trajectories Press it on The probability weighted summation / integration under represents the trajectory Obedience Strategy The expected value of the probability distribution of Represents the learnable parameters of the control model, and the learnable parameters include at least: convolution weight value and convolution bias value; Indicates that the current parameter The policy function is used to output a control action based on the current state and the context sequence. The probability distribution of Indicates the number of time steps of the target task; Indicates the current moment; represents the discount factor, , used to control the degree of attention paid to rewards in the future; Indicates time A combination of state, action, and reward at a moment.

[0130] In one embodiment, obtaining a target control action corresponding to a current state output by a robot control model includes: reading a historical task execution trajectory and preset prior knowledge; the historical task execution trajectory includes the current state, control action, and reward acquisition of multiple time steps; and inputting the current state of the robot at the current moment, the historical task execution trajectory, and the prior knowledge into a strategy function to generate the target control action at the current moment.

[0131] For example, the target task is to jump from a high altitude and land stably.

[0132] The robot's current status may be:

[0133]

[0134] in, Indicates the position angle of the first joint (unit: rad), Indicates the position angle of the second joint, represents the angular velocity of the first joint, Indicates the rotation angle (roll angle) of the body around the x-axis, Indicates the rotation angle (pitch angle) of the body around the y-axis.

[0135] Read the historical trajectory data of the robot's interaction with the simulation environment over a period of time, and extract relevant trajectory segments to construct a context sequence .

[0136] Assume that the historical mission trajectory is as follows:

[0137] Time step 1: The state is standing, the action is to take a small step forward, and the reward is 0.5;

[0138] Time step 2: The state is jumping, the action is a strong kick, and the reward is 0.7;

[0139] Assumed prior knowledge It can be understood as auxiliary information provided in advance that can be used directly to guide model reasoning and can be associated with different target tasks. In this task, the system's preset priori prompts include: 1) Joint 1 needs to be stretched quickly (>0.8 rad / s) to obtain vertical speed during takeoff; 2) The pitch angle is adjusted to be close to 0 in the air to keep the body horizontal; 3) Joint 2 is bent ( 2≈-0.3 rad) to absorb impact and avoid roll angle fluctuations.

[0140] The current state of the robot at the current moment, the historical task execution trajectory and prior knowledge are input into the strategy function to generate the target control action at the current moment.

[0141] Policy Function Expressed as:

[0142]

[0143] The context learning model is based on the current state and historical trajectory , using prior information in the context and a posteriori feedback to generate real-time control signals =[Joint 1 torque, Joint 2 torque] = [−0.1, −0.5], Joint 1 brakes slightly (-0.1) to maintain balance, and Joint 2 bends strongly (-0.5) to reach the target angle Here, the posterior feedback refers to the feedback signal obtained after the target control action is executed, which can include the reward signal, the execution result (whether the action is successfully completed), the change of the sensor reading (such as whether the distance is reduced), etc. =0.8 (landing stability reward), the execution result is successful and stable landing, and the task is completed.

[0144] In order to implement the context-based learning universal motion control method driven by a reinforcement learning coach in the embodiment of the present application, the embodiment of the present application also provides a context-based learning universal motion control system driven by a reinforcement learning coach, please refer to Figure 2 , the system comprises:

[0145] The structural parameter sampling module 201 is used to receive predefined structural parameters of multiple robots and randomly sample them within a preset parameter space to generate multiple robots with different structures; the structural parameters include: structural length, degree of freedom distribution, number of degrees of freedom, joint working angles, and dynamic parameters;

[0146] A reinforcement learning coach training module 202 is used to define target tasks and establish a simulation environment, allowing the robot to interact with the environment within the simulator, and to train the reinforcement learning coach corresponding to each target task based on the reward function defined by the target task;

[0147] The task simulation and trajectory acquisition module 203 is used to use the reinforcement learning coach to guide the robot to perform preset target tasks and record the task execution trajectory corresponding to each target task; the preset target task includes a preset target task definition, state space, action space, and reward function; the task execution trajectory includes the current state, control action, and reward obtained at each time step;

[0148] A noise perturbation module 204 is configured to add different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory to obtain the task execution trajectory after being perturbed by Gaussian noise;

[0149] The context modeling and strategy training module 205 is configured to extract a fixed-length context sequence from the task execution trajectory using a general context learning framework, perform cross-task meta-training with the goal of maximizing the expected cumulative reward under each target task, and optimize the context learning model parameters to obtain a pre-trained robot control model. The context learning model is configured to output the target control action in the current state based on the context sequence.

[0150] The robot current state acquisition module 206 inputs the acquired current state of the robot into the pre-trained robot control model;

[0151] an action reasoning module 207 for obtaining a target control action corresponding to the current state output by the pre-trained robot control model;

[0152] The action execution module 208 is used to determine a target action execution instruction based on the target control action, and the target action execution instruction is used to control the robot to execute the target control action.

[0153] In one embodiment, the context sequence is represented as:

[0154]

[0155] in, Indicates the current time context sequence; Indicates the step index; Indicates the The status of the step; Indicates the The action of the step; Indicates the Step reward; Indicates the The status of the step; Indicates the context window length.

[0156] In one embodiment, adding Gaussian noise to the current state and control action at each time step in the task execution trajectory includes:

[0157]

[0158] in, Indicates the current moment after adding Gaussian noise Status; Represents the original current moment Status; represents the state noise, , that is, the state noise has a mean of 0 and a variance of Normal distribution; Indicates the current moment after adding Gaussian noise Action; Represents the original current moment Action; represents motion noise, , that is, the action noise has a mean of 0 and a variance of Normal distribution.

[0159] In one embodiment, the general context learning framework processes the context sequence through a self-attention mechanism, and the specific calculation formula is:

[0160]

[0161] in, Indicates output attention; Represents the query matrix, which stores the query vector generated based on the current state; represents the key matrix formed by embedding the context sequence vector; denotes a similarity score matrix, used to calculate the similarity between the current input and the historical context; denotes the dimension of the key vector, used to prevent gradient explosion caused by too large vector dimension, and scale the inner product; denotes a value matrix composed of the context sequence vector embedding.

[0162] In an embodiment, the training objective is to maximize the expected cumulative reward under each target task, and the specific expression is:

[0163]

[0164] wherein, denotes the cumulative reward expected to be obtained when completing the target task under the policy with parameters ; is the sum / integral weighted by the probability of under , and represents the expected value when the trajectory obeys the probability distribution of the policy ; denotes the learnable parameters of the control model, including at least: convolution weight values, convolution bias values; denotes the policy function currently determined by the parameters , used to output the probability distribution of the control action according to the given current state and context sequence; denotes the time step of the target task; denotes the current time; denotes the discount factor , used to control the degree of attention to the reward in the future time; denotes the combination of state, action and reward at time .

[0165] In an embodiment, the action inference module is specifically configured to:

[0166] read historical task execution trajectories and preset prior knowledge; the historical task execution trajectories include current states, control actions and obtained rewards at multiple time steps;

[0167] input the current state of the robot at the current time, the historical task execution trajectories and the prior knowledge into the policy function together to generate the target control action at the current time.

[0168] It should be noted that the above embodiment provides a reinforcement learning coach-driven context-learning universal motion control system. When performing motion control, the division of the above-mentioned program modules is only used as an example. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the reinforcement learning coach-driven context-learning universal motion control system provided in the above embodiment and the reinforcement learning coach-driven context-learning universal motion control method embodiment are of the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0169] Based on the hardware implementation of the above program modules, in order to implement the context-based learning general motion control method driven by reinforcement learning coach provided in the embodiment of the present application, the embodiment of the present application also provides an electronic device, such as Figure 3 As shown, the electronic device 300 includes:

[0170] CPU 301, memory 302 and input / output interface 303;

[0171] The memory 302 is a temporary storage memory or a permanent storage memory;

[0172] The central processing unit 301 is configured to communicate with the memory 302 and execute instruction operations in the memory 302 to perform the method described in the first aspect of the embodiment of the present application and any specific implementation of the first aspect.

[0173] Of course, in actual application, the various components in the electronic device 300 are coupled together through the bus system 304. It can be understood that the bus system 304 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 304 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 3 Various buses are labeled as bus system 304 .

[0174] The memory 302 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device 300. Examples of such data include: any computer program used to operate on the electronic device 300.

[0175] It is understandable that when the processor in the electronic device described above executes the computer program, it can also implement the functions of the various units in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, and one or more modules / units are stored in the memory and executed by the processor to complete the various embodiments of the present application. One or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device. For example, the computer program can be divided into the various units in the above-mentioned electronic device, and each unit can implement the specific functions described in the above-mentioned corresponding electronic device.

[0176] Electronic devices may be computing devices such as desktop computers, laptops, PDAs, and cloud servers. Electronic devices may include, but are not limited to, processors and memory. Those skilled in the art will appreciate that processors and memory are merely examples of electronic devices and do not limit the scope of electronic devices. Electronic devices may include more or fewer components, or combinations of certain components, or different components. For example, electronic devices may also include input / output devices, network access devices, buses, and the like.

[0177] A processor can be a central processing unit (CPU), other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of an electronic device, connecting all parts of the entire electronic device using various interfaces and lines.

[0178] The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application required for the function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0179] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the embodiment of the present application and any specific implementation of the first aspect is executed.

[0180] An embodiment of the present application also provides a computer program product having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, it is used to implement the method described in the first aspect of the embodiment of the present application and any specific implementation method of the first aspect.

[0181] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0183] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0184] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

Claims

1. A context-based general motion control method driven by a reinforcement learning coach, characterized by: include: Receive predefined structural parameters of multiple robots and perform random sampling within the preset parameter space to generate multiple robots with different structures; The structural parameters include: structural size length, degree of freedom distribution, number of degrees of freedom, joint working angle, and dynamic parameters; Define the target tasks and establish a simulation environment, allowing the robot to interact with the environment within the simulator and train a reinforcement learning coach for each target task based on the reward function defined by the target task; A reinforcement learning coach is used to guide the robot to perform preset target tasks and record the task execution trajectory corresponding to each target task; the preset target task includes a preset target task definition, state space, action space, and reward function; the task execution trajectory includes the current state, control action, and reward obtained at each time step; Adding different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory to obtain the task execution trajectory after Gaussian noise disturbance; A general context learning framework is used to extract a fixed-length context sequence from the task execution trajectory. Cross-task meta-training is performed with the training objective of maximizing the expected cumulative reward under each target task to optimize the context learning model parameters and obtain a pre-trained robot control model. The context learning model is used to output the target control action in the current state based on the context sequence. Inputting the obtained current state of the robot into the pre-trained robot control model; Obtaining a target control action corresponding to the current state output by the pre-trained robot control model; Based on the target control action, a target action execution instruction is determined, and the target action execution instruction is used to control the robot to execute the target control action.

2. The method according to claim 1, characterized in that The context sequence is represented as: in, Indicates the current time context sequence; Indicates the step index; Indicates the The status of the step; Indicates the The action of the step; Indicates the Step reward; Indicates the The status of the step; Indicates the context window length.

3. The method according to claim 1, characterized in that The adding of Gaussian noise to the current state and control action of each time step in the task execution trajectory includes: in, Indicates the current moment after adding Gaussian noise Status; Represents the original current moment Status; represents the state noise, , that is, the state noise has a mean of 0 and a variance of Normal distribution; Indicates the current moment after adding Gaussian noise Action; Represents the original current moment Actions; represents motion noise, , that is, the action noise has a mean of 0 and a variance of Normal distribution.

4. The method according to claim 1, wherein The general context learning framework processes the context sequence through the self-attention mechanism. The specific calculation formula is: in, Indicates output attention; Represents the query matrix, which stores the query vector generated based on the current state; represents the key matrix formed by embedding the context sequence vector; Represents the similarity score matrix, which is used to calculate the similarity between the current input and the historical context; Indicates the dimension of the key vector, which is used to prevent gradient explosion caused by excessive vector dimension and to scale the inner product; Represents the value matrix composed of the context sequence vector embedding.

5. The method according to claim 1, wherein The training goal is to maximize the expected cumulative reward under each target task. The specific expression is: in, Indicates that the parameter is Strategy The cumulative reward expected to be obtained when the target task is completed; In all possible trajectories Press it on The probability weighted summation / integration under represents the trajectory Obedience Strategy The expected value of the probability distribution of Represents the learnable parameters of the control model, and the learnable parameters include at least: convolution weight value and convolution bias value; Indicates that the current parameter The policy function is used to output a control action based on the current state and the context sequence. The probability distribution of Indicates the number of time steps of the target task; Indicates the current moment; represents the discount factor, , used to control the degree of attention paid to rewards in the future; Indicates time A combination of state, action, and reward at a given moment.

6. The method according to claim 1, characterized in that The obtaining of a target control action corresponding to the current state output by the robot control model includes: Read historical task execution trajectories and preset prior knowledge; the historical task execution trajectories include the current state, control actions, and rewards obtained at multiple time steps; The current state of the robot at the current moment, the historical task execution trajectory and the prior knowledge are input into the strategy function to generate the target control action at the current moment.

7. A context-learning general motion control system driven by a reinforcement learning coach, characterized in that include: The structural parameter sampling module is used to receive the predefined structural parameters of multiple robots and perform random sampling within the preset parameter space to generate multiple robots with different structures; The structural parameters include: structural size length, degree of freedom distribution, number of degrees of freedom, joint working angle, and dynamic parameters; The reinforcement learning coach training module is used to define the target task and establish a simulation environment, allowing the robot to interact with the environment within the simulator and train the reinforcement learning coach corresponding to each target task based on the reward function defined by the target task; The task simulation and trajectory acquisition module is used to use the reinforcement learning coach to guide the robot to perform preset target tasks and record the task execution trajectory corresponding to each target task; the preset target task includes the preset target task definition, state space, action space and reward function; the task execution trajectory includes the current state, control action and reward obtained at each time step; a noise perturbation module, configured to add different degrees of Gaussian noise to the current state and control action of each time step in the task execution trajectory, thereby obtaining the task execution trajectory after being perturbed by Gaussian noise; A context modeling and policy training module is used to extract a fixed-length context sequence from the task execution trajectory using a general context learning framework. The module performs cross-task meta-training with the goal of maximizing the expected cumulative reward under each target task to optimize the context learning model parameters and obtain a pre-trained robot control model. The context learning model is used to output the target control action in the current state based on the context sequence. A robot current state acquisition module inputs the acquired current state of the robot into the pre-trained robot control model; an action reasoning module, configured to obtain a target control action corresponding to the current state output by the pre-trained robot control model; The action execution module is used to determine a target action execution instruction based on the target control action, and the target action execution instruction is used to control the robot to execute the target control action.

8. An electronic device, characterized in that: include: CPU, memory and input / output interfaces; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the context-based learning general motion control method driven by a reinforcement learning coach according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the reinforcement learning coach-driven context-based learning general motion control method according to any one of claims 1 to 6 is performed.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the reinforcement learning coach-driven context-based learning general motion control method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Human indoor behavior intention reasoning method and system based on deep inverse reinforcement learning, electronic equipment and storage medium

    CN122020424A