Method and device for performing dynamic object operation on quadruped robot

By adopting a hierarchical architecture method in the four-legged robot, skill indexes and low-level control commands are generated, and the coordinated control problem of dynamic object operation and complex terrain crossing is solved, and efficient operation in complex environments is achieved.

CN120056132AActive Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510493022.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-30
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing four-legged robots have problems that are difficult to effectively solve in the coordinated control of dynamic object operations and complex terrain crossing, which limits their application potential in complex environments.

Method used

Using a hierarchical architecture including high-level policy networks and low-level skill networks, the high-level policy network generates skill indexes and low-level control commands based on the robot's ontology perception data, target object locations and user instructions, and then generates the robot's target joint position through the low-level skill network, achieving flexible control of robot movement and target object operation.

Benefits of technology

Dynamic object operation under complex terrain is realized, the problem of coordinated control of dynamic object operation and terrain adaptability in the prior art is overcome, and the operation capability of four-legged robots in complex environments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056132A_ABST
    Figure CN120056132A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for a quadruped robot to perform dynamic object operation, relates to the technical field of quadruped robot control, and aims to realize dynamic object operation of the quadruped robot on a complex terrain. The method comprises the following steps: acquiring robot body sensing data, a target object position and a user instruction; generating a skill index and a low-layer control command through a high-layer strategy network according to the robot body sensing data, the target object position and the user instruction; wherein the skill index is used for indicating a selected target low-layer skill, and the low-layer control command represents a control command of the target low-layer skill; generating a target joint position of the robot through a low-layer skill network according to the skill index and the low-layer control command; and controlling the robot to operate the target object on the target terrain according to the target joint position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of quadruped robot control, and particularly to a method and device for a quadruped robot to perform dynamic object manipulation. Background Art

[0002] In the task of quadruped robot dynamic object manipulation, motion control methods based on reinforcement learning have made significant progress in recent years. Related technologies mainly achieve forelimb object manipulation in two ways, namely, static object manipulation is completed through forelimb end-effector trajectory tracking, and end-to-end deep reinforcement learning is used to train the forelimb end-effector to complete dynamic object manipulation tasks; these methods mainly adopt a single-layer policy network architecture for operating static objects on flat terrain.

[0003] In terms of hierarchical reinforcement learning architectures, existing methods mainly include a combined architecture of a task-agnostic low-level controller and a task-related high-level planner, and a multi-skill integration architecture based on residual learning. However, these methods are mainly designed for single-task scenarios and adopt fixed skill-switching logics, making it difficult to effectively solve the cooperative control problem of dynamic object manipulation and complex terrain traversal, thus limiting the application potential of quadruped robots in complex environments. Therefore, there is an urgent need for a cooperative control method that can balance dynamic object manipulation and terrain adaptability. Summary of the Invention

[0004] In view of the above problems, embodiments of this application provide a method and device for a quadruped robot to perform dynamic object manipulation, so as to overcome or at least partially solve the above problems.

[0005] In a first aspect of the embodiments of this application, a method for a quadruped robot to perform dynamic object manipulation is disclosed. The method includes: Obtain robot body perception data, the position of the target object, and a user instruction, where the user instruction represents the specified speed of the target object; Generate a skill index and a low-level control command through a high-level policy network according to the robot body perception data, the position of the target object, and the user instruction; where the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move with a first amplitude, a second skill for moving with a second amplitude, a third skill for controlling the robot to move on flat terrain, and a fourth skill for moving on complex terrain; the low-level control command represents the control command of the target low-level skill; Generate the target joint positions of the robot through a low-level skill network according to the skill index and the low-level control command; Control the robot to manipulate the target object on the target terrain according to the target joint positions.

[0006] In a second aspect of the embodiments of the present application, a device for a quadruped robot to perform dynamic object manipulation is disclosed. The device includes: A first acquisition module, configured to acquire robot body perception data, target object position, and user instructions, where the user instructions represent the specified speed of the target object; A first generation module, configured to generate a skill index and low-level control commands through a high-level policy network based on the robot body perception data, the target object position, and the user instructions; wherein, the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move by a first amplitude, a second skill for moving by a second amplitude, a third skill for controlling the robot to move on flat terrain, and a fourth skill for moving on complex terrain; the low-level control commands represent the control commands of the target low-level skill; A second generation module, configured to generate the target joint positions of the robot through a low-level skill network based on the skill index and the low-level control commands; A first operation module, configured to control the robot to manipulate the target object on the target terrain according to the target joint positions.

[0007] In a third aspect of the embodiments of the present application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method for a quadruped robot to perform dynamic object manipulation described in the first aspect of the embodiments of the present application are implemented.

[0008] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for a quadruped robot to perform dynamic object manipulation described in the first aspect of the embodiments of the present application are implemented.

[0009] In a fifth aspect of the embodiments of the present application, a computer program product is disclosed, including a computer program. When the computer program is executed by a processor, the steps of the method for a quadruped robot to perform dynamic object manipulation described in the first aspect of the embodiments of the present application are implemented.

[0010] The embodiments of the present application have the following advantages: In the embodiments of the present application, a hierarchical architecture including a high-level policy network and a low-level skill network is adopted. The high-level policy network generates a skill index and a low-level control command according to the acquired perception data of the robot body, the position of the target object, and the user instruction. The skill index is used to indicate the selected target low-level skill, and the low-level control command represents the control command of the target low-level skill. Since the low-level skills include a first skill for operating the target object to move by a first amplitude and a second skill for moving by a second amplitude, as well as a third skill for controlling the robot to move on a flat terrain and a fourth skill for moving on a complex terrain, the high-level policy network can perform dynamic policy adjustment according to the real-time terrain features and the state of the target object, and select a suitable control strategy. Then, the low-level skill network generates the target joint positions of the robot according to the skill index and the low-level control command, and controls the robot to operate on the target object on the target terrain according to the target joint positions, so as to realize flexible control of the robot movement and the target object operation based on the dynamically adjusted control strategy. Therefore, through this method, dynamic object operation under complex terrains can be achieved. Description of the Drawings

[0011] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 is a flowchart of the steps of a method for a quadruped robot to perform dynamic object operation provided by the embodiments of the present application; Figure 2 is a schematic diagram of a hierarchical architecture of a high-level policy network and a low-level skill network provided by the embodiments of the present application; Figure 3 is a schematic diagram of the training curve of DSF-PO and standard PPO provided by the embodiments of the present application; Figure 4 is a schematic diagram of the comparison result of the operation success rate on different terrains provided by the embodiments of the present application; Figure 5 is a schematic diagram of the evaluation result of the cross-terrain object operation performance provided by the embodiments of the present application; Figure 6 is a schematic diagram of the usage frequency of low-level skills on different terrains provided by the embodiments of the present application; Figure 7 is a schematic diagram of the actual deployment on different terrains provided by the embodiments of the present application; Figure 8It is a schematic structural diagram of a device for a quadruped robot to perform dynamic object manipulation provided by an embodiment of the present application; Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0013] To make the above objects, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] In the related art, the related technology mainly realizes forelimb object manipulation through two methods: one is to complete static object manipulation through forelimb foot-end trajectory tracking (for example, the leg trajectory interpolation method); the other is to use end-to-end deep reinforcement learning to train the forelimb foot-end to complete dynamic object manipulation tasks (for example, the DribbleBot system and the DexDribbler system); among them, the DribbleBot system realizes the football dribbling ability on flat ground through large-scale parallel training in a simulation environment and combined with domain randomization technology; the DexDribbler system introduces a feedback control term and a context-assisted state estimator to improve the dribbling stability.

[0015] In terms of the hierarchical reinforcement learning architecture, the existing methods mainly include the combined architecture of a task-agnostic low-level controller and a task-related high-level planner, and the multi-skill integration architecture based on residual learning. However, these methods are mostly designed for a single task scenario and adopt a fixed skill switching logic, which is difficult to effectively solve the cooperative control problem of dynamic object manipulation and complex terrain crossing, and limits the application potential of quadruped robots in complex environments.

[0016] Therefore, the related technology has the following limitations: (1) Terrain adaptability limitation: For end-to-end reinforcement learning systems such as DribbleBot and DexDribbler, their policy networks are only optimized for flat terrains during the training process. When facing rough terrains, due to the dynamic coupling of the movement mode and the dribbling mode of the quadruped robot, a single policy is difficult to handle both modes at the same time. In addition, rough terrains will make it difficult for the end-to-end reinforcement learning policy to obtain effective reward signals in the initial stage, which will lead to sparse reward signals and a significant decrease in the policy exploration efficiency.

[0017] (2) Lack of skill coordination mechanism: The traditional hierarchical architecture adopts a fixed skill switching logic and cannot dynamically adjust strategies according to real-time terrain features and the motion state of the sphere. This leads to target conflicts between the two subtasks of movement control and object operation in complex terrain scenarios, reducing the overall task success rate.

[0018] (3) Insufficient optimization of the hybrid action space: When dealing with the discrete-continuous hybrid action space, the existing hierarchical reinforcement learning methods adopt a uniform gradient update strategy, resulting in the update of key skill parameters being interfered by the gradients of inactive skills and significantly prolonging the training convergence time.

[0019] To overcome the limitations of related technologies, the embodiments of the present application provide a method for a quadruped robot to perform dynamic object operation. The method is based on a dynamic skill coordination framework of hierarchical reinforcement learning, which can dynamically adjust strategies according to real-time terrain features and the state of the target object, select appropriate control strategies, and flexibly control the movement of the robot and the operation of the target object based on the dynamically adjusted control strategies. In addition, by combining a gradient optimization mechanism, the efficient dynamic object operation ability of the quadruped robot in complex terrain scenarios is further improved.

[0020] The following will describe the method for a quadruped robot to perform dynamic object operation in the embodiments of the present application with reference to the accompanying drawings, respectively through Section 1.1 Method for a quadruped robot to perform dynamic object operation, Section 1.2 Method for generating skill indices and low-level control commands, Section 1.3 High-level policy network training method, and Section 1.4 Experimental result performance analysis.

[0021] 1.1 Method for a quadruped robot to perform dynamic object operation: Refer to Figure 1 As shown, Figure 1 is a flowchart of the steps of a method for a quadruped robot to perform dynamic object operation provided by the embodiments of the present application. As Figure 1 shown, a method for a quadruped robot to perform dynamic object operation provided by the embodiments of the present application may include steps S110 to S140: Step S110: Obtain the perception data of the robot body, the position of the target object, and the user instruction, where the user instruction represents the specified speed of the target object.

[0022] Step S120: Generate a skill index and a low-level control command based on the robot body perception data, the target object position, and the user instruction through a high-level policy network; wherein, the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move by a first amplitude, a second skill for moving by a second amplitude, a third skill for controlling the robot to move on flat terrain, and a fourth skill for moving on complex terrain; the low-level control command represents the control command of the target low-level skill.

[0023] Step S130: Generate the target joint positions of the robot based on the skill index and the low-level control command through a low-level skill network. Step S140: Control the robot to operate the target object on the target terrain according to the target joint positions.

[0024] In the embodiments of the present application, the robot refers to a quadruped robot, and the target object refers to a dynamic object to be operated. For example, in the scenario where the robot kicks a ball, the target object can be the ball. Considering the dynamic target object operation task on complex terrain, relying solely on a single movement or operation skill cannot achieve good performance. Therefore, a hierarchical architecture including a high-level policy network and a low-level skill network is adopted. The low-level skills of the low-level skill network are mainly responsible for specific target object operation and robot movement control, while the high-level policy network is responsible for achieving smooth switching and coordination between different low-level skills.

[0025] Specifically, the dynamic target object operation task on complex terrain is regarded as a long-time domain task composed of a movement sub-task and a movement operation sub-task. The movement operation sub-task focuses on operating the target object when approaching the target object, while the movement sub-task focuses on approaching the target object or crossing complex terrain when far from the target object. Therefore, four low-level policies (low-level skills), namely the first skill, the second skill, the third skill, and the fourth skill, are used in the low-level skill network, and each low-level policy is applicable to a specific scenario. These low-level policies are trained end-to-end through deep reinforcement learning and used as the low-level skill network, which can be flexibly combined for more complex tasks and actual deployment.

[0026] Among them, the first skill and the second skill are used to operate the target object. Operating the target object means that the robot can operate through its front limbs to enable the quadruped robot to move the target object in a given direction and speed. The first skill and the second skill are applied to operating the target object with different amplitudes (for example, the first amplitude can be that the robot places the target object beside itself, and the second amplitude is that the robot operates the target object away from itself), enabling the robot to effectively achieve the operation of the target object on different terrains. The third skill and the fourth skill are used for the robot to move. Considering that the operation of the target object on complex terrains poses higher requirements on the robot's movement ability and requires better adaptability to slopes and rugged terrains, and these conditions cannot be fully covered in the training of operating the target object. Therefore, skills specifically for the robot to move are introduced. The robot is controlled to move quickly on flat terrains through the third skill and move steadily on complex terrains through the fourth skill.

[0027] The high-level policy network is responsible for achieving smooth switching and coordination among different low-level skills. Specifically, the high-level policy network can generate corresponding low-level actions (that is, generate skill indices and low-level control commands) based on the observation data (i.e., the robot's own perception data, the position of the target object, and the user's instruction), so that the low-level skill network can flexibly control the movement of the robot and the operation of the target object based on the corresponding low-level actions.

[0028] In the specific implementation, first, step S110 is executed to obtain the robot's own perception data, the position of the target object, and the user's instruction. Among them, the robot's own perception data can include: the gravity unit vector, the robot's joint positions and speeds, the low-level actions at the previous moment (i.e., the skill indices and low-level control commands at the previous moment), and the global yaw angle. The robot's own perception data is obtained through the robot's own sensors. The position of the target object refers to the specific position of the target object in the environment, and the position of the target object can be given artificially. The user's instruction represents the specified speed of the target object, that is, the user's instruction specifies the speed of the target object in the global coordinate system, and the user's instruction can be given by the user according to the task requirements.

[0029] Secondly, step S120 is executed to generate a skill index and a low-level control command, that is, to generate corresponding low-level actions, through a high-level policy network based on the robot body perception data, the target object position, and the user instruction. The skill index is used to indicate the selected target low-level skill. For example, if the selected target low-level skill is the first skill, the skill index is the index corresponding to the first skill; if the selected target low-level skill is the second skill, the skill index is the index corresponding to the second skill. The low-level control command represents the control command of the target low-level skill and is used to control the robot to move or control the robot to operate on the target object. The control commands corresponding to different target low-level skills are different. For example, if the target low-level skill is the second skill, the low-level control command is the control command corresponding to the second skill; if the target low-level skill is the third skill, the low-level control command is the control command corresponding to the third skill.

[0030] Next, step S130 is executed to generate the target joint positions of the robot through a low-level skill network based on the skill index and the low-level control command. In some embodiments, the observation space of the low-level skill network includes the low-level control command, part of the robot body perception data, the time reference variable, and the feature task information (such as the target object position). Furthermore, the low-level skill network generates the target joint positions according to the skill index and the low-level control command, and in combination with the observation data in other observation spaces, at the target frequency. For example, the low-level skill network can generate 12-dimensional target joint positions at a frequency of 50 Hz.

[0031] Finally, step S140 is executed to control the robot to operate on the target object on the target terrain according to the target joint positions; where the target terrain can be flat ground, uphill, downhill, rough terrain, going downstairs, etc. Specifically, the target joint positions can be input into a PD (Proportional-Differential) controller to drive the motors of the robot to achieve operating on the target object on the target terrain.

[0032] It can be understood that, in order to balance the switching frequency in the future, the high-level policy network runs at a lower inference frequency than the low-level skill network. For example, if the inference frequency of the low-level skill network is 50 Hz, the inference frequency of the high-level policy network is lower than that of the low-level skill network. For example, the high-level policy network can run at an inference frequency of 10 Hz.

[0033] Through the above implementation process, a hierarchical architecture including a high-level policy network and a low-level skill network is adopted. The high-level policy network generates a skill index and a low-level control command based on the robot's body perception data, the position of the target object, and the user instruction obtained. Among them, the skill index is used to indicate the selected target low-level skill, and the low-level control command represents the control command of the target low-level skill. Since the low-level skills include a first skill for operating the target object to move by a first amplitude and a second skill for moving by a second amplitude, as well as a third skill for controlling the robot to move on flat terrain and a fourth skill for moving on complex terrain, the high-level policy network can perform dynamic policy adjustment according to the real-time terrain features and the state of the target object, and select an appropriate control strategy. Furthermore, the low-level skill network generates the target joint positions of the robot according to the skill index and the low-level control command, and controls the robot to operate on the target object on the target terrain according to the target joint positions, so as to realize flexible control of the robot's movement and the operation of the target object based on the dynamically adjusted control strategy. Therefore, through this method, dynamic object operation under complex terrain can be achieved.

[0034] 1.2 Method for generating a skill index and a low-level control command: Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a method for a quadruped robot to perform dynamic object operation. In this method, the high-level policy network includes: a context-assisted estimation network and a high-level action network.

[0035] The step of "generating a skill index and a low-level control command through the high-level policy network according to the robot's body perception data, the position of the target object, and the user instruction" in the above step S120 may specifically include sub-steps S120-1 to step S120-2: Step S120-1: The context-assisted estimation network performs external state prediction according to the robot's body perception data and the position of the target object to obtain context information, where the context information includes: terrain parameters, robot posture, and the motion state of the target object.

[0036] Step S120-2: The high-level action network performs low-level action generation according to the context information, the robot's body perception data, the position of the target object, and the user instruction to obtain the skill index and the low-level control command.

[0037] Wherein, when the target low-level skill selected as indicated by the skill index is the first skill or the second skill, the low-level control command is an operation command represented by two dimensions, and the operation command includes the horizontal direction speed and the vertical direction speed of the target object; when the target low-level skill selected as indicated by the skill index is the third skill or the fourth skill, the low-level control command is a movement command represented by three dimensions, and the movement command includes the horizontal direction speed and the vertical direction speed of the robot, and the rotation information of the robot around its own z-axis.

[0038] In the embodiments of the present application, the context-assisted estimation network is used for external state prediction, and the context-assisted estimation network can be trained on the real state provided by the simulator through supervised learning. After obtaining the context information through the context-assisted estimation network, the context information, the robot body perception data, the target object position, and the user instruction are used as the input information of the high-level action network together. The high-level action network can generate corresponding low-level actions based on the input information, that is, the action space of the high-level action network includes a skill index and a low-level control command.

[0039] Considering that the dimensions of the required low-level instructions for the target object operation skill and the robot movement skill are different, therefore, in the embodiments of the present application, the output dimension of the low-level control command is set to the sum of these individual dimensions; since the low-level skills in the same category (for example, the first skill and the second skill, the third skill and the fourth skill) have behavioral consistency, the same part of the low-level command can be used. Therefore, the low-level control command is defined as a five-dimensional vector, wherein the first two dimensions are used for the target object operation skill, and the operation command is represented by these two dimensions ( ), and the operation command includes the horizontal direction speed of the target object and the vertical direction speed ; the last three dimensions are used for the movement skill, and the movement command is represented by these three dimensions ( ), and the movement skill includes the horizontal direction speed of the robot and the vertical direction speed , and the rotation information of the robot around its own z-axis .

[0040] Furthermore, the high-level action network includes a shared feature extractor, a skill selector, and a command generator; in step S120-2, "generating low-level actions according to the context information, the robot body perception data, the target object position, and the user instruction through the high-level action network to obtain the skill index and the low-level control command" specifically includes steps A1 to A3: Step A1: The shared feature extractor extracts features from the context information, the robot body perception data, the target object position, and the user instruction to obtain an extraction result.

[0041] Step A2: The skill selector generates a normalized four-dimensional classification distribution based on the extraction result and samples a skill index from the four-dimensional classification distribution.

[0042] Step A3: The command generator generates a five-dimensional continuous vector as the mean of a normal distribution based on the extraction result and samples a low-level control command from the normal distribution.

[0043] In the embodiment of the present application, the shared feature extractor can be a three-layer MLP (Multilayer Perceptron) with an ELU (Exponential Linear Unit) activation function, which is used to extract features from the input information of the high-level action network. The action space of the high-level action network contains discrete skill indexes, including 4 categories. Therefore, in order for the high-level action network to generate discrete skill indexes, the skill selector includes a linear layer and a softmax function (i.e., the normalized exponential function). Thus, a normalized four-dimensional classification distribution can be generated based on the skill selector, that is, the distribution followed by the skill index. By sampling the four-dimensional classification distribution, the skill index can be obtained.

[0044] The action space of the high-level action network contains continuous low-level control commands. Therefore, in order for the high-level action network to generate continuous low-level control commands, the command generator includes a linear layer and a tanh activation function (i.e., the hyperbolic tangent function). Based on the command generator, a five-dimensional continuous vector with a range of can be output as the mean of the normal distribution to sample the low-level control command therefrom.

[0045] Through the above implementation process, a skill index can be generated according to the real-time terrain features and the target object state to indicate the selected target low-level skill, and a control command for the target low-level skill can be generated. Moreover, the low-level control command is defined by five dimensions in the action space of the high-level action network. For different types of low-level skills, different-dimensional low-level control commands are generated for control. In this way, flexible control of the robot movement and the target object operation is achieved.

[0046] As Figure 2 shown, Figure 2It is a schematic diagram of a hierarchical architecture of a high-level policy network and a low-level skill network provided by an embodiment of the present application. Among them, the high-level policy network includes a context-assisted estimation network and a high-level action network, and the high-level action network includes a shared feature extractor, a skill selector, and a command generator. The context-assisted estimation network performs external state prediction based on the robot's own perception data and the position of the target object to obtain context information, and inputs the context information, the robot's own perception data, the position of the target object, and the user instruction into the high-level action network. The shared feature extractor in the high-level action network performs feature extraction, generates a skill index through the skill selector, and generates a low-level control command through the command generator, so as to achieve smooth switching and coordination between different low-level skills; while the low-level skill network realizes the operation of the target object and the movement control of the robot according to the corresponding target low-level skill and the low-level control command branch. Therefore, dynamic object operation in complex terrains is realized based on this architecture.

[0047] 1.3 High-level policy network training method: In the embodiment of the present application, considering that in the training of the high-level policy network, due to the asymmetry of skill execution, the standard Proximal Policy Optimization (PPO) algorithm may not be optimal. For example, when the skill index corresponds to the first skill or the second skill, only the first two dimensions of the low-level control command are used to operate the target object, but all dimensions will generate gradient updates. This mechanism has two main problems: 1) The unused command dimensions (for example, the movement command when operating the target object) will still generate gradients, interfering with the optimization process and reducing the convergence speed; 2) The high-level policy network wastes the ability to optimize these irrelevant parameters, reducing the overall learning efficiency.

[0048] To solve the above problems, the embodiment of the present application provides a reinforcement learning method for the high-level policy network. This method uses a novel form of surrogate loss function, namely Dynamic Skill-Focused Policy Optimization (DSF-PO), which dynamically adjusts the optimization weights according to the probability of skill selection to ensure that the high-level policy network only optimizes the relevant command dimensions corresponding to the selected skills. The training method of the high-level policy network will be described in the following two parts (1.3.1 and 1.3.2).

[0049] 1.3.1 Theoretical derivation of DSF-PO: First, define the high-level action network as , and the high-level action network has two output heads, a discrete skill selector and a continuous command generator; among them, the skill selector is represented as , and the skill selector outputs the skill index as , used to indicate the selected target low-level skill; the command generator represents , and the command generator represents The output command vector is the union of all skill commands, that is , where is the command corresponding to skill , and the dimension is . This modeling corresponds to the hierarchical learning problem: once the skill index is selected, the low-level skill network only processes the corresponding command , and discards other commands.

[0050] Specifically, the four-dimensional categorical distribution that the low-level skill index follows can be expressed as:

[0051] Among them, the probability of selecting the low-level skill is which can be expressed as:

[0052] Among them, corresponds to the unnormalized logits of the low-level skill.

[0053] The selected low-level skill The corresponding low-level command follows a multivariate normal distribution and can be expressed as:

[0054] Therefore, the high-level action network can be decomposed into:

[0055] Among them, is the indicator function, indicating that only the command vector corresponding to the low-level skill d is selected .

[0056] Therefore, the skill focus weight value is defined as:

[0057] Among them, the skill focus weight value represents the probability of selecting the low-level skill in the state , so the importance ratio in the original Proximal Policy Optimization (PPO) algorithm can be rewritten as:

[0058] Among them, Indicates the probability of selecting the target low-level skill in the current state , Indicates the probability of selecting the target low-level skill in the previous training , Indicates the current low-level control command , Indicates the low-level control command of the previous training.

[0059] Therefore, the agent loss function of DSF-PO can be expressed as:

[0060] Among them, is the advantage estimation at time , is the clipping parameter of PPO.

[0061] The policy gradient generated by it can be expressed as:

[0062] It can be seen from this agent loss function that the greater the probability that the current policy tends to select the low-level skill k, the greater the gradient weight of the corresponding command parameter. In this way, based on the DSF-PO method, the high-level policy network is trained, and the optimization weight can be dynamically adjusted according to the probability of low-level skill selection, ensuring that the high-level policy network only optimizes the relevant command dimensions corresponding to the selected skills, and guaranteeing the performance of the high-level policy network.

[0063] 1.3.2 Training process of the high-level policy network based on DSF-PO: Combined with the above embodiments, in one implementation manner, the embodiments of the present application also provide a method for a quadruped robot to perform dynamic object operation. In this method, the high-level policy network is trained based on the above DSF-PO to solve the problems of skill switching and training stability in the high-level policy network. Specifically, the high-level policy network is obtained through reinforcement learning according to the following steps S210 to step S240: Step S210: Determine the current terrain difficulty and the current user instruction required for the current training.

[0064] Specifically, in order to train the robot's target object operation ability in a complex environment, five terrains are designed, namely flat ground, uphill, downhill, rough terrain, and going down stairs, and the difficulty of the terrain is controlled. The terrain difficulty refers to the steepness and ruggedness of going down stairs, uphill, downhill, and rough terrain. At the beginning of each training, the current terrain difficulty required for the current training is determined, and the robot is randomly placed on one of the above terrains that meets the current terrain difficulty for training. The current user instruction is the specified target object speed, which can be randomly given by the user or given according to the training situation of the robot.

[0065] Step S220: Generate an action under the current terrain difficulty according to the current robot body perception data, the current target object position, and the current user instruction. The action includes a current skill index and a current low-level control command.

[0066] Among them, the current robot body perception data is obtained through the robot's body sensors. At the beginning of training, after randomly placing the robot on the terrain with the corresponding difficulty, the robot is initialized simultaneously, that is, initialized with a random yaw angle, and its joint angles are randomly initialized around the standard posture, and then the current robot body perception data is determined through the body sensors. The current target object position can be a position where the target object is randomly placed within a radius of N meters (for example, 2 meters) around the robot.

[0067] During training, the current robot body perception data, the current target object position, and the current user instruction are used as the input data of the high-level policy network and input into the high-level policy network for processing. After the shared feature extractor in the high-level policy network performs feature extraction, the skill selector generates a normalized four-dimensional classification distribution according to the extraction result, samples a current skill index from the four-dimensional classification distribution, the command generator generates a five-dimensional continuous vector as the mean of the normal distribution according to the extraction result, and samples a current low-level control command from the normal distribution, so as to obtain an action under the current terrain difficulty.

[0068] Exemplarily, the action under the current terrain difficulty The specific sampling process can be expressed as:

[0069]

[0070] Among them, the target low-level skill corresponding to the current skill index is one of the first skill, the second skill, the third skill, and the fourth skill, and the current low-level control command is the specific control command corresponding to the target low-level skill.

[0071] Step S230: Determine the current skill focus weight value according to the action, and determine the importance ratio according to the current skill focus weight value. The skill focus weight value represents the probability of selecting the target low-level skill in the current state, and the skill focus weight value is positively correlated with the importance ratio.

[0072] In the embodiments of the present application, the probability of selecting a target low-level skill in the current state can be determined according to the probability distribution followed by the skill index. The probability of selecting a target low-level skill in the current state is used as the current skill focus weight value. That is, the greater the probability of selecting a target low-level skill in the current state, the greater the current skill focus weight value; the smaller the probability of selecting a target low-level skill in the current state, the smaller the current skill focus weight value.

[0073] Specifically, determining the current skill focus weight value according to the action and determining the importance ratio according to the current skill focus weight value includes: determining the probability of selecting a target low-level skill in the current state according to the four-dimensional classification distribution followed by the skill index, and defining the probability of selecting a target low-level skill in the current state as the current skill focus weight value; determining the importance ratio according to the ratio of the probability of selecting a target low-level skill in the current state to the probability of selecting a target low-level skill in the previous training, the ratio of the current low-level control command to the low-level control command in the previous training, and the current skill focus weight value.

[0074] Exemplarily, the four-dimensional classification distribution followed by the low-level skill index can be expressed as:

[0075] Among them, the probability of selecting a low-level skill is which can be expressed as:

[0076] Among them, corresponds to the unnormalized logits of the skill.

[0077] Therefore, for the low-level skill k, the skill focus weight value in the state is expressed as:

[0078] Among them, the first skill and the second skill share a subset of the command output, that is , where is the command set related to the skill . Therefore, the importance ratio can be expressed as:

[0079] Among them, represents the probability of selecting the target low-level skill in the current state, represents the probability of selecting the target low-level skill in the previous training The probability represents the current low-level control command , and represents the low-level control command of the previous training.

[0080] Step S240: Determine the surrogate loss value according to the importance ratio, and update the parameters of the high-level action network in the high-level policy network according to the surrogate loss value. When the training end condition is met, the trained high-level policy network is obtained.

[0081] In the embodiment of the present application, determining the surrogate loss value according to the importance ratio may be determined according to the importance ratio and the reward value, where the reward value is calculated according to the current action and the current state. During training, the parameters of the context-aided estimation network in the high-level policy network are frozen, and only the parameters of the high-level action network in the high-level policy network are updated. Specifically, the gradient of the command parameter is determined according to the surrogate loss value, and then the parameters of the high-level action network in the high-level policy network are updated according to the gradient. Among them, the training end condition may be that the number of training times reaches the training times threshold, or the performance of the high-level policy network meets the preset requirements.

[0082] Through the above implementation process, the probability of selecting the target low-level skill in the current state is used as the current skill focus weight value. Therefore, according to the current skill focus weight value, the importance ratio can be determined, and the optimization weight can be dynamically adjusted according to the probability of skill selection, ensuring that the high-level policy network only optimizes the relevant command dimensions corresponding to the selected target low-level skill 1.3.2.1 Reward items in reinforcement learning: In some embodiments, "determining the surrogate loss value according to the importance ratio" in the above step S240 may specifically include sub-steps S240-1 to step S240-3: Step S240-1: Calculate the reward value corresponding to the reward item according to the current state and the current action; wherein, the current state at least includes the real speed of the robot and the real speed of the target object, and the reward item includes: encouraging the robot to maintain balance, encouraging the robot to approach and face the target object, encouraging the robot to continuously output the same skill index, rewarding the actual speed of the target object to be close to the expected speed, and encouraging the robot to use the first skill or the second skill when approaching the target object.

[0083] Step S240-2: Calculate the advantage estimation according to the reward value, and the advantage estimation represents the goodness or badness of the current action.

[0084] Step S240-3: Determine the surrogate loss value according to the advantage estimation and the importance ratio.

[0085] In the embodiments of the present application, in order to minimize the attention of the high-level policy network to the low-level skills of the robot, as few reward items as possible are used during training. The reward items mainly include the above-mentioned 5 categories; among them, encouraging the robot to maintain balance includes the gravity projection amount; encouraging the robot to approach and face the target object includes: the distance between the robot and the target object, and the yaw angle alignment; encouraging the robot to continuously output the same skill index includes: using a temporally consistent skill index and using a temporally varying skill index; rewarding the actual speed of the target object to be close to the desired speed includes: the target object speed norm, the target object speed angle, and the target object speed error; encouraging the robot to use the first skill or the second skill when approaching the target object includes near-target object operations.

[0086] To ensure stable training, a bounded summation function is used to aggregate all reward items, and an exponential kernel is used for weighted integration to obtain the advantage estimate (i.e., the advantage estimate is the goodness or badness of the current action compared to the average policy). Among them, the expression and weight value of each reward item are shown in Table 1.

[0087] Table 1 Expression and weight value of reward items

[0088] Among them, is the projection magnitude of the gravity vector on the x-y plane, is the exponential factor, is the target object coordinate, is the coordinate of the right front thigh joint of the quadruped robot, is the exponential factor, is the yaw angle difference between the quadruped robot and the target object, is the yaw angle difference between the quadruped robot and the desired speed of the target object, is the skill index at time i, is the skill index at time i-1, is the skill index at time t, is the skill index at time t-1, is the exponential factor, is the desired speed of the target object, is the actual speed of the target object, is the actual speed angle of the target object, is the desired speed angle of the target object, is the exponential factor.

[0089] Finally, the agent loss function can be expressed as:

[0090] Among them, is the time Advantage estimation is the clipping parameter for PPO.

[0091] Through the above implementation process, as few reward terms as possible are used during training, and the advantage estimation is calculated using the reward values of the above five types of reward terms, and then the agent loss value is calculated, so that the high-level policy should be able to reduce the attention to the low-level actions of the robot.

[0092] 1.3.2.1 Curriculum learning to determine the current terrain difficulty and the current user instruction: In some embodiments, the step of "determining the current terrain difficulty and the current user instruction required for the current training" in the above step S210 may specifically include sub-steps S210-1 to step S210-2: Step S210-1: Determine the joint distribution of the current training user instruction and the terrain difficulty by using the method of curriculum learning, and the joint distribution represents the sampling range of the terrain difficulty and the user instruction for the current training; wherein, the terrain difficulty is used to control the steepness and ruggedness of going downstairs, uphill, downhill and rugged terrain.

[0093] Step S210-2: Sample the joint distribution to obtain the current terrain difficulty and the current user instruction.

[0094] In the embodiment of the present application, since the high-level policy network is trained from scratch, the use of curriculum learning can significantly improve the stability of the training process. Therefore, terrain curriculum learning is used to help the robot adapt to complex terrains. Specifically, the terrain difficulty is used as a parameter of the terrain curriculum to control the steepness and ruggedness of going downstairs, uphill, downhill and rugged terrain respectively; at the same time, instruction curriculum learning is used to improve the robot's ability to follow a wider range of user instructions.

[0095] Specifically, during the sampling process of the kth training episode, the user instruction and the joint distribution of the terrain difficulty t is , and the distribution of the user instruction is multiplied by the distribution of the terrain difficulty to obtain the joint distribution . During the training process, the box adaptive curriculum update rule based on the reward method can be adopted. After obtaining the joint distribution of the current training user instruction and the terrain difficulty, the current terrain difficulty and the current user instruction are obtained by sampling the joint distribution. Specifically, the initial distribution is initialized to a uniform distribution , and in each training episode k, samples are independently drawn from the distribution to obtain the current user instruction and the terrain difficulty t.

[0096] In some alternative embodiments, the method updates the joint distribution according to the following steps: calculating a first reward value for the robot to follow the user's instructions, and calculating a second reward value for the robot to cross the current terrain; updating the distribution of the user's instructions obedience according to the first reward value, and updating the distribution of the terrain difficulty obedience according to the second reward value; obtaining an updated joint distribution of the user's instructions and the terrain difficulty according to the updated distribution of the user's instructions obedience and the updated distribution of the terrain difficulty obedience, and the updated joint distribution of the user's instructions and the terrain difficulty is used to determine the terrain and user's instructions required for the next training. Exemplarily, the calculation methods of the first reward value and the first reward value are shown in Table 2.

[0097] Table 2 Calculation methods of reward values

[0098] Wherein, is the speed error of the target object, and are set thresholds.

[0099] Specifically, updating the distribution of the user's instructions obedience according to the first reward value, and updating the distribution of the terrain difficulty obedience according to the second reward value includes: when the first reward value represents that the robot completes the task of following the user's instructions under the current joint distribution, expanding the distribution of the user's instructions obedience to the neighboring area; when the second reward value represents that the robot completes the task of crossing the current terrain under the current joint distribution, expanding the distribution of the terrain difficulty obedience to the neighboring area.

[0100] That is to say, if the robot successfully completes the task within the current distribution area, the sampling distribution is expanded to the neighboring area, thereby increasing the difficulty of the task. Therefore, the update rules for the distribution of the user's instructions obedience and the distribution of the terrain difficulty obedience can be expressed as:

[0101] Wherein, represents the user's instruction distribution for the (k + 1)-th training, represents the terrain difficulty distribution for the (k + 1)-th training; the neighborhood of the user's instruction and the neighborhood of the terrain difficulty have an increased sampling probability density, and the neighborhood is defined as: .

[0102] Through the above implementation process, the terrain difficulty and user instructions for each training are determined through curriculum learning, enabling the high-level policy network to gradually learn complex terrains to help the robot gradually adapt to complex terrains. At the same time, curriculum learning of instructions is carried out to improve the robot's ability to follow a wider range of user instructions.

[0103] In summary, the method for a quadruped robot to perform dynamic object manipulation provided by the embodiments of this application, based on a hierarchical reinforcement learning-based dynamic skill coordination framework and combined with a gradient optimization mechanism, realizes the efficient dynamic object manipulation ability of the quadruped robot in complex terrain scenarios. Specifically, this application can solve the following problems: (1) Solved the problem of dynamic operation failure in complex terrains: enabling the quadruped robot to simultaneously meet the dual-modal task requirements of stable movement control and dynamic object manipulation in rough terrain scenarios, and overcoming the problem of sparse reward signals caused by the terrain-dribbling coupling effect in existing end-to-end policies.

[0104] (2) Solved the problem of the lack of a multi-skill dynamic coordination mechanism: constructed an adaptive policy switching mechanism based on real-time environment perception, solved the problem of target conflicts caused by the fixed skill switching logic in traditional hierarchical architectures, and realized the dynamic priority adjustment of the movement and manipulation subtasks.

[0105] (3) Solved the problem of low optimization efficiency in the hybrid action space: this method dynamically adjusts the optimization weights according to the probability of skill selection, ensuring that the policy network only optimizes the relevant command dimensions corresponding to the selected skills. Therefore, the coupling interference between discrete skill selection and continuous command adjustment is eliminated through this gradient update strategy, improving the training convergence speed and stability of hierarchical reinforcement learning in hybrid action space scenarios.

[0106] 1.4 Experimental Result Performance Analysis: In this section, taking the robot's operation of a sphere as an example, the learning performance of training the high-level policy using Proximal Policy Optimization (PPO) is demonstrated through experiments. Both simulation and training are carried out on the Isaac Gym platform equipped with an NVIDIA RTX 3090 Ti. At the same time, an ablation study is also conducted to evaluate the performance improvement of the DSF-PO method proposed in the embodiments of this application compared with the standard PPO algorithm. The two methods use the same hyperparameters, and the only difference is the calculation method of the loss function.

[0107] As Figure 3 shown, Figure 3 is a schematic diagram of the training curves of DSF-PO and the standard PPO provided by the embodiments of this application, Figure 3The performance of two methods in 12,000 iterations of training is shown. The shaded area represents the standard deviation of multiple experiments. The total reward value and the duration of a single episode (episode length) are used as evaluation metrics. From the results, it can be seen that the PPO algorithm with the DSF-PO loss function achieved higher reward values and longer episode lengths, and its performance steadily improved during the training process, while the standard PPO algorithm encountered a performance bottleneck earlier. This indicates that DSF-PO improves the exploration ability by dynamically optimizing the command dimension according to the skill selection probability, effectively avoiding premature convergence and making the learning of the policy more robust.

[0108] To evaluate the terrain traversal ability of the trained policy, the proposed method is compared with DribbleBot and DexDribbler. Since the publicly available weight models of the aforementioned methods were trained on the Unitree Go1 robot, the URDF model of the Go1 robot is also used in this scheme for fair testing. All policies use the same control commands, and 100 tests are conducted for each method on each terrain respectively. The success is defined as the robot being able to reach the end boundary of the terrain. As Figure 4 shown, Figure 4 is a schematic diagram of the comparison results of the operation success rates on different terrains provided by an embodiment of this application, Figure 4 which shows the success rates of DribbleBot (DB), DexDribbler (DD) and this scheme evaluated on various terrains: flat ground, uphill (slope = 0.1), downhill (slope = -0.1), rough terrain and going downstairs (height = 0.05m, width = 0.5m). It can be seen that all methods achieved a 100% success rate on flat ground; however, in more challenging terrains, the performance differences among the methods were significant. The method proposed in this application was superior to the comparison methods in tests such as uphill (92%), downhill (97%), rough terrain (80%) and going downstairs (78%). These results highlight the stronger adaptability and robustness of this scheme under diverse terrain conditions.

[0109] As Figure 5 shown, Figure 5 is a schematic diagram of the evaluation results of the cross-terrain object operation performance provided by an embodiment of this application, Figure 5 in which (a) is a schematic diagram of the trajectory of the robot dribbling through five terrains in sequence, with each terrain being 10 m long, Figure 5 and (b) in it is the visualization of different low-level skill calls, where each thin vertical line represents one call, Figure 5Visualization of the speed magnitude and direction of the (c-d) ball. Specifically, the ball control ability of the robot was evaluated on five terrains: going down stairs, going down slopes, rough terrain, going uphill, and flat ground. Given a fixed instruction, the robot successfully completed the task within 121 seconds. On most terrains, the robot was able to maintain a stable movement direction. Only on rough terrain, due to external disturbances, there was a certain deviation in the movement direction. Figure 5 The (b) in [reference] shows the situation where the high-level policy invokes low-level skills under different terrain conditions. The results show that the high-level policy uses the ball control skill more frequently and can adaptively switch between different ball control skills according to terrain changes. In contrast, the movement skill is called less frequently and is mainly used for repositioning and attitude adjustment after the robot separates from the football.

[0110] To systematically analyze the differences in the use of low-level skills on different terrains, 10,000 steps of data were collected on each terrain, and the skill indices selected by the high-level policy were counted. The statistical distribution results are as Figure 6 shown, Figure 6 The numbers in [figure] represent the proportion of the use frequency of each low-level skill on the given terrain. This reveals significant behavioral patterns of the high-level policy under different terrain conditions: 1) Flat ground: The high-level policy mainly selects the second skill , because the training environment of the second skill matches the flat ground environment, thus ensuring the stability of ball control. 2) Going down stairs: The high-level policy uses the movement skill (the third skill or the fourth skill) more frequently to help the robot handle the discontinuous features on the terrain. 3) Going uphill and going downhill: The skill distributions are almost the same, indicating that the control strategies for the two terrain environments are highly consistent. 4) Rough terrain: The high-level policy uses the fourth skill more to adapt to the irregular ground and maintain stability.

[0111] In addition, actual deployment experiments were carried out on different terrains. The indoor environment included going uphill, going downhill, and going down stairs, while the outdoor environment consisted of complex terrains. In this section, the Unitree Go2 quadruped robot was used for the actual deployment experiment. An additional downward-mounted fisheye camera with a 240° field of view was installed on the robot's head. All policy inferences were completed on the NVIDIA Jetson Orin NX platform carried by the robot. To estimate the position of the football, we trained a YOLOv11 network for ball detection. The data from the front camera and the downward-looking fisheye camera were fused to improve the football positioning accuracy. Using the fisheye equidistant projection model, the angular information in the image was converted into distance, and two methods were used to calculate the two-dimensional coordinates of the football: one method used the actual diameter information of the ball; the other method was based on the distance and angle measured by the camera. Finally, these estimates were fused through Kalman filtering to further improve the positioning accuracy.

[0112] In the strategy training and evaluation phase, the parameters used are PD controller; while in mobile skills Different PD control parameters are set in All low-level and high-level strategies are directly deployed on the physical robot using zero-shot transfer.

[0113] Specifically, real machine evaluation was carried out on four different terrains, such as Figure 7 As shown, Figure 7 This is a schematic diagram of actual deployment on different terrains provided by an embodiment of the present application. The uphill, downhill and downstairs scenes are all built indoors. In the downstairs scene, each step is 50 cm wide and 5 cm high. The cross-terrain scene includes irregular ground, raised curbs, smooth slate roads, and deformable gravel areas.

[0114] At the beginning of each experiment, the robot and the football were initialized and placed on a flat ground. The goal was to successfully push the football to the other end of the field while ensuring that the robot maintained balance. High-level instructions were issued in real time through the remote control. Five tests were conducted on each terrain. The statistical results are shown in Table 3. The results show that the method of the embodiment of the present application achieved the highest success rate in all four terrains, reflecting excellent generalization ability. Further experimental analysis shows that the high-level strategy of the embodiment of the present application enables the robot to dynamically adjust the walking and ball control skills according to the terrain changes, that is, on a flat ground, the robot maintains a stable ball control gait with only a small amount of adjustment; in a slope environment, the robot will actively adjust the coordination of the limbs and the body posture to offset the slope effect and stabilize the ball control; in the downstairs scene, the robot first kicks the football to the next step, then carefully follows the football forward, and adjusts the gait in time to maintain stability and balance; in the rugged outdoor terrain, the robot actively adjusts the stride and standing width, flexibly switches between different skills, effectively crosses obstacles and maintains a stable ball control state.

[0115] The above experimental results show that the method of dynamic object manipulation by a quadruped robot proposed in the embodiment of the present application achieves successful ball control performance of the robot in an actual environment, and is highly consistent with the behavior in a simulation environment, verifying the effectiveness of this solution in actual deployment.

[0116] Table 3 Real-world dribbling performance evaluation

[0117] The present application also provides a device for a quadruped robot to manipulate dynamic objects, referring to Figure 8 As shown, Figure 8It is a schematic structural diagram of a device for a quadruped robot to perform dynamic object operation provided by an embodiment of the present application. The device includes: A first acquisition module 810, configured to acquire robot body perception data, target object position, and user instructions, where the user instructions represent the specified speed of the target object; A first generation module 820, configured to generate a skill index and a low-level control command through a high-level policy network according to the robot body perception data, the target object position, and the user instructions; wherein, the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move by a first amplitude and a second skill for moving by a second amplitude, as well as a third skill for controlling the robot to move on a flat terrain and a fourth skill for moving on a complex terrain; the low-level control command represents the control command of the target low-level skill; A second generation module 830, configured to generate the target joint position of the robot through a low-level skill network according to the skill index and the low-level control command; A first operation module 840, configured to control the robot to operate on the target object on the target terrain according to the target joint position.

[0118] An embodiment of the present application also provides an electronic device. Refer to Figure 9 , Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 9 shown, the electronic device 900 includes: a memory 910 and a processor 920. The memory 910 is communicatively connected to the processor 920 through a bus. A computer program is stored in the memory 910, and the computer program can run on the processor 920, thereby implementing the steps of the method for a quadruped robot to perform dynamic object operation described in the embodiments of the present application.

[0119] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for a quadruped robot to perform dynamic object operation described in the embodiments of the present application are implemented.

[0120] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method for a quadruped robot to perform dynamic object operation described in the embodiments of the present application are implemented.

[0121] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0122] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods and apparatuses according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.

[0123] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0124] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0125] The above has introduced in detail a method and an apparatus for a quadruped robot to perform dynamic object operations provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for a quadruped robot to manipulate dynamic objects, characterized in that: The method comprises: Acquire robot proprioceptive perception data, a target object position, and a user instruction, wherein the user instruction represents a specified target object speed; Generate a skill index and a low-level control command through a high-level strategy network according to the robot body perception data, the target object position and the user instruction; wherein the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move according to a first amplitude and a second skill for moving according to a second amplitude, and a third skill for controlling the robot to move on a flat terrain and a fourth skill for moving on a complex terrain; the low-level control command represents the control command of the target low-level skill; generating a target joint position of the robot according to the skill index and the low-level control command through a low-level skill network; The robot is controlled to operate the target object on the target terrain according to the target joint position.

2. The method according to claim 1, characterized in that The high-level strategy network includes: a context-assisted estimation network and a high-level action network; generating a skill index and a low-level control command according to the robot body perception data, the target object position and the user instruction through the high-level strategy network, including: The context-assisted estimation network predicts external states according to the robot proprioception data and the position of the target object, and obtains context information, wherein the context information includes: terrain parameters, robot posture, and motion state of the target object; Generate low-level actions through the high-level action network according to the context information, the robot body perception data, the target object position and the user instruction to obtain the skill index and the low-level control command; Among them, when the skill index indicates that the selected target low-level skill is the first skill or the second skill, the low-level control command is an operation command represented by two dimensions, and the operation command includes the horizontal speed and the vertical speed of the target object; when the skill index indicates that the selected target low-level skill is the third skill or the fourth skill, the low-level control command is a movement command represented by three dimensions, and the movement command includes the horizontal speed and the vertical speed of the robot, and the rotation information of the robot around its own z-axis.

3. The method according to claim 2, characterized in that The high-level action network includes a shared feature extractor, a skill selector and a command generator; the high-level action network generates low-level actions according to the context information, the robot body perception data, the target object position and the user instruction to obtain the skill index and the low-level control command, including: Performing feature extraction on the context information, the robot proprioception data, the target object position and the user instruction by the shared feature extractor to obtain an extraction result; Generating a normalized four-dimensional classification distribution according to the extraction result by the skill selector, and sampling from the four-dimensional classification distribution to obtain the skill index; The command generator generates a five-dimensional continuous vector as the mean of the normal distribution according to the extraction result, and obtains the low-level control command by sampling from the normal distribution.

4. The method according to any one of claims 1 to 3, characterized in that: The high-level policy network is obtained by reinforcement learning according to the following steps: Determine the current terrain difficulty and current user instructions required for the current training; Generate an action under the current terrain difficulty according to the current robot proprioception data, the current target object position and the current user instruction, wherein the action includes a current skill index and a current low-level control command; Determine a current skill focus weight value according to the action, and determine an importance ratio according to the current skill focus weight value, wherein the skill focus weight value represents a probability of selecting a target low-level skill in a current state, and the skill focus weight value is positively correlated with the importance ratio; The proxy loss value is determined according to the importance ratio, and the parameters of the high-level action network in the high-level policy network are updated according to the proxy loss value, and a trained high-level policy network is obtained when the training end condition is met.

5. The method according to claim 4, characterized in that Determining a current skill focus weight value according to the action, and determining an importance ratio according to the current skill focus weight value, including: Determine the probability of selecting the target low-level skill in the current state according to the four-dimensional classification distribution obeyed by the skill index, and define the probability of selecting the target low-level skill in the current state as the current skill focus weight value; The importance ratio is determined based on the ratio of the probability of selecting the target low-level skill in the current state to the probability of selecting the target low-level skill in the last training, the ratio of the current low-level control command to the low-level control command in the last training, and the current skill focus weight value.

6. The method according to claim 4, characterized in that Determining a proxy loss value according to the importance ratio includes: Calculate a reward value corresponding to a reward item according to the current state and the current action; wherein the current state includes at least the real speed of the robot and the real speed of the target object, and the reward items include: encouraging the robot to maintain balance, encouraging the robot to approach and face the target object, encouraging the robot to continuously output the same skill index, rewarding the actual speed of the target object to be close to the expected speed, and encouraging the robot to use the first skill or the second skill when approaching the target object; Calculating an advantage estimate according to the reward value, wherein the advantage estimate represents the goodness of the current action; A proxy loss value is determined based on the advantage estimate and the importance ratio.

7. The method according to claim 4, characterized in that Determines the current terrain difficulty and current user instructions required for the current session, including: The joint distribution of the current training user instructions and the terrain difficulty is determined by curriculum learning, and the joint distribution represents the sampling range of the current training terrain difficulty and the user instructions; wherein the terrain difficulty is used to control the steepness and ruggedness of going down stairs, uphill, downhill and rugged terrain; The joint distribution is sampled to obtain the current terrain difficulty and the current user instruction.

8. The method according to claim 7, characterized in that The method further comprises: Calculating a first reward value for the robot following a user instruction, and calculating a second reward value for the robot traversing a current terrain; updating the distribution of user instruction compliance according to the first reward value, and updating the distribution of terrain difficulty compliance according to the second reward value; According to the distribution of updated user instructions and the distribution of updated terrain difficulty, a joint distribution of updated user instructions and terrain difficulty is obtained, and the joint distribution of updated user instructions and terrain difficulty is used to determine the terrain and user instructions required for the next training.

9. The method according to claim 8, characterized in that The method includes updating the distribution of user instruction compliance according to the first reward value, and updating the distribution of terrain difficulty compliance according to the second reward value, including: When the first reward value represents that the robot completes the task of following the user's instruction under the current joint distribution, expanding the distribution of the user's instruction compliance to an adjacent area; When the second reward value represents that the robot completes the task of crossing the current terrain under the current joint distribution, the distribution of the terrain difficulty is expanded to the adjacent area.

10. A device for a quadruped robot to manipulate dynamic objects, characterized in that: The device comprises: A first acquisition module is used to acquire robot body perception data, target object position and user instructions, wherein the user instructions represent a specified target object speed; A first generation module is used to generate a skill index and a low-level control command according to the robot body perception data, the target object position and the user instruction through a high-level strategy network; wherein the skill index is used to indicate the selected target low-level skill, and the low-level skills include a first skill for operating the target object to move according to a first amplitude and a second skill for moving according to a second amplitude, and a third skill for controlling the robot to move on a flat terrain and a fourth skill for moving on a complex terrain; the low-level control command represents the control command of the target low-level skill; A second generation module is used to generate a target joint position of the robot according to the skill index and the low-level control command through a low-level skill network; The first operating module is used to control the robot to operate the target object on the target terrain according to the target joint position.

Citation Information

Patent Citations

  • Quadruped robot multi-skill motion control method and system and medium

    CN113110442A

  • Motion control method for quadruped robot with damaged legs

    CN118915802A

  • Robot control system and method, storage medium, controller and robot

    CN118927246A

  • Interaction system of quadruped robot, planning control method and medium

    CN119148751A

  • Model training method and related device

    WO2023246819A1