Robot humanoid impedance control method and system based on multimodal signal fusion
Through the robotic human impedance control method of multimodal signal fusion, the problem of low control accuracy of traditional smart hands is solved, and the precise grasp of objects of different stiffness by smart hands is realized, especially the protection of fragile objects.
Patent Information
- Application Number
- CN202310154443.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-02-23
AI Technical Summary
In the prior art, traditional smart hand control methods use unknown object dynamics models to identify, resulting in low contact force control accuracy and difficult to reach human-like impedance adjustment level.
The robot human impedance control method based on multimodal signal fusion is adopted, including the imitation learning stage, the reinforcement learning stage and the behavior control stage, obtaining multimodal signals through sensors, using closed-loop kinematics algorithm and motion redirection method to train the Gaussian hybrid model, constructing a reinforcement learning model, and realizing human-like impedance control of robot joint position and contact force.
It improves the precise controllability of the contact force and relative position of the skilled hands when grabbing objects of different stiffness, especially the protection of fragile objects, and improves the control accuracy.
Smart Images

Figure CN116300436B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot control, and in particular relates to a robot humanoid impedance control method and system based on multimodal signal fusion. Background Art
[0002] With the successful application of robotics in industrial environments, the application of service robots in daily life has attracted widespread attention. Multi-finger dexterous hands, as robotic end-effectors, can be used for humanoid dexterous manipulation. A safe and stable grip force regulation strategy is key to ensuring humanoid manipulation of multi-finger dexterous hands. When dealing with objects of varying stiffness in real life, dexterous hands require not only excellent position control accuracy but also precise and anthropomorphic control of the contact force between the fingertips and the object.
[0003] Currently, traditional dexterous hand control methods rely on identifying the dynamic model of an unknown object and manually adjusting the contact force and position by setting an impedance relationship between them. These methods often employ simplified ideal models, resulting in low contact force control accuracy and difficulty achieving human-like impedance control. Biomechanical research has shown that humans can subconsciously adjust muscle stiffness to suit specific tasks under the control of the central nervous system when manipulating objects. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a robot humanoid impedance control method and system based on multimodal signal fusion, which solves the problem of low contact force control accuracy caused by dynamic model identification in the prior art.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] The robot humanoid impedance control method based on multimodal signal fusion includes imitation learning stage, reinforcement learning stage and behavior control stage; wherein,
[0007] In the imitation learning phase, sensors are first used to obtain multimodal signals from human demonstrations, and several physically meaningful parameters related to the task are learned through a closed-loop kinematics algorithm. These parameters are then transferred to the robot task using a motion redirection method, used to train a Gaussian mixture model, and Gaussian mixture regression is used to obtain the trajectory profile and stiffness profile composed of dynamic motion primitives.
[0008] In the reinforcement learning phase, a reinforcement learning model is constructed. The simulation sample data obtained from the environmental dynamics model and the trajectory profile and stiffness profile generated in the imitation learning phase are used to train the reinforcement learning model, and an online adjustment strategy for the task stiffness profile is obtained.
[0009] In the behavior control stage, the task stiffness profile is provided to the robot in an adaptive and optimal manner, and human-like impedance control is performed on the robot's joint position, joint angle, and contact force.
[0010] The multimodal signals include hand posture signals and tactile signals collected by data gloves, visual signals of hand behavior and arm posture collected by cameras, and arm electromyography signals collected by electromyography sensors.
[0011] The physically meaningful parameters related to the task include finger contact force, finger joint angles, Cartesian coordinates of contact points and arm joint angles when grasping various objects.
[0012] The motion redirection is based on the posture mapping of the human hand and arm to obtain the posture of the robot's mechanical arm and manipulator; wherein, the posture of the human hand and arm is obtained by the forward kinematic equations and sensor measurements of the given human hand and arm model, and the posture of the robot's mechanical arm and manipulator is obtained by scaling and transforming the posture of the human hand and arm.
[0013] The trajectory profile is obtained by the robot's manipulator arm and manipulator through the robot's inverse kinematics equation; the stiffness profile includes the finger stiffness profile and the arm stiffness profile. Among them, the finger stiffness profile is the change curve of the parameters obtained by calculating the finger contact force and the Cartesian coordinates of the contact point during the task execution, and the arm stiffness profile is estimated by the electromyography signal.
[0014] The dynamic motion primitive is expressed by the following formula
[0015]
[0016]
[0017] Where τ is the time constant, k is the stiffness coefficient, c is the damping coefficient, x is the Cartesian coordinate of the current position, x0 is the Cartesian coordinate of the initial position, g is the Cartesian coordinate of the target position, v is the velocity of the robot end, s is the current system state, and f(s) is a nonlinear function composed of a Gaussian mixture model, obtained by Gaussian mixture regression.
[0018] The reinforcement learning model includes a state space and an action space, wherein the state space includes the absolute value and next-moment increment of the tactile contact force, the absolute value and next-moment increment of the finger joint angle, and the absolute value and next-moment increment of the Cartesian position of the tactile contact; the action space includes finger movement, finger reopening indication, end effector posture adjustment, lifting and grasping stiffness.
[0019] Apply the following formula to obtain the optimization objective of the reinforcement learning model:
[0020]
[0021] Among them, ρ0 is the probability distribution of the initial state, π θ is a parameterized reinforcement learning policy network for robot control, a t is the action space, s t is the state space, is the reinforcement learning policy network before parameterized update for robot control, clip(·) is the cutoff function, ∈ is the hyperparameter used to control the cutoff range, is an estimator of the advantage function, is the prediction function of the state value.
[0022] The humanoid impedance control refers to the control of the expected trajectory of the robot arm and the impedance relationship between the robot arm and the object through the reinforcement learning model and the finite state machine. It is expressed by the following formula:
[0023]
[0024] in, is the hand closing range, is a measure of how wide the hand is closed, is the grab impedance,
[0025] A robot humanoid impedance control system based on multimodal signal fusion includes a robot body and a controller, wherein the robot body's manipulator is a five-finger dexterous hand, each finger has three joints, and the entire hand has a total of 15 joints. A joint angle sensor is installed at each joint to measure the rotation angle of each joint, and three-dimensional force sensors are installed at the fingertips of the five fingers; the controller executes the method to control the manipulator.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] 1. The present invention mainly obtains demonstrations of humans grasping objects of different stiffness through sensors, learns the stiffness geometry of humans performing tasks of grasping objects of different stiffness, uses motion redirection methods to transfer it to robot tasks, encodes the stiffness geometry related to different online tasks through imitation learning methods, and then trains the reinforcement learning model through the stiffness profile reproduced by imitation learning to improve the online adaptability of the controller.
[0028] 2. During the task execution, the reinforcement learning model adjusts the impedance model parameters online. By adjusting the impedance relationship between the contact force and relative position of objects of different stiffness and the dexterous hand, the dexterous hand can achieve precise control of the contact force and relative position during the grasping process of objects of different stiffness, especially fragile objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is an overall flow chart of the robot humanoid impedance control method and system based on multimodal signal fusion of the present invention.
[0030] Figure 2 This is a flow chart of the online control mode of the human-simulated impedance control system of the present invention. DETAILED DESCRIPTION
[0031] The structure and working process of the present invention will be further described below with reference to the accompanying drawings.
[0032] The present invention primarily uses sensors to capture demonstrations of humans grasping objects of varying stiffness, learning the stiffness geometry of these tasks. This is then transferred to robotic tasks using motion redirection methods. Imitation learning is then used to encode the stiffness geometry associated with these tasks online. The resulting stiffness profiles are then used to train a reinforcement learning model, improving the controller's online adaptability. During task execution, the reinforcement learning model adjusts impedance model parameters online. By adjusting the impedance relationship between the contact force and relative position of objects of varying stiffness and the dexterous hand, the dexterous hand achieves precise control of the contact force and relative position during grasping of objects of varying stiffness, particularly fragile objects.
[0033] The robot humanoid impedance control method based on multimodal signal fusion includes imitation learning stage, reinforcement learning stage and behavior control stage; wherein,
[0034] In the imitation learning phase, sensors are first used to obtain multimodal signals from human demonstrations, and several physically meaningful parameters related to the task are learned through a closed-loop kinematics algorithm. These parameters are then transferred to the robot task using a motion redirection method, used to train a Gaussian mixture model, and Gaussian mixture regression is used to obtain the trajectory profile and stiffness profile composed of dynamic motion primitives.
[0035] In the reinforcement learning phase, a reinforcement learning model is constructed. The simulation sample data obtained from the environmental dynamics model and the trajectory profile and stiffness profile generated in the imitation learning phase are used to train the reinforcement learning model, and an online adjustment strategy for the task stiffness profile is obtained.
[0036] In the behavior control stage, the task stiffness profile is provided to the robot in an adaptive and optimal manner, and human-like impedance control is performed on the robot's joint position, joint angle, and contact force.
[0037] Specific embodiments, such as Figure 1 、 Figure 2 As shown,
[0038] The robot humanoid impedance control method based on multimodal signal fusion includes imitation learning stage, reinforcement learning stage and behavior control stage; wherein,
[0039] In the imitation learning stage, first, sensors are used to obtain multimodal signals from human demonstrations, and a number of physically meaningful parameters related to the task are learned through a closed-loop kinematic algorithm; the parameters are transferred to the robot task using a motion redirection method to train a Gaussian mixture model, and the trajectory profile and stiffness profile composed of dynamic motion primitives are obtained through Gaussian mixture regression; the multimodal signals include human hand posture signals and tactile signals collected by data gloves, visual signals of human hand behavior and arm posture collected by cameras, and arm electromyography signals collected by electromyography sensors.
[0040] The physical parameters related to the task include but are not limited to the finger contact force {f t}、Finger joint angle The Cartesian coordinates of the contact point {c t} and the joint angle of the arm wait.
[0041] The motion redirection is based on the position of the human hand and arm Mapping to obtain the position of the robot's manipulator arm and manipulator To solve the geometric and dynamic differences between the human body and the robot; among them, the posture of the human hand and arm It is obtained by the forward kinematic equations and sensor measurements of a given human hand and arm model. and Obtained, the robot's robotic arm and manipulator pose It is obtained by scaling and transforming the posture of the human hand and arm to fit the motion space of the robotic arm and manipulator.
[0042] The trajectory profile The robot's manipulator and manipulator are obtained through the robot's inverse kinematics equation; the stiffness profile includes the finger stiffness profile and arm stiffness profile Among them, the finger stiffness profile The curve of the parameters obtained by calculating the finger contact force and the Cartesian coordinates of the contact point during the task execution is shown in Figure 2. The arm stiffness profile is obtained by the electromyographic signal. Estimated;
[0043] The arm stiffness profile is obtained by optimizing the F-norm of the following expression:
[0044]
[0045] in, is the pseudo-inverse of the Jacobian matrix of the arm, Fex It is the external force applied to the human hand. is the stiffness of the end of the arm.
[0046] The finger joint angles are obtained through a closed-loop inverse kinematics algorithm, and the used imitation learning model includes a Gaussian mixture model, a Gaussian mixture regression and a dynamic motion primitive.
[0047] The Gaussian mixture model is as follows:
[0048]
[0049] Where x is a variable and p(x) is the joint probability distribution of the variable x. i , μ i ,∑ i represents the prior probability, mean, and covariance of the i-th Gaussian component, K is the total number of Gaussian components, and K is obtained by the statistical method of the Bayesian Information Criterion.
[0050] The goal of the offline training phase of the imitation learning model is to maximize the log-likelihood function equation with respect to the model parameters, as follows
[0051]
[0052] Here, N refers to the total number of data points in the training dataset.
[0053] The imitation learning model uses the following expectation maximization process to iteratively update the model parameters until convergence:
[0054]
[0055]
[0056]
[0057]
[0058] The dynamic motion primitive is expressed by the following formula
[0059]
[0060]
[0061] Where τ is the time constant, k is the stiffness coefficient, c is the damping coefficient, x is the Cartesian coordinate of the current position, x0 is the Cartesian coordinate of the initial position, g is the Cartesian coordinate of the target position, v is the velocity of the robot end, s is the current system state, and f(s) is a nonlinear function composed of a Gaussian mixture model, obtained by Gaussian mixture regression.
[0062] In the reinforcement learning phase, a reinforcement learning model is constructed. The simulation sample data obtained from the environmental dynamics model and the trajectory profile and stiffness profile generated in the imitation learning phase are used to train the reinforcement learning model, and the online adjustment data of the task stiffness profile is obtained.
[0063] The reinforcement learning model includes a state space and an action space, wherein the state space includes the absolute value of the tactile contact force and the next moment increment Absolute value of finger joint angle and the next moment increment Absolute Cartesian position of tactile contact and the next moment increment
[0064] Specifically, The tactile sensor readings for the last 20 time points, The increment of the tactile sensor readings between the last 20 time points. is the joint angle value of the last 20 time points, is the incremental value of the joint angle at the last 20 time points, The readings of the contact point positions at the last 20 time points, is the change in the contact point position at the last 20 time points.
[0065] The action space includes the hand closure range Finger reopen indication Dexterous hand expected posture lift and grip stiffness
[0066] Apply the following formula to obtain the optimization objective of the reinforcement learning model:
[0067]
[0068] Among them, ρ0 is the probability distribution of the initial state, π θ is a parameterized reinforcement learning policy network for robot control, clip(·) is the truncation function, ∈ is a hyperparameter used to control the truncation range, is an estimator of the advantage function, is the prediction function of the state value.
[0069] Furthermore, the reinforcement learning model realizes humanoid impedance control of the robot through a learning control strategy during the interaction with the environmental dynamics model, and the environmental dynamics model is realized in the computer system through the pybullet software package.
[0070] The humanoid impedance control refers to the control of the expected trajectory of the robot arm and the impedance relationship between the robot arm and the object through the reinforcement learning model and the finite state machine. It is expressed by the following formula:
[0071]
[0072] in, is the measurement of finger closure width, is the grab impedance,
[0073] In the behavior control stage, the task stiffness profile is provided to the robot in an adaptive and optimal manner, and human-like impedance control is performed on the robot's joint position, joint angle, and contact force.
[0074] A robot humanoid impedance control system based on multimodal signal fusion includes a robot body and a controller. The robot body has a five-finger dexterous hand, each finger has three joints, and the entire hand has a total of 15 joints. A joint angle sensor is installed at each joint to measure the rotation angle of each joint, and three-dimensional force sensors are installed at the fingertips of the five fingers. The controller executes the robot humanoid impedance control method based on multimodal signal fusion to control the robot arm.
[0075] By collecting kinematic and dynamic information of humans grasping objects with different stiffness, we transfer the human stiffness adjustment strategy to the robot control system through imitation learning, making it exhibit human-like impedance behavior.
[0076] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.
[0077] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0078] The above process of obtaining sample data includes human body demonstration, that is, motion clips of real people grasping objects.
[0079] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.
[0080] Those skilled in the art should understand that they can implement variations by combining the prior art and the above embodiments. Such variations do not affect the essence of this solution and are not described in detail here.
[0081] It should be understood that this solution is not limited to the specific implementation methods described above. Devices and structures not described in detail should be understood to be implemented in a common manner in the art. Any person skilled in the art can, without departing from the scope of this solution, use the methods and technical content disclosed above to make many possible changes and modifications to this solution, or modify it into equivalent embodiments with equivalent changes, without affecting the essence of this solution. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this solution without departing from the content of this solution are still within the scope of protection of this solution.
Claims
1. A robot humanoid impedance control method based on multimodal signal fusion, characterized by: It includes the imitation learning stage, reinforcement learning stage and behavior control stage; among them, In the imitation learning phase, sensors are first used to obtain multimodal signals from human demonstrations, and several physically meaningful parameters related to the task are learned through a closed-loop kinematic algorithm. These parameters are then transferred to the robot task using a motion redirection method, used to train a Gaussian mixture model, and the trajectory profile and stiffness profile composed of dynamic motion primitives are obtained through Gaussian mixture regression. The multimodal signals include hand posture and tactile signals collected by data gloves, visual signals of hand behavior and arm posture collected by cameras, and arm electromyography signals collected by electromyography sensors. The physically meaningful parameters related to the task include finger contact force, finger joint angles, Cartesian coordinates of contact points, and arm joint angles when grasping various objects. In the reinforcement learning phase, a reinforcement learning model is constructed and trained using simulation sample data obtained from the environmental dynamics model and the trajectory profile and stiffness profile generated in the imitation learning phase to obtain an online adjustment strategy for the task stiffness profile. The reinforcement learning model includes a state space and an action space. The state space includes the absolute value and next-moment increment of the tactile contact force, the absolute value and next-moment increment of the finger joint angle, and the absolute value and next-moment increment of the Cartesian position of the tactile contact. The action space includes finger movement, finger reopening indication, end-effector posture adjustment, lifting, and grasping stiffness. The optimization objective of the reinforcement learning model is obtained using the following formula: , in, is the probability distribution of the initial state, is a parameterized reinforcement learning policy network for robot control, is the action space, is the state space, , is a parameterized pre-update reinforcement learning policy network for robot control, is the truncation function, is a hyperparameter used to control the cutoff range, is an estimator of the advantage function, is the prediction function of the state value; In the behavior control stage, the task stiffness profile is provided to the robot in an adaptive and optimal manner, and humanoid impedance control is performed on the robot's joint positions, joint angles, and contact forces.
2. The method for controlling humanoid impedance of a robot based on multimodal signal fusion according to claim 1, characterized in that: The motion redirection is based on the posture mapping of the human hand and arm to obtain the posture of the robot's mechanical arm and manipulator; wherein the posture of the human hand and arm is obtained by the forward kinematic equations and sensor measurements of the given human hand and arm model, and the posture of the robot's mechanical arm and manipulator is obtained by scaling and transforming the posture of the human hand and arm.
3. The robot humanoid impedance control method based on multimodal signal fusion according to claim 1, characterized in that: The trajectory profile is obtained by the robot's manipulator arm and manipulator through the robot's inverse kinematics equation; the stiffness profile includes the finger stiffness profile and the arm stiffness profile. Among them, the finger stiffness profile is the change curve of the parameters obtained by calculating the finger contact force and the Cartesian coordinates of the contact point during the task execution, and the arm stiffness profile is estimated by the electromyography signal.
4. The method for controlling humanoid impedance of a robot based on multimodal signal fusion according to claim 3, characterized in that: The dynamic motion primitive is expressed by the following formula , in, is the time constant, is the stiffness coefficient, is the damping coefficient, is the Cartesian coordinate of the current position, is the Cartesian coordinate of the initial position, is the Cartesian coordinate of the target position, is the velocity of the robot's end, is the current system status, A nonlinear function composed of Gaussian mixture models, obtained through Gaussian mixture regression.
5. The method for controlling humanoid impedance of a robot based on multimodal signal fusion according to claim 1, wherein: The humanoid impedance control refers to the control of the expected trajectory of the robot arm and the impedance relationship between the robot arm and the object through the reinforcement learning model and the finite state machine. It is expressed by the following formula: , in, To indicate finger movement, is a measure of how wide the hand is closed, is the grab impedance, , is the grasping stiffness, To adjust the hyperparameters.
6. A robotic humanoid impedance control system based on multimodal signal fusion, characterized by: The invention comprises a robot body and a controller, wherein the manipulator of the robot body is a five-finger dexterous hand, each finger consists of three joints, and the whole hand has a total of 15 joints. A joint angle sensor is installed at each joint to measure the rotation angle of each joint, and three-dimensional force sensors are installed at the fingertips of the five fingers; the controller executes the robot humanoid impedance control method based on multimodal signal fusion as described in any one of claims 1 to 5 to control the robot humanoid impedance.
Citation Information
Patent Citations
Control system and method for learning variable impedance
CN108153153A
Impedance control imitation learning training method beyond expert demonstration
CN113641099A
Cited By
Skill acquisition and migration method for dexterous hand operation
CN120901995A