A robot variable impedance skill learning method based on a conservative extended dynamic system
Patent Information
- Application Number
- CN202611186163.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-18
AI Technical Summary
1.解决现有保守扩展动态系统框架中旋转与平移刚度阻尼耦合、调控灵活性不足的问题,同时建立刚度矩阵特性与系统存在性、无源性之间的系统理论联系,为技能学习提供明确的约束依据;
1.实现平移-旋转解耦调控,理论体系完备:本发明构建平移-旋转解耦的保守扩展动态系统框架,二者共享虚拟任务进度坐标并配置独立标量阻尼系数,打破传统耦合框架的限制,可针对旋转与平移的不同物理特性独立调控阻抗参数。同时从刚度角度建立系统理论:证明刚度矩阵的恰当性与对称性是保守扩展动态系统存在的充分必要条件,刚度矩阵的一致正定性是变阻抗控制系统无源性的充分条件,明确了刚度可编码性与系统稳定性的核心判定依据。
Smart Images

Figure CN122769997A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot compliant control and teaching learning technology, specifically to a robot variable impedance skill learning method based on Conservative Extended Dynamical Systems (CEDS). Background Technology
[0002] In unstructured environments, robots face challenges related to geometric constraints and uncertainties in physical properties (such as stiffness and damping) when performing contact-rich tasks (e.g., inserting into holes, grinding, polishing). Controllers with fixed impedance parameters struggle to balance interaction safety and task accuracy. Variable Impedance Control (VIC), as an effective solution, allows robots to dynamically adjust stiffness and damping parameters based on task progress and system state, thereby improving overall task performance. However, stiffness variations introduce additional energy changes into the system, potentially leading to unstable behavior; therefore, the design of VIC systems requires comprehensive consideration of both task performance and system stability.
[0003] Current research on variable impedance skill design for tasks involving a wide range of tasks mainly falls into the following directions: The first category is instruction-based methods for acquiring variable impedance skills. These methods collect human demonstration data (including position, velocity, contact force, etc.) and extract variable impedance curves using imitation learning techniques. Common techniques include: least squares fitting, directly fitting the desired stiffness and damping curves from the demonstration data; Gaussian mixture models, probabilistically modeling the demonstration data to generate smooth impedance parameter trajectories; dynamic motion primitives, adjusting impedance parameters while encoding motion trajectories; and a combination of imitation learning and reinforcement learning, further optimizing skill performance through reinforcement learning based on initial imitation learning. These methods primarily focus on optimizing task performance and typically neglect theoretical guarantees of system stability. The second category is the variable impedance stability control method based on energy pools. To ensure the stability of variable impedance control, energy pool technology is widely used to guarantee the passivity of the system. Its core idea is to decompose stiffness into constant and time-varying components; actively recover dissipated energy using a virtual energy pool to compensate for energy changes caused by stiffness variations; and ensure the stability of the system during free motion and interaction with any passive environment, without requiring an explicit model of the environmental dynamics. The main limitation of this method is that once the energy pool is depleted, the system can only realize the constant component of the stiffness curve and cannot accurately execute the desired variable impedance skill. The third category is variable impedance control methods based on dynamic systems (DS). Dynamic systems (DS) have mature applications in robot motion generation. In recent years, researchers have extended DS to the field of variable impedance control: encoding the desired task trajectory through DS and encoding state-related stiffness curves through its Jacobian matrix; achieving symmetric and attractive stiffness through local modifications to the Jacobian matrix, thus improving trajectory tracking performance; and combining passive theory to ensure the stability of variable impedance control. However, existing DS-based variable impedance methods mainly have the following limitations: stiffness depends only on the robot state and cannot capture changes related to task progress; damping coupling of rotational and translational motions is in the same matrix, making independent design difficult; and it mainly relies on kinematic demonstrations (position and velocity), without incorporating contact force information, limiting its applicability in contact-rich tasks. The fourth category is stable variable impedance learning methods based on potential functions. These methods guarantee stability by modeling elastic forces as gradients of potential functions. Key research includes: Khansari-Zadeh et al. proposed the VIC framework for state-dependent stiffness and damping, deriving sufficient conditions for stability; Jin et al. used a potential function based on the neural network SNEUM to learn error-dependent variable stiffness curves with formal stability guarantees. These methods can learn state-dependent variable stiffness skills, but they also suffer from problems such as stiffness not changing with task progress and rotational / translational coupling. The fifth category is extended dynamical systems and their application in variable impedance. To enhance the encoding capability of DS (Dynamic Stress Distributor) for complex trajectories, related research has proposed Extended Dynamical Systems (EDS). By introducing virtual coordinates representing task progress, the system can encode complex motions such as self-intersecting trajectories. Furthermore, by imposing conservative conditions on EDS, Conservative Extended Dynamical Systems (CEDS) are formed to encode variable impedance skills that depend simultaneously on task progress and robot state. However, the existing framework still has shortcomings such as coupling of rotational and translational stiffness damping, lack of system passivity analysis and proof, lack of a theoretical connection between stiffness characteristics and system performance, and lack of a systematic demonstration data learning method.
[0004] In summary, existing technologies have not yet been able to simultaneously achieve multiple objectives within a unified framework, such as learning variable impedance skills from demonstrations, adaptive stiffness changes with task progress and status, independent control of rotational and translational impedance, passive guarantee of skill execution, and incorporating contact force information to improve the adaptability of contact tasks. Therefore, they are insufficient to meet the actual needs of complex contact tasks such as precision assembly.
[0005] Therefore, those skilled in the art urgently need to develop a robot impedance-based skill learning method based on conservative extended dynamic systems. Summary of the Invention
[0006] In view of the above-mentioned deficiencies of the prior art, the present invention at least solves the following technical problems: 1. To address the issues of insufficient control flexibility and coupling of rotational and translational stiffness damping in the existing conservative extended dynamic system framework, and to establish a theoretical connection between stiffness matrix characteristics and system existence and passivity, providing a clear constraint basis for skill learning; 2. To address the problem that existing variable impedance learning methods cannot inherently guarantee stiffness characteristics at the network architecture level, and to overcome the limitations of post-verification or external constraints, to achieve end-to-end learning of passive variable impedance skills, thereby ensuring system stability from the bottom layer. 3. To address the issue that existing dynamic system impedance methods do not incorporate contact force information, enabling learned impedance skills to reflect real interactive characteristics and improving adaptability and success rate in handling a wide range of contact tasks; 4. To address the problem that existing stiffness models rely solely on robot state and cannot match the differentiated needs of different task stages, this paper proposes to achieve joint adaptive adjustment of stiffness based on task progress and robot state, thereby improving cross-task generalization capability.
[0007] To achieve the above objectives, this invention discloses a robot variable impedance skill learning method based on a conservative extended dynamic system, the method comprising the following steps: S1: Establish a Cartesian space dynamics model for the robot, and decompose the generalized state of the robot's end effector into mutually independent translational and rotational states; S2: Construct a translation-rotation decoupled CEDS framework; the framework includes a translation CEDS and a rotation CEDS that share the same virtual task progress coordinates, and the translation CEDS and rotation CEDS are configured with independent scalar damping coefficients respectively; S3: Collect physical guidance demonstration data containing kinematic and contact force information, infer the target velocity field based on the robot dynamics equations, and construct the CEDS training dataset; S4: A potential function for a conservative extended dynamic system is constructed using a Partially Input Convex Neural Network (PICNN). The inherent structure of the network ensures that the potential function is strictly convex with respect to the robot state, so that the corresponding stiffness matrix naturally satisfies symmetry and uniform positive definiteness. The potential function takes both the robot state and the virtual task progress coordinates as inputs, so that the stiffness depends on both the robot state and the task progress. S5: Train a conservative extended dynamic system based on the training dataset, substitute the trained velocity field output into the decoupled variable impedance control law, and generate robot joint torque control commands.
[0008] Furthermore, in step S2, both the translational CEDS and the rotational CEDS include the state velocity field equation and the virtual task progress evolution equation, with the following expressions: Translate CEDs:
[0009] Rotating CEDs:
[0010] in, The translation state vector, Let the rotation state vector be... For virtual task progress coordinates; For the translational velocity field, For rotational velocity field; , For state-schedule coupling terms, For nominal progress evolution items; translational CEDS and rotational CEDS maintain progress synchronization through shared virtual task progress coordinates and virtual force adjustments; The expression is as follows:
[0011] in, Indicates the task execution cycle, parameters This determines the strength of the coupling between the robot's state and the task progress; It is a positive real number used to determine The location of the equilibrium point.
[0012] Furthermore, in step S2, the decoupling variable impedance control law includes a generalized control force equation and a virtual progress adjustment force equation, expressed as follows:
[0013] in, The generalized control force applied to the controller, This is for the robot's gravity compensation. The rotational scalar damping coefficient is... The translational scalar damping coefficient; This represents the actual rotational speed of the robot's end effector. This represents the actual translational speed of the robot's end effector. This is a virtual progress adjustment force used to synchronize the virtual task progress evolution rate of the translational and rotational CEDS. , For state-schedule coupling terms, For the nominal schedule evolution item, This represents the actual rate of change of the virtual task progress coordinates.
[0014] Furthermore, in step S3, the specific steps for constructing the training dataset are as follows: S31: Collect demonstration data through physical guidance, and record timestamps, end position, end attitude, and raw readings of force / torque sensors; S32: Preprocess the raw force / torque data, remove the influence of end-load gravity, and perform coordinate transformation to obtain the generalized external contact force. ; S33: Perform low-pass filtering and numerical differentiation on the position and attitude data to obtain the actual velocity and actual acceleration of the robot end effector; S34: Based on the robot dynamics equations, assuming the control force equals the guiding force, the target velocity field of CEDS is derived by reverse calculation, and its expression is:
[0015] in, For the robot's inertia matrix, The matrix of Coriolis force and centrifugal force. For the generalized acceleration of the robot's end effector, It is an identity matrix.
[0016] Further, in step S4, the potential function is decomposed into a state-schedule coupled component and a pure schedule component, expressed as:
[0017] in, The potential function components, which depend on both the robot's state and the task progress, are learned by a partially input convex neural network. The potential function component that depends only on the task progress is obtained by analytical integration of the nominal progress evolution function; For component identification, Corresponding rotational component, Corresponding translation component.
[0018] Further, in step S4, the state-progress coupling potential function component is composed of the superposition of the PICNN learning potential function and the baseline stiffness potential function, and its expression is:
[0019] Among them, the baseline stiffness potential function The expression is:
[0020] in, The average reference trajectory corresponding to the task progress is learned by a multilayer perceptron network. The baseline stiffness matrix is diagonally positive definite; the baseline stiffness potential function is used to ensure that the stiffness matrix is not lower than a preset threshold and to maintain consistent positive definiteness.
[0021] Furthermore, in step S4, the partially input convex neural network uses the robot state as the convex input and the virtual task progress coordinates as the non-convex input; the convex path of the network uses non-negative weights and a convex non-decreasing activation function to ensure that the output potential function is strictly convex with respect to the robot state, and its Hessian matrix, i.e., the stiffness matrix, naturally satisfies symmetry and uniform positive definiteness.
[0022] Furthermore, the virtual task progress coordinates are constructed as follows:
[0023] in, The current task execution time. The nominal total task duration, The scaling factor is the state-schedule coupling strength; the nominal evolution law of the virtual task schedule coordinates is a piecewise function, which evolves at a constant rate during the task execution phase and enters the convergence phase after the task ends and asymptotically approaches the final value.
[0024] Furthermore, in step S5, the training of the conservative extended dynamic system adopts a two-stage training method: Phase 1: Train the multilayer perceptron network to learn the average reference trajectory under different task progresses, and optimize the network parameters by minimizing the trajectory fitting error; Phase 2: The joint demonstration dataset and the constraint dataset are used to train the input convex neural network, and the network parameters are optimized using a composite loss function. The composite loss function includes velocity field fitting loss and task cycle constraint loss, which are used to align the demonstration speed and ensure the nominal task duration, respectively.
[0025] Furthermore, the stiffness matrix corresponding to the variable impedance control is a block diagonal structure, with the rotational stiffness block and the translational stiffness block being independent of each other; the appropriateness and symmetry of the stiffness matrix are necessary and sufficient conditions for the existence of the conservative extended dynamic system, and the uniform positive definiteness of the stiffness matrix is a sufficient condition for the passivity of the variable impedance control system.
[0026] This invention achieves at least the following beneficial technical effects: 1. Achieving Decoupled Translation-Rotation Control with a Complete Theoretical System: This invention constructs a conservative extended dynamic system framework with translation-rotation decoupling. Both translation and rotation share virtual task progress coordinates and are configured with independent scalar damping coefficients, breaking the limitations of traditional coupled frameworks. Impedance parameters can be independently controlled based on the different physical characteristics of rotation and translation. Simultaneously, a system theory is established from a stiffness perspective: it is proven that the appropriateness and symmetry of the stiffness matrix are necessary and sufficient conditions for the existence of the conservative extended dynamic system, and the uniform positive definiteness of the stiffness matrix is a sufficient condition for the passivity of the variable impedance control system. The core criteria for determining stiffness codedability and system stability are clarified.
[0027] 2. The architecture inherently guarantees passivity while balancing stability and task performance: This invention employs a partially input convex neural network to construct the potential function, treating the appropriateness, symmetry, and consistent positive definiteness of the stiffness matrix as inherent structural characteristics of the network. This inherently guarantees the system's passivity from the architectural level without requiring post-implementation verification or modification. In a square shaft-hole assembly experiment with a 0.14mm gap, this method achieved a success rate of 93.33% in 30 trials, significantly outperforming the GMM-based CEDS (coupled) method (0%) and the SNEUM-based VIC method (80%); the trajectory reproduction error (SEA) was 3.054cm. 2 The DTWD (341.534 cm) and force reproduction error were the lowest among the three comparison methods, and stable interactive control and high-precision skill reproduction were achieved at the same time.
[0028] 3. Incorporating Contact Force Information Significantly Improves Interaction Adaptability: This invention simultaneously collects kinematic and contact force data during physical-guided demonstrations, and constructs a training dataset by inversely deducing the target velocity field based on the robot's dynamics equations, ensuring that the learned variable impedance skills align with real-world interaction requirements. Ablation experiments verify that the success rate of the scheme incorporating contact force training is 93.33%, while the success rate of the scheme without contact force is 0%, demonstrating that contact force information plays a decisive role in learning effective variable impedance skills. The translational stiffness learned with contact force is approximately one order of magnitude higher than that without contact force, exhibiting task-adaptive directional stiffness at each assembly stage, and maintaining a continuous contact force of approximately 22N during the insertion stage to ensure alignment.
[0029] 4. Adaptive Task Progress Control and Excellent Generalization Ability: This invention uses virtual task progress coordinates to make stiffness dependent on both robot state and task progress. Stiffness changes are precisely aligned with the task progress, exhibiting differentiated directional stiffness characteristics in each stage of assembly—approach, alignment, rotation, and insertion—matching the task requirements of different stages. This skill can be successfully generalized to assembly tasks with different geometries, achieving an 86.67% success rate for assembling round holes (0.04mm gap) and a 90.00% success rate for assembling triangular holes (0.28mm gap), demonstrating good cross-task versatility.
[0030] 5. Strong robustness and engineering applicability: The skills learned in this invention exhibit good robustness against visual errors, calibration errors, and kinematic errors, achieving an overall success rate of 83.2% in 125 tests with different positional errors. The solution is compatible with mainstream industrial robot hardware and control architectures, facilitates demonstration and data acquisition, and enables efficient network inference. It requires no special customized equipment and possesses excellent conditions for industrial application. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the framework of the robot variable impedance skill learning method based on conservative extended dynamic systems according to the present invention. Figure 2 This is a schematic diagram of the CEDS architecture based on PICNN of the present invention; Figure 3 This is a schematic diagram of the data acquisition for the pin hole assembly based on physical guidance according to the present invention, where 3(a) is the physical guidance process in the assembly task and 3(b) is the workflow for constructing the dataset from the recorded data. Figure 4 This is a schematic diagram of the overall experimental scenario for the shaft and hole assembly task of this invention; Figure 5 This is a schematic diagram illustrating the establishment of the coordinate system of this invention and the dimensions of the square axis and square hole. Detailed Implementation
[0032] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0033] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0034] This invention proposes a robot impedance-changing skill learning method based on a conservative extended dynamic system. The method comprises five core steps: establishing a Cartesian space dynamic model of the robot, constructing a translation-rotation decoupled conservative extended dynamic system framework, building a training dataset containing contact force information, constructing a potential function using a partially input convex neural network, training the system, and generating control commands. It can be applied to various contact-rich robot tasks. The overall process is as follows: Figure 1 As shown, the method includes the following steps: S1: Establish a Cartesian space dynamics model for the robot, and decompose the generalized state of the robot's end effector into mutually independent translational and rotational states; S2: Construct a conservative extended dynamic system CEDS framework with translation-rotation decoupled; the framework includes a translation CEDS and a rotation CEDS that share the same virtual task progress coordinates, and the translation CEDS and rotation CEDS are configured with independent scalar damping coefficients respectively; S3: Collect physical guidance demonstration data containing kinematic and contact force information, infer the target velocity field based on the robot dynamics equations, and construct the CEDS training dataset; S4: The potential function of the conservative extended dynamic system is constructed using PICNN. The inherent structure of the network ensures that the potential function is strictly convex with respect to the robot state, so that the corresponding stiffness matrix naturally satisfies symmetry and uniform positive definiteness. The potential function takes both the robot state and the virtual task progress coordinates as input, so that the stiffness depends on both the robot state and the task progress. S5: Train a conservative extended dynamic system based on the training dataset, substitute the trained velocity field output into the decoupled variable impedance control law, and generate robot joint torque control commands.
[0035] Example 1
[0036] This embodiment verifies the effectiveness of the proposed translation-rotation decoupled CEDS-based VIC frame. In traditional EDS / VIC frames, the stiffness and damping of rotational and translational motions are coupled in the same generalized velocity field and controlled by the same eigenvalue along the EDS direction, making it impossible to independently design the magnitude of rotational and translational damping. This embodiment verifies the crucial role of decoupling design in improving the adaptability of contact tasks through comparative experiments in a square shaft-hole assembly task.
[0037] This embodiment implements a CEDS-based variable impedance control framework with translation-rotation decoupling. The specific implementation process is as follows: First, a Cartesian space dynamics model of the robot is established, and the dynamic equations are:
[0038] The generalized state of a robot end effector is defined as a combination of attitude and position components, where the attitude components can be parameterized using exponential coordinates or Euler angles; the position components correspond to the translational state, and the attitude components correspond to the rotational state. The dynamic equations of a rigid robot in Cartesian space include the inertia matrix, the Coriolis force and centrifugal force matrices, the gravity compensation term, and the generalized control force applied by the controller and the contact force from the external environment.
[0039] Subsequently, independent rotation CEDS and translation CEDS are constructed, both sharing the same virtual task progress coordinates: The expression for rotating CEDS is:
[0040] The expression for CEDS translation is:
[0041] in, The translation state vector, Let the rotation state vector be... For virtual task progress coordinates; For the translational velocity field, For rotational velocity field; , For state-schedule coupling terms, The nominal schedule evolution item; translational CEDS and rotational CEDS maintain schedule synchronization through shared virtual task schedule coordinates and virtual force adjustments; nominal schedule evolution item The expression is as follows:
[0042] in, Indicates the task execution cycle, parameters This determines the strength of the coupling between the robot's state and the task progress. It is a positive real number used to determine The location of the equilibrium point.
[0043] The rotational CEDS includes a rotational velocity field and a schedule evolution equation. The schedule evolution consists of the sum of rotation-related coupling terms and the nominal schedule term. The translational CEDS includes a translational velocity field and a schedule evolution equation. The schedule evolution consists of the sum of translation-related coupling terms and the nominal schedule term. Both CEDS share the same virtual schedule coordinates, and rotation and translation are synchronized through virtual force adjustments. The virtual schedule coordinate is defined as the current time divided by the product of the scaling factor and the total duration. The task starts with a schedule of zero and ends when the schedule reaches "one divided by the square root of the scaling factor," where the scaling factor controls the state-schedule coupling strength.
[0044] Based on this, a decoupled variable impedance control law is designed, using independent scalar damping coefficients to control the rotating CEDS and the translational CEDS respectively. The control law expression is as follows:
[0045] in, The generalized control force applied to the controller, This is for the robot's gravity compensation. The rotational scalar damping coefficient is... The translational scalar damping coefficient; This represents the actual rotational speed of the robot's end effector. This represents the actual translational speed of the robot's end effector. This is a virtual progress adjustment force used to synchronize the virtual task progress evolution rate of the translational and rotational CEDS. , For state-schedule coupling terms, For the nominal schedule evolution item, This represents the actual rate of change of the virtual task progress coordinates.
[0046] The control law consists of two parts: the first part is the robot's gravity compensation term; the second part is the damping term, where rotational damping is the rotational damping coefficient multiplied by the rotational speed deviation, and translational damping is the translational damping coefficient multiplied by the translational speed deviation. The calculation of virtual forces also incorporates the progress coupling deviations of rotation and translation, weighted by the damping coefficients of both. This control law generates an isotropic damping structure within each subsystem, while allowing different damping magnitudes for rotation and translation.
[0047] Further derivation yields the stiffness matrix of the block-diagonal structure, where the rotational and translational stiffness blocks are independent. Each rotational and translational stiffness block is obtained by multiplying its corresponding damping coefficient by the negative partial derivative of the velocity field with respect to the state. The stiffness depends on both the robot's state and the task progress, achieving adaptive variable impedance control during task phases. Theoretical verification shows that the appropriateness and symmetry of the stiffness matrix are necessary and sufficient conditions for the existence of CEDS. The appropriateness condition requires that the partial derivatives of each element of the stiffness matrix with respect to the state satisfy a symmetric relationship; the uniform positive definiteness of the stiffness matrix (i.e., all eigenvalues are greater than a certain positive constant) is a sufficient condition for the passivity of the CEDS-based variable impedance control.
[0048] In the comparative experiment of square shaft-hole assembly, the success rate of the VIC method based on decoupled CEDS of this invention was 93.33% (28 out of 30 tests were successful), while the success rate of the comparative method based on GMM-based CEDS (coupling) was 0% (all 30 tests failed), verifying the key role of decoupled design in adaptability to contact tasks. Specific assembly success rate statistics are shown in Table 1: Table 1 Comparison of assembly success rates using different methods
[0049] Example 2
[0050] This embodiment verifies the effectiveness of the CEDS learning architecture based on PICNN proposed in this invention. The core innovation of this invention lies in treating the appropriateness, symmetry, and uniform positive definiteness of the stiffness matrix as inherent structural characteristics of the PICNN network rather than as ex-post constraints, thus naturally guaranteeing the passivity of the learned variable impedance skills. The overall network architecture and computational process are as follows: Figure 2 As shown, the specific implementation process is as follows: First, the potential function of CEDS is decomposed, and the expression of the total potential function is:
[0051] in, The state-progress coupling component, which depends on both the robot's state and the task's progress, is learned by PICNN. The nominal schedule evolution term is obtained by analyzing the integral of the nominal schedule evolution function, which is a pure schedule component that depends only on the task progress. The definition and expression are the same as the aforementioned general scheme; For component identification, Corresponding rotational component, Corresponding translation component.
[0052] PICNN is a neural network with special structural constraints. Its output is convex with respect to convex inputs and remains non-convex with respect to non-convex inputs. In this embodiment, the robot state is used as the convex input, and the virtual task progress coordinates are used as the non-convex input. The network includes non-convex paths (u-paths) and convex paths (z-paths). The non-convex paths use a conventional neural network structure, while the convex paths ensure the convexity of the output with respect to convex inputs through constraints of non-negative weights and convex non-decreasing activation functions. The two paths interact through the Hadamard product. PICNN's... The layer architecture is defined as follows:
[0053] in, Network layer number; For the first Feature vectors of non-convex paths in layers. For the first Feature vectors of the convex path; , , , , , , This is the weight matrix for the corresponding path, where the weights for convex paths are all non-negative. , , For the corresponding bias vector; For the first The activation function of the layer is used, and the convex path uses a convex and non-decreasing activation function; This represents the Hadamard product (element-by-element multiplication). The input is a non-convex coordinate system, i.e., the virtual task progress coordinates. The input is a convex shape, representing the robot's state. Output the potential function value to the network.
[0054] Through this structural constraint, the output potential function is strictly convex with respect to the robot state, and its Hessian matrix, i.e., the stiffness matrix, naturally satisfies symmetry and uniform positive definiteness.
[0055] The expression for the baseline stiffness potential function is:
[0056] in, The average reference trajectory corresponding to the task progress is learned by a multi-layer perceptron (MLP) network; The baseline stiffness matrix is diagonally positive definite. The baseline stiffness potential function is used to ensure that the stiffness matrix is not lower than a preset threshold, so as to avoid insufficient robustness to disturbances such as friction and modeling errors due to excessively low learned CEDS stiffness, and further maintain consistent positive definiteness.
[0057] The state-schedule coupled potential function component is composed of the superposition of the PICNN learning potential function and the baseline stiffness potential function, and its expression is:
[0058] Leveraging the system's conservative properties, the velocity field of CEDS is given by the negative gradient of the total potential function with respect to the state, and the progress coupling term is given by the negative gradient of the total potential function with respect to the progress. These gradients are calculated through automatic differentiation, and the nominal progress term is directly calculated from its analytical expression. Before training, the input is normalized: each dimension of the robot state is independently normalized to [-1,1] using min-max scaling; the virtual task progress coordinates are normalized to [0,1] using a smooth saturation function, enhancing the robustness and stability of network training.
[0059] The training process adopts a two-stage training approach: Phase 1: Train the MLP network to learn the average reference trajectory under different task progresses, and optimize the network parameters by minimizing the trajectory fitting error. The MLP adopts a single hidden layer structure (1D input layer, 32D hidden layers, 6D output layer) and uses the sigmoid activation function.
[0060] Phase 2: PICNN is trained using a joint demonstration dataset and a constraint dataset, with network parameters optimized using a composite loss function. This composite loss function includes a velocity field fitting loss and a task cycle constraint loss, used to align the demonstration velocity and ensure the nominal task duration, respectively. Both losses have a weight of 1. PICNN's non-convex paths employ a single hidden layer structure (1D input layer, 32D hidden layers, 32D output), while convex paths employ a dual hidden layer structure (3D input layer, 24D hidden layers each, 1D output). The convex paths use the softplus activation function, and the output layer uses an identity mapping.
[0061] In the square shaft-hole assembly experiment, the success rate of the PICNN-based CEDS method was 93.33% (28 out of 30 successful trials), with an average assembly time of 9.900 seconds and a standard deviation of 0.074 seconds. Regarding trajectory reproduction error, the swept error area (SEA) was 3.054 square centimeters, and the dynamic time warping distance (DTWD) was 341.534 centimeters, both superior to the GMM-based CEDS and SNEUM-based VIC methods. In terms of force reproduction error, the dynamic time warping distance was 15709.2 Newtons, also the lowest among the three methods, achieving optimal demonstration and reproduction results while ensuring system stability.
[0062] Example 3
[0063] This embodiment verifies the effectiveness of the proposed CEDS demonstration dataset construction method that incorporates contact force information. Existing DS-based variable impedance methods only utilize kinematic demonstration data (position and velocity) and do not incorporate contact force / torque information. This invention extracts both kinematic and contact force data from the physically guided demonstration and inversely derives the target velocity field of the CEDS using the robot dynamics equations. The dataset acquisition and construction process is as follows: Figure 3 As shown, 3(a) is the physical guidance process in the assembly task, and 3(b) is the workflow for building a dataset from the recorded data. The specific implementation process is as follows: First, demonstration data is collected through physical guidance (drag-and-drop teaching). The operator drags the robot's end effector to complete the target task. During the demonstration, the robot's dynamics satisfy the force balance relationship.
[0064] That is, the sum of the robot's inertial force, Coriolis force, and gravity is equal to the sum of the operator's guiding force and the external contact force.
[0065] In the square shaft-hole assembly task, the square shaft (side length 20.00mm) is mounted on the force / torque sensor, and the square hole (side length 20.14mm) is fixed to the worktable. During data acquisition, it is necessary to ensure that the operator does not touch the output side of the force / torque sensor to avoid introducing additional measurement errors.
[0066] Each demonstration collects timestamps, end effector position, end effector attitude, and raw readings from the force / torque sensors. The raw data is preprocessed: the influence of the end effector load gravity is subtracted from the raw force / torque data, and a coordinate transformation is performed to obtain the generalized external contact force. Low-pass filtering and numerical differentiation are performed on the position and attitude data to obtain the actual velocity and actual acceleration of the robot end effector.
[0067] This embodiment uses the aforementioned translation-rotation decoupled CEDS framework, with the nominal progress evolution term... The definition and expression are the same as the aforementioned general scheme. Based on the robot dynamics equations, assuming the control force equals the guiding force during the demonstration, the target velocity field of CEDS is derived by reverse calculation, and its expression is:
[0068] in, For the robot's inertia matrix, The matrix of Coriolis force and centrifugal force. For the generalized acceleration of the robot's end effector, It is an identity matrix.
[0069] The target speed is equal to the actual speed plus the inverse of the damping coefficient and the result of the combined calculation of inertial force, Coriolis force, and external contact force, which ultimately constructs the complete training dataset.
[0070] To verify the role of contact force, two sets of ablation comparison datasets were constructed: one set included contact force in the calculation of the target velocity field, and the other set ignored contact force and directly used the actual velocity as the target velocity field. The remaining data processing methods for the two sets were completely identical. The ablation experiment results showed that the success rate of the method including contact force training was 93.33% (drag-and-drop teaching), while the success rate of the method without contact force was 0% (drag-and-drop teaching). The translational stiffness of the method with contact force training was about one order of magnitude higher than that without contact force training. It exhibited task-adaptive directional stiffness characteristics in each stage of assembly. During the insertion stage, it could maintain a continuous contact force of about 22N along the insertion direction to ensure alignment. The system without contact force, on the other hand, lost contact during the rotation stage due to insufficient stiffness, resulting in assembly failure.
[0071] Example 4
[0072] This embodiment verifies the effectiveness of the present invention in making stiffness dependent on both robot state and task progress simultaneously through virtual task progress coordinates. The requirements for impedance characteristics differ significantly at different stages of a contact task (free approach, contact search, alignment, insertion), and state-dependent models cannot capture these task-process-related changes. The specific implementation process is as follows: First, construct virtual task progress coordinates, normalize the demonstration trajectory by time, and the coordinate expression is:
[0073] in, The current task execution time. The nominal total task duration, The scaling factor for state-schedule coupling strength controls the coupling strength between state and schedule. The task starts at schedule 0 and reaches a schedule of... The event will end at that time.
[0074] Nominal progress evolution item The aforementioned general piecewise function form is adopted. The actual schedule evolution rate is jointly determined by the nominal schedule evolution term and the state-schedule coupling term, that is:
[0075] When the robot's state deviates from the average reference trajectory, the schedule coupling term generates an additional rate of change, causing the task schedule to adjust adaptively, thus forming a two-way coupling between state and schedule. The smaller the scaling factor, the smaller the ratio of the schedule coupling term to the nominal schedule term, and the weaker the influence of the state on the schedule evolution.
[0076] By evaluating the trajectory reproduction error (dynamic time warping distance and sweep error area) under different scaling factors, the optimal coupling strength parameter is selected. As the scaling factor increases, the reproduction error first decreases rapidly and then gradually stabilizes. In this embodiment, the scaling factor is selected as 0.001 and the progress convergence time constant is set to 0.2 seconds.
[0077] Through this mechanism, stiffness can change simultaneously with the robot's state and task progress, exhibiting differentiated stiffness characteristics at different stages of the task: During the non-contact phase, the stiffness is at its maximum along the insertion direction, allowing the shaft to stably approach and establish contact with the hole. During the lateral alignment stage, the stiffness in the insertion direction and the lateral direction are enhanced. The former maintains the contact force, while the latter overcomes the friction force to ensure the motion trajectory. During the rotational alignment phase, the direction of maximum stiffness is perpendicular to the rotation axis to ensure stable contact during rotation. The maximum stiffness during the insertion phase is along the insertion direction to ensure stable insertion.
[0078] This skill can be successfully generalized to assembly tasks with different geometries. The assembly success rate is 86.67% for round holes (0.04mm gap) and 90.00% for triangular holes (0.28mm gap), demonstrating good cross-task versatility.
[0079] Example 5
[0080] This embodiment conducts a physical experiment on a real six-DOF robot platform to perform square shaft-hole assembly, verifying that the CEDS-based variable impedance skill learning method can enable the robot to learn and successfully execute precision assembly tasks with rich contact, while maintaining high success rate and system stability (passivity). The overall experimental scenario is as follows: Figure 4 As shown, the coordinate system and workpiece dimensions are as follows: Figure 5 As shown, the specific implementation process is as follows: The experiment used a 6-DOF Chinrobo CRB-7 collaborative robot equipped with joint torque sensors; it was operated via a Beckhoff CX2030 industrial controller with an EtherCAT communication frequency of 4kHz; and an ATI Gamma IP65 force / torque sensor was installed at the end effector. The experimental workpiece consisted of a square shaft with a side length of 20.00 mm and a square hole with a side length of 20.14 mm, with a single-sided gap of 0.14 mm.
[0081] The experiment first identified the robot's dynamic parameters using an iterative method. Based on the identification results, the inertia matrix, Coriolis force matrix, and gravity vector were calculated, and model-free compensation was performed for joint friction. Subsequently, six demonstration trajectories were collected through physical guidance. For each trajectory, time, position, attitude, and force / torque data were recorded, and all trajectories were normalized to a duration of 10 seconds.
[0082] The PICNN-based CEDS is trained using the aforementioned two-stage method, with the nominal progress evolution term... The definitions and expressions are the same as the aforementioned general scheme; the key parameter settings are as follows: rotational damping coefficient 50N. m s / rad², translational damping coefficient 800N s / m²; the rotational baseline stiffness matrix is 40 times the identity matrix N. m / rad, the translation baseline stiffness matrix is 100 times the identity matrix N / m; the state-progress coupling strength scaling factor is 0.001, and the progress convergence time constant is 0.2 seconds. After training, the velocity field output by CEDS is substituted into the decoupled variable impedance control law to generate robot joint torque control commands to execute assembly tasks; the force sensor monitors the contact force in real time and automatically adjusts the speed command when it exceeds the safety threshold.
[0083] Experimental results showed that 28 out of 30 assembly tests were successful, achieving a success rate of 93.33%. The average assembly time for successful tests was 9.900 seconds, with a standard deviation of 0.074 seconds. In 125 tests with different positional errors, the overall success rate reached 83.2%, verifying the good robustness of the learned skill against visual errors, calibration errors, and kinematic errors. This skill can be further generalized to assembly tasks with different geometries, such as round shaft-hole and triangular shaft-hole assembly. The success rate for round shaft-hole assembly (0.04mm gap) was 86.67%, and the success rate for triangular shaft-hole assembly (0.28mm gap) was 90.00%.
[0084] After completing the above core implementation steps, this solution possesses significant advantages in engineering implementation and industrial application. At the hardware level, all involved hardware and software are industrial standard configurations, requiring no special customized equipment and can be directly deployed on existing industrial robot platforms. At the data acquisition level, demonstration data is collected through physical-guided drag-and-drop teaching, without relying on remote operation equipment or professional programming skills; frontline workers can complete the operation after simple training. At the computing resource level, PICNN adopts a lightweight network structure with a small number of parameters and fast inference speed, enabling real-time operation on industrial controllers without the need for expensive computing resources such as GPUs. At the system integration level, the solution outputs the CEDS velocity field and stiffness matrix, which can be seamlessly integrated with existing impedance control architectures. The control law structure is simple and compatible with existing robot control software stacks.
[0085] This invention can be widely applied to multiple industrial scenarios: in precision assembly scenarios, it can meet the precision assembly requirements such as inserting shafts into holes with extremely small gaps, and is suitable for high-end manufacturing fields such as electronics, automobiles, and aerospace; in surface treatment scenarios, it can adaptively adjust the stiffness of different processing stages to achieve constant contact force control for tasks such as grinding and polishing; in human-machine collaboration scenarios, it can learn impedance adjustment strategies from human demonstrations to adapt to the flexible assembly requirements of customized production.
[0086] From a market perspective, the global industrial robot market is projected to reach $24 billion to $30 billion by 2026, while the Chinese industrial robot market is expected to reach $9 billion by 2034. Variable impedance control, as a core technology for improving robot compliance and adaptability, has a clear demand for industrialization. In summary, this invention boasts high technological maturity, low hardware barriers, and well-defined application scenarios, possessing favorable conditions for industrialization.
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0088] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A robot variable impedance skill learning method based on conservative extended dynamic systems, characterized in that, Includes the following steps: S1: Establish a Cartesian space dynamics model for the robot, and decompose the generalized state of the robot's end effector into mutually independent translational and rotational states; S2: Construct a conservative extended dynamic system CEDS framework with translation-rotation decoupled; the framework includes a translation CEDS and a rotation CEDS that share the same virtual task progress coordinates, and the translation CEDS and rotation CEDS are configured with independent scalar damping coefficients respectively; S3: Collect physical guidance demonstration data containing kinematic and contact force information, infer the target velocity field based on the robot dynamics equations, and construct the CEDS training dataset; S4: The potential function of the conservative extended dynamic system is constructed using PICNN. The inherent structure of the network ensures that the potential function is strictly convex with respect to the robot state, so that the corresponding stiffness matrix satisfies symmetry and uniform positive definiteness. The potential function takes both the robot state and the virtual task progress coordinates as input, so that the stiffness depends on both the robot state and the task progress. S5: Train CEDS based on the training dataset, substitute the trained velocity field output into the decoupled variable impedance control law, and generate robot joint torque control commands.
2. In the robot variable impedance skill learning method based on conservative extended dynamic systems according to claim 1, in step S2, both the translational CEDS and the rotational CEDS include the state velocity field equation and the virtual task progress evolution equation, the expressions of which are as follows: Translation of CEDS: Rotating CEDS: in, The translation state vector, Let the rotation state vector be... For virtual task progress coordinates; For the translational velocity field, For rotational velocity field; , For state-schedule coupling terms, For nominal progress evolution items; translational CEDS and rotational CEDS maintain progress synchronization through shared virtual task progress coordinates and virtual force adjustments; The expression is as follows: in, Indicates the task execution cycle, parameters This determines the strength of the coupling between the robot's state and the task progress; It is a positive real number used to determine The location of the equilibrium point.
3. In the robot variable impedance skill learning method based on a conservative extended dynamic system according to claim 1, in step S2, the decoupled variable impedance control law includes a generalized control force equation and a virtual progress adjustment force equation, expressed as: in, The generalized control force applied to the controller, This is for the robot's gravity compensation. The rotational scalar damping coefficient is... The translational scalar damping coefficient; This represents the actual rotational speed of the robot's end effector. This represents the actual translational speed of the robot's end effector. This is a virtual progress adjustment force used to synchronize the virtual task progress evolution rate of the translational and rotational CEDS. , For state-schedule coupling terms, For the nominal schedule evolution item, This represents the actual rate of change of the virtual task progress coordinates.
4. In the robot variable impedance skill learning method based on conservative extended dynamic systems according to claim 1, the specific steps for constructing the training dataset in step S3 are as follows: S31: Collect demonstration data through physical guidance, and record timestamps, end position, end attitude, and raw readings of force / torque sensors; S32: Preprocess the raw force / torque data, remove the influence of end-load gravity, and perform coordinate transformation to obtain the generalized external contact force. ; S33: Perform low-pass filtering and numerical differentiation on the position and attitude data to obtain the actual velocity and actual acceleration of the robot end effector; S34: Based on the robot dynamics equations, assuming the control force equals the guiding force, the target velocity field of CEDS is derived by reverse calculation, and its expression is: in, For the robot's inertia matrix, The matrix of Coriolis force and centrifugal force. For the generalized acceleration of the robot's end effector, It is an identity matrix.
5. In the robot variable impedance skill learning method based on a conservative extended dynamic system according to claim 1, in step S4, the potential function is decomposed into a state-progress coupling component and a pure progress component, expressed as: in, The potential function components, which depend on both the robot's state and the task progress, are learned by a partially input convex neural network. The potential function component that depends only on the task progress is obtained by analytical integration of the nominal progress evolution function; For component identification, Corresponding rotational component, Corresponding translation component.
6. In the robot variable impedance skill learning method based on conservative extended dynamic systems according to claim 5, in step S4, the state-progress coupling potential function component is composed of the superposition of the PICNN learning potential function and the baseline stiffness potential function, and its expression is: in, Baseline stiffness potential function The expression is: in, The average reference trajectory corresponding to the task progress is learned by a multilayer perceptron network. The baseline stiffness matrix is diagonally positive definite; the baseline stiffness potential function is used to ensure that the stiffness matrix is not lower than a preset threshold and to maintain consistent positive definiteness.
7. In the robot variable impedance skill learning method based on conservative extended dynamic system according to claim 1, in step S4, the partially input convex neural network takes the robot state as convex input and the virtual task progress coordinates as non-convex input; the convex path of the network adopts non-negative weights and convex non-decreasing activation functions to ensure that the output potential function is strictly convex with respect to the robot state, and its Hessian matrix, i.e., stiffness matrix, naturally satisfies symmetry and uniform positive definiteness.
8. The robot variable impedance skill learning method based on conservative extended dynamic systems according to claim 1, wherein the virtual task progress coordinates are constructed as follows: in, The current task execution time. The nominal total task duration, The scaling factor is the state-schedule coupling strength; the nominal evolution law of the virtual task schedule coordinates is a piecewise function, which evolves at a constant rate during the task execution phase and enters the convergence phase after the task ends and asymptotically approaches the final value.
9. The robot variable impedance skill learning method based on a conservative extended dynamic system according to claim 1, wherein in step S5, the training of the conservative extended dynamic system adopts a two-stage training method: Phase 1: Train the multilayer perceptron network to learn the average reference trajectory under different task progresses, and optimize the network parameters by minimizing the trajectory fitting error; Phase 2: The joint demonstration dataset and the constraint dataset are used to train the input convex neural network, and the network parameters are optimized using a composite loss function. The composite loss function includes velocity field fitting loss and task cycle constraint loss, which are used to align the demonstration speed and ensure the nominal task duration, respectively.
10. The robot variable impedance skill learning method based on conservative extended dynamic system according to claim 1, wherein the stiffness matrix corresponding to the variable impedance control is a block diagonal structure, and the rotational stiffness block and the translational stiffness block are independent of each other.