Demonstration video-based cross-modality robot learning method and system

CN122550647BActive Publication Date: 2026-09-22SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611048070.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-22
Estimated Expiration
2046-07-15

AI Technical Summary

Technical Problem

[0004]在机械臂模仿学习领域,现有技术往往是利用数据集训练的多样本学习模式,然而,该方法存在以下技术缺陷:其一,该方法依赖演示轨迹数据增强和模型训练,部署至新构型机械臂时需重新采集数据并微调模型参数,跨形态迁移成本高、周期长;其二,在物体被遮挡时仅能依赖模型基于历史数据预测未来轨迹,预测误差会随遮挡时间呈指数级累积;其三,扩散模型对毫米级微位移信号不敏感,容易丢失按压、旋转、拧动等精细操作的关键运动特征

Benefits of technology

在本发明中,实时计算物体遮挡置信度与位移尺度因子,生成自适应动态融合权重,引导机械臂在物体主导与手部引导之间平滑过渡,当物体遮挡置信度低于遮挡判定阈值或物体位移尺度小于微位移判定阈值时,通过自适应动态融合权重使融合目标轨迹由手部轨迹主导,解决物体遮挡断裂与微位移捕捉难题,显著提升了模仿学习的鲁棒性与精细操作能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550647B_ABST
    Figure CN122550647B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of mechanical arm imitation learning, in order to solve the problem that the existing object occlusion or micro displacement is difficult to capture, resulting in inaccurate planning of the mechanical arm, a cross-modal mechanical arm imitation learning method and system based on demonstration video are proposed, based on human demonstration video, the object motion trajectory sequence and the operator hand trajectory sequence are extracted synchronously; the object occlusion confidence and the displacement scale factor are calculated in real time, and the adaptive dynamic fusion weight is generated; based on the adaptive dynamic fusion weight, the current object position, the current hand position and the hand speed are weighted and fused to obtain the fused target trajectory; the fused target trajectory is mapped from the human operation space to the target mechanical arm joint space, and the control instruction sequence executed by the target mechanical arm is generated combined with the kinematics parameters of the target mechanical arm to reproduce the demonstration operation. The present application solves the problem of object occlusion breakage and micro displacement capture, and significantly improves the robustness and fine operation ability of imitation learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of robotic arm imitation learning, and particularly relates to a cross-morphological robotic arm imitation learning method and system based on demonstration videos. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Imitation learning based on human demonstration videos, as an efficient means of human-machine skill transfer, has been widely studied and used to solve robot dexterity tasks due to its intuitiveness and low-cost data acquisition characteristics.

[0004] In the field of robotic arm imitation learning, existing technologies often utilize multi-sample learning models trained on datasets. However, this method has the following technical drawbacks: First, it relies on demonstration trajectory data augmentation and model training. When deployed to a new robotic arm configuration, data needs to be re-collected and model parameters fine-tuned, resulting in high costs and long cycles for cross-morphological transfer. Second, when an object is occluded, it can only rely on the model to predict future trajectories based on historical data, and the prediction error accumulates exponentially with the occlusion time. Third, the diffusion model is not sensitive to millimeter-level micro-displacement signals and easily loses key motion features of fine operations such as pressing, rotating, and twisting.

[0005] Zero-shot imitation learning techniques also exist in this field, which typically extract the motion trajectory of objects from demonstration videos and directly drive the robotic arm to execute the movements after coordinate system transformation. However, this type of method has the following technical drawbacks: First, when objects in the demonstration video are frequently obscured by hands, tools, or background clutter, the object trajectory extraction may experience breaks, jumps, or systematic deviations, leading to inaccurate motion planning for the robotic arm. Second, for fine operations with extremely small displacements, such as pressing switches or rotating knobs, the changes in the visual features of the object are subtle, making it difficult for traditional trajectory extraction algorithms to capture effective motion signals, thus preventing the robotic arm from completing micro-displacement tasks. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, this invention provides a cross-morphological robotic arm imitation learning method and system based on demonstration videos. By adaptively and dynamically fusing weights, the robotic arm smoothly transitions between object-led and hand-led modes, solving the problems of object occlusion and breakage and micro-displacement capture, and significantly improving the robustness and precision operation capabilities of imitation learning.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a cross-morphological robotic arm imitation learning method based on demonstration videos, including: Based on human demonstration videos, the sequence of human motion trajectories and the sequence of operator hand trajectories are extracted simultaneously; Real-time calculation of object occlusion confidence and displacement scale factor to generate adaptive dynamic fusion weights; An adaptive dynamic weighted fusion mechanism is adopted to weight and fuse the current object position, hand position, and hand velocity to generate the fused target trajectory. When the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory dominates the fused target trajectory through adaptive dynamic fusion weight. The target trajectory is mapped from the human operating space to the joint space of the target robotic arm, and the kinematic parameters of the target robotic arm are combined to generate a sequence of control commands to be executed by the target robotic arm in order to reproduce the demonstration operation.

[0008] Secondly, the present invention provides a cross-morphological robotic arm imitation learning system based on demonstration videos, comprising: The extraction module is configured to simultaneously extract the body motion trajectory sequence and the operator's hand trajectory sequence based on the human demonstration video; The dynamic fusion weight module is configured to: calculate the object occlusion confidence and displacement scale factor in real time, and generate adaptive dynamic fusion weights. The target trajectory fusion module is configured to: adopt an adaptive dynamic weighted fusion mechanism to weight and fuse the current object position, hand position, and hand velocity to generate a fused target trajectory; when the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory will dominate the fused target trajectory through adaptive dynamic fusion weights; The learning module is configured to: map the fused target trajectory from the human operating space to the target robotic arm joint space, and generate a sequence of control commands to be executed by the target robotic arm in combination with the kinematic parameters of the target robotic arm, so as to reproduce the demonstration operation.

[0009] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.

[0010] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.

[0011] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0012] The above one or more technical solutions have the following beneficial effects: In this invention, the occlusion confidence and displacement scale factor of an object are calculated in real time to generate adaptive dynamic fusion weights, which guide the robotic arm to smoothly transition between object-led and hand-led operation. When the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is less than the micro-displacement judgment threshold, the adaptive dynamic fusion weights make the fusion target trajectory dominated by the hand trajectory, solving the problems of object occlusion breakage and micro-displacement capture, and significantly improving the robustness and fine operation capability of imitation learning.

[0013] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0014] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0015] Figure 1 This is a flowchart of the cross-morphological robotic arm imitation learning method based on a demonstration video in an embodiment of the present invention; Figure 2 This is a schematic diagram of the dual-trajectory dynamic adaptive fusion module and the velocity feedforward compensation structure in an embodiment of the present invention; Detailed Implementation It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0016] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0017] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0018] Example 1 This embodiment discloses a cross-morphological robotic arm imitation learning method based on demonstration videos. It parses complex human demonstration videos into occlusion-resistant, high-precision fused trajectories and achieves cross-configuration transfer through morphology-independent mapping. This method decouples the imitation learning task into two stages: dual-trajectory dynamic fusion and zero-shot kinematic mapping. The fusion stage allows the system to adaptively adjust the object and hand trajectory weights based on visual confidence at a higher level, driven by a smooth switching mechanism. The mapping stage consists of analytical geometric constraints, allowing adaptation to new robotic arms without neural network fine-tuning.

[0019] like Figure 2As shown, this embodiment introduces a dynamic adaptive weighted fusion module based on occlusion confidence and displacement scale, aiming to solve the failure problem of a single visual source in occlusion and micro-displacement scenarios. This embodiment adopts a data processing architecture of "parallel extraction - spatiotemporal alignment - dynamic fusion". During the fusion process, the reliability index of the object trajectory and the task displacement scale are calculated in real time to generate continuously differentiable fusion weights, guiding the robotic arm to smoothly transition between object-led and hand-led modes. In addition, cross-morphological mapping in this embodiment is completed through analytical kinematic constraints. Using the Cartesian trajectory output from the upper-layer fusion as a condition, combined with the target robotic arm parameters, the joint space commands are directly solved.

[0020] Combination Figure 1 The method in this embodiment specifically includes the following processes: Step 1: Based on the human demonstration video, simultaneously extract the object motion trajectory sequence and the operator's hand trajectory sequence, and perform spatiotemporal alignment between the object motion trajectory sequence and the operator's hand trajectory sequence.

[0021] Specifically, the process begins by acquiring a sequence of video demonstrations of human operation. On one hand, an instance segmentation model is used to extract body masks and calculate the centroid 3D coordinate sequence. On the other hand, the HaMeR model was used to extract 21 3D key points of the hand. The key point sequences of the fingertips or wrists were smoothed by Kalman filtering to obtain the original hand trajectory. .

[0022] Due to the frame rate difference and coordinate system offset between the two trajectories, spatiotemporal alignment preprocessing is required. First, based on the original video frame rate, cubic spline interpolation is used to fill in missing frames in the object trajectory sequence and hand trajectory sequence caused by occlusion or inference delay. Then, the two sequences are resampled to a unified control frequency. First, an isochronous time step sequence is obtained. Second, based on the pre-calibrated camera intrinsic and extrinsic transformation matrices, the object pixel coordinates are back-projected to the world coordinate system. Simultaneously, the normalized hand coordinates output by HaMeR are mapped to the same physical reference system through scale factor and rigid body transformation, completing dimensional unification and origin calibration, and outputting the aligned synchronization sequence. and The superscript t indicates the t-th control time; T indicates the total number of control times; This indicates the trajectory of the object at time t. This represents the hand trajectory at time t.

[0023] Step 2: An adaptive dynamic weighted fusion mechanism is adopted to weight and fuse the current object position, hand position, and hand velocity to generate the fused target trajectory. When the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory is made dominant by the adaptive dynamic fusion weight.

[0024] Specifically, this embodiment introduces a dynamic adaptive weight generator for trajectory fusion and completes task execution based on this dynamic adaptive weight.

[0025] The weights for all time steps, based on real-time calculated occlusion confidence and displacement scale, represent a high-level understanding of the current operational state. This design facilitates the adaptation of the corresponding imitation learning system to highly dynamic occlusion environments or fine-grained micro-displacement operation scenarios.

[0026] At time t, acquire the aligned object trajectory. With hand trajectory First, calculate the object occlusion confidence at control time t. ,in, The visibility ratio of the object at time t is controlled by the ratio of the current mask area to the reference area. To assess the confidence level of the model output at control time t, the displacement of the object at adjacent time steps is calculated. ;in, This indicates that t controls the trajectory of the object at time t. This represents the trajectory of the object at control time t-1; This represents the L2 norm.

[0027] The detection model takes a single frame from a human demonstration video as input and outputs a segmentation mask of the target object and its confidence score. The detection model can be from the YOLO series, such as YOLOv8, YOLOv9, and YOLOv10.

[0028] Secondly, an adaptive weight allocation function based on smooth switching logic is constructed. :

[0029]

[0030] in, For the standard Sigmoid function, This is the sensitivity adjustment coefficient. The threshold for determining occlusion. The threshold for determining micro-displacement; Let t be the confidence level of object occlusion at control time t; This indicates the displacement of an object at adjacent time steps. Let t be the weight corresponding to the hand trajectory at time t. The weights corresponding to the object's trajectory at time t are used to control the time.

[0031] This function guarantees that... or When any condition is met, The smoothness approaches 1, achieving a seamless transition to hand-centric trajectories.

[0032] Next, the aligned hand position sequence is temporally differentiald and a cutoff frequency is applied. A second-order Butterworth low-pass filter (LPF) is used to filter out high-frequency detection noise, resulting in a smooth end-effector velocity vector. .

[0033] Finally, the fused target trajectory is generated. :

[0034] in, This is the velocity feedforward compensation coefficient, and the velocity feedforward term... It provides motion trend prediction during occlusion or micro-displacement phases, effectively suppressing end effector jitter in robotic arms. Let t be the velocity vector of the hand's end effector at time t.

[0035] When the object occlusion confidence is lower than the occlusion determination threshold or the object displacement scale is smaller than the micro-displacement determination threshold, the weight corresponding to the hand trajectory is increased, so that the fused target trajectory switches to hand trajectory dominance; when the object occlusion confidence is higher than the occlusion determination threshold and the object displacement scale is greater than the micro-displacement determination threshold, the weight corresponding to the object trajectory is dynamically increased, so that the fused target trajectory switches to object trajectory dominance.

[0036] This embodiment proposes a dual-trajectory dynamic adaptive fusion method to achieve occlusion-resistant trajectory reconstruction. It captures both global object pose (object-driven) and local hand interaction intent (hand-guided), avoiding the trajectory fragmentation and suboptimal solution problems inherent in traditional single-visual-source-dependent methods. The hierarchical perception architecture decouples complex feature extraction from smooth trajectory generation, significantly reducing signal processing complexity in occluded scenarios.

[0037] First, the concept of adaptive dynamic fusion weights is embedded in the fusion decision layer to complete the allocation of dominance. This allows the system to learn the current trajectory switching strategy from a higher-level perspective based on real-time occlusion confidence and displacement scale. In addition, the Sigmoid smooth transition mechanism ensures weight switching between two time steps, allowing the macroscopic transport and microscopic interaction motion modes to mutually promote each other.

[0038] This embodiment presents a vision-based, autonomous, and adaptive velocity feedforward compensation method for fine-grained tasks. The underlying trajectory enhancement is achieved through a feedforward compensation framework, composed of end-effector velocity vectors extracted via low-pass filtering. This guides the control system to autonomously select the compensation intensity for the current state, maximizing its tracking accuracy. During sequence trajectory generation, an intelligent velocity compensation mechanism effectively suppresses position lag and detection noise, dynamically improving the success rate of the operation. The underlying trajectory output is achieved through weighted position fusion and velocity feedforward compensation, enabling the robotic arm to learn how to generate continuous control commands under the conditions of upper-level dynamic weights and spatiotemporal alignment information.

[0039] Step 3: Map the fused target trajectory from the human operation space to the joint space of the target robotic arm, and generate a sequence of control commands to be executed by the target robotic arm in combination with the kinematic parameters of the target robotic arm to reproduce the demonstration operation.

[0040] Specifically, a morphology-independent analytical kinematic mapping module, i.e., a morphology-independent mapping network, is constructed to integrate the fused Cartesian trajectories. The morphology-independent analytical kinematics mapping module maps the human operating space to the target robotic arm joint space. It employs a purely analytical computational architecture and does not contain any trainable neural network parameters. Therefore, when deployed to heterogeneous robotic arms with different degrees of freedom and link layouts, there is no need to re-collect demonstration data or perform model fine-tuning.

[0041] DH parameter table based on target robotic arm Construct homogeneous transformation matrices between adjacent links, and then obtain the total transformation relationship from the base coordinates to the end effector coordinate system:

[0042] Where n represents the degrees of freedom of the robotic arm. The DH parameter represents the total homogeneous transformation matrix of the robotic arm, describing the pose transformation from the base coordinate system (0 coordinate system, the fixed base of the robotic arm) to the end effector (n coordinate system, the end effector of the robotic arm). Let represent the link length of the i-th mechanical link from the base of the robotic arm. This represents the link torsion angle of the i-th mechanical link. This represents the link offset of the i-th mechanical link. This represents the rotation angle of the i-th joint. R(Q) is the joint angle vector; R(Q) is the end-effector attitude rotation matrix; p(Q) is the end-effector position vector, and the superscript T indicates transpose.

[0043] Next, inverse kinematics is solved in real time for the fused Cartesian target pose. Find satisfaction joint angle vector Where the superscript t represents the t-th control time; To merge target trajectories The corresponding target transformation matrix is ​​derived from the fused target trajectory. The 3D target position and 3D target pose are analytically constructed from the fused target trajectory. Separate the target attitude angle and target position vector. Then, the target attitude angle is converted into a target attitude rotation matrix. Finally, the combination and The target homogeneous transformation matrix is ​​obtained. :

[0044] For industrial robotic arms with six degrees of freedom or less that meet the Pieper criterion, the analytical inverse kinematics method is adopted. The inverse kinematics of the position of the first three joints and the inverse kinematics of the posture of the last three joints are solved by the wrist decomposition strategy. The calculation delay is less than 1ms, which meets the 30Hz real-time control requirement.

[0045] For a seven-DOF redundant robotic arm or a robotic arm with a special configuration, the Newton-Raphson numerical inverse method is used, and the iterative formula is as follows:

[0046] in, The Moore-Penrose pseudoinverse of the Jacobian matrix. Let be the error vector between the current pose and the target pose. This is the joint angle vector of the robotic arm at the k-th iteration, which represents the joint state at the current iteration step.

[0047] When multiple solutions exist for the inverse kinematics, the minimum joint displacement criterion is used to select the optimal solution to ensure continuity.

[0048] in, This is the set of all inverse solutions corresponding to the current pose. This is the joint angle vector.

[0049] To avoid impact vibrations in the robotic arm, fifth-order polynomial interpolation is used on the discrete joint angle sequence obtained from the original inverse kinematics solution to achieve third-order continuity of position, velocity, and acceleration.

[0050] in, For the i-th joint in interpolation time The continuous joint angle positions at the given location represent the continuous trajectory to be solved during the interpolation process. The time interval between adjacent keyframes. ~ The coefficients of the fifth-order polynomial corresponding to the i-th joint are uniquely determined by a system of linear equations consisting of the position, velocity, and acceleration boundary conditions of the keyframes at both ends of the interpolation interval. Physical hard constraints on the motors are applied to the interpolated joint motion parameters.

[0051] in, This represents the instantaneous angular velocity of the i-th joint at time t. This represents the instantaneous angular acceleration of the i-th joint at time t. This represents the maximum permissible angular velocity of the i-th joint, determined by the performance parameters of the robotic arm's joint motors. This represents the maximum permissible angular acceleration of the i-th joint, which is determined by the performance parameters of the robotic arm's joint motor.

[0052] If the constraints are exceeded, a time scaling method is used to extend the interpolation interval, ensuring all parameters meet the motor drive capability while maintaining the trajectory shape. Finally, the smoothed joint angle sequence is discretized at a control frequency of 30Hz to generate control commands executable by the underlying servo driver.

[0053] This embodiment introduces a shape-independent analytical kinematic mapping module for cross-domain trajectory planning, and generates instructions for robotic arms with different configurations based on this mapping relationship. The analytical mapping process based on the kinematic parameters of the target robotic arm represents a high-level understanding of the geometric constraints of the human demonstration space and the robot joint space. This design facilitates the adaptation of the imitation learning system corresponding to this embodiment to heterogeneous robotic arm platforms with different degrees of freedom and different link layouts, or to highly complex industrial environments involving multi-model collaboration, achieving truly zero-shot rapid deployment.

[0054] In this embodiment, the underlying control of the robotic arm is driven by the fused target trajectory. Based on the hand guidance signal output by the upper dynamic weight, the decision is made based on continuous spatiotemporal alignment information, so that the capability of the basic control strategy characterizes the robustness against occlusion or the precision of micro-operation.

[0055] Example 2 The purpose of this embodiment is to provide a cross-morphological robotic arm imitation learning system based on demonstration videos, including: The extraction module is configured to simultaneously extract the body motion trajectory sequence and the operator's hand trajectory sequence based on the human demonstration video; The dynamic fusion weight module is configured to: calculate the object occlusion confidence and displacement scale factor in real time, and generate adaptive dynamic fusion weights. The target trajectory fusion module is configured to: adopt an adaptive dynamic weighted fusion mechanism to weight and fuse the current object position, hand position, and hand velocity to generate a fused target trajectory; when the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory will dominate the fused target trajectory through adaptive dynamic fusion weights; The learning module is configured to: map the fused target trajectory from the human operating space to the target robotic arm joint space, and generate a sequence of control commands to be executed by the target robotic arm in combination with the kinematic parameters of the target robotic arm, so as to reproduce the demonstration operation.

[0056] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.

[0057] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0058] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0059] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.

[0060] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0061] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.

[0062] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.

[0063] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0064] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.

[0065] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0066] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A cross-morphological robotic arm imitation learning method based on demonstration videos, characterized in that, include: Based on human demonstration videos, the sequence of human motion trajectories and the sequence of operator hand trajectories are extracted simultaneously; Real-time calculation of object occlusion confidence and object displacement scale factor to generate adaptive dynamic fusion weights; The adaptive dynamic fusion weights are specifically as follows: in, The weights corresponding to the hand trajectory at time t are used to control the movement. Let t be the weight corresponding to the object's trajectory at time t. For the standard Sigmoid function, and These are the occlusion sensitivity coefficient and the displacement scale sensitivity coefficient, respectively. The threshold for determining occlusion. The threshold for determining micro-displacement. Let t be the confidence level of object occlusion at time t. The displacement of the object in adjacent time steps; in, Let t control the visibility ratio of objects at time t; To detect the confidence level of the model output at control time t; This indicates that t controls the trajectory of the object at time t. This represents the trajectory of the object at control time t-1; Represents the L2 norm; , These are the aligned object trajectory and the hand trajectory, respectively. An adaptive dynamic weighted fusion mechanism is adopted to weight and fuse the current object position, hand position, and hand velocity to generate the fused target trajectory. When the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory dominates the fused target trajectory through adaptive dynamic fusion weight. The target trajectory is mapped from the human operating space to the joint space of the target robotic arm, and the kinematic parameters of the target robotic arm are combined to generate a sequence of control commands to be executed by the target robotic arm in order to reproduce the demonstration operation.

2. The cross-morphological robotic arm imitation learning method based on demonstration videos as described in claim 1, characterized in that, When the object occlusion confidence is higher than the occlusion judgment threshold and the object displacement scale is greater than the micro-displacement judgment threshold, the weight corresponding to the object trajectory is dynamically increased, so that the fused target trajectory is dominated by the object trajectory.

3. The cross-morphological robotic arm imitation learning method based on demonstration videos as described in claim 1, characterized in that, The object occlusion confidence is determined based on the ratio of the object mask area to the complete reference mask area in frame t, and the object detection confidence; the object displacement scale is determined based on the object displacement in adjacent time steps.

4. The cross-morphological robotic arm imitation learning method based on demonstration videos as described in claim 1, characterized in that, Based on adaptive dynamic fusion weights, the current object position, current hand position, and hand velocity are weighted and fused to obtain the fused target trajectory, specifically: in, and These represent the aligned object position and hand position at time t, respectively. For velocity feedforward compensation coefficient, Let t control the hand speed at time t. The weights corresponding to the object trajectory at time t control the time. The weights corresponding to the hand trajectory at time t are used to control the movement.

5. The cross-morphological robotic arm imitation learning method based on demonstration videos as described in claim 1, characterized in that, Also includes: Spatiotemporal alignment of the object's motion trajectory sequence with the operator's hand trajectory sequence is performed, specifically as follows: Cubic spline interpolation is used to fill in missing frames in the object trajectory sequence and hand trajectory sequence caused by occlusion or inference delay, and the object trajectory sequence and hand trajectory sequence are resampled to a unified control frequency; Based on the camera intrinsic and extrinsic transformation matrices, the object pixel coordinates are back-projected to the world coordinate system, and the normalized coordinates of the hand are mapped to the same world coordinate system through rigid body transformation, thus completing the unification of dimensions and origin calibration.

6. A cross-morphological robotic arm imitation learning system based on demonstration videos, characterized in that: include: The extraction module is configured to simultaneously extract the body motion trajectory sequence and the operator's hand trajectory sequence based on the human demonstration video; The dynamic fusion weight module is configured to: calculate the object occlusion confidence and displacement scale factor in real time, and generate adaptive dynamic fusion weights; the adaptive dynamic fusion weights are specifically as follows: in, The weights corresponding to the hand trajectory at time t are used to control the movement. Let t be the weight corresponding to the object's trajectory at time t. For the standard Sigmoid function, and These are the occlusion sensitivity coefficient and the displacement scale sensitivity coefficient, respectively. The threshold for determining occlusion. The threshold for determining micro-displacement. Let t be the confidence level of object occlusion at time t. The displacement of the object in adjacent time steps; in, Let t control the visibility ratio of objects at time t; To detect the confidence level of the model output at control time t; This indicates that t controls the trajectory of the object at time t. This represents the trajectory of the object at control time t-1; Represents the L2 norm; , These are the aligned object trajectory and the hand trajectory, respectively. The target trajectory fusion module is configured to: adopt an adaptive dynamic weighted fusion mechanism to weight and fuse the current object position, hand position, and hand velocity to generate a fused target trajectory; when the object occlusion confidence is lower than the occlusion judgment threshold or the object displacement scale is smaller than the micro-displacement judgment threshold, the hand trajectory will dominate the fused target trajectory through adaptive dynamic fusion weights; The learning module is configured to: map the fused target trajectory from the human operating space to the target robotic arm joint space, and generate a sequence of control commands to be executed by the target robotic arm in combination with the kinematic parameters of the target robotic arm, so as to reproduce the demonstration operation.

7. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-5.

9. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Deep learning-based billion-pixel video image stitching method and system

    CN121213346A

  • Method for learning operation track of robot from human demonstration video and computer

    CN121267889A