Humanoid robot teleoperation method and system based on digital twinning

By using digital twin technology and an unsupervised neural network redirection system, high-precision motion mapping and real-time status monitoring for teleoperation of humanoid robots were achieved, solving the problems of unnatural motion mapping and insufficient safety in existing technologies, and improving the naturalness and safety of teleoperation.

CN122077618APending Publication Date: 2026-05-26HUAZHONG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-02-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing humanoid robot teleoperation technology suffers from problems such as unnatural motion mapping, low accuracy, lack of intuitive status feedback and adaptive collision judgment, making it difficult to perform tasks safely and reliably in complex environments.

Method used

An unsupervised neural network redirection system based on digital twins is adopted to map human motion data into robot motion data. Combined with a communication system, a data service system, and a collision avoidance module, the system enables real-time visualization monitoring and adaptive adjustment of the robot's state.

Benefits of technology

It improves the naturalness and accuracy of teleoperation, enhances the immersive experience and intuitiveness of operation, improves the safety and reliability of the system, and supports flexible switching between online teleoperation and offline simulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122077618A_ABST
    Figure CN122077618A_ABST
Patent Text Reader

Abstract

The invention discloses a digital twinning-based humanoid robot teleoperation method and system, and belongs to the technical field of digital twinning and humanoid robots, and the system comprises a physical body, a twinning body, a redirection system, a communication system and a data service system. Human body action data are mapped into robot action data through an unsupervised neural network of a redirection system, the robot action data are transmitted to a physical body through a communication system to be executed, twin bodies synchronously display action states in real time, a data service system completes storage, filtering and format conversion, high-precision, visual and safe man-machine heterogeneous whole-body teleoperation is achieved, and the operation efficiency is improved. Contact judgment and abnormal early warning are supported, and the complex scene operation capacity of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital twin and humanoid robot technology, and particularly relates to a method and system for teleoperating a humanoid robot based on digital twin. Background Technology

[0002] In recent years, with the rise of the embodied intelligence concept and the development of humanoid robot technology, teleoperation technology has become an important bridge for promoting the intelligent evolution of robots. Through the "human-in-the-loop" interaction mode, operators can transmit their perception, decision-making, and action capabilities to the robot in real time, enabling it to perform tasks safely and controllably in complex real-world environments. However, existing humanoid robot teleoperation technology still faces the following challenges:

[0003] First, due to the significant differences in kinematic structure between humans and humanoid robots (i.e., heterogeneity), achieving natural, coordinated, and accurate motion mapping remains challenging. Existing methods mostly rely on inverse kinematics solutions or retargeting strategies based on optimization models. While inverse kinematics solutions can obtain robot joint information by aligning the end effector pose, the inverse solution process may result in no solution or multiple solutions. On the other hand, optimization model-based methods often require precise calibration and training of human-robot paired motion data, resulting in high data acquisition costs and limited generalization ability, making it difficult to adapt to diverse natural human movements.

[0004] Secondly, traditional teleoperation systems lack intuitive and comprehensive status feedback on the robot's execution process. Operators can typically only monitor the robot's status through limited sensor data or two-dimensional images, making it difficult to perceive its overall motion posture, joint load, and potential environmental interaction risks in real time, which affects the accuracy and safety of operation.

[0005] Furthermore, existing teleoperation systems often lack real-time judgment and adaptive adjustment mechanisms for the robot-environment contact state during operation. In confined or unstructured workspaces, robots may collide with surrounding objects due to forced motion mapping, leading to equipment damage or mission failure.

[0006] Digital twin technology, by constructing virtual mappings of physical entities, offers a new approach to achieving comprehensive, visual monitoring and prediction of robot states. However, how to deeply integrate digital twins with humanoid robot teleoperation, achieving high-fidelity, low-latency virtual-real synchronization while combining intelligent algorithms to solve heterogeneous motion mapping and safety issues, still requires further research.

[0007] Therefore, there is an urgent need in this field for a humanoid robot teleoperation solution that can integrate digital twin visualization monitoring and intelligent motion redirection to improve the intuitiveness of operation, the accuracy of motion mapping, and the overall safety of the system. Summary of the Invention

[0008] This invention proposes a teleoperation method and system for humanoid robots based on digital twins to solve the problems existing in the prior art.

[0009] To achieve the above objectives, the present invention provides a method for teleoperating a humanoid robot based on a digital twin, comprising:

[0010] Acquire human motion data;

[0011] The human motion data is mapped to robot motion data by using the unsupervised neural network of the redirection system.

[0012] The robot's motion data is transmitted to the physical body via a communication system;

[0013] Based on the robot's motion data, the virtual model is driven to synchronously display the motion state of the physical object;

[0014] The data service system stores, filters, and converts the human motion data and robot motion data, and then forwards the converted data to the physical entity and the twin entity through the communication system.

[0015] Optionally, acquiring the state information of the physical body includes: acquiring the humanoid robot's own state information and environmental information through the sensing module; and controlling the humanoid robot to perform actions based on the robot's motion data through the control module.

[0016] Optionally, the synchronous display of the driving virtual model includes: displaying the operating status information of the humanoid robot through a visual status monitoring interface.

[0017] Optionally, mapping human motion data to robot motion data via a redirection system includes: acquiring human motion data via a motion data acquisition module; and mapping the human motion data to robot motion data via a redirection network.

[0018] Optionally, the redirection network includes two encoders with identical structures and a decoder; the encoders are used to map human motion data and robot motion data to a shared latent space, respectively; the decoder is used to decode features in the shared latent space into robot motion data.

[0019] The present invention also provides a humanoid robot teleoperation system based on digital twins, comprising:

[0020] A physical entity, the physical entity including a sensing module for acquiring entity state information and a control module for performing actions;

[0021] A twin, the twin comprising a virtual model for mapping the posture of a physical body and a visual monitoring interface for displaying the operating status;

[0022] A redirection system, which is built on a neural network, is used to convert human motion data into robot motion data;

[0023] A communication system for transmitting data between modules;

[0024] A data service system is used to store, filter, and convert the human motion data and robot motion data.

[0025] Optionally, the sensing module is arranged on the head and chest of the humanoid robot to acquire environmental images and joint encoder information.

[0026] Optionally, the virtual model maps the posture of the physical object in real time at a 1:1 scale and supports both online and offline operating modes.

[0027] Optionally, the redirection system includes a motion data acquisition module and a redirection network; the motion data acquisition module is used to acquire human motion data; the redirection network is used to map the human motion data to robot motion data, and the redirection network is trained using a joint loss function that includes triplet loss, reconstruction loss and potential consistency loss.

[0028] Optionally, the redirection network includes two encoders with identical structures and a decoder; the encoders are used to map human motion data and robot motion data to a shared latent space, respectively; the decoder is used to decode features in the shared latent space into robot motion data.

[0029] Compared with the prior art, the present invention has the following advantages and technical effects:

[0030] This invention achieves high-precision, high-generalization motion mapping between the heterogeneous motion structures of the human body and robots through a redirection system based on unsupervised neural networks. This significantly improves the naturalness and accuracy of teleoperation without relying on difficult-to-obtain paired training data. By constructing a 1:1 synchronized digital twin with the physical robot, real-time 3D visualization and monitoring of the robot's motion state are achieved, enhancing the immersive experience and intuitive interaction. The system integrates contact state judgment and self-collision avoidance mechanisms, enabling real-time perception and avoidance of robot-environment interaction risks, improving the safety and reliability of teleoperation in complex scenarios. Multimodal time-series data is simultaneously collected and stored during teleoperation, providing a high-quality data foundation for subsequent robot imitation learning and autonomous skill training. The modular system architecture based on ROS supports flexible switching between online teleoperation and offline simulation verification, improving system scalability and engineering applicability. Attached Figure Description

[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0032] Figure 1 This is a schematic diagram of the teleoperation system structure for a humanoid robot based on digital twins in this embodiment;

[0033] Figure 2 This is a schematic diagram of the remote operating system operation process in this embodiment;

[0034] Figure 3 This is a schematic diagram illustrating the acquisition of human motion data in this embodiment;

[0035] Figure 4 This is a schematic diagram of the redirected neural network training in this embodiment;

[0036] Figure 5 This is a schematic diagram of the communication system framework in this embodiment;

[0037] Figure 6 This is a functional diagram of the data service system in this embodiment. Detailed Implementation

[0038] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0039] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0040] Example 1

[0041] like Figure 1 As shown, this embodiment provides a humanoid robot teleoperation system based on digital twins, including:

[0042] A physical entity, the physical entity including a sensing module for acquiring entity state information and a control module for performing actions;

[0043] A twin, the twin comprising a virtual model for mapping the posture of a physical body and a visual monitoring interface for displaying the operating status;

[0044] A redirection system, which is built on a neural network, is used to convert human motion data into robot motion data;

[0045] A communication system for transmitting data between modules;

[0046] A data service system is used to store, filter, and convert the human motion data and robot motion data.

[0047] Furthermore, the physical body is a 27-DOF humanoid robot body with joints throughout its entire body, excluding the head and finger degrees of freedom. Specifically, the robot has 1 degree of freedom in its waist, located in the torso for rotational movement relative to the lower limbs in the vertical direction; 14 degrees of freedom in the upper limbs, 7 degrees of freedom per arm; 12 degrees of freedom in the lower limbs, and 6 degrees of freedom per leg.

[0048] Furthermore, the physical body includes a perception module and a control module. The perception module is arranged on the head and chest of the humanoid robot to acquire environmental images and joint encoder information. The perception module mainly consists of visual sensors and joint motor encoders arranged on the robot's head, chest, and other positions to acquire information about the robot's surrounding environment and its own state. The control module includes a collision avoidance module and a command following module to achieve reasonable mapping of human posture while ensuring the robot's own safety.

[0049] Furthermore, the self-collision avoidance module is used to ensure the safe operation of the robot. This module relies on the contact state judgment function of the redirection system to perceive the contact state between the robot's limbs and the surrounding environment in real time. When a collision is detected, it can automatically limit the relevant joints to avoid the robot from performing forced redirection, thereby preventing damage to the limbs or the environment. The instruction following module is used to drive the robot to move along the predetermined trajectory processed by the self-collision avoidance module.

[0050] Optionally, the twin includes a virtual model and a visual status monitoring interface. The virtual model is a high-fidelity model running in real time in the Webots simulation environment, which can realize 1:1 real-time mapping of the robot's running posture. It can present the robot's whole-body teleoperation in a more intuitive and realistic way while ensuring stable and accurate human-robot motion redirection. The visual status monitoring interface is a numerical representation of the twin.

[0051] Furthermore, the virtual model includes two operating modes: online and offline. In the online mode, the robot and the virtual model run synchronously, while in the offline mode, the virtual model runs independently. The offline mode can be used to test the teleoperation mapping effect and play back recorded robot actions, thereby improving the visualization effect of teleoperation data.

[0052] Optionally, the redirection system is a key module for ensuring the similarity between human and robot limbs during teleoperation. Its core is a redirection strategy based on neural networks, which can ensure high precision and good robustness in the redirection of heterogeneous motions between the human body and the robot under various postures.

[0053] Furthermore, the redirection strategy is an unsupervised strategy with higher generalization ability. Its core idea is cross-domain similarity measurement. First, a similarity index is defined to capture the similarity between human and robot limbs. Then, two encoders trained using a reconstruction error loss function are used. , and a decoder , Transmitting human body movements into a shared potential space. Transmit the aforementioned robot actions into the shared potential space;

[0054] Furthermore, the cross-domain similarity index is the quaternion information of each limb of the human body and robot relative to the pelvis, and the human posture and robot posture are respectively represented by... and Indicate, and through and Construct a triple, with the anchor set to the human body's limb movements processed by the encoder and decoder. positive is the same as The limb movements of robots with high similarity negative is related to Low similarity Their similarity measurement formula is:

[0055] ;

[0056] in, Let be the quaternion of each limb of the human body relative to pelvis. Let be the quaternion of each limb of the robot relative to pelvis;

[0057] Optionally, the reconstruction error loss function is used to learn the retargeting function through the constructed triples:

[0058] ;

[0059] Furthermore, both the encoder and decoder adopt an optimized multi-layer perceptron (OMLP) structure. The main body of the network consists of one input mapping layer, four residual-enhanced hidden layers, and one output mapping layer. Each hidden layer integrates a residual module and a layer normalization structure, and introduces PReLU as an activation function to enhance nonlinear expressive power and training stability. The shared latent space dimension is set to 512.

[0060] Furthermore, the loss function comprises triplet loss, standard reconstruction loss, and potential consistency loss;

[0061] The triplet loss aims to identify highly similar samples in the potential shared space. and To bring samples with low similarity closer together and By pushing away, similar clustering can be achieved, but different separations can be achieved, resulting in triplet loss. as follows:

[0062] ;

[0063] in, The compensation coefficient;

[0064] The standard reconstruction loss is used to realize the hidden variables in the shared latent space. Standard reconstruction loss is achieved by directly mapping the decoder to robot data. for:

[0065] ;

[0066] The potential consistency loss is used to achieve a direct mapping from human motion to robot motion. for:

[0067] ;

[0068] The above redirection employs an end-to-end training strategy, enabling the encoder to learn the shared representation space of human and robot actions in an unsupervised manner. Then, the decoder directly reconstructs the robot joint information from the shared representation space, with a total loss of [missing information]. for:

[0069] ;

[0070] in, and These are the weight coefficients during the training process;

[0071] Optionally, the communication system, based on ROS topic communication, uses a network approach for data transmission and sharing. The entire communication network includes a central hub and four nodes: a data processing and relay center (Speedgoat), a human information node (Xsens), a robot information node (27-DOF humanoid robot), a virtual model node (Webots), and a visual status monitoring interface (App Designer) node. Its detailed functions and communication flow include:

[0072] The data processing and relay center uses the Speedgoat real-time control platform to receive, process and forward the running data of each node. The device can import and run real-time models based on Simulink, realize data interaction through the Ethernet interface, and obtain data from human information nodes and robot information nodes through the ROS subscription module.

[0073] The human information node is used to acquire human motion data and send the collected information to the system's data processing and relay center through the ROS publishing module to realize the real-time transmission and sharing of human motion information.

[0074] Furthermore, the human motion information acquired by the Xsens motion capture device is transmitted to the Linux system via a Socket UDP stream on a Windows computer through MNV Studio, with MNV Studio acting as the server and the Linux system acting as the client.

[0075] Furthermore, the Linux client creates a ROS workspace, compiles the xsens_mnv_ros ROS package, and sends data to other ROS endpoints via ROS topics. The transmitted data is the global quaternion LinkState of each limb.

[0076] Furthermore, an end-to-end redirection network is constructed. After subscribing to LinkState data, the global quaternion is converted into a quaternion relative to the pelvis and input into the neural network model in 6D pose form. At the same time, the generated robot joint information is published through the ROS topic.

[0077] The robot information node is used to obtain information about 27 robot joints from the redirection system, and to publish the actual motion information in the form of ROS topics.

[0078] The virtual model node is used to obtain robot motion information from Speedgoat and control the robot in Webots to perform the same motion through ROS.

[0079] The visual status monitoring interface node is used to monitor the robot's operating status in real time, including key parameters such as the torque and rotation angle of each motor. When abnormal information is detected, the system will automatically issue an alarm prompt, such as abnormal floating base data, abnormal motor temperature or rotation angle, abnormal sensor data, etc. Through this node, the operator can grasp the robot's operating status at the first time and take corresponding measures in a timely manner to ensure the safe and stable operation of the system.

[0080] Furthermore, the visual status monitoring interface is AppDesigner running on the Windows computer connected to Speedgoat. This interface is a GUI interface, and the data processed in Speedgoat can be read in real time by setting a timer.

[0081] Optionally, the data service system is deployed on the Speedgoat platform running Simulink, and is used to analyze, store and convert the motion data of human and robot at a specific frequency. The stored data is time-series data, which can be used for subsequent humanoid robot anthropomorphic operation training and teleoperation operation status analysis, etc.

[0082] Furthermore, the motion data analysis results are visualized through a GUI interface built on App Designer. The interface components include a Speedgoat connection button, a system running status indicator, a virtual-real interaction system switch, a modal switching knob, a data display text box, a data curve analysis area, an instrument monitoring area, a data save button, and a sub-boundary for virtual-real interaction analysis and real-time stress analysis of components.

[0083] Another objective of this invention is to provide a humanoid robot teleoperation remapping system that integrates digital twin technology. This system can convert human motion data into corresponding robot motion data in real time and transmit it to the robot via a highly convenient and visual teleoperation method. In the future, it can leverage the dynamic balancing technology of a general-purpose humanoid robot controller to learn and execute multi-task motion strategies in complex scenarios.

[0084] Example 2

[0085] like Figure 2 As shown, this embodiment provides a teleoperation method for a humanoid robot based on digital twins, including:

[0086] Acquire human motion data;

[0087] The human motion data is mapped to robot motion data by using the unsupervised neural network of the redirection system.

[0088] The robot's motion data is transmitted to the physical body via a communication system;

[0089] Based on the robot's motion data, the virtual model is driven to synchronously display the motion state of the physical object;

[0090] The data service system stores, filters, and converts the human motion data and robot motion data, and then forwards the converted data to the physical entity and the twin entity through the communication system.

[0091] Specifically, the following steps are included:

[0092] S201: Initialize the remote operating system, start each hardware device and software module in sequence, and complete the startup of Xsens motion capture device, humanoid robot, Webots virtual simulation environment and Speedgoat real-time control platform, including the startup of each ROS node and the publication of topics, thereby completing the initial startup of the entire human-machine-twin system and providing a stable operating foundation for subsequent human motion mapping and remote control.

[0093] S202: Calibrate the human motion capture equipment and set up the data stream. The calibration of the human motion capture equipment refers to calibrating the Xsens motion capture equipment using MVNStudio software, including sensor initialization, zero-bias correction, and posture reference adjustment, to ensure high-precision acquisition of human joint data. The calibration method involves first positioning the device in T-Pose, and then performing a series of movements such as standing and walking. Setting up the data stream refers to establishing a data stream channel through MVN Studio software after calibration, transmitting the human motion data from the Windows end to the Linux end in real time via a UDP Socket Server.

[0094] S203: The redirection system inputs human motion and outputs robot motion information. The main implementation method of this process is to pre-train an unsupervised neural network model that accurately maps human motion to robot motion. The input of the model is 126-dimensional 6D limb poses, and the output is 27-dimensional robot joint angles. The 6 hidden layers have a dimension of 128. It includes two processes: training and prediction. The training phase requires training and testing sets composed of robot data and human data. The prediction phase only requires real-time input of 126-dimensional 6D limb poses. Finally, the generated robot joint angles are published to the physical body and data service system in real time through ROS topics, providing technical support for stable and efficient human-robot motion redirection.

[0095] S204: Speedgoat acquires and processes the robot's rotation angle information, which consists of 27 joint angles output by the redirecting neural network. Speedgoat uses steady-state algorithms such as wavelet smoothing, moving average, and outlier removal to remove and filter jitter, noise, or spike data in real time. At the same time, it sets a safety threshold for the command change rate to avoid sudden changes in joint angles that could cause the robot to perform violent movements, thereby generating mechanical shock or hardware overload risks. The processed data is then synchronously sent to the robot controller, virtual model controller, and visual status monitoring interface by setting ROS topics and signal tags.

[0096] S205: The processed data will be subscribed to by the robot controller and the virtual model controller respectively through ROS topics, and will be acquired in real time by the visual status monitoring interface through a timer mechanism; at the same time, Speedgoat continuously listens for status commands from the robot, the virtual model, and the monitoring interface, and dynamically adjusts the running status of the virtual model according to information changes, so as to achieve state synchronization and consistency between the virtual and physical systems, and ensure that the virtual model and the physical robot maintain collaborative operation and consistent feedback.

[0097] Based on the specific implementation of the above process, the following will be combined with Figures 3-6 In a detailed explanation, one implementation of step S202 requires the Xsens motion capture device to first perform position calibration under T-Pose using MNV Studio running on a Windows computer. The calibration process includes three actions: T-Pose holding, walking, and standing.

[0098] Figure 3 The diagram of the human body on the left shows the IMU locations of the Xsens motion capture device, which has a total of 18 IMUs. It fully expresses the movements of the human torso and various limbs. The data is transmitted to the receiver via WIFI and then to the Window. The latency of the whole process is negligible.

[0099] After the posture calibration is completed in step S202, human movements can be transmitted to the Linux system via MNV Studio in Socket UDP format data stream, with MNV Studio acting as the server and the Linux system acting as the client.

[0100] Furthermore, in step S202, the ROS workspace needs to be configured on the Ubuntu side to receive human motion information transmitted from MNV Studio. xsens_mnv_ros is the ROS package provided by the official website, which can be directly used to receive motion data transmitted from MNV Studio. The data content includes global quaternions LinkState for 18 limbs and several calculated human joint information JointState. In subsequent embodiments of this invention, only LinkState can be used. Thus, step S202 is completed.

[0101] As one implementation of the redirection system, step S203 processes the LinkState obtained in step S202 by relative rotation, and then converts it into information of each joint of the robot through a redirection neural network model. The redirection network model adopts an unsupervised learning structure that does not require human-robot paired motion data. The unsupervised network consists of two encoders with the same structure and a decoder. The encoders are used to extract the latent variable features of human motion data and robot motion data respectively, so that the two are comparable in the same latent space. The decoder is used to reconstruct the latent variables and decode the encoded human motion latent variables into the corresponding robot joint driving angles. Through the joint loss function based on robot joint parameters and shared latent space, the accurate redirection and consistency optimization of motion between human and robot are achieved.

[0102] To obtain a highly accurate and reliable retargeting neural network model, step S203 includes a training phase and a prediction phase for the neural network model. The specific setup method for the training phase is as follows:

[0103] The training dataset for constructing the neural network model is provided in this embodiment using the AMASS open-source human motion dataset and the corresponding robot motion data generated by optimization methods. This avoids the difficulty and cost of relying on a large amount of real machine data collection. The human motion data is represented by parameters based on the SMPL-X model, and the robot motion data consists of 27 joint angles. Subsequently, the human pose, robot joint angles and their corresponding latent space representations are organized into triplet samples to support unsupervised domain alignment learning and improve the generalization ability and stability of human-machine pose mapping.

[0104] Furthermore, the construction scheme for the triplet data employs a cross-domain similarity measure, that is, comparing the quaternion information of the human body and the robot relative to their own pelvis, with human pose and robot pose respectively using... and Indicate, and through and Construct a triple, with the anchor set to the human body's limb movements processed by the encoder and decoder. positive is the same as The limb movements of robots with high similarity negative is related to Low similarity Their similarity measurement formula is: ,in Let be the quaternion of each limb of the human body relative to pelvis. Let be the quaternion of each limb of the robot relative to pelvis;

[0105] In this embodiment, as Figure 4 As shown, As a reference anchor for human movement, To achieve robot limb movements with a high degree of similarity to the anchor, Human body movements that have low similarity to the anchor;

[0106] The similarity metric is only one method for constructing triples. The final triple data consists of the complete SMPL-X model data and robot joint angles obtained through the similarity metric.

[0107] Furthermore, the human motion data input into the shared latent space is represented using a 6D pose representation, that is, the Body_pose in the SMPL-X model is converted from axis-angle form to a 6D rotational representation to avoid singularity issues and enhance model training stability and convergence performance, such as... Figure 4 As shown, in this embodiment, the redirecting neural network consists of two encoders with identical structures and a decoder, wherein the first encoder... The second encoder is used to encode human motion data into a shared latent space. Decoder is used to map robot motion data to the same latent space. It is responsible for decoding the shared latent features back to the robot's 27 joint angle output, so as to achieve a unified representation and consistent mapping between human posture and robot posture;

[0108] Furthermore, the encoder and decoder are composed of OMLPs with the same structure. Specifically, the human motion encoder... The input dimension is 126, and the output latent feature dimension is 512; Robot motion encoder The input dimension is 27, and the output latent feature dimension is 512; the decoder The input dimension is 512, and the output dimension is 27. The network structure consists of one input mapping layer, four hidden layers containing residual enhancement mechanisms, and one output mapping layer. Each hidden layer integrates residual connections and layer normalization structures, and uses the PReLU activation function to improve nonlinear expressiveness and training convergence stability. The final shared latent space dimension is set to 512.

[0109] Furthermore, the loss function comprises triplet loss, standard reconstruction loss, and potential consistency loss, wherein the triplet loss is... , The compensation coefficient is the standard reconstruction loss. The potential consistency loss is The total loss function is The weight coefficients during the training process and For ease of distinction and representation, Figure 4 Chinese express , express ;

[0110] Furthermore, Figure 4 middle The corresponding propagation loop is used to ensure the robot encoder and decoder Consistency during the training process;

[0111] In this embodiment, the training data consists of 600,000 triples, with an Epoch of 300 and a Batch size of 256.

[0112] Furthermore, step S203, the prediction phase, adopts the results obtained during the training phase. and The final retargeting neural network model, used during teleoperation, is input to real-time acquired 6D human pose data. This pose data is a LinkState that has undergone relative rotation processing. Encode into a shared potential space, and then... It can directly output the corresponding 27 robot joint angles;

[0113] Furthermore, the 27 joint angles obtained are published to Speedgoat as a ROS topic, thus completing step S203.

[0114] Step S204 is performed in the Simulink model running Speedgoat. In this embodiment, Subscribe and Publish modules are added to Simulink Real-Time. Speedgoat can automatically initialize ROS nodes during the running process to achieve ROS communication.

[0115] Furthermore, step S204 performs smoothing, filtering, and outlier removal on the received joint angles. First, the pre-detection module removes outliers such as NaN, Inf, and those exceeding physical limits, while recording the number of consecutive outliers for safety assessment. Then, the moving average module is used to suppress jitter in the joint angle data, and wavelets are used to process high-frequency noise in order to filter out spikes and jitter while maintaining the motion trend.

[0116] Furthermore, after filtering, each joint angle also goes through a rate limiting module to control the angle change within a set safety threshold to prevent sudden changes from causing mechanical impact or hardware overload to the robot.

[0117] In this embodiment, in addition to the above-mentioned work, the output signals of each module are marked in order to enable visual status monitoring and subsequent data changes recorded in Speedgoat. Finally, the processed robot motion data is sent to the robot controller and the virtual model controller through two Publish modules respectively.

[0118] Furthermore, Simulink also includes a Subscribe module for receiving motion information from the robot and virtual model, and for comparing the running status in real time to facilitate error analysis and correction. This completes step S204.

[0119] In step S205, the robot controller and the virtual model controller have already encapsulated the sensor and joint motor module into the same interface, which will not be described again here;

[0120] The status monitoring interface runs on the Windows system connected to Speedgoat. It is built as a GUI interface by MATLAB App Designer and achieves real-time data interaction with the Simulink model in Speedgoat through timers and related APIs.

[0121] In this embodiment, by configuring multiple interface components, functions such as Speedgoat device connection and startup, real-time numerical and curve monitoring, online and offline mode switching, data recording and saving, instrumentation status display, stress and health status analysis of key components, and system abnormal indicator alarms are realized.

[0122] Furthermore, in this embodiment, the timer frequency of the visual state machine monitoring interface is set to 100Hz, and an abnormal diagnosis indication mechanism is configured. When the robot or virtual model experiences abnormal situations such as disconnection or communication delay exceeding the threshold, the status indicator light will switch to flashing red. At the same time, the abnormal prompt information will be output in real time through the status label component so as to promptly notify the operator to take corresponding measures, thereby ensuring the safety and stability of the remote operating system.

[0123] This invention also proposes a neural network-based motion retargeting strategy for achieving high-precision mapping of human actions to humanoid robot actions. This strategy employs a dual-encoder—shared latent space—single-decoder structure to perform cross-domain unified modeling of human and robot motion features, and achieves attitude retargeting between heterogeneous systems without requiring paired teaching data. The method maintains high generalization ability and robustness under complex human-machine structural differences, effectively improving the accuracy of robot motion imitation and teleoperation consistency, thereby significantly enhancing the naturalness and user experience of humanoid robots in remote operation scenarios.

[0124] The various embodiments described in this specification employ... Figure 2 The description follows a progressive approach, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to accordingly.

[0125] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for teleoperating a humanoid robot based on digital twins, characterized in that, include: Acquire human motion data; The human motion data is mapped to robot motion data by using the unsupervised neural network of the redirection system. The robot's motion data is transmitted to the physical body via a communication system; Based on the robot's motion data, the virtual model is driven to synchronously display the motion state of the physical object; The data service system stores, filters, and converts the human motion data and robot motion data, and then forwards the converted data to the physical entity and the twin entity through the communication system.

2. The method according to claim 1, characterized in that, The acquisition of the physical body's state information includes: acquiring the humanoid robot's own state information and environmental information through the sensing module; and controlling the humanoid robot to perform actions based on the robot's motion data through the control module.

3. The method according to claim 1, characterized in that, The synchronous display of the driving virtual model includes: displaying the operating status information of the humanoid robot through a visual status monitoring interface.

4. The method according to claim 1, characterized in that, The process of mapping human motion data to robot motion data through a redirection system includes: acquiring human motion data through a motion data acquisition module; and mapping the human motion data to robot motion data through a redirection network.

5. The method according to claim 4, characterized in that, The redirection network includes two identical encoders and one decoder; the encoders are used to map human motion data and robot motion data to a shared latent space, respectively; the decoder is used to decode features in the shared latent space into robot motion data.

6. A humanoid robot teleoperation system based on digital twins, characterized in that, include: A physical entity, the physical entity including a sensing module for acquiring entity state information and a control module for performing actions; A twin, the twin comprising a virtual model for mapping the posture of a physical body and a visual monitoring interface for displaying the operating status; A redirection system, which is built on a neural network, is used to convert human motion data into robot motion data; A communication system for transmitting data between modules; A data service system is used to store, filter, and convert the human motion data and robot motion data.

7. The system according to claim 6, characterized in that, The sensing module is positioned on the head and chest of the humanoid robot to acquire environmental images and joint encoder information.

8. The system according to claim 6, characterized in that, The virtual model maps the posture of the physical object in real time at a 1:1 scale and supports both online and offline operating modes.

9. The system according to claim 6, characterized in that, The redirection system includes a motion data acquisition module and a redirection network; the motion data acquisition module is used to acquire human motion data; the redirection network is used to map the human motion data into robot motion data, and the redirection network is trained using a joint loss function that includes triplet loss, reconstruction loss and potential consistency loss.

10. The system according to claim 9, characterized in that, The redirection network includes two identical encoders and one decoder; the encoders are used to map human motion data and robot motion data to a shared latent space, respectively; the decoder is used to decode features in the shared latent space into robot motion data.