Immersive robot teleoperation method and system

By building multiple models to process the robot's perspective images and obtain the operator's posture control information in real time, the problem of lack of feedback when controlling the robot with VR equipment is solved, immersive three-dimensional perspective feedback and high-precision control are achieved, and the operator's control ability and task execution effect are improved.

CN120578299BActive Publication Date: 2025-10-03HANGZHOU YUSHU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511075241.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-03
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

In the existing technology, the body data collected by VR devices to control robots lacks an effective and timely feedback mechanism, resulting in users being unable to promptly know the robot's execution status, poor operational intuitiveness, difficulty in performing complex and delicate operational tasks, and low control accuracy.

Method used

By constructing a first-person mapping model, an interactive data acquisition model, a remote operation processing model, an execution perception model, and an immersive perspective simulation model, the robot's perspective image is acquired and processed, the operator's posture control information is obtained in real time, and converted into control instructions in the robot coordinate system, and immersive three-dimensional perspective feedback is generated to achieve robot posture control and motion image acquisition.

Benefits of technology

It enables the operator to accurately and timely understand the robot's execution status, improves control accuracy and operational intuitiveness, enables the operator to perform complex and delicate operational tasks, and improves control capabilities and task execution effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578299B_ABST
    Figure CN120578299B_ABST
Patent Text Reader

Abstract

The present invention discloses an immersive robot teleoperation method and system, which belongs to the field of robot teleoperation technology. Existing teleoperation schemes lack an effective and timely feedback mechanism, have poor intuitive operation, low control accuracy, and are difficult to perform complex and delicate operation tasks. An immersive robot teleoperation method of the present invention, by constructing a first-person mapping model, an interactive data acquisition model, a teleoperation processing model, an execution perception model, and an immersive perspective simulation model, can map the operator's posture and movements in real time, and generate a real-time responsive three-dimensional image that can be displayed in a VR device, forming an effective and timely first-person perspective immersive feedback mechanism, which can make the operator perform tasks as if they were in person, improve the immersion, intuitiveness and naturalness of the interaction, thereby significantly improving the control ability and task execution effect, improving the operation efficiency and accuracy, and thus being able to perform complex and delicate operation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an immersive robot teleoperation method and system, and belongs to the technical field of robot teleoperation. Background Art

[0002] A Chinese patent application (publication number: CN119347781A) discloses a method for remotely controlling a robot. The method comprises: collecting first position data of key points on a user's body using a VR device; then performing coordinate conversion on the first position data to obtain second position data of the first position data in a robot coordinate system; determining joint control parameters of the robot based on the second position data; and finally controlling the robot based on the joint control parameters. The aforementioned application uses a VR device to collect position data of key points on the body, which can make the collected body data more accurate; converting the body data collected by the VR device into the robot coordinate system achieves the purpose of controlling the robot based on the body data collected by the VR device.

[0003] The above solution controls the robot based on body data collected by VR equipment, but lacks an effective and timely feedback mechanism, which results in users being unable to promptly know the robot's execution status. As a result, the operation is unintuitive and difficult to perform complex and delicate operation tasks.

[0004] Furthermore, the above solutions and existing technologies do not disclose how to convert the operator's posture control information into control instructions in the robot coordinate system, resulting in low control accuracy of robot teleoperation and affecting the promotion and use of robot teleoperation solutions.

[0005] The information disclosed in this Background Art is only for understanding the background of the present inventive concept and therefore it may include information that does not constitute prior art. Summary of the Invention

[0006] In response to the above problem or one of the above problems, an object of the present invention is to provide an immersive robot remote operation method and system, which obtains the robot perspective image by constructing a first-person mapping model, an interactive data acquisition model, a remote operation processing model, an execution perception model, and an immersive perspective simulation model, and processes the robot perspective image to obtain a mapping image that can be presented in the operator's VR device; then, the operator's posture control information is obtained in real time, and the operator's posture control information can be accurately converted into control instructions in the robot coordinate system; then, based on the control instructions, the robot's posture is controlled, and the robot motion image is collected, thereby generating a real-time responsive three-dimensional image that can be displayed in the VR device, so that the operator can obtain an immersive three-dimensional perspective, forming an effective and timely first-person perspective immersive feedback mechanism, so that the operator can accurately, comprehensively and timely know the robot's execution status, and the operation is intuitive, which can effectively improve the control accuracy, and thus can be competent for complex and delicate operation tasks.

[0007] In response to the above problem or one of the above problems, the second object of the present invention is to provide an immersive robot remote operation method and system. By providing a three-dimensional first-person perspective image and combining it with natural posture control, the operator can perform the task as if he were in person, thereby improving the immersiveness, intuition, intuitiveness and naturalness of the interaction, thereby significantly improving the control ability and task execution effect, improving operational efficiency and accuracy, and then realizing real-time synchronous motion control, so that the robot can quickly and smoothly follow the operator's movements.

[0008] To achieve one of the above purposes, the first technical solution of the present invention is:

[0009] An immersive robot teleoperation method, comprising the following:

[0010] Based on the pre-built first-person mapping model, the robot's perspective image is acquired and processed to obtain a mapping image that can be presented in the operator's VR device;

[0011] Through the pre-built interactive data acquisition model, the operator's posture control information is obtained in real time;

[0012] Using a pre-built teleoperation processing model, the operator's posture control information is converted into control instructions in the robot coordinate system; this includes the following:

[0013] Obtain the operator's posture control information and perform base coordinate system transformation on it to obtain the initial posture agreement;

[0014] Perform alignment mapping processing for different initial posture conventions to obtain alignment data;

[0015] Extract standard data structures for inverse kinematics and control planning from the alignment data;

[0016] Generate control instructions in the robot coordinate system based on standard data structure for robot control;

[0017] Using a pre-built execution perception model, the robot's posture is controlled based on control instructions, and images of the robot's movements are collected.

[0018] Using a pre-built immersive perspective simulation model, the robot motion image and mapping image are coupled and processed to generate real-time responsive three-dimensional images that can be displayed in VR devices, allowing the operator to obtain an immersive three-dimensional perspective.

[0019] The present invention constructs a first-person mapping model, an interactive data acquisition model, a remote operation processing model, an execution perception model, and an immersive perspective simulation model to obtain a robot perspective image, and processes the robot perspective image to obtain a mapping image that can be presented in the operator's VR device; then the operator's posture control information is obtained in real time, and the operator's posture control information can be accurately converted into control instructions in the robot coordinate system; based on the control instructions, the robot's posture is controlled, and the robot's motion image is collected, thereby generating a real-time responsive three-dimensional image that can be displayed in the VR device, so that the operator can obtain an immersive three-dimensional perspective, forming an effective and timely first-person perspective immersive feedback mechanism, so that the operator can accurately, comprehensively and timely know the robot's execution status, and the operation is intuitive, which can effectively improve the control accuracy, so that it can be competent for complex and delicate operation tasks, and facilitate the promotion and use of robot remote operation solutions.

[0020] Furthermore, the present invention provides a three-dimensional first-person perspective image and combines it with natural posture control, which allows the operator to perform tasks as if he or she is personally there, thereby improving the immersiveness, intuition, intuitiveness and naturalness of the interaction, thereby significantly improving the control ability and task execution effect, improving operational efficiency and accuracy, and thus realizing real-time synchronous motion control, so that the robot can quickly and smoothly follow the operator's movements. The solution is scientific, reasonable and feasible.

[0021] As preferred technical measures:

[0022] The method for obtaining a robot-viewpoint image based on a pre-built first-person mapping model and processing the robot-viewpoint image to obtain a mapping image that can be presented in the operator's VR device is as follows:

[0023] Obtain the VR registration information of the VR device used by the operator and the corresponding robot registration information;

[0024] The server establishes a three-party network transmission channel based on VR registration information and robot registration information;

[0025] Using a three-party network transmission channel, the server receives the robot's perspective images collected by the robot's binocular camera in real time;

[0026] Process the robot's view image to obtain the robot's binocular camera image stream;

[0027] Using a three-party network transmission channel, the robot's binocular camera image stream is transmitted to the operator's VR device to obtain a mapping image that can be presented in the operator's VR device.

[0028] As preferred technical measures:

[0029] The method for obtaining the operator's posture control information in real time through the pre-built interactive data acquisition model is as follows:

[0030] Capture the operator's motion information, which includes the operator's head posture information, left and right hand posture information, and handle operation information; the left and right hand posture information includes finger joint position information and wrist posture information;

[0031] Construct wrist transformation relationship based on the world reference system and the wrist local reference system;

[0032] Based on the wrist transformation relationship, the left and right hand posture information is transformed into the wrist local reference system to obtain the finger joint data;

[0033] The head posture information, finger joint data and handle operation information are stored in a dictionary type format to form posture control information.

[0034] As preferred technical measures:

[0035] The method for obtaining the operator's posture control information and performing base coordinate transformation to obtain the initial posture agreement is as follows:

[0036] Acquiring operator posture control information, which at least includes left wrist posture, left handle posture, right wrist posture, right handle posture, finger joint posture and head posture;

[0037] According to the specific transformation matrix and based on the similarity transformation principle, the base coordinate system transformation of the posture control information is performed to obtain the posture output information based on the robot coordinate system convention;

[0038] The posture output information is converted into an initial posture convention to obtain an initial posture convention, where the initial posture convention at least includes an initial posture convention required by the robot arm and an initial posture convention required by the robot finger.

[0039] As preferred technical measures:

[0040] The specific transformation matrix is ​​constructed as follows:

[0041] Based on the Open Extended Reality standard coordinate system convention, the z-axis is set to the back direction, the y-axis is set to the up direction, and the x-axis is set to the right direction;

[0042] Based on the robot coordinate system convention, the z-axis is set to the upward direction, the y-axis is set to the left direction, and the x-axis is set to the forward direction;

[0043] Based on the coordinate axis transformation relationship between the open extended reality standard coordinate system and the robot coordinate system, a specific transformation matrix is ​​constructed.

[0044] As preferred technical measures:

[0045] Based on the standard data structure, the method for generating control instructions in the robot coordinate system is as follows:

[0046] The first step is to obtain the standard data structure, which includes the data structure and the type structure;

[0047] The data structure includes head posture and joint posture; the type structure includes gesture status information;

[0048] The second step is to load the overall model of a robot, whose format is the model file under the robot operating system standard;

[0049] The third step is to determine the position error and rotation error between the robot's current state and the target state based on the head posture, joint posture and gesture state information;

[0050] The fourth step is to construct a total cost function based on the position error and rotation error. The optimization variable of the total cost function is the joint motor position.

[0051] Step 5: Solve the total cost function to obtain the joint motor position solution;

[0052] Step 6: Based on the joint motor position solution, the feedforward torque vector of the robot joint is calculated;

[0053] Step 7: Determine the joint motor position based on the feedforward torque vector;

[0054] Step 8: Perform real-time smoothing on the joint motor position to obtain the joint motor position angle;

[0055] The ninth step is to process the joint motor position angle and generate control instructions in the robot coordinate system.

[0056] As preferred technical measures:

[0057] The method for controlling the robot's posture and collecting the robot's motion images using the pre-built execution perception model based on control instructions is as follows:

[0058] receiving a control instruction, which includes at least a target joint angle and a torque instruction;

[0059] Based on the target joint angle and torque instructions, the local controller is called to drive the motor to perform the action, realizing the coordinated control of the arms and hands; and the status information of each motor is collected, including angle, speed, current and force feedback information;

[0060] Real-time perception of camera image data streams from the head and wrist to obtain robot motion images.

[0061] As preferred technical measures:

[0062] The method for using a pre-built immersive perspective simulation model to couple the robot motion image and the mapping image to generate a real-time responsive 3D image that can be displayed in a VR device is as follows:

[0063] Obtain robot motion images collected synchronously in real time by the robot's binocular camera;

[0064] The robot motion images are left-right image pairs with natural parallax;

[0065] Performing distortion correction processing on the left and right image pairs to obtain left and right corrected images;

[0066] The left and right corrected images are subjected to layered mask isolation processing and coupled with the mapped image to obtain a real-time responsive three-dimensional image that can be output to the left and right display screens of the VR device.

[0067] As preferred technical measures:

[0068] The method for performing distortion correction processing on the left and right image pairs to obtain the left and right corrected images is as follows:

[0069] Normalize the left and right image pairs to obtain several undistorted normalized coordinates;

[0070] Calculate the squared radial distance between the undistorted normalized coordinates;

[0071] Inversely estimate all distortion coefficients using a method of minimizing image point projection errors, wherein the distortion coefficients include radial distortion coefficients and tangential distortion coefficients;

[0072] Based on the Brown-Conradie distortion model, the undistorted normalized coordinates, the squared radial distance value, the radial distortion coefficient and the tangential distortion coefficient are processed to obtain the distorted image coordinates;

[0073] The distorted image coordinates are summed up to obtain the left and right corrected images.

[0074] To achieve one of the above purposes, the second technical solution of the present invention is:

[0075] An immersive robotic teleoperation system for humanoid robots or dual-arm robotic arms, using a modular architecture including a VR interaction terminal, a teleoperation algorithm processing unit, and a robot execution and perception terminal;

[0076] The input data of the VR interaction end is the camera image stream data of the robot execution and perception end; the output data is the posture data of the wearable device, or the posture data and all button data of the VR device's matching handle;

[0077] The remote operation algorithm processing unit uses the output data of the VR interactive end as input data and obtains the output data of the robot execution and perception end in real time. Based on the input data, it obtains the action data of the robot execution device. It is equipped with an inverse kinematics module, an image transceiver module, and a data recording module.

[0078] The inverse kinematics module is used to achieve a complete mapping from posture data and all key data to the joint motor position angles that the robot can execute;

[0079] The image transceiver module is used to receive camera image stream data from the robot execution and perception end, and forward the camera image stream data to the VR device end for rendering in real time;

[0080] The data recording module is used to record the image data, status data and control instructions during the remote operation in real time;

[0081] The input data of the robot execution and perception end is the action data of the remote operation algorithm processing unit; the output data is the camera image stream data on the robot and the status data of the robot execution device.

[0082] Compared with the existing technical solutions, the present invention has the following beneficial effects:

[0083] The present invention acquires robot perspective images by constructing a first-person mapping model, an interactive data acquisition model, a remote operation processing model, an execution perception model, and an immersive perspective simulation model, and processes the robot perspective images to obtain mapping images that can be presented in the operator's VR device; then, the operator's posture control information is acquired in real time, and the operator's posture control information can be accurately converted into control instructions in the robot coordinate system; based on the control instructions, the robot's posture is controlled, and the robot's motion image is acquired, thereby generating a real-time responsive three-dimensional image that can be displayed in the VR device, so that the operator can obtain an immersive three-dimensional perspective, forming an effective and timely first-person perspective immersive feedback mechanism, so that the operator can accurately, comprehensively, and timely know the robot's execution status, and the operation is intuitive, which can effectively improve the control accuracy, and thus can be competent for complex and delicate operation tasks.

[0084] Furthermore, the present invention provides a three-dimensional first-person perspective image and combines it with natural posture control, which allows the operator to perform tasks as if he or she is personally there, thereby improving the immersiveness, intuition, intuitiveness and naturalness of the interaction, thereby significantly improving the control ability and task execution effect, improving operational efficiency and accuracy, and thus realizing real-time synchronous motion control, so that the robot can quickly and smoothly follow the operator's movements. The solution is scientific, reasonable and feasible. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 The figure is a flow chart of the immersive robot teleoperation method of the present invention. DETAILED DESCRIPTION

[0086] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of this application. The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims.

[0087] like Figure 1 As shown, the first specific embodiment of the immersive robot teleoperation method of the present invention is:

[0088] An immersive robot teleoperation method, comprising the following:

[0089] Based on the pre-built first-person mapping model, the robot's perspective image is acquired and processed to obtain a mapping image that can be presented in the operator's VR device;

[0090] Through the pre-built interactive data acquisition model, the operator's posture control information is obtained in real time;

[0091] Using the pre-built teleoperation processing model, the operator's posture control information is converted into control instructions in the robot coordinate system;

[0092] Using a pre-built execution perception model, the robot's posture is controlled based on control instructions, and images of the robot's movements are collected.

[0093] Using a pre-built immersive perspective simulation model, the robot motion image and mapping image are coupled and processed to generate real-time responsive three-dimensional images that can be displayed in VR devices, allowing the operator to obtain an immersive three-dimensional perspective.

[0094] A second specific embodiment of the immersive robot teleoperation method of the present invention:

[0095] An immersive robot teleoperation method comprises the following steps:

[0096] Step 1: Obtain the operator's posture control information in real time through the pre-built interactive data acquisition model;

[0097] Step 2: Using the pre-built teleoperation processing model, the operator's posture control information is converted into control instructions in the robot coordinate system;

[0098] Step 3: Using the pre-built execution perception model, the robot's posture is controlled based on the control instructions, and the execution image captured by the camera on the robot is uploaded to the server;

[0099] Step 4: Use the pre-built immersive perspective simulation model to process the execution image so that the operator can obtain an immersive three-dimensional perspective.

[0100] A specific embodiment of an immersive robot teleoperation system using the method of the present invention:

[0101] An immersive robot teleoperation system based on VR devices, designed for humanoid robots or dual-arm robotic arms, adopts a modular architecture and mainly includes a VR interaction terminal, a teleoperation algorithm processing unit, and a robot execution and perception terminal.

[0102] The VR interactive terminal is built based on the interactive data collection model, and its working principle is as follows:

[0103] The input data is the camera image stream data of the robot execution and perception end, which can also be forwarded by the remote operation algorithm processing unit; the output data is the head and finger tracking (including wrist) posture data of the user wearing the device, or the posture and all button data of the handle device supporting the VR device.

[0104] The teleoperation algorithm processing unit is built based on the teleoperation processing model. It includes an inverse kinematics module, an image transceiver module, and a data recording module. Its working principle is as follows:

[0105] The output data of the VR interaction end is used as input data, and the output data of the robot execution and perception end is obtained in real time.

[0106] Based on the input data, the action data of the robot execution device is obtained; and the camera image stream data of the robot execution and perception end is forwarded to the VR interaction end.

[0107] The inverse kinematics module is used to achieve a complete mapping from the left and right wrist postures and finger postures input by the data analysis module to the joint motor position angles that can be executed by the robot.

[0108] The image transceiver module is used to receive image streams from the robot's execution and perception ends, and forward them to the VR device end for rendering in real time.

[0109] The data recording module is used to record images, status, control instructions and other data during the teleoperation process in real time and save them in a structured manner as high-quality data sets.

[0110] The robot execution and perception side is built based on the execution perception model, and its working principle is as follows:

[0111] The input data is the action data of the teleoperation algorithm processing unit; the output data is the image stream data of all cameras on the robot and the status data of the robot's execution equipment.

[0112] The present invention utilizes a VR headset (which may include a handle device) to enable the operator to remotely control the robot naturally and intuitively from an immersive first-person perspective. At the same time, the system has a built-in data acquisition and storage mechanism to synchronously record the control and perception data during the remote operation process for subsequent AI model training.

[0113] In this embodiment, the method of using the VR interactive terminal to allow the operator to "enter" the robot's body and control its movements is as follows:

[0114] The operator uses a VR device to interact with the robot. The VR device is an Apple headset or a ByteDance headset.

[0115] Based on the browser WebXR environment, the VR device accesses the local server of the teleoperation algorithm processing unit through the Hypertext Transfer Protocol Secure HTTPS, thereby establishing a full-duplex real-time communication protocol WebSocket channel.

[0116] The VR interactive terminal captures head posture, left and right hand (including wrist) posture data, or raw controller data in real time. Except for finger joints, all data is the original data collected by the device. The reference coordinate system is the world coordinate system under the positioning of the VR device system. All data is stored in a dictionary type, which includes the following key-value data structure:

[0117] Head posture: , left wrist pose: , right wrist pose: , the left finger joint position set: , right finger joint position set: , left finger joint posture set: , right finger joint posture set: .in, Special Euclidean group (SE), represents the set of real numbers, It is a special orthogonal group (SO).

[0118] All knuckle data above All are located in the local reference system of the corresponding wrist. The data in this local reference system is the original data in the world reference system. , calculated by the coordinate transformation formula, the expression of the coordinate transformation formula is as follows:

[0119]

[0120] Where, is an index, and its meaning is consistent with the hand tracking interface XRHand definition in the WebXR Device API standard. When the index is 0, the original data Represented as wrist joint data , which is expressed as follows:

[0121]

[0122] Hand gestures, including left-hand pinch strength: , right hand pinching strength: , left fist strength: , right fist strength: , pinch the left and right hands together to form a fist. The fist state is obtained by Boolean function To express it, the expression is as follows:

[0123]

[0124] in, Indicates the pinching state. Indicates the fist state, 0 means the state has not been reached, and 1 means the state has been reached.

[0125] The handle pose includes the left handle pose: And the right handle posture: .

[0126] The handle buttons include the left trigger button: , right trigger button: , left grip key: , right grip key: , Left stick: (Unit circle, ), right joystick: (Unit circle, ).

[0127] Define all buttons as boolean functions , which is expressed as follows:

[0128]

[0129] Among them, A and B represent the two main buttons on the right handle, and X and Y represent the two main buttons on the left handle. and Respectively represent the press buttons of the left and right trigger keys; 0 means the key is not activated, and 1 means the key is activated.

[0130] Furthermore, the left and right image pairs with natural parallax are collected synchronously in real time through the binocular camera. After distortion correction, they are output to the left and right display screens of the VR device through layered mask isolation, obtaining a natural and immersive image. The parallax characteristics of the human eye are used to form an immersive 3D view effect with depth perception, achieving the effect of rendering from a first-person perspective, so that the operator feels as if "entering" the robot's body to control its movements.

[0131] The distortion correction process is based on the Brown-Conrady distortion model, which is mathematically expressed as follows:

[0132]

[0133]

[0134]

[0135] in: is the undistorted normalized coordinate; is the coordinate of the distorted image; is the square of the radial distance; is the radial distortion coefficient; is the tangential distortion coefficient.

[0136] The model parameters are obtained through a calibration process of multiple frames of standard checkerboard images, and all distortion coefficients are inversely estimated by minimizing the image point projection error to ensure correction accuracy and geometric consistency.

[0137] In this embodiment, the layered mask isolation display process is as follows:

[0138] Obtain the distortion-corrected left and right binocular image pair , which is expressed as follows:

[0139]

[0140] Where x is the pixel width of the image, y is the pixel height of the image, W is the maximum pixel width of the image, and H is the maximum pixel height of the image.

[0141] The left view , right view It can be expressed as:

[0142]

[0143]

[0144] In this embodiment, the teleoperation algorithm processing unit runs on the user's computer or embedded platform. It can convert and correct the raw pose data output by the VR interactive terminal to obtain valid data for robot control. Because the spatial reference system used by VR devices (OpenXR coordinate system) and the spatial coordinate system used by robots (Robot coordinate system) are significantly different, the core functions of the teleoperation algorithm processing unit include the following:

[0145] The original pose information output by the VR interactive terminal is transformed into the base coordinate system to obtain the initial pose; then the alignment mapping between different initial pose conventions is completed to obtain alignment data; and then from the alignment data, a standard data structure that can be used for subsequent inverse kinematics and control planning is extracted.

[0146] In this embodiment, the original posture information output by the VR interaction terminal is transformed into a base coordinate system to obtain the initial posture as follows:

[0147] The first step is to convert the left and right wrist / handle postures, finger joint postures and head postures output by the VR interaction terminal into For the sake of convenience, the left wrist / handle posture data are uniformly named as , the right wrist / handle pose data is uniformly named ; Name the left finger joint pose data as , the right finger joint pose data is named All of this data serves as input to the remote operation algorithm processing unit. They are all based on the OpenXR coordinate system convention of the open extended reality standard, with the z-axis set to the back direction, the y-axis set to the up direction, and the x-axis set to the right direction.

[0148] The data format output by the teleoperation algorithm processing unit is based on the robot coordinate system convention, with the z-axis set to the upward direction, the y-axis set to the left direction, and the x-axis set to the forward direction.

[0149] The second step is to complete the base coordinate system transformation through the similarity transformation principle to obtain the output data based on the Robot coordinate system convention. , which is calculated as follows:

[0150]

[0151] Where, Represents all input data under the OpenXR coordinate system convention of the open extended reality standard, including 、 、 、 、 ; Is a specific transformation matrix. The input data is multiplied by the specific transformation matrix , and right-multiply by a specific transformation matrix After the inverse of the Robot coordinate system, the output data is obtained. .

[0152] The third step is to convert the initial posture convention. The wrist / controller pose based on the Robot coordinate system is multiplied by the specific correction matrix to transform it to the initial posture convention required by the specific robot arm. The process is shown in the following formula:

[0153]

[0154]

[0155] in, is the left wrist / controller specific correction matrix, Correction matrix specific to the right wrist / controller.

[0156] The finger joint pose based on the Robot coordinate system is multiplied by the specific correction matrix to transform it to the initial pose required by the specific robot finger. The process is shown in the following formula:

[0157]

[0158]

[0159] in, is the specific correction matrix for the left finger joints, is the specific correction matrix for the right finger joints.

[0160] Adjust the reference coordinate system to obtain the wrist / handle pose, which includes the following:

[0161] Under the Robot coordinate system convention, the input left and right wrist / handle pose data uses the world system positioned by the VR device system as the reference coordinate system. , , Left wrist / controller , right wrist / controller , head posture The position component of the wrist / handle pose is converted from the world coordinate system to the head local coordinate system as shown in the following formula:

[0162]

[0163]

[0164] in, is the position component of the left wrist / controller in the head local coordinate system, is the position component of the right wrist / controller in the head's local coordinate system.

[0165] To ensure the correct planning of the subsequent inverse kinematics module, a fixed translation operation is required to convert the origin of the coordinate system from the head to the waist of the robot. The expression is as follows:

[0166]

[0167]

[0168] The translation value t can be adjusted according to the actual body size of the user and the robot. is the position component of the left wrist / controller in the robot waist coordinate system, is the position component of the right wrist / controller in the robot waist coordinate system.

[0169] In this embodiment, the data matrix also needs to be validated, which includes the following:

[0170] Before each posture transformation, the error-proof matrix update function can be used to check whether the input matrix is ​​non-singular (that is, whether it is a legal transformation). The processing method is as follows:

[0171] Get the previous frame matrix previous_matrix, the current matrix current_matrix, and the initial matrix init_matrix, and input them into the error-proofing matrix update function to obtain the determinant det of the current matrix.

[0172] When the determinant det is not finite or is close to zero (such as less than or equal to 0.000001), it is determined whether the previous frame matrix is ​​valid. If it is valid, the previous frame matrix is ​​returned; otherwise, the initial legal matrix init_matrix is ​​returned.

[0173] When the determinant det is finite and the determinant det is not close to zero, returns the current matrix.

[0174] If the posture is invalid (such as VR data loss or freezing), the initial legal posture or the last valid posture will be used to continue output to ensure system stability and control continuity.

[0175] In this embodiment, the standard data structure includes a data structure TeleData and a type structure TeleStateData.

[0176] The data structure TeleData includes head posture, wrist / controller posture, finger joint position and posture, and hand gesture posture.

[0177] The head pose is a head pose matrix, which is constructed using a floating-point array to represent the head pose. The wrist / controller pose includes the left arm pose matrix and the right arm pose matrix, which are constructed using a floating-point array. The finger joint position and pose include the left finger joint position set, the right finger joint position set, the left finger joint pose set, and the right finger joint pose set. Each joint position set is constructed using a floating-point array, and each row is a 3D joint point position. Each joint pose set is constructed using a floating-point array to represent the rotation matrix of each joint. The hand gesture pose includes the left hand pinch strength, the right hand pinch strength, the left hand trigger pull depth, and the right hand trigger pull depth, which are represented by floating-point type data structures.

[0178] The TeleStateData structure is a structure with gesture and button status information, which contains all gesture / button status information. Gesture status information includes left hand pinch state, left hand fist state, left hand fist strength, right hand pinch state, right hand fist state, and right hand fist strength.

[0179] The left hand pinch state is represented by a Boolean data structure, which indicates whether it is pinched; the left hand fist state is represented by a Boolean data structure, which indicates whether the fist is clenched; the left hand fist strength is represented by a floating-point data structure, and its value range is 0.0~1.0; the right hand pinch state is represented by a Boolean data structure; the right hand fist state is represented by a Boolean data structure; the right hand fist strength is represented by a floating-point data structure, and its value range is 0.0~1.0.

[0180] The controller button status includes the left trigger button status, left grip button status, left grip pull strength, left joystick button status, left joystick input vector, left A button status, left B button status, right trigger button status, right grip button status, right grip pull strength, right joystick button status, right joystick input vector, right A button status, and right B button status.

[0181] The left trigger key state is represented by a Boolean data structure; the left grip key state is represented by a Boolean data structure; the left grip pull strength is represented by a floating-point data structure, and its value range is 0.0~1.0; the left joystick button state is represented by a Boolean data structure; the left joystick input vector is a two-dimensional vector; the left A key state is represented by a Boolean data structure; the left B key state is represented by a Boolean data structure; the right trigger key state is represented by a Boolean data structure; the right grip key state is represented by a Boolean data structure; the right grip pull strength is represented by a floating-point data structure, and its value range is 0.0~1.0; the right joystick button state is represented by a Boolean data structure; the right joystick input vector is a two-dimensional vector; the right A key state is represented by a Boolean data structure; the right B key state is represented by a Boolean data structure.

[0182] Through this module, the raw VR input is reliably converted into standard format data that can be used for robot control, and this format data can be directly input into the subsequent inverse kinematics module.

[0183] In this embodiment, the inverse kinematics module supports dual-arm and dual-hand joint control through a hierarchical structure and a unified process. The main steps are as follows:

[0184] The first step is the model preparation stage, which is to load the robot model as a whole. The format is the URDF model file in XML format under the Robot Operating System ROS standard.

[0185] Simplify or select the joints related to the arms and hands in the model file through program code to form a sub-model with reversible solution;

[0186] The second step is the target input stage, which inputs the left and right wrist / handle pose data output by the data analysis module; and the left and right finger joint pose data output by the data analysis module;

[0187] The third step is the inverse kinematics solution phase, where we construct an optimization problem. The optimization terms include the position and rotation errors between the current and target positions, as well as regularization terms and smoothing terms to enhance robustness. The overall cost function formula is shown below, where the optimization variable is the joint motor position q:

[0188]

[0189]

[0190] Among them, p To obtain the end position function based on the current joint motor position; The position component of the left and right wrist / handle pose data output by the input data parsing module; To obtain the end posture function based on the current joint motor position; The rotation component of the left and right wrist / handle pose data output by the input data parsing module; ; is the joint motor position at the last iteration; and are real constants, representing the lower and upper limits of the corresponding joint motor positions in the URDF model file. The first two terms of the cost function are the position and rotation error terms, the third is the regularization term, and the fourth is the smoothing term. The expressions below the cost function are the constraints.

[0191] Using the optimization algorithm library, the above optimization problem is automatically solved based on the nonlinear optimizer IPOPT to obtain the joint motor position solution.

[0192] The fourth step is to calculate the feedforward torque. The reverse Newton-Euler method (RNEA) is used to calculate the corresponding torque feedforward to assist in robot control. The calculation formula is as follows:

[0193]

[0194] in, is the feedforward torque vector; is the current joint motor position; is the joint velocity; is the joint acceleration; is the inertia matrix; are the Coriolis force and centrifugal force terms; is the gravity term.

[0195] The fifth step is filtering. The single Euro filter algorithm is used to smooth the joint motor position in real time, suppressing jitter while ensuring response speed. The calculation formula is as follows:

[0196]

[0197]

[0198]

[0199]

[0200]

[0201] in, is the current input joint motor position; The joint motor position input for the previous frame; is the output signal after current filtering; Indicates the time interval between the current frame and the previous frame; is the derivative of the original input (estimated velocity); is the derivative after low-pass filtering, which represents a smooth estimate of the rate of change; is the derivative filter smoothing coefficient, which is determined by the derivative cutoff frequency Decide; is the cutoff frequency of the derivative filter; It is a first-order low-pass filtering operation; The dynamic cutoff frequency of the main filter changes adaptively with the speed of signal change; is the speed response coefficient, which is used to adjust the response sensitivity; The minimum cutoff frequency of the main filter (the smaller the better); The smoothing factor for the main signal filter.

[0202] In this embodiment, the method by which the data recording module records the image, status, control instructions and other data during the teleoperation process is as follows:

[0203] Step 1: Initialize the task settings, create a new recording task directory, and generate task metadata (such as version, time, task description, etc.).

[0204] Step 2: Cache and write data, writing multimodal data such as images, control status, and actions into a memory queue in chronological order. A background thread continuously reads the queue contents and writes them asynchronously to the local file system to ensure that the main process is not blocked and data is written to disk in real time.

[0205] Step 3: Data is structured and saved. During the write process, all frame data is organized into a lightweight JSON file, supporting clear data organization and time alignment. This also requires support for playback, allowing for visual playback of recorded results for action reproduction, effect verification, and pre-training model verification.

[0206] In this embodiment, the robot execution and perception end includes the following:

[0207] The system receives target joint angle and torque commands from the teleoperation algorithm processing unit, then calls the local controller to drive the motors to execute the movements, achieving coordinated control of both arms and hands. It also streams camera image data from the head, wrist, and other parts of the body in real time for sensory feedback, and provides status information such as the angle, speed, current, and force feedback of each motor for subsequent closed-loop control and data recording.

[0208] Therefore, the present invention uses a VR headset to provide a three-dimensional first-person perspective, combined with natural gesture control, so that the operator can perform tasks as if they were in person, thereby improving the immersiveness, intuition, intuitiveness and naturalness of the interaction. Compared with the traditional remote operation framework, VR-based control significantly improves the control ability and task execution effect. At the same time, the present invention can effectively improve operational efficiency and accuracy. Through the combination of high-refresh-rate VR visual feedback and low-latency communication links, real-time synchronous motion control is achieved, allowing the robot to quickly and smoothly follow the operator's movements. The system's built-in motion planning algorithm ensures stable and reliable robot movement, effectively reducing the risk of misoperation due to network delays.

[0209] Furthermore, the present invention enables full-process data recording and utilization, integrating a one-stop data platform that simultaneously collects all control and perception data during teleoperation. The collected high-quality data can be directly used to train robot motion strategies and AI models, accelerating the robot's learning and autonomy, and improving the efficiency of subsequent task execution.

[0210] Furthermore, this invention effectively enhances safety and applicability, allowing operators to complete tasks without entering hazardous environments, reducing personal safety risks. The system is user-friendly with regard to both operational techniques and usage scenarios, and can be applied to scientific research simulations, industrial field operations, disaster relief, and other fields. Furthermore, by providing complete operational data records, this invention can also be used for remote training, playback and review, and task optimization, possessing significant engineering and societal value.

[0211] A specific application example of the method of the present invention is as follows:

[0212] This example builds an immersive teleoperation system in a real environment using the Apple Vision Pro headset and the Unitree G1 humanoid robot as carriers. The system process is as follows:

[0213] In step 1, the operator wears an Apple Vision Pro headset, opens a browser and enters a URL, accesses a local server, and establishes full-duplex, real-time WebSocket communication based on the Hypertext Transfer Protocol (HTTPS). The local server runs on a PC with the Ubuntu 22.04 operating system or a Jetson embedded system module, deploying the automatic remote control operation avp_teleoperate project to achieve data capture, transformation, inverse kinematics solution, control command sending, and image forwarding (taken from the camera on the Unitree G1).

[0214] Step 2: Use the VR interactive terminal to obtain the head, wrist / handle, finger joint posture / button data in real time and publish it to the teleoperation algorithm processing unit; at the same time, receive the distortion-corrected robot binocular three-color camera image stream in real time and render it into an immersive 3D view.

[0215] In step 3, the teleoperation algorithm processing unit receives and corrects the posture data published by the VR interactive terminal, converting it into a standard data structure in the robot coordinate system. Using this standard data, the inverse kinematics of the terminal are solved using the IPOPT solver in the nonlinear optimization and automatic differentiation tool CasADi. Based on the solution, the feedforward torque compensation (RNEA) is solved, and all results are finally filtered using a single Euro filter to obtain smooth control data. This final processed control data (including joint poses and torques) is sent to the robot's underlying controller (PD) via the robot operating system (ROS) and the Unitree SDK.

[0216] Step 4: Based on the robot's execution and perception terminals, Unitree G1 is controlled to receive and execute dual-arm control commands, while simultaneously uploading data from the robot's head and wrist, along with camera images, to the server. Optionally, the "record" parameter is used to enable recording and call the data recording submodule to save the multimodal teleoperation data (images, control status, gestures, joints, etc.) in a structured format (JSON) for use in AI training and action reproduction.

[0217] In step 5, the operator obtains an immersive 3D perspective through the Apple Vision Pro headset. The dual-arm robot responds to head perspective switching and natural gestures in real time, with smooth operation and low-latency feedback. At the same time, the recorded high-quality annotated data can be used for reinforcement learning model training, task playback, or operation review.

[0218] A server embodiment using the method of the present invention:

[0219] A server comprising:

[0220] one or more processing units;

[0221] a storage device for storing one or more programs;

[0222] When the one or more programs are executed by the one or more processing units, the one or more processing units implement the above-mentioned immersive robot teleoperation method.

[0223] The storage device is an internal memory, an external memory, a cache memory or other special memory. The processing unit has the signal processing capability and can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, an off-the-shelf programmable gate array or other programmable logic device.

[0224] An embodiment of a device applying the method of the present invention:

[0225] An electronic device is provided with a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium. When the program is executed by a processing unit, the above-mentioned immersive robot remote operation method is implemented.

[0226] Computer-readable storage media refers to physical media that can store computer-readable data, instructions, or programs. These media must meet the core characteristic of being "computer-readable" (i.e., data exists in the form of electrical, magnetic, optical, or other signals that can be converted into binary information that can be processed by a computer using appropriate equipment). These physical media include magnetic storage media, optical storage media, semiconductor storage media, or other storage media.

[0227] The model in this application is an object that objectively describes the morphological structure with the help of physical or virtual representation. The object is not equal to the physical body and is not limited to physical and virtual. It can be a data processing function, software program, processing mode, usage method, operation method, workflow, application process, electronic hardware, circuit module, processing system, system imitation or simulation object.

[0228] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and any modifications or equivalent replacements that do not deviate from the spirit and scope of the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. An immersive robot teleoperation method, characterized by: Includes the following: Based on the pre-built first-person mapping model, the robot's perspective image is acquired and processed to obtain a mapping image that can be presented in the operator's VR device; Through the pre-built interactive data acquisition model, the operator's posture control information is obtained in real time; Using the pre-built teleoperation processing model, the operator's posture control information is converted into control instructions in the robot coordinate system; It includes the following: Obtain the operator's posture control information and perform base coordinate system transformation on it to obtain the initial posture agreement; Perform alignment mapping processing for different initial posture conventions to obtain alignment data; Extract standard data structures for inverse kinematics and control planning from the alignment data; Generate control instructions in the robot coordinate system based on standard data structure for robot control; Using a pre-built execution perception model, the robot's posture is controlled based on control instructions, and images of the robot's movements are collected. Using a pre-built immersive perspective simulation model, the robot's motion image and mapping image are coupled to generate real-time responsive 3D images that can be displayed in VR devices, allowing the operator to obtain an immersive 3D perspective. The method for obtaining a mapping image that can be presented in the operator's VR device is as follows: Obtain the VR registration information of the VR device used by the operator and the corresponding robot registration information; The server establishes a three-party network transmission channel based on VR registration information and robot registration information; Using a three-party network transmission channel, the server receives the robot's perspective images collected by the robot's binocular camera in real time; Process the robot's view image to obtain the robot's binocular camera image stream; Using a three-party network transmission channel, the robot's binocular camera image stream is transmitted to the operator's VR device to obtain a mapping image that can be presented in the operator's VR device.

2. The immersive robot teleoperation method according to claim 1, wherein: The method for obtaining the operator's posture control information in real time through the pre-built interactive data acquisition model is as follows: Capture the operator's motion information, which includes the operator's head posture information, left and right hand posture information, and handle operation information; the left and right hand posture information includes finger joint position information and wrist posture information; Construct wrist transformation relationship based on the world reference system and the wrist local reference system; Based on the wrist transformation relationship, the left and right hand posture information is transformed into the wrist local reference system to obtain the finger joint data; The head posture information, finger joint data and handle operation information are stored in a dictionary type format to form posture control information.

3. The immersive robot teleoperation method according to claim 1, wherein: The method for obtaining the operator's posture control information and performing base coordinate transformation to obtain the initial posture agreement is as follows: Acquiring operator posture control information, which at least includes left wrist posture, left handle posture, right wrist posture, right handle posture, finger joint posture and head posture; According to the specific transformation matrix and based on the similarity transformation principle, the base coordinate system transformation of the posture control information is performed to obtain the posture output information based on the robot coordinate system convention; The posture output information is converted into an initial posture convention to obtain an initial posture convention, where the initial posture convention at least includes an initial posture convention required by the robot arm and an initial posture convention required by the robot finger.

4. The immersive robot teleoperation method according to claim 3, wherein: The specific transformation matrix is ​​constructed as follows: Based on the Open Extended Reality standard coordinate system convention, the z-axis is set to the back direction, the y-axis is set to the up direction, and the x-axis is set to the right direction; Based on the robot coordinate system convention, the z-axis is set to the upward direction, the y-axis is set to the left direction, and the x-axis is set to the forward direction; Based on the coordinate axis transformation relationship between the open extended reality standard coordinate system and the robot coordinate system, a specific transformation matrix is ​​constructed.

5. The immersive robot teleoperation method according to claim 1, wherein: Based on the standard data structure, the method for generating control instructions in the robot coordinate system is as follows: The first step is to obtain the standard data structure, which includes the data structure and the type structure; The data structure includes head posture and joint posture; the type structure includes gesture status information; The second step is to load the overall model of a robot, whose format is the model file under the robot operating system standard; The third step is to determine the position error and rotation error between the robot's current state and the target state based on the head posture, joint posture and gesture state information; The fourth step is to construct a total cost function based on the position error and rotation error. The optimization variable of the total cost function is the joint motor position. Step 5: Solve the total cost function to obtain the joint motor position solution; Step 6: Based on the joint motor position solution, the feedforward torque vector of the robot joint is calculated; Step 7: Determine the joint motor position based on the feedforward torque vector; Step 8: Perform real-time smoothing on the joint motor position to obtain the joint motor position angle; The ninth step is to process the joint motor position angle and generate control instructions in the robot coordinate system.

6. The immersive robot teleoperation method according to claim 1, wherein: The method for controlling the robot's posture and collecting the robot's motion images using the pre-built execution perception model based on control instructions is as follows: receiving a control instruction, which includes at least a target joint angle and a torque instruction; Based on the target joint angle and torque instructions, the local controller is called to drive the motor to perform the action, realizing the coordinated control of the arms and hands; and the status information of each motor is collected, including angle, speed, current and force feedback information; Real-time perception of camera image data streams from the head and wrist to obtain robot motion images.

7. The immersive robot teleoperation method according to claim 1, wherein: The method for using a pre-built immersive perspective simulation model to couple the robot motion image and the mapping image to generate a real-time responsive 3D image that can be displayed in a VR device is as follows: Obtain robot motion images collected synchronously in real time by the robot's binocular camera; The robot motion images are left-right image pairs with natural parallax; Performing distortion correction processing on the left and right image pairs to obtain left and right corrected images; The left and right corrected images are subjected to layered mask isolation processing and coupled with the mapped image to obtain a real-time responsive three-dimensional image that can be output to the left and right display screens of the VR device.

8. The immersive robot teleoperation method according to claim 7, wherein: The method for performing distortion correction processing on the left and right image pairs to obtain the left and right corrected images is as follows: Normalize the left and right image pairs to obtain several undistorted normalized coordinates; Calculate the squared radial distance between the undistorted normalized coordinates; Inversely estimate all distortion coefficients using a method of minimizing image point projection errors, wherein the distortion coefficients include radial distortion coefficients and tangential distortion coefficients; Based on the Brown-Conradie distortion model, the undistorted normalized coordinates, the squared radial distance value, the radial distortion coefficient and the tangential distortion coefficient are processed to obtain the distorted image coordinates; The distorted image coordinates are summed up to obtain the left and right corrected images.

9. An immersive robotic teleoperation system, characterized by: An immersive robot teleoperation method according to any one of claims 1 to 8 is applied, which is oriented to a humanoid robot or a dual-arm robotic arm and adopts a modular architecture, including a VR interaction terminal, a teleoperation algorithm processing unit, and a robot execution and perception terminal; The input data of the VR interaction end is the camera image stream data of the robot execution and perception end; the output data is the posture data of the wearable device, or the posture data and all button data of the VR device's matching handle; The remote operation algorithm processing unit uses the output data of the VR interactive end as input data and obtains the output data of the robot execution and perception end in real time. Based on the input data, it obtains the action data of the robot execution device. It is equipped with an inverse kinematics module, an image transceiver module, and a data recording module. The inverse kinematics module is used to achieve a complete mapping from posture data and all key data to the joint motor position angles that the robot can execute; The image transceiver module is used to receive camera image stream data from the robot execution and perception end, and forward the camera image stream data to the VR device end for rendering in real time; The data recording module is used to record the image data, status data and control instructions during the remote operation in real time; The input data of the robot execution and perception end is the action data of the remote operation algorithm processing unit; the output data is the camera image stream data on the robot and the status data of the robot execution device.

Citation Information

Patent Citations

  • Teleoperation method and device of robot, terminal equipment and storage medium

    CN119347781A

  • Humanoid robot remote control system and method based on virtual reality technology

    CN119501953A

Cited By

  • An immersive robotic teleoperation method, system, device, and medium

    CN122547239A