A wireless teleoperation method and system based on multi-modal feedback
By fusing real-time images and data using a multimodal feedback method, VR visuals are generated and control data is analyzed. This solves the problems of unintuitive interaction and limited feedback in wireless robot remote control technology, enabling efficient and safe remote robot operation and transparent fee settlement.
Patent Information
- Application Number
- CN202511445567.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing wireless robot remote control technology suffers from problems such as unintuitive interaction, limited feedback, and expensive equipment, making it difficult to achieve precise control and widespread application.
A multimodal feedback method is adopted to generate real-time VR images by integrating real-time images, target detection information, target tracking information, voxel information and early warning information collected by the robot. Combined with the analysis of head data, hand data and chassis motion data of user order tasks, the robot is controlled to perform operation actions. At the same time, a task fee calculation and settlement mechanism is introduced.
It significantly enhances the operator's sense of immersion and presence, enables multi-dimensional and precise control of the robot, reduces operational difficulty and learning costs, improves task execution efficiency and system security, and optimizes resource scheduling and cost settlement transparency.
Smart Images

Figure CN120921399B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a wireless remote control method and system based on multi-modal feedback, and belongs to the technical field of intelligent robots. BACKGROUND
[0002] In the current era of rapid technological development, 5G technology provides a solid guarantee for real-time data transmission with its high speed, low latency and large capacity; VR technology continues to innovate, enabling highly realistic immersive experiences; and robot technology has made significant breakthroughs in mechanical design, intelligent control and other aspects, with stronger environmental adaptability and operational precision. The vigorous development of these technologies has made it possible for remote control robot assistance systems to gradually move from theoretical concepts to practical applications in the technical field, opening up new paths for solving complex tasks in various fields and meeting the needs of special groups.
[0003] The inventor found that existing wireless robot remote control technology has many shortcomings: first, the interaction is not intuitive, most systems rely only on VR and handle control, lack immersive perception, and operators have difficulty accurately perceiving the remote environment; second, the feedback is single and has high latency, lacks visual or tactile feedback, and fine control is difficult to achieve; third, the system is expensive, often requiring expensive VR equipment and specialized control hardware, which is not conducive to popular application. SUMMARY
[0004] The purpose of the present application is to overcome the shortcomings of the prior art and provide a wireless remote control method and system based on multi-modal feedback, which solves the problems of non-intuitive interaction, single feedback and expensive equipment in the prior art.
[0005] To achieve the above-mentioned purpose / to solve the above-mentioned technical problems, the present application adopts the following technical solutions:
[0006] A wireless remote control method based on multi-modal feedback applied to a cloud server, the method comprising:
[0007] Fusing real-time images, target detection information, target tracking information, voxel information and early warning information collected by the robot at the control terminal to generate a VR real-time picture;
[0008] Real-time analysis of head data, hand data and chassis motion data for controlling the robot according to the user's order task and the VR real-time picture;
[0009] Determining the upper limb motion trajectory of the robot through the hand data;
[0010] According to the upper limb movement trajectory, the head data, the hand data, and chassis movement data, control the robot to perform the operation action of the order task, wherein: the upper limb movement trajectory is used to control the joint movement and posture of the robot arm, the hand data is used to realize the operation action such as grabbing, rotating and placing of the robot hand, the head data is used to control the rotation and observation angle of the robot head, and the chassis movement data is used to control the forward movement, backward movement and turning of the robot chassis, so as to realize the accurate operation of the order task.
[0011] Optionally, the method further comprises:
[0012] Determining the cost of the order task;
[0013] According to the operation time and difficulty coefficient of the order task, determining the settlement cost of the operation end from the cost of the order task.
[0014] Optionally, the real-time image, target detection information, target tracking information, voxel information and early warning information collected by the robot comprise:
[0015] Collecting real-time images of the environment where the robot is currently located;
[0016] Pretreating the real-time images to obtain pretreated image data;
[0017] Inputting the pretreated image data into a visual neural network model to obtain target detection information, target tracking information, voxel information and early warning information.
[0018] Optionally, the pretreating the real-time images to obtain pretreated image data comprises:
[0019] Performing frame-by-frame slicing operation on the real-time images to decompose the continuous video stream into a time-ordered single-frame image sequence to obtain RGB images and depth images;
[0020] The RGB images are sequentially subjected to color balance processing based on the gray world assumption, then subjected to bilateral filtering to suppress noise and retain edge information, and finally subjected to resolution adjustment and pixel value normalization;
[0021] The depth images are subjected to depth denoising and smoothing processing, then subjected to interpolation scaling and spatial alignment with the RGB frames at the pixel level, and subjected to normalization processing to obtain the pretreated image data.
[0022] Optionally, the inputting the pretreated image data into the visual neural network model to obtain the target detection information, the target tracking information, the voxel information and the early warning information comprises:
[0023] extracting multi-scale feature structures from the preprocessed RGB image and the depth image;
[0024] performing spatial alignment and weighted fusion of the extracted multi-scale feature structures and depth features corresponding to the RGB features to generate fused features;
[0025] identifying a candidate region in the fused features, outputting a bounding box coordinate, a class label, and a confidence score, and dynamically adjusting a detection threshold according to brightness of the candidate region and texture complexity of the candidate region; wherein the brightness is obtained by converting the RGB image into an HSV color space to extract a brightness component and calculating a brightness mean value of the candidate region, and the texture complexity is obtained by performing an edge detection operator on the RGB image to obtain an edge binary image and counting a proportion of edge pixel points in a total number of pixel points of the candidate region;
[0026] performing pixel-level probability prediction on the candidate region and obtaining a binary mask through threshold segmentation to label a foreground region and obtain target detection information;
[0027] generating depth information based on the fused features, converting the depth information into a voxel block, and generating voxel information, wherein the voxel information is used to present a three-dimensional spatial structure;
[0028] calculating a three-dimensional position and a minimum safety distance of an obstacle using the voxel information in combination with a current pose of a robot to obtain early warning information;
[0029] obtaining target tracking information based on a detection result output by target detection in combination with a tracking state of a historical frame.
[0030] Optionally, the dynamically adjusting the detection threshold according to the brightness of the candidate region and the texture complexity of the candidate region comprises:
[0031] when the brightness mean value is less than a first preset value and the proportion of edge pixel points is less than a first preset percentage, setting the confidence threshold value as a first target value;
[0032] when the brightness mean value is greater than a second preset value and the proportion of edge pixel points is greater than a second preset percentage, setting the confidence threshold value as a second target value;
[0033] when the brightness mean value is greater than the first preset value and less than the second preset value, and the proportion of edge pixel points is greater than the first preset percentage and less than the second preset percentage, setting the confidence threshold value as a third target value.
[0034] Optionally, the upper limb movement trajectory is inversely solved according to the hand data, comprising:
[0035] The wrist IMU based on the operating terminal captures three-axis acceleration, three-axis angular velocity data, and changes in finger joint flexion state and applied pressure based on hand data;
[0036] The triaxial acceleration and triaxial angular velocity data, combined with the finger joint flexion state and the applied pressure changes, are input into the upper limb motion inverse model to output the upper limb motion trajectory. The upper limb motion trajectory includes the joint angles of the elbow flexion and extension and forearm rotation, shoulder internal rotation and / or external rotation and lifting degrees of freedom.
[0037] The second aspect: a wireless remote control system based on multimodal feedback, comprising:
[0038] The cloud server is used to receive user order tasks, send the order tasks to the operation terminal, determine the cost of the user order tasks, and determine the settlement fee of the operation terminal from the cost calculated for the order tasks based on the operation time and difficulty coefficient of the order tasks.
[0039] The vision processing module is used to receive real-time images of the robot's current environment, obtain target detection information, target tracking information, voxel information and warning information based on the real-time images, and send the real-time images, target detection information, target tracking information, voxel information and warning information to the control terminal. After fusion at the control terminal, a VR real-time image is generated.
[0040] On the operating end, based on the order task and the real-time VR screen received by the control terminal, the head data, hand data and chassis motion control data of the robot are analyzed in real time and sent to the control terminal.
[0041] The control terminal is used to receive real-time VR images and send the hand movement data, chassis movement data, and head data simulated by the gyroscope built into the control terminal to the robot.
[0042] The robot is used to perform the control actions of the order placement task based on head data, hand movement data and chassis movement data sent by the receiving control terminal.
[0043] Optionally, the operating terminal includes a pedal control panel, a VR stand, and a remote control glove;
[0044] The pedal control panel is used to identify pressing actions in the forward, backward, left, and right directions, as well as to detect foot rotation actions, thereby controlling the robot's directional angle.
[0045] The VR bracket trigger control terminal uses the attitude information generated by the built-in gyroscope to synchronously control the rotation of the robot's head;
[0046] The remote control glove is used to achieve coordinated control of the robot's hand movements and upper limb motion trajectories, as well as to provide feedback on the pressure of the robot's hand.
[0047] Optionally, the visual processing module is further configured to collect the temperature of the target object and send the collected temperature of the target object to the control terminal, so that the control terminal displays the temperature in the VR picture.
[0048] Compared with the prior art, the present application has the following beneficial effects:
[0049] The control terminal can generate a highly realistic VR real-time picture by fusing the real-time image collected by the robot, the target detection information, the target tracking information, the voxel information and the early warning information, thereby significantly improving the immersion and the sense of presence of the operator, and making the remote operation like on-site operation. In combination with the analysis of the head data, the hand data and the chassis motion data of the user's ordered task, multi-dimensional accurate control of the robot upper limbs, hands, head and chassis can be realized, the naturalness and coordination of the control action are ensured, and the problems of non-intuitive interaction and single feedback in the prior art are solved. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 The figure shows the structure schematic diagram of the remote control equipment of the embodiment of the present application;
[0051] Figure 2 The figure shows the flow chart of the cloud server of the embodiment of the present application;
[0052] Figure 3 The figure shows the overall method flow schematic diagram of the embodiment of the present application;
[0053] Figure 4 The figure shows the multi-modal perception and control processing flow schematic diagram of the robot of the embodiment of the present application;
[0054] Figure 5 The figure shows the upper limb motion inverse solution model flow chart of the embodiment of the present application;
[0055] Figure 6 The figure shows the overall structure schematic diagram of the pedal control disc of the embodiment of the present application;
[0056] Figure 7 The figure shows the exploded structure schematic diagram of the pedal control disc of the embodiment of the present application.
[0057] In the figure: 1, pedal; 2, rotating shaft; 3, shell; 4, cover plate; 5, ball; 21, cloud server; 22, robot; 23, operation end; 24, VR support; 25, control terminal; 26, remote control glove; 27, pedal control disc; 28, visual processing module. DETAILED DESCRIPTION
[0058] In order to make the technical means, creative features, purposes and effects realized by the present application easy to understand, the present application will be further described in conjunction with specific embodiments.
[0059] In the description of the present application, it needs to be understood that the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" and the like are only for the purpose of description and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second" and the like can be explicitly or implicitly included one or more. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0060] In the description of the present application, it needs to be understood that unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood through specific circumstances.
[0061] Embodiment 1, reference Figure 2 As shown, a wireless remote control method based on multi-modal feedback is disclosed, applied to a cloud server 21, the method comprises:
[0062] Step 11, after the real-time image, target detection information, target tracking information, voxel information and early warning information collected by the robot 22 are fused in the control terminal 25, a VR real-time picture is generated;
[0063] Step 12, according to the user's order task and the VR real-time picture, the head data, hand data and chassis motion data for controlling the robot 22 are analyzed in real time;
[0064] Step 13, the upper limb motion trajectory of the robot 22 is determined through the hand data;
[0065] Step 14, according to the upper limb movement trajectory, head data, hand data and chassis movement data, control the robot to perform the operation action of the order task, wherein: the upper limb movement trajectory is used to control the joint movement and posture of the robot arm, the hand data is used to realize the operation action such as grabbing, rotating and placing of the robot hand, the head data is used to control the rotation and observation angle of the robot head, and the chassis movement data is used to control the forward, backward and turning of the robot chassis, so as to realize the accurate operation of the order task.
[0066] By fusing the real-time images collected by the fusion robot 22, target detection information, target tracking information, voxel information and early warning information, the control terminal 25 can generate a highly realistic VR real-time picture, thereby significantly improving the immersion and presence of the operator, making the remote operation like on-site operation. Combined with the analysis of the head data, hand data and chassis movement data by the user order task, multi-dimensional accurate control of the robot upper limb, hand, head and chassis can be realized, ensuring the naturalness and coordination of the operation action. In particular, the combination of upper limb movement trajectory and hand data not only enables fine operations such as grabbing, rotating and placing, but also ensures accurate reproduction of joint movement and posture, thereby improving the execution efficiency and reliability of complex tasks. At the same time, the fused early warning information can effectively prompt potential risks, reduce the occurrence of misoperation and collision, and improve the safety and robustness of the system. Overall, this method reduces the operation difficulty and learning cost, so that even non-professionals can quickly get started and achieve efficient, safe and accurate task execution in a remote environment.
[0067] As shown in Figure 3 , the method of the present embodiment further comprises:
[0068] determining the cost of the order task; determining the settlement cost of the operation end 23 from the cost of the order task according to the operation time and difficulty coefficient of the order task. Specifically, it is described as follows:
[0069] By introducing the task cost calculation and settlement mechanism in the system, the cost can be dynamically determined according to the task operation time and difficulty coefficient, so as to ensure that the task cost matches the actual operation workload. This automatic billing method reduces manual intervention and accounting errors, improves the fairness and transparency of cost settlement. At the same time, this mechanism provides the operation end with reasonable remuneration proportional to the task difficulty and time consumption, which helps to encourage the operation personnel to improve the operation efficiency and service quality. Combined with the unified management of the cloud server, the system can realize real-time billing and settlement for a large number of tasks, thereby optimizing resource scheduling and task allocation and improving overall operation efficiency. Ultimately, this mechanism not only improves the user's use satisfaction, but also enhances the enthusiasm of the operation end and the sustainability of the system.
[0070] In the specific implementation process of the embodiment, the real-time image, the target detection information, the target tracking information, the voxel information and the early warning information collected by the robot include:
[0071] collecting a real-time image of an environment in which the robot is currently located;
[0072] preprocessing the real-time image to obtain image data after preprocessing;
[0073] inputting the image data after preprocessing into a visual neural network model to obtain the target detection information, the target tracking information, the voxel information and the early warning information.
[0074] In the specific implementation process of the embodiment, the preprocessing of the real-time image to obtain the image data after preprocessing includes:
[0075] extracting a multi-scale feature structure of the image from the RGB image and the depth image after preprocessing;
[0076] performing spatial alignment and weighted fusion of the extracted feature structure to generate fusion features;
[0077] identifying a candidate target region in the fusion features, outputting a bounding box coordinate, a class label and a confidence score, and dynamically adjusting a detection threshold according to brightness and texture complexity; the brightness information of the candidate region is obtained by converting the RGB image into an HSV color space to extract a brightness component and calculating a brightness mean value of the region; the texture complexity information of the candidate region is obtained by performing an edge detection operator on the RGB image to obtain an edge binary image and counting a proportion of edge pixel points in a total number of pixel points of the candidate region;
[0078] performing pixel-level probability prediction on the candidate target region and obtaining a binary mask through threshold segmentation to label a foreground region, to obtain the target detection information;
[0079] combining the detection information output based on the target detection with a tracking state of a historical frame to obtain the target tracking information.
[0080] generating depth information based on the fusion features, the depth information including a smooth disparity map and a pseudo-color depth map, and further converting the depth information into a voxel block to generate voxel information, the voxel information being used to present a three-dimensional spatial structure, and using the voxel information in combination with a current pose of the robot to calculate a three-dimensional position and a minimum safety distance of an obstacle to obtain the early warning information;
[0081] The embodiment can significantly improve the accuracy and stability of target recognition by performing multi-level preprocessing and feature extraction on real-time images. The extraction of multi-scale feature structures can capture target information of different sizes and complexities in the image, making the system more sensitive to small targets and local details. The spatial alignment and weighted fusion of RGB features and depth features result in fused features, which fully integrate visual information and spatial information, improving the accuracy of target detection and enhancing the robustness in varying light conditions or occlusion.
[0082] By dynamically adjusting the detection threshold based on the brightness and texture complexity of the candidate target region, the system can adapt to different environmental lighting conditions and background complexity, reducing the probability of missed detection and false detection. Pixel-level probability prediction and binary mask generation allow accurate labeling of foreground targets, providing high-quality target information for subsequent operations. Combined with the tracking state of historical frames, target tracking can achieve continuous tracking and state updating of targets, ensuring the real-time response capability of the system in dynamic environments, thereby improving the accuracy and smoothness of remote control.
[0083] Based on the depth information generated from the fused features, the depth information is converted into voxel blocks to generate voxel information, thereby constructing a three-dimensional spatial structure. This can intuitively and accurately reconstruct the three-dimensional morphology of the environment and targets. Voxel information can not only present the spatial layout of complex scenes, but also assist in calculating the three-dimensional position and minimum safety distance of obstacles, achieving safe constraints on the robot's operation path. The calculation of the minimum safety distance and the generation of warning information help to real-time prompt potential collision risks, improving the safety and robustness of the system. Overall, the embodiment can achieve high-precision, high-robustness, and high-safety target detection, tracking, and environment perception based on multi-modal information fusion and voxel information three-dimensional reconstruction, providing reliable and comprehensive data support for remote robot control.
[0084] As shown in Figure 4 The embodiment further illustrates that the visual neural network model adopts a shared backbone network structure, and three prediction heads are set in parallel on the fused features extracted by the structure, respectively completing target detection and mask generation, target tracking, and warning information calculation. This realizes multi-task cooperative processing based on unified features, and outputs target detection information, target tracking information, voxel information, and warning information.
[0085] Specifically, the visual neural network model obtains RGB images and depth images from a camera in real time, inputs the images into a shared backbone network structure after frame-by-frame slicing to extract multi-scale features, and performs spatial alignment and weighted fusion on the multi-scale RGB features and depth features to obtain fused features. The images are preprocessed before being input to ensure the accuracy and robustness of the subsequent model. The image preprocessing includes color correction based on the gray world assumption, i.e., calculating the average values of each channel in the image and adjusting them to be consistent to reduce the impact of light source changes; after color correction, bilateral filtering is applied to the image to remove high-frequency noise while preserving edge features through spatial Gaussian kernel and pixel value Gaussian kernel; then image scaling is performed to match the input size of the detection model, and normalization (pixel value scaling to [0, 1]) is performed on the RGB channels. The depth image is simultaneously filled with invalid values and smoothed in the preprocessing stage, and its output is used for subsequent voxel information construction and early warning information calculation.
[0086] In the detection prediction head, candidate target regions are identified based on the fused features and through the anchor box mechanism, and the boundary box coordinates, class label and confidence score are output, and NMS (Non-Maximum Suppression) is used to suppress redundant boxes. The system also designs a dynamic threshold adjustment mechanism: taking the global mean of the brightness component in the image HSV space as the brightness indicator, and taking the edge point proportion extracted by the edge detection operator as the texture complexity reference, automatically down-regulating to the first target value in weak light or weak texture scenes, and up-regulating to the second target value when the light is good and the structure is rich, so as to balance the detection sensitivity and false detection rate. The detection result outputs the boundary box [xmin, ymin, xmax, ymax], and further calculates the target environment occlusion score (based on the edge point proportion), and through hierarchical weight correction, the accuracy is enhanced, providing prior information for the subsequent prediction head.
[0087] On this basis, the mask prediction takes the result target candidate region of the detection prediction head as the prior input, uses the U-Net type decoding structure to output the pixel-level probability map, and performs foreground / background segmentation with a threshold of 0.5. The final result is rendered as a semi-transparent green layer (α=0.4), which is superimposed on the original image through OpenCV to realize real-time highlighting of the target region, and serves as the target detection information.
[0088] In the tracking prediction head, the target state [x, y, vx, vy] is updated based on the fused features and combined with the Kalman filter prediction-correction framework. The detection frame of each frame is input as an observation value, and the association with the historical trajectory is completed through IOU (threshold 0.3). If there is no match for 5 consecutive frames, the trajectory is removed, and the target tracking information is generated. To deal with small targets at a long distance, the system automatically performs 2x upsampling (bicubic interpolation) when the detected target box area is less than 10% of the image area, and marks the enlarged state with a green thick border on the interface to maintain tracking stability and visibility. At the same time, the system maintains historical state (position, speed and motion trend), and enhances the local feature channel of the target when occluded, thereby improving the robustness of long sequence tracking.
[0089] In the depth-warning prediction head, the depth information is obtained based on the fused feature output disparity map and pseudo-color depth map, and further converted into voxel blocks to generate voxel information, thereby constructing a three-dimensional spatial structure. Combined with the current pose data of the robot, the fusion modeling is completed and the three-dimensional position and minimum safety distance of the obstacle are calculated in real time to generate warning information. Specifically, the depth image pixels are back-projected to a three-dimensional coordinate system and organized into local voxel blocks for spatial modeling and collision detection; the point cloud in the ±10° range in front of the robot is filtered and the nearest distance is taken as the obstacle distance. When the obstacle distance is less than 0.5m, the first level warning (yellow frame 2Hz flicker) is triggered; when the obstacle distance is less than 0.1m, the second level warning (red frame 5Hz flicker) is triggered. The final detection result, mask layer, depth pseudo-color map and warning marker are uniformly encoded and compressed, and pushed to the mobile terminal at 60fps. The bandwidth is optimized through the UDP protocol to ensure real-time visual feedback and safety prompt.
[0090] The mobile terminal simultaneously collects the operator's head pose using a gyroscope and an accelerometer, and calculates the pitch angle and yaw angle using adaptive Kalman filtering. In the prediction stage, the angle is updated based on the integral of angular velocity, and in the update stage, the acceleration inclination is corrected, and the noise covariance R is dynamically adjusted during rapid rotation to improve response speed. The final attitude angle is mapped to the view angle control signal, which is transmitted to the robot servo motor in real time through UDP. The entire interaction process is based on the prediction-correction closed-loop compensation mechanism, which makes the end-to-end control delay ≤50ms, and ensures smooth switching of the view angle without sudden changes under the condition of 60fps.
[0091] In the specific implementation process of the embodiment, the detection threshold is dynamically adjusted according to the brightness and texture complexity, which includes:
[0092] The RGB image of the candidate region is converted to the HSV color space to extract the brightness component, and the average brightness of the region is calculated. At the same time, an edge detection operator is performed on the region to obtain an edge binary image, and the proportion of edge pixels to the total number of pixels in the candidate region is calculated as a texture complexity indicator;
[0093] When the average brightness is less than the first preset value and the edge pixel ratio is less than the first preset percentage, the confidence threshold is set to the first target value; in the specific implementation process of the embodiment, when the average brightness is less than 80 and the edge pixel ratio is less than 5%, the confidence threshold is set to 0.6;
[0094] When the average brightness is greater than the second preset value and the edge pixel ratio is greater than the second preset percentage, the confidence threshold is set to the second target value; in the specific implementation process of the embodiment, when the average brightness is greater than 160 and the edge pixel ratio is greater than 15%, the confidence threshold is set to 0.8;
[0095] In addition to the above two cases, when the average brightness is greater than the first preset value and less than the second preset value, and the edge pixel ratio is greater than the first preset percentage and less than the second preset percentage, the confidence threshold is set to the third target value, which is set to 0.7, so as to realize dynamic threshold adjustment under different light conditions and scene complexity.
[0096] As shown in Figure 5 In the specific implementation process of the embodiment, the robot performs inverse solution processing on the upper limb motion trajectory based on the hand data, and the specific process is referred to Figure 5 Upper limb motion inverse solution model flow chart:
[0097] First, three-axis acceleration, three-axis angular velocity data captured by the wrist IMU based on the operating end according to the hand data, and finger joint bending state and applied pressure change are obtained, wherein the hand data includes hand state data (hand bending degree), pressure distribution data (applied pressure change), and wrist motion data (three-axis acceleration, three-axis angular velocity) of the wrist IMU;
[0098] Subsequently, the motion feature extraction module (LSTM network) is used to model the time sequence features of the wrist motion data, and the hand state data and the pressure distribution data are processed through the operation intention analysis module (MLP network). After the above two types of features are integrated through the feature fusion processing module, they are input into the joint angle prediction module to obtain the estimation result of the operator's upper limb joint angle, including the angle information of the elbow flexion and extension, forearm rotation (pronation / retroversion), shoulder internal rotation / external rotation, and lifting, etc. Degree of freedom.
[0099] In the prediction process, the upper limb motion inverse solution model adopts a deep neural network structure composed of a long short-term memory network (LSTM) and a multi-layer perceptron (MLP). The input is a sequence of wrist IMU data sampled continuously for 0.5 seconds (50 frames) and the pressure distribution of the fingertips and palm center. The output is the shoulder and elbow joint angles. To ensure the rationality of the prediction, a biomechanical constraint module (including angle range and joint speed limit) is introduced. In addition to the mean square error (MSE), a regularization term based on the kinematic constraint is also introduced in the loss function. The model is trained using the Adam optimizer (adaptive moment estimation optimizer). The training data comes from a variety of hand-upper limb coordinated action samples labeled by a high-precision motion capture system, covering common scenarios such as grasping, twisting, lifting, and holding, totaling 100,000 sequence samples.
[0100] In real-time operation, hand data is read at a certain frequency and input into the upper limb motion inverse solution model through a sliding window to predict joint angles and action intentions. The joint angle prediction results are converted into control instructions for the robot's shoulder and elbow joints through a motion planning module and combined with the hand end pose control signal to achieve coordinated control of the robot arm and hand. Action intention prediction is used to anticipate the operator's future action trends, providing feedforward information for the robot control system to improve its response speed and prediction ability during execution. To avoid control oscillation caused by high dynamic actions, a first-order low-pass filter is introduced at the output end for smoothing, and a Kalman filter is used for state fusion, significantly improving the stability and robustness of prediction and control.
[0101] Embodiment 2 discloses a wireless remote control system based on multi-modal feedback, comprising:
[0102] A cloud server 21 is configured to receive a user's order task, send the order task to an operation terminal, settle the cost of the order task for the user, and settle the cost for the operation terminal from the cost settled by the user according to the operation time and difficulty coefficient of the order task;
[0103] A visual processing module 28 is configured to receive real-time images of the environment in which the robot is currently located, obtain target detection information, target tracking information, voxel information, and warning information from the real-time images, and send the real-time images, target detection information, target tracking information, voxel information, and warning information to a control terminal. The control terminal 25 generates a VR real-time picture after fusion.
[0104] The operation terminal 23 synchronously completes the task operation according to the order task combined with the VR real-time picture received by the control terminal 25, and generates head data, hand data, and chassis motion control data for controlling the robot by real-time analysis of the operation input.
[0105] The control terminal 25 is used for receiving the VR real-time picture and sending the hand movement data, chassis movement data and head data simulated by the built-in gyroscope of the control terminal to the robot;
[0106] The robot 22 is used for moving according to the head data, hand movement data and chassis movement data sent by the control terminal to realize accurate control of the task target and complete the ordered task.
[0107] As shown in the figure, the operation end 23 comprises a VR support 24, a control terminal 25, a remote control glove 26 and a pedal control panel 27. Figure 1
[0108] As shown in the figure, in the specific implementation process of the embodiment, the pedal control panel 27 comprises a pedal 1, a rotating shaft 2, a shell 3 and a cover plate 4, the pedal 1 is fixedly connected with the rotating shaft 2 to form an integral body, the integral body is inserted into the opening end of the shell 3 during installation, so that the sphere 5 at the lower end of the rotating shaft 2 is embedded into the spherical support of the shell 3 to form a universal support, and the bottom of the sphere 5 is provided with a protrusion; the rotating shaft 2 is arranged in a cross shape, and the cross-shaped shaft end is in contact with the protrusion at the bottom of the sphere 5. Figures 6-7
[0109] During use, the operator applies pressure to the pedal 1 in the front-rear or left-right direction by the foot, so that the pedal 1 and the sphere 5 below it produce tilting movement, the protrusion at the bottom of the sphere 5 pushes the cross-shaped shaft to swing in the corresponding direction, thereby outputting control signals in the front-rear and left-right directions; when the operator applies a rotating force to the pedal 1, the pedal 1 drives the sphere to rotate around the central shaft, thereby realizing the rotating control function. The pedal control panel 27 is designed with foot interaction as the core, the inside of the pedal control panel 27 triggers micro switches or sensing elements through the sphere 5, which is used for identifying the stepping actions in the front, rear, left and right directions, and generating corresponding translation control signals accordingly, the sphere 5 cooperates with the rotating sensing device to detect the rotating action of the foot, thereby realizing the control of the direction angle of the robot, and the VR support 24 triggers the attitude information generated by the built-in gyroscope of the control terminal, which is used for synchronously controlling the rotation of the head of the robot.
[0110] During use, the user applies pressure to the pedal 1 in the front-rear or left-right direction by the foot, so that the pedal 1 and the sphere 5 produce tilting movement, and the protrusion at the bottom of the sphere 5 sequentially presses the buttons in the corresponding direction, thereby realizing the identification of the translation direction. When the user applies a rotating force (for example, the foot rotates along an arc), the sphere can rotate around the central shaft, and the rotating sensing structure at the bottom captures the action, thereby generating continuous direction angle (Yaw axis) control signals. The single sphere realizes multi-dimensional interaction of translation and rotation, so that the operation is more intuitive and smooth, and has high sensitivity and good stability.
[0111] In terms of signal acquisition, the pedal control disc is internally integrated with a signal acquisition circuit for obtaining the foot action signals of the operator. Among them, the control in the front, back, left and right directions is triggered by buttons and outputs digital signals, and the rotating action is collected through a rotating sensing structure to obtain the angle change. In order to improve the anti-interference ability, the signal is subjected to de-bouncing processing and low-pass filtering processing to filter out the slight shaking of the foot and external electromagnetic interference. The signal analyzed by the microcontroller is transmitted to the control terminal at a frequency of 50Hz through the low-power Bluetooth module, and the transmission data includes the translation direction signal and the rotation angular velocity, and is accompanied by timestamp information to ensure the synchronicity and integrity of the signal.
[0112] After the control terminal receives the above-mentioned signal, the translation direction signal and the rotation signal are respectively mapped into standardized robot motion instructions. Specifically, the translation direction signal is mapped into forward / backward speed and lateral speed, and the rotation signal is mapped into angular velocity, thereby forming a two-dimensional plane motion state parameter, including forward speed vx, lateral speed vy and steering angular velocity ω.
[0113] In order to avoid sudden actions caused by small signal fluctuations, the control terminal introduces a threshold filtering mechanism to the received signal and adopts exponential moving average smoothing processing to enhance the stability of the motion instructions. The finally generated motion instructions are sent to the robot body through a wireless network, and the robot receives and analyzes the instructions and drives the motor to execute the corresponding action. Based on the closed-loop control strategy, the robot body combines the feedback of the built-in sensors (including encoders and gyroscopes) to correct the execution result in real time, thereby realizing high-precision control. The delay of the whole signal acquisition and transmission to execution is controlled within 100ms to ensure the real-time requirement of the system.
[0114] The VR support 24 triggers the attitude information generated by the built-in gyroscope of the control terminal, which is used to synchronously control the rotation of the robot head;
[0115] The remote control gloves 26 are used to realize the coordinated control of the robot hand action and the upper limb motion trajectory, and to feedback the pressure of the robot hand.
[0116] Through the control terminal 25 and in cooperation with the VR support 24, the chassis motion control is realized in combination with the pedal control disc 27, and the remote control gloves 26 are used to capture the tactile and action feedback, and the visual, action and other multi-modal information are fused and displayed, forming a low-cost, low-delay and immersive remote control system, which effectively solves the problem of non-intuitive interaction and single feedback in the prior art.
[0117] By controlling the terminal 25 in cooperation with the VR support 24, a low-cost immersive control interface is constructed, which supports head pose control and can complete environment observation and target positioning in real time following the operator's view angle, and the rendering visualization of visual information is completed by the control terminal, which significantly improves the intuitiveness and immersion of remote control.
[0118] In the specific implementation of this embodiment, the control terminal 25, i.e., the mobile phone, serves as the core device of the operation terminal 23, such as... Figure 3 As shown, the control terminal 25 is responsible for integrating head data, hand data, and chassis motion control data, while sending instructions from the operator terminal 23 to the robot 22. The control terminal 25 serves as both a data display platform and a signal relay station, undertaking multiple functions including command generation, data processing, and feedback presentation. This control terminal supports various input methods, including video feeds and feedback information. The control terminal maintains a real-time connection with the robot via efficient wireless communication, ensuring seamless integration of command transmission, feedback reception, and voice signal transmission. It also supports rapid data processing and display to provide a smooth operating experience. The entire design emphasizes low cost and ease of use, enabling operators to interact with the remote robot in an intuitive way, making it suitable for various home service scenarios.
[0119] In this embodiment, the control terminal is equipped with a VR stand for easy operation and use, achieving an effect equivalent to expensive VR display devices. The control terminal connects to the robot via a wireless network, fusing real-time images, target detection information, target tracking information, voxel information, and warning information sent by the vision processing module 28. This allows for real-time visual feedback on the control terminal, displaying the robot's perspective of the environment. During operation, the control terminal supports the operator's perspective adjustment function, capturing head movements through the phone's built-in gyroscope and converting the data into synchronized commands from the robot's perspective. This design ensures the operator can naturally observe the task environment, enhancing the intuitiveness of operation. To optimize transmission efficiency, the method employs compression technology to process video data, ensuring image clarity and smoothness while reducing network resource consumption.
[0120] In this specific implementation, the remote control glove 26 has 10 pressure sensors (distributed at the fingertips and palm area, with a range of 0–50N and a resolution of 0.1N) at its front end to accurately sense the gripping contact force, capture the operator's movements, and simulate the interaction between the robot and the environment. A six-axis IMU (Inertial Measurement Unit) is embedded in the wrist of the remote control glove 26 to continuously track changes in wrist posture. For example, when the robot contacts an object, the operator can sense the corresponding resistance or surface characteristics.
[0121] The remote control glove 26 is connected to the control terminal in a wireless manner, data transmission is fast and stable, and the VR picture of the tactile feedback is highly consistent with the actual operation. At the same time, the built-in wrist IMU can be used to sense the wrist posture in real time, the approximate motion state of the operator's upper arm joint is inferred by inputting the hand data into the upper limb motion inverse solution model, and the motion trajectory of the robot's upper limb is obtained, so as to realize the cooperative control of hand action and upper arm posture, and further improve the freedom and naturalness of robot operation.
[0122] In the specific implementation process of the embodiment, the robot captures RGB color pictures and depth pictures through a binocular camera, and the RGB color pictures and the depth pictures are subjected to target detection and depth estimation through a visual neural network model to help the robot to locate and track targets and prevent the robot from colliding.
[0123] The remote control glove 26 is transmitted to the control terminal through the Bluetooth protocol, and the average communication delay is less than 10 milliseconds. The control terminal is not only used for generating tactile feedback signals, but also based on the wrist IMU time sequence data and the hand state, the dynamic posture of the operator's elbow and shoulder is inferred in real time through the upper limb motion inverse solution model, and the complete upper limb motion data is constructed. The upper limb motion inverse solution model adopts a time sequence deep learning structure (such as a mixed network of LSTM-MLP), supports an online inference rate of 100 Hz, and encodes the prediction results and the hand control signal into action instructions of the robot end together, which is transmitted to the robot wirelessly, and the robot executes the action instructions to control the robot to perform the operation action of the single task.
[0124] In use, the embodiment:
[0125] The operator establishes a connection with the robot through the APP of the control terminal, and realizes low-delay communication by using the UDP protocol. The connection process is initiated by the robot end, which continuously sends 30 data packets at an interval of 5 milliseconds, each containing a sequence number and a timestamp, and the mobile phone end immediately returns an acknowledgement packet after receiving. The robot end calculates the network bandwidth (obtained by dividing the size of the acknowledgement packet by the time difference) and the round-trip delay RTT (obtained by the difference between the sending and returning timestamps) accordingly, and calculates the packet loss rate by detecting the proportion of missing sequence numbers. If the bandwidth is lower than 20 Mbps, the sending window is fixedly set to 5 packets to avoid congestion; if the packet loss rate exceeds 5%, a forward error correction mechanism is enabled, and 1 redundant packet (redundancy rate 10%) is inserted in every 9 valid packets by using XOR (exclusive OR operation), so that the receiving end can recover the single packet loss. The entire connection establishment process is completed within 100 milliseconds, and after the connection is successful, the network status indicators, including the bandwidth, delay and packet loss rate, are displayed on the mobile phone APP interface, and the system enters the ready state. In this stage, the mobile phone end is only responsible for decoding and displaying the video transmitted by the robot end, and continuously sends the operator's instructions, such as head motion control instructions, to the robot to ensure the two-way synchronization of communication.
[0126] First step: Start the platform, and establish a connection between the robot and the control terminal by wireless means. The user wears a VR support on the head and a teleoperation glove on the hand, and places both feet above the foot control panel 27. The teleoperation glove and the foot control panel respectively establish a communication connection with the control terminal through Bluetooth.
[0127] Second step: The robot end acquires real-time RGB images and depth images through binocular cameras, processes image data in real time to realize target detection, target tracking and depth estimation, and transmits the data packets of the picture and processing results to the control terminal in a low-delay and high-bandwidth manner through the server. The control terminal accepts the picture and processing result data packets, superimposes the processing results on the original picture in the form of a semi-transparent information layer, etc., so that the operator can obtain an immersive perception experience of the first-person perspective on the control terminal. The gyroscope built-in the control terminal collects the head movement information of the operator in real time and converts it into robot head control instructions to realize visual angle linkage control, i.e. "the human head moves, the robot head moves", further enhancing the natural interaction and perception sensitivity of remote operation.
[0128] Third step: The user controls the foot control panel through the feet to realize forward and backward, left and right direction tilt control, as well as clockwise and counterclockwise rotation control, and the related motion data is sent to the control terminal through Bluetooth.
[0129] Fourth step: The user can freely bend the fingers and rotate the wrist of the teleoperation glove worn by the user, and the related motion information is transmitted to the control terminal through Bluetooth.
[0130] Fifth step: The control terminal receives multi-channel control signals from the foot control panel and the teleoperation glove, performs fusion processing, and transmits the instructions to the robot end in real time through the server. The robot end analyzes the control instructions to realize the overall movement of the robot (forward, backward, rotation) and the fine operation of the dexterous hand. At the same time, based on the wrist IMU (Inertial Measurement Unit) data, the robot end realizes more accurate arm motion mapping through the upper limb motion inverse solution model to ensure "human motion, robot motion".
[0131] Sixth step: When the robot dexterous hand contacts an external object, the built-in sensor collects real-time tactile information, i.e. force feedback information, and at the same time collects the surface temperature information of the object, and transmits them to the control terminal. The control terminal visualizes the tactile feedback information and temperature information, and transmits them to the feedback module of the teleoperation glove through Bluetooth, so that the glove produces corresponding simulation feedback, enabling the operator to obtain a real interactive experience.
[0132] Seventh step: After the task is completed, the user disconnects the system connection, turns off the robot, removes the VR support and teleoperation glove, and completes the teleoperation process.
[0133] The operation flow of the platform when accepting tasks is as follows Figure 2 The detailed description is divided into the following five stages:
[0134] 1. Task initiation stage: the user initiates a task request on the platform through the mobile phone and inputs specific task requirements (such as cleaning the floor and arranging items). The platform quickly processes the request and assigns the task to the operator. The task information is transmitted to the background storage through the wireless network.
[0135] 2. Operation preparation stage: after receiving the task notification, the operator wears the visual and tactile feedback device and establishes a Bluetooth connection with the visual and tactile feedback device and the treadle control base through the control terminal.
[0136] 3. Real-time control stage: the operator observes the environment at the robot end in real time through the visual feedback and controls the movement by using the treadle control. The treadle control device is inspired by the working principle of traditional remote sensing devices and uses a spherical linkage trigger structure as the core sensing method. The operator applies pressure in the forward, backward, left, and right directions on the treadle with the feet, causing the sphere to tilt and press the corresponding direction button, achieving linear translation control of the robot in the up-down and left-right directions. When the operator applies a rotating force to the treadle (such as rotating the feet along a circular arc), the sphere rotates around the central axis, and the rotating sensing structure at the bottom captures the action and generates a direction angle control signal, achieving rotation control of the robot. This design makes translation and rotation control intuitive and smooth, responsive and stable.
[0137] 4. Feedback cycle stage: the robot transmits real-time environmental data and interaction information, and the operator perceives the environment through visual and tactile feedback to ensure smooth task progress. The feedback process is fast and efficient, ensuring the continuity and accuracy of task execution.
[0138] 5. Task completion stage: after completing the task, the operator ends the task, disconnects the platform connection, and records the operation data for subsequent reference or optimization.
[0139] The above is only the preferred embodiment of the present application. It should be noted that for ordinary skilled persons in the technical field, without departing from the technical principles of the present application, several improvements and modifications can be made, which should also be considered as the protection scope of the present application.
Claims
1. A wireless remote control method based on multimodal feedback, characterized in that, The method applied to a cloud server comprises: fusing real-time images, target detection information, target tracking information, voxel information and early warning information collected by a robot at a control terminal to generate a VR real-time picture; real-time analyzing head data, hand data and chassis motion data for controlling the robot according to a user-side ordering task and the VR real-time picture; determining an upper limb motion trajectory of the robot through the hand data; controlling the robot to perform a control action of the ordering task according to the upper limb motion trajectory, the head data, the hand data and the chassis motion data, wherein the upper limb motion trajectory is used to control joint motion and posture of a robot arm, the hand data is used to realize grasping, rotating and placing operation actions of a robot hand, the head data is used to control rotation and observation angle of a robot head, and the chassis motion data is used to control forward movement, backward movement and steering of a robot chassis, thereby realizing accurate control of the ordering task; the real-time images, the target detection information, the target tracking information, the voxel information and the early warning information collected by the robot comprise: collecting real-time images of an environment where the robot is currently located; preprocessing the real-time images to obtain image data after preprocessing; inputting the image data after preprocessing into a visual neural network model to obtain the target detection information, the target tracking information, the voxel information and the early warning information; the inputting the image data after preprocessing into the visual neural network model to obtain the target detection information, the target tracking information, the voxel information and the early warning information comprises: extracting multi-scale feature structures from the RGB image and the depth image after preprocessing; performing spatial alignment and weighted fusion of RGB features and depth features corresponding to the RGB features of the extracted multi-scale feature structures to generate fusion features; identifying a candidate region in the fusion features, outputting a bounding box coordinate, a class label and a confidence score, and dynamically adjusting a detection threshold according to brightness of the candidate region and texture complexity of the candidate region; wherein the brightness is obtained by converting the RGB image into an HSV color space to extract a brightness component and calculating a brightness mean value of the candidate region, and the texture complexity is obtained by performing an edge detection operator on the RGB image to obtain an edge binary image and counting a proportion of edge pixel points in a total number of pixel points of the candidate region; performing pixel-level probability prediction on the candidate region and obtaining a binary mask through threshold segmentation, labeling a foreground region to obtain the target detection information; obtaining target tracking information based on detection information output by target detection in combination with a tracking state of a historical frame; generating depth information based on the fusion features, converting the depth information into a voxel block to generate voxel information, wherein the voxel information is used to present a three-dimensional spatial structure; calculating a three-dimensional position and a minimum safety distance of an obstacle using the voxel information in combination with a current pose of the robot to obtain early warning information.
2. The multi-modal feedback based wireless tele-manipulation method of claim 1, wherein, The method further comprises: determining a cost of the ordering task; determining a settlement cost of an operation end from the cost of the ordering task according to an operation time and a difficulty coefficient of the ordering task.
3. The multi-modal feedback based wireless tele-manipulation method of claim 1, wherein, The real-time image is preprocessed to obtain preprocessed image data, including: Performing frame-by-frame slicing operation on the real-time image to decompose the continuous video stream into a sequence of time-ordered single-frame images to obtain RGB images and depth images; The RGB images are sequentially subjected to color balance processing based on the gray world assumption, then subjected to bilateral filtering to suppress noise and retain edge information, and finally subjected to resolution adjustment and pixel value normalization; The depth images are subjected to depth denoising and smoothing, then subjected to interpolation scaling and spatial alignment with the RGB frames at the pixel level, and subjected to normalization to obtain the preprocessed image data.
4. The multi-modal feedback based wireless tele-manipulation method of claim 1, wherein, The detection threshold is dynamically adjusted according to the brightness of the candidate region and the texture complexity of the candidate region, including: When the brightness mean value is less than a first preset value and the edge pixel ratio is less than a first preset percentage, the confidence threshold is set to a first target value; When the brightness mean value is greater than a second preset value and the edge pixel ratio is greater than a second preset percentage, the confidence threshold is set to a second target value; When the brightness mean value is greater than the first preset value and less than the second preset value, and the edge pixel ratio is greater than the first preset percentage and less than the second preset percentage, the confidence threshold is set to a third target value.
5. The multi-modal feedback based wireless tele-manipulation method of claim 1, wherein, The upper limb motion trajectory is inversely solved from the hand data, including: Based on the three-axis acceleration and three-axis angular velocity data captured by the wrist IMU of the operation end, and the bending state and applied pressure change of the finger joints; The three-axis acceleration, three-axis angular velocity data, bending state and applied pressure change of the finger joints are input into the upper limb motion inverse solution model to output the upper limb motion trajectory, wherein the upper limb motion trajectory includes the joint angles of elbow flexion and extension, forearm rotation, shoulder internal and / or external rotation, and lifting freedom.
6. A wireless remote control system based on multimodal feedback, characterized in that, It includes: A cloud server configured to receive a user order task, send the order task to an operation end, determine a cost of the order task, and determine a settlement cost of the operation end from the cost of the order task according to an operation time and a difficulty coefficient of the order task; A visual processing module configured to receive real-time images of an environment in which a robot is currently located, obtain target detection information, target tracking information, voxel information, and early warning information from the real-time images, and send the real-time images, the target detection information, the target tracking information, the voxel information, and the early warning information to a control terminal to generate a VR real-time picture after fusion in the control terminal; An operation end configured to analyze head data, hand data, and chassis motion control data for controlling the robot in real time based on the order task and the VR real-time picture received by the control terminal, and send the head data, the hand data, and the chassis motion control data to the control terminal; A control terminal configured to receive the VR real-time picture, and send the hand motion data, the chassis motion data, and the head data simulated by a gyroscope built-in in the control terminal to the robot; A robot configured to perform a control action of the order task according to the head data, the hand motion data, and the chassis motion data received from the control terminal.
7. The multi-modal feedback based wireless tele-manipulation system of claim 6, wherein, The operation end includes a pedal control disc, a VR support, and a remote control glove. The pedal control disc is used for identifying the pressing actions in front, rear, left and right directions, and detecting the rotating action of the feet, so as to realize the control of the direction angle of the robot. The VR support triggers the posture information generated by the built-in gyroscope of the control terminal, and is used for synchronously controlling the rotation of the head of the robot. The remote control gloves are used for realizing the cooperative control of the hand action and the upper limb movement track of the robot, and feeding back the pressure of the hand of the robot.
8. The multi-modal feedback based wireless tele-manipulation system of claim 6, wherein, The visual processing module is also used for collecting the temperature of the target object, and sending the collected temperature of the target object to the control terminal, so that the control terminal displays the temperature in the VR picture.
Citation Information
Patent Citations
Immersive robot teleoperation method and system
CN120578299A