Humanoid robot teleoperation method based on multi-modal data fusion and related devices

By using multimodal data fusion and spatiotemporal alignment processing, control and correction commands are generated, solving the problem of response deviation in teleoperation and improving the accuracy, real-time performance, and stability of teleoperation.

CN121132639BActive Publication Date: 2026-08-25广州里工实业有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511257571.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-08-25
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing teleoperation technologies often rely on single or a few modal data, resulting in response bias, low operational accuracy, and insufficient robustness. Furthermore, the fusion of multimodal data does not fully utilize complementarity, thus failing to effectively improve operational stability.

Method used

The robot collects the operator's brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals through a multimodal acquisition device. After data preprocessing and spatiotemporal alignment, the data is input into a multimodal fusion model to generate control commands. When a response deviation is detected, a correction command is generated to correct the robot's movements.

Benefits of technology

It improves the accuracy and real-time performance of teleoperation, solves the problem of time asynchrony of multimodal data through spatiotemporal alignment processing, and enhances the stability of teleoperation and the robot's autonomous learning ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121132639B_ABST
    Figure CN121132639B_ABST
Patent Text Reader

Abstract

The application discloses a human-shaped robot remote operation method based on multi-modal data fusion and related equipment, and the method comprises the following steps: collecting original multi-modal data of an operator through a multi-modal acquisition device; performing data preprocessing on the original multi-modal data through a controller to obtain candidate multi-modal data, performing spatio-temporal alignment processing on the candidate multi-modal data through the controller to obtain target multi-modal data; inputting the target multi-modal data into a multi-modal fusion model through the controller to output fusion features; generating a control instruction for a human-shaped robot based on the fusion features through the controller, so that the human-shaped robot executes remote operation according to the control instruction; if a response deviation occurs in the remote operation, generating a correction instruction based on the fusion features through the controller, so that the human-shaped robot executes the correction instruction to correct the motion trajectory of the remote operation. The application can improve the accuracy, real-time performance and stability of remote operation, and can be widely applied to the technical field of robot control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and in particular to a method and related equipment for teleoperating humanoid robots based on multimodal data fusion. Background Technology

[0002] With the development of humanoid robot technology, teleoperation has become a research hotspot. Related teleoperation technologies often rely on single or limited modal data (such as vision or motion capture), resulting in problems such as response bias, low operational accuracy, and insufficient robustness. For example, teleoperation based solely on visual signals is easily affected by ambient lighting, and systems relying solely on motion capture struggle to reflect the details of the operator's intentions. During teleoperation, due to environmental interference, differences in equipment performance, and other factors, there is often a response bias between the robot's actual actions and the operator's intentions. If this bias is not corrected in time, it can lead to operational errors. Related technologies often address this problem by using simple compensation with single or limited modal data. Furthermore, when considering the fusion of multimodal data, they often involve simple stitching, which fails to effectively improve operational stability under scenarios with response bias.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0004] The embodiments of this application aim to at least partially address one of the technical problems in the related art. Therefore, the main objective of the embodiments of this application is to propose a teleoperation method and related equipment for humanoid robots based on multimodal data fusion, which can improve the accuracy, real-time performance, and stability of teleoperation.

[0005] To achieve the above objectives, one aspect of this application proposes a teleoperation method for a humanoid robot based on multimodal data fusion, the method comprising the following steps: The operator's raw multimodal data is collected by a multimodal acquisition device and sent to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals and motion signals; The controller performs data preprocessing on the original multimodal data to obtain candidate multimodal data, and then performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into the multimodal fusion model and outputs fused features. The controller generates control commands for the humanoid robot based on the fused features and sends the control commands to the humanoid robot, so that the humanoid robot can perform teleoperation according to the control commands. If the controller detects a response deviation in the teleoperation of the humanoid robot, the controller generates a correction instruction based on the fused features and sends the correction instruction to the humanoid robot. The correction instruction includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction instruction to correct the teleoperation trajectory.

[0006] In some embodiments, the method further includes: The controller obtains feedback data corresponding to the humanoid robot executing the correction command; The humanoid robot's motion model is trained by the controller in combination with the target multimodal data and the feedback data. When the motion model meets the model convergence condition, the trained motion model is obtained. When the smoothness score of the autonomous operation of the trained action model is greater than the preset score threshold, the system switches to model control. When the smoothness score or autonomous operation response deviation of the trained action model is less than or equal to the preset score threshold, the system switches to multimodal teleoperation control.

[0007] In some embodiments, the step of performing spatiotemporal alignment processing on the candidate multimodal data through the controller to obtain time-synchronized target multimodal data includes: The controller obtains the acquisition timestamp corresponding to each candidate modality data in the candidate multimodal data; The controller calibrates the acquisition timestamps corresponding to different types of candidate modal data based on a preset high-precision time reference. The controller performs interpolation processing on the calibrated candidate multimodal data to obtain the time-synchronized target multimodal data.

[0008] In some embodiments, the multimodal fusion model includes a feature extraction layer, a dynamic weighting layer, and a fusion output layer. The step of inputting the target multimodal data into the multimodal fusion model via the controller and outputting fused features includes: The controller inputs the target multimodal data into the multimodal fusion model, and the feature extraction layer extracts the target modal features corresponding to each target modal data in the target multimodal data; wherein, the target modal features include operation intention features and environmental interaction features; Through the dynamic weighting layer, corresponding weights are assigned to each target modal feature based on the signal-to-noise ratio of each target modal feature and the correlation between each target modal feature and the current operation task; The fusion output layer performs weighted fusion of the target modal features based on the weights of each target modal feature, and outputs the fused feature.

[0009] In some embodiments, generating control commands for the humanoid robot based on the fused features via the controller includes: The controller inputs the fused features into a preset command to generate a model and outputs joint motion control parameters. The joint motion control parameters satisfy the joint motion range constraints of the humanoid robot. The joint motion control parameters include joint angle, joint velocity, and joint acceleration. The joint velocity is obtained by smoothing and correcting the motion trend of the visual signal in the fused features. The preset instruction generation model generates control instructions for the humanoid robot based on the joint motion control parameters.

[0010] In some embodiments, if the controller detects a response deviation in the teleoperation of the humanoid robot, the controller generates a correction instruction based on the fused features, including: The controller dynamically outputs the current operational intent features and the robot's actual action features through the multimodal fusion model. The controller calculates the degree of matching between the current operational intent features and the actual action features of the robot. If the matching degree is lower than the preset matching threshold, it is determined that the teleoperation of the humanoid robot has a response deviation; When the teleoperation of the humanoid robot results in a response deviation, the controller generates the correction command based on the fused features.

[0011] To achieve the above objectives, another aspect of this application proposes a humanoid robot teleoperation device based on multimodal data fusion, the device comprising the following modules: The data acquisition module is used to acquire the operator's raw multimodal data through a multimodal acquisition device and send the raw multimodal data to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals; The spatiotemporal alignment module is used to preprocess the original multimodal data through the controller to obtain candidate multimodal data, and to perform spatiotemporal alignment processing on the candidate multimodal data through the controller to obtain time-synchronized target multimodal data. The feature fusion module is used to input the target multimodal data into the multimodal fusion model through the controller and output fused features; The instruction generation module is used to generate control instructions for the humanoid robot based on the fusion features through the controller, and send the control instructions to the humanoid robot so that the humanoid robot can perform teleoperation according to the control instructions; The deviation correction module is used to generate a correction instruction based on the fusion features if the controller detects a response deviation in the teleoperation of the humanoid robot, and send the correction instruction to the humanoid robot. The correction instruction includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction instruction to correct the teleoperation trajectory.

[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0014] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0015] The embodiments of this application include at least the following beneficial effects: This application provides a teleoperation method and related equipment for humanoid robots based on multimodal data fusion. This method collects the operator's raw multimodal data using a multimodal acquisition device and sends the raw multimodal data to a controller. The raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals. The controller preprocesses the raw multimodal data to obtain candidate multimodal data, and performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into a multimodal fusion model and outputs fusion features. The controller generates control commands for the humanoid robot based on the fusion features and sends the control commands to the humanoid robot, enabling the humanoid robot to perform teleoperation according to the control commands. If the controller detects a response deviation in the humanoid robot's teleoperation, it generates correction commands based on the fusion features and sends the correction commands to the humanoid robot. The correction commands include correction trajectory information and differentiated operation guidance information, enabling the humanoid robot to execute the correction commands to correct the teleoperation trajectory. This application's embodiments, by integrating multimodal data such as brain modality, electromyography, tactile sensation, vision, and motion, fully utilize the advantages of each modality to improve the accuracy and real-time performance of teleoperation; by using spatiotemporal alignment processing, the asynchronous nature of multimodal data is resolved, ensuring data alignment in the time dimension; when a response deviation is detected in the teleoperation of the humanoid robot, by generating corrected trajectory information and differentiated operation guidance information, the operator can understand the robot's action status in a timely manner, thereby further improving the stability of teleoperation. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the steps of a teleoperation method for a humanoid robot based on multimodal data fusion provided in an embodiment of this application. Figure 2 This is a schematic diagram of the humanoid robot teleoperation system architecture based on multimodal data fusion provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the teleoperation method for a humanoid robot based on multimodal data fusion provided in an embodiment of this application. Figure 4 This is a schematic diagram of the multimodal fusion model structure provided in the embodiments of this application; Figure 5 This is a visual interface diagram of the response deviation scenario provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the humanoid robot teleoperation device based on multimodal data fusion provided in the embodiments of this application; Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application.

[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] With the development of humanoid robot technology, teleoperation and autonomous training have become research hotspots. Related teleoperation technologies often rely on single or limited modal data (such as vision or motion capture), resulting in problems such as response bias, low operational accuracy, and insufficient robustness. For example, teleoperation based solely on visual signals is easily affected by ambient lighting, and systems relying solely on motion capture struggle to reflect the details of the operator's intentions.

[0022] During teleoperation, due to factors such as environmental interference and differences in equipment performance, there is often a response deviation between the robot's actual actions and the operator's intentions. If this deviation is not corrected in time, it may lead to operational errors. Related technologies often address this issue by using simple compensation with single-modal data, failing to fully utilize the complementarity of multimodal data. Furthermore, the fusion of multimodal data is often a simple stitching process, neglecting the task relevance and signal-to-noise ratio differences among the various modalities, thus failing to effectively improve operational stability under response deviation scenarios.

[0023] In robot training, traditional methods often use preset trajectories or single-modal feedback, resulting in poor adaptability of the trained models, making it difficult to cope with complex dynamic scenarios.

[0024] In view of this, embodiments of this application provide a method and related equipment for teleoperating a humanoid robot based on multimodal data fusion. This method collects the operator's raw multimodal data using a multimodal acquisition device and sends the raw multimodal data to a controller. The raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals. The controller preprocesses the raw multimodal data to obtain candidate multimodal data, and performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into a multimodal fusion model and outputs fusion features. The controller generates control commands for the humanoid robot based on the fusion features and sends the control commands to the humanoid robot, enabling the humanoid robot to perform teleoperation according to the control commands. If the controller detects a response deviation in the humanoid robot's teleoperation, it generates correction commands based on the fusion features and sends the correction commands to the humanoid robot. The correction commands include correction trajectory information and differentiated operation guidance information, enabling the humanoid robot to execute the correction commands to correct the teleoperation trajectory. This application's embodiments, by integrating multimodal data such as brain modality, electromyography, tactile sensation, vision, and motion, fully utilize the advantages of each modality to improve the accuracy and real-time performance of teleoperation; by using spatiotemporal alignment processing, the asynchronous nature of multimodal data is resolved, ensuring data alignment in the time dimension; when a response deviation is detected in the teleoperation of the humanoid robot, by generating corrected trajectory information and differentiated operation guidance information, the operator can understand the robot's action status in a timely manner, thereby further improving the stability of teleoperation.

[0025] The humanoid robot teleoperation method based on multimodal data fusion provided in this application relates to the field of robot control technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited thereto. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the humanoid robot teleoperation method based on multimodal data fusion, but is not limited to the above forms.

[0026] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0027] Please see Figure 1 , Figure 1 This is an optional flowchart of a humanoid robot teleoperation method based on multimodal data fusion provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.

[0028] Step S101: Collect the operator's raw multimodal data through a multimodal acquisition device and send the raw multimodal data to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals; Optionally, the raw multimodal data includes brain modal signals, electromyographic signals, tactile signals, visual signals, and motion signals.

[0029] In this embodiment, the operator's brain modality signals, electromyographic signals, tactile signals, visual signals and motion signals are collected by a multimodal acquisition device (including EEG device, EMG sensor, force feedback glove, binocular camera, IMU), and the collected multimodal data is sent to the controller via 5G communication.

[0030] In practice, the specific content and process of multimodal data acquisition are as follows: (1) Brain modal signals: The operator's brain signals were collected using an EEG (electroencephalography) device at a sampling frequency of 250 Hz, reflecting the operator's motor intentions (such as the intention to "grab" or "move"). (2) Electromyographic signals: Electromyographic signals are collected by attaching an electromyographic sensor to the operator's arm muscle activity signals (sampling frequency 1000Hz) to reflect the intensity of muscle activity and movement trend; (3) Tactile signals: The operator wears force feedback gloves to collect pressure and friction signals of the hand in contact with the object (sampling frequency 500Hz), which are sensory data of the operator's interaction with the environment; (4) Visual signals: Images of the operator's limb movements are captured using a binocular camera (30fps) to extract the operator's posture features; (5) Action signal: The operator's joint angle, speed (sampling frequency 100Hz) and other data are collected by the inertial measurement unit (IMU) to directly reflect the operator's action state.

[0031] Step S102: The controller performs data preprocessing on the original multimodal data to obtain candidate multimodal data, and performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. In some embodiments, the step of performing spatiotemporal alignment processing on candidate multimodal data by a controller to obtain time-synchronized target multimodal data may include: obtaining the acquisition timestamps corresponding to each candidate modal data in the candidate multimodal data by a controller; calibrating the acquisition timestamps corresponding to different types of candidate modal data by a controller based on a preset high-precision time reference; and interpolating the calibrated candidate multimodal data by a controller to obtain time-synchronized target multimodal data.

[0032] After collecting multimodal data (before performing spatiotemporal alignment processing), this embodiment of the application will preprocess the multimodal data through the data processing unit in the controller to obtain candidate multimodal data. The preprocessing step aims to improve the quality of the original data and provide reliable input for subsequent spatiotemporal alignment and feature fusion. The preprocessing steps specifically include: (1) Noise removal: adopting a dedicated denoising method for the characteristics of different modal data. For example: wavelet transform is used to remove power frequency interference (50Hz / 60Hz) and electromyography artifacts from EEG signals; low-frequency motion artifacts and high-frequency noise are filtered through a bandpass filter (20-500Hz) for EMG signals; and Gaussian filtering is used to remove image noise from visual signals and image enhancement is used to improve the clarity of limb contours; (2) Data standardization: converting each modal data to a uniform level. For example: normalizing physiological signals such as EEG and EMG (mapping to the [-1,1] interval); and normalizing the joint angle (unit: degrees) and velocity (unit: degrees / ) of motion signals. (2) Perform order-of-magnitude calibration to avoid the impact of unit differences on subsequent feature extraction; (3) Outlier handling: Detect and remove outliers in each modal data (such as jump values ​​caused by instantaneous equipment failure) through the 3σ criterion, and fill them with adjacent valid data interpolation to ensure data continuity.

[0033] After preprocessing the multimodal data, spatiotemporal alignment is performed on the candidate multimodal data obtained from the preprocessing to obtain the target multimodal data for time synchronization. The specific implementation process of spatiotemporal alignment is as follows: First, extract the timestamps of each modality (such as the timestamps of the EEG device). electromyography (EMG) sensors (etc.); then, using a high-precision clock as a reference, each timestamp is calibrated. The calibration calculation formula is: ,in, For the calibrated timestamp, This is the original timestamp. The time deviation of this mode is measured in advance through synchronous experiments; finally, linear interpolation is performed on the calibrated data, and the sampling frequency is uniformly set to 100Hz to ensure that the sampling frequency of each mode data is consistent in the time dimension, thus ensuring data alignment in the time dimension.

[0034] Step S103: Input the target multimodal data into the multimodal fusion model through the controller, and output the fused features; Optionally, the multimodal fusion model includes a feature extraction layer, a dynamic weighting layer, and a fusion output layer.

[0035] The multimodal fusion model is a functional component of the data processing unit within the controller. The controller's data processing unit relies on algorithmic models to perform deep processing of multimodal data. The multimodal fusion model (including a feature extraction layer, a dynamic weighting layer, and a fusion output layer), as the core algorithm set for converting "target multimodal data into fused features," is deployed within the data processing unit, utilizing its computing resources to perform feature extraction, dynamic weight adjustment, and fusion calculation. In essence, the data processing unit represents the hardware and basic processing framework, while the multimodal fusion model is the algorithmic logic running on the data processing unit; the two work together to complete the data fusion processing task in this stage.

[0036] In some embodiments, step S103 may include: inputting target multimodal data into a multimodal fusion model via a controller; extracting target modal features corresponding to each target modal data in the target multimodal data via a feature extraction layer; wherein, the target modal features include operation intention features and environmental interaction features; assigning corresponding weights to each target modal feature based on the signal-to-noise ratio of each target modal feature and the correlation between each target modal feature and the current operation task via a dynamic weighting layer; and performing weighted fusion of each target modal feature based on the weights of each target modal feature via a fusion output layer to output fused features.

[0037] In practice, the feature extraction layer, dynamic weight layer, and fusion output layer in the multimodal fusion model are implemented as follows: (1) Feature extraction layer, used to extract the operational intention features and environmental interaction features of time-synchronized multimodal data. The extracted features include the following, among which the intention tendency features of EEG signals and the movement force features of EMG signals are more representative: 1) Electroencephalogram (EEG) signals: Features were extracted using a convolutional neural network (CNN) to obtain EEG features with a dimension of 128. (Including intentional tendencies); 2) Electromyographic signals: Time-domain and frequency-domain features were extracted using wavelet transform to obtain 64-dimensional electromyographic features. (Including characteristics of the force of the movement); 3) Tactile signals: Environmental interaction features such as pressure intensity and friction distribution are extracted through a fully connected layer to obtain 32-dimensional tactile features. ; 4) Visual signals: Movement features such as limb posture and motion trajectory are extracted using a skeleton extraction algorithm to obtain 24-dimensional posture features. It is used to capture the operator's limb spatial position and movement trends; 5) Motion signals: Dynamic features such as joint angles and motion speed are directly extracted from IMU data to obtain 16-dimensional motion features. It directly reflects the details of the operator's actions.

[0038] (2) Dynamic weighting layer: Based on the signal-to-noise ratio (data quality) of each modality data and the correlation between each modality feature and the current operation task (feature effectiveness), corresponding weights are assigned to each modality feature, increasing the weight ratio of EEG signals and tactile signals in response deviation scenarios. Signal-to-noise ratio (SNR) is an inherent property of multimodal data (such as raw EEG and EMG signals). It represents factors like the power frequency interference intensity of EEG signals and the proportion of motion artifacts in EMG signals. It is a quantifiable indicator during the data acquisition and preprocessing stages (e.g., calculated through signal power spectrum analysis). The feature extraction layer extracts features based on the raw data, while the dynamic weighting layer adjusts weights by combining the SNR of the raw data (reflecting data quality) and the relevance of the features to the task (reflecting feature effectiveness). Both serve as the basis for weight adjustment.

[0039] For the current task, it refers to the task target information extracted in real time from the operator's multimodal data (after feature extraction). For example: the intention tendency features of EEG signals (such as the intention of "grabbing" or "moving") directly indicate the task type; the limb posture features of visual signals (such as the trajectory of the hand reaching towards the object) and the joint movement features of motion signals (such as the tendency of arm flexion and extension) help to clarify the task details; the multimodal fusion model integrates the above single-modal features through the feature extraction layer and parses the current task (such as "grabbing a specific object" or "moving to the target location") in real time.

[0040] First, calculate the task relevance score for each feature. Task relevance score The specific calculation formula is as follows:

[0041] in, For the Sigmoid function, , For learnable parameters, The features of the i-th mode are given.

[0042] Then, based on task relevance scoring Calculate the weights using the following formula:

[0043] in, Let be the weight of the i-th mode, and n=5 be the total number of modes. The higher the relevance score, the greater the corresponding weight.

[0044] In a normal scenario, the weights of each modality are as follows: , , , , .

[0045] In the response bias scenario (matching degree < 0.7), the weights of various modalities are as follows: , , , , (Increase the weight of the intention-environment interaction modality).

[0046] Specifically, in scenarios where teleoperation response deviation is detected (i.e., the matching degree between the robot's actual actions and the operator's intention is lower than a preset threshold), the weight ratio of EEG signals and tactile signals in the fusion features is dynamically increased. For example, in normal scenarios, the weight of EEG signals is 0.2 and that of tactile signals is 0.1; in response deviation scenarios, their weights are increased to 0.35 and 0.15, respectively. This is because EEG signals directly reflect the operator's original intention (unaffected by action execution errors), while tactile signals provide real-time feedback on the operator's interaction with the environment (such as whether grip strength is sufficient). Increasing their weights enhances the fusion features' ability to represent "true intention" and "environmental interaction," thereby more efficiently correcting deviations.

[0047] (3) Fusion output layer: fusion features The calculation formula is as follows:

[0048] Step S104: The controller generates control instructions for the humanoid robot based on the fusion features and sends the control instructions to the humanoid robot so that the humanoid robot can perform teleoperation according to the control instructions; In some embodiments, the step of generating control commands for a humanoid robot based on fused features by a controller may include: inputting the fused features into a preset command generation model through the controller, and outputting joint motion control parameters; wherein the joint motion control parameters satisfy the joint motion range constraints of the humanoid robot, and the joint motion control parameters include joint angle, joint velocity, and joint acceleration, and the joint velocity is obtained by smoothing and correcting the motion trend of the visual signal in the fused features; and generating control commands for the humanoid robot based on the joint motion control parameters through the preset command generation model.

[0049] The instruction generation model belongs to the controller's data processing unit. On one hand, the controller's data processing unit is the core carrier for executing complex algorithms and model calculations. The instruction generation model relies on its computing resources to analyze and compute the fused features (such as deriving joint motion control parameters based on the fused features and then generating control commands). On the other hand, after preprocessing, spatiotemporal alignment, and fusion, multimodal data is integrated into the instruction generation model within the same data processing unit. This avoids delays and errors in cross-module transmission, ensuring the efficiency and accuracy of control command generation. Therefore, the instruction generation model runs within the controller's data processing unit, serving as the algorithm module for "converting fused features into control commands," collaboratively completing the control logic for the humanoid robot.

[0050] In the specific implementation, the first step is to fuse the features. Input instructions to generate a model (such as an LSTM network), and output the angles of each joint of the robot. The parameters are velocity v and acceleration a, where the velocity parameter is smoothed based on the motion trend of the visual signal (the upper limit of velocity is reduced when the trend changes abruptly). Then, standardized, multi-system compatible control commands are generated based on these parameters, such as the shoulder joint angle when grasping an object. =30°, elbow joint speed .

[0051] Step S105: If the controller detects a response deviation in the teleoperation of the humanoid robot, the controller generates a correction instruction based on the fusion features and sends the correction instruction to the humanoid robot. The correction instruction includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction instruction to correct the teleoperation trajectory.

[0052] In some embodiments, if the controller detects a response deviation in the teleoperation of the humanoid robot, the step of generating a correction instruction based on the fusion features by the controller may include: dynamically outputting the current operation intention features and the robot's actual action features through a multimodal fusion model in the controller; calculating the matching degree between the current operation intention features and the robot's actual action features by the controller; if the matching degree is lower than a preset matching threshold, determining that the teleoperation of the humanoid robot has a response deviation; and generating a correction instruction based on the fusion features by the controller when the teleoperation of the humanoid robot has a response deviation.

[0053] The specific implementation process for calculating the matching degree between the current operational intent features and the robot's actual action features through the controller is as follows: (1) Eigenvector preparation: 1) Operational Intent Feature Vector: The multimodal fusion model extracts key information reflecting the operator's operational intent from multimodal data such as EEG signals, EMG signals, tactile signals, visual signals, and motion signals, and integrates them into a high-dimensional feature vector. For example, intentional tendency features in EEG signals (such as EEG pattern encoding corresponding to thoughts like "grab" and "put down") and action force features in EMG signals (such as numerical values ​​corresponding to the degree of muscle exertion) are processed by the feature extraction layer to form the operational intent feature vector; 2) Robot's Actual Motion Feature Vector: Various sensors equipped on the robot (such as joint angle sensors, velocity sensors, etc.) collect real-time data on the robot's actual motion during operation (such as the angles, velocities, accelerations, etc. of each joint). This data is processed and encoded to form the robot's actual motion feature vector.

[0054] (2) Cosine similarity calculation of matching degree: Cosine similarity measures the degree of similarity between two vectors by calculating the cosine value of the angle between them. The formula for calculating cosine similarity is:

[0055] in, The feature vector representing the operational intent. This represents the feature vector of the robot's actual actions. It is the dot product of two vectors. and These are the magnitudes of the two vectors, respectively.

[0056] When two vectors are in the same direction, the angle between them is . A cosine value of 1 means that the robot's actual action perfectly matches the operator's intention, i.e., the matching degree is 1; when the two vectors are in completely opposite directions, the angle between them is... A cosine value of -1 indicates that the robot's actual actions are completely contrary to the operator's intentions.

[0057] In practical applications, the matching degree ranges from [-1, 1]. The closer the value is to 1, the higher the degree of matching between the robot's actual actions and the operator's intentions; the closer the value is to -1, the lower the degree of matching. In this embodiment, a preset matching threshold is set (e.g., 0.7, which is determined based on a large number of experiments and actual operation scenarios). When the calculated matching degree is lower than this threshold, it is determined that the teleoperation of the humanoid robot has a response deviation.

[0058] In some embodiments, the method may further include: acquiring feedback data corresponding to the humanoid robot executing correction instructions through a controller; training the humanoid robot's motion model by combining the target multimodal data and the feedback data through the controller; obtaining the trained motion model when the motion model meets the model convergence condition; switching to model control when the motion smoothness score of the trained motion model's autonomous operation is greater than a preset score threshold; and switching to multimodal teleoperation control when the motion smoothness score or autonomous operation response deviation of the trained motion model's autonomous operation is less than or equal to a preset score threshold.

[0059] Autonomous operation response deviation specifically refers to the deviation between the actual actions of a robot and the intended intent / standard trajectory when the robot autonomously performs a task using a trained motion model. Specifically, the robot's actions are monitored when it switches to "model control" (autonomously performing a task using a trained motion model). For example, suppose the robot should smoothly grasp an object, but instead its movements are jerky or its position deviates; this deviation in "the robot autonomously performing a task using a trained motion model" is called "autonomous operation response deviation." The autonomous operation response deviation is used to determine whether it is necessary to switch back to user-controlled multimodal teleoperation.

[0060] In practical implementation, a multimodal fusion model outputs the current operation intention features and the robot's actual action features in real time, calculating the matching degree between the intention features and action features. If the matching degree is lower than a preset threshold, it is determined that there is a response deviation in the teleoperation. Specifically, when a response deviation is detected (matching degree < 0.7): based on the fused features... Generate a correction trajectory (including a continuous adjustment sequence of joint angles and speeds) and generate differentiated operation guidance information: on the display, the correction trajectory is marked with a green dashed line, and the original trajectory is marked with a gray solid line, and the deviation reason prompt is displayed simultaneously (such as "weak tactile signal, please increase grip strength").

[0061] Steps S101 to S105 as illustrated in this embodiment involve: acquiring the operator's raw multimodal data using a multimodal acquisition device and sending the raw multimodal data to a controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals; preprocessing the raw multimodal data using the controller to obtain candidate multimodal data, and then performing spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data; inputting the target multimodal data into a multimodal fusion model using the controller and outputting fusion features; generating control commands for the humanoid robot based on the fusion features and sending the control commands to the humanoid robot to enable the humanoid robot to perform teleoperation according to the control commands; if the controller detects a response deviation in the humanoid robot's teleoperation, generating correction commands based on the fusion features and sending the correction commands to the humanoid robot, the correction commands include correction trajectory information and differentiated operation guidance information to enable the humanoid robot to execute the correction commands to correct the teleoperation trajectory. This application's embodiments, by integrating multimodal data such as brain modality, electromyography, tactile sensation, vision, and motion, fully utilize the advantages of each modality to improve the accuracy and real-time performance of teleoperation; by using spatiotemporal alignment processing, the asynchronous nature of multimodal data is resolved, ensuring data alignment in the time dimension; when a response deviation is detected in the teleoperation of the humanoid robot, by generating corrected trajectory information and differentiated operation guidance information, the operator can understand the robot's action status in a timely manner, thereby further improving the stability of teleoperation.

[0062] Furthermore, by employing a dynamic weighting mechanism to adjust the weights of each modality (especially in response bias scenarios), the effectiveness of the fused features was enhanced; and by combining multimodal data with execution feedback to train the robot model, the robot's autonomous learning ability and environmental adaptability were improved.

[0063] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.

[0064] The embodiments of this application aim to solve the problems in related technologies such as poor robustness of single-modal teleoperation, low efficiency of multimodal data fusion (without considering task relevance and signal-to-noise ratio), insufficient adaptability of robot training, and weak operational stability under response deviation scenarios. This application embodiment collects multimodal data, including brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals from the operator, using a multimodal acquisition device. A spatiotemporal alignment processing mechanism is used to calibrate the timestamps of different modal data (based on a preset high-precision time reference and calibration formula), and interpolation is used to unify the sampling frequency of each modality, resulting in spatiotemporally aligned multimodal data. This solves the problem of time asynchrony in multimodal data and provides high-quality input data for subsequent feature extraction and fusion of the fusion model. The feature extraction layer of the multimodal fusion model extracts the operational intent features and environmental interaction features of the spatiotemporally aligned multimodal data. These features are then fused through the dynamic weight layer and fusion output layer of the multimodal fusion model to generate fused features. Robot control commands are generated based on these fused features to achieve real-time teleoperation of the humanoid robot. When a deviation in the humanoid robot's teleoperation response is detected, the control commands are dynamically corrected based on the fused features, and differentiated operation guidance information is generated. Simultaneously, the multimodal data and feedback data after the robot executes control commands are used to train the robot's motion model to optimize its autonomous operation capabilities. This application embodiment improves the accuracy and real-time performance of teleoperation through multimodal data fusion, enhances operational stability in complex scenarios through dynamic correction mechanisms, and solves the problem of insufficient robustness of single-modal data in complex scenarios. The humanoid robot teleoperation method based on multimodal data fusion provided in this application embodiment is particularly suitable for response deviation correction scenarios in complex environments.

[0065] Please see Figure 2 , Figure 2 This is a schematic diagram of the humanoid robot teleoperation system architecture based on multimodal data fusion provided in an embodiment of this application, such as... Figure 2 As shown, the humanoid robot teleoperation system based on multimodal data fusion mainly comprises three modules: a multimodal acquisition device, a controller, and a humanoid robot. The multimodal acquisition device includes an EEG device, an EMG sensor, a force feedback glove, a camera, and an IMU sensor. This device is used to acquire the operator's brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals. The humanoid robot executes control commands and provides feedback on the execution data. The controller communicates with both the multimodal acquisition device and the humanoid robot via 5G communication. The controller includes a data processing unit and a deviation detection module. The data processing unit's functions include spatiotemporal alignment, feature fusion, command generation, deviation correction (including trajectory optimization and guidance information generation), and model training.

[0066] Please see Figure 3 , Figure 3 This is a flowchart illustrating the teleoperation method for humanoid robots based on multimodal data fusion provided in this application embodiment, as shown below. Figure 3 As shown, the overall implementation process of the humanoid robot teleoperation method based on multimodal data fusion includes: Step 1, Multimodal Data Acquisition: Collect the operator's brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals through multimodal acquisition devices (including EEG devices, EMG sensors, force feedback gloves, binocular cameras, and IMUs), and send the acquired multimodal data to the controller via 5G communication.

[0067] In practice, the specific content and process of multimodal data acquisition are as follows: (1) Brain modal signals: The operator's brain signals were collected using an EEG (electroencephalography) device at a sampling frequency of 250 Hz, reflecting the operator's motor intentions (such as the intention to "grab" or "move"). (2) Electromyographic signals: Electromyographic signals are collected by attaching an electromyographic sensor to the operator's arm muscle activity signals (sampling frequency 1000Hz) to reflect the intensity of muscle activity and movement trend; (3) Tactile signals: The operator wears force feedback gloves to collect pressure and friction signals of the hand in contact with the object (sampling frequency 500Hz), which are sensory data of the operator's interaction with the environment; (4) Visual signals: Images of the operator's limb movements are captured using a binocular camera (30fps) to extract the operator's posture features; (5) Action signal: The operator's joint angle, speed (sampling frequency 100Hz) and other data are collected by the inertial measurement unit (IMU) to directly reflect the operator's action state.

[0068] It should be noted that the robot's data (such as visual feedback from the robot and sensor data of the actions performed) belongs to the subsequent "feedback data" and is only used in combination with the operator's multimodal data when training the robot's motion model.

[0069] After acquiring multimodal data (but before performing the spatiotemporal alignment processing in step two), this embodiment of the application preprocesses the multimodal data, specifically including: (1) Noise Removal: Dedicated noise removal methods are adopted for different modal data. For example: wavelet transform is used to remove power frequency interference (50Hz / 60Hz) and electromyography artifacts from EEG signals; low-frequency motion artifacts and high-frequency noise are filtered out from electromyography signals through bandpass filters (20-500Hz); and Gaussian filtering is used to remove image noise from visual signals and image enhancement is used to improve the clarity of limb contours. (2) Data standardization: Convert the data of each modality to a uniform scale. For example, normalize physiological signals such as EEG and EMG (map to the [-1,1] interval); calibrate the scale of joint angle (unit: degrees) and velocity (unit: degrees / second) of motion signals to avoid the impact of unit differences on subsequent feature extraction. (3) Outlier handling: Outliers in each modal data (such as jump values ​​caused by instantaneous equipment failure) are detected and removed by the 3σ criterion, and adjacent valid data are interpolated to fill in the gaps to ensure data continuity.

[0070] The above preprocessing steps aim to improve the quality of the raw data and provide reliable input for subsequent spatiotemporal alignment and feature fusion.

[0071] Step 2, Spatiotemporal alignment processing: After receiving the multimodal data sent by the multimodal acquisition device, the controller's internal data processing unit performs data preprocessing on the multimodal data, and then performs spatiotemporal alignment processing on the preprocessed multimodal data to obtain time-synchronized multimodal data.

[0072] In the specific implementation, the spatiotemporal alignment process is as follows: First, extract the timestamps of each modality's data (such as the timestamps of EEG devices). electromyography (EMG) sensors (etc.); then, using a high-precision clock as a reference, each timestamp is calibrated. The calibration calculation formula is: ,in, For the calibrated timestamp, This is the original timestamp. The time deviation of this mode is measured in advance through synchronous experiments; finally, linear interpolation is performed on the calibrated data, and the sampling frequency is uniformly set to 100Hz to ensure that the sampling frequency of each mode data is consistent in the time dimension, thus ensuring data alignment in the time dimension.

[0073] Step 3, Feature Fusion: The controller's data processing unit inputs the time-synchronized multimodal data into a preset multimodal fusion model and outputs fused features.

[0074] Please see Figure 4 , Figure 4 This is a schematic diagram of the multimodal fusion model structure provided in the embodiments of this application, as shown below. Figure 4 As shown, the multimodal fusion model includes a feature extraction layer, a dynamic weight layer (including a weight calculation module), and a fusion output layer. In its specific implementation, the feature extraction layer of the multimodal fusion model extracts the operational intent features and environmental interaction features of the time-synchronized multimodal data. Then, the dynamic weight layer and fusion output layer of the multimodal fusion model perform feature fusion to generate fused features.

[0075] In practice, the feature extraction layer, dynamic weight layer, and fusion output layer in the multimodal fusion model are implemented as follows: (1) Feature extraction layer, used to extract the operational intention features and environmental interaction features of time-synchronized multimodal data. The extracted features include the following, among which the intention tendency features of EEG signals and the movement force features of EMG signals are more representative: 1) Electroencephalogram (EEG) signals: Features were extracted using a convolutional neural network (CNN) to obtain EEG features with a dimension of 128. (Including intentional tendencies); 2) Electromyographic signals: Time-domain and frequency-domain features were extracted using wavelet transform to obtain 64-dimensional electromyographic features. (Including characteristics of the force of the movement); 3) Tactile signals: Environmental interaction features such as pressure intensity and friction distribution are extracted through a fully connected layer to obtain 32-dimensional tactile features. ; 4) Visual signals: Movement features such as limb posture and motion trajectory are extracted using a skeleton extraction algorithm to obtain 24-dimensional posture features. It is used to capture the operator's limb spatial position and movement trends; 5) Motion signals: Dynamic features such as joint angles and motion speed are directly extracted from IMU data to obtain 16-dimensional motion features. It directly reflects the details of the operator's actions.

[0076] (2) Dynamic weighting layer: Based on the signal-to-noise ratio (data quality) of each modality data and the correlation between each modality feature and the current operation task (feature effectiveness), corresponding weights are assigned to each modality feature, increasing the weight ratio of EEG signals and tactile signals in response deviation scenarios. First, calculate the task relevance score for each feature. Task relevance score The specific calculation formula is as follows:

[0077] in, For the Sigmoid function, , For learnable parameters, The features of the i-th mode are given.

[0078] Then, based on task relevance scoring Calculate the weights using the following formula:

[0079] in, Let be the weight of the i-th mode, and n=5 be the total number of modes. The higher the relevance score, the greater the corresponding weight.

[0080] In a normal scenario, the weights of each modality are as follows: , , , , .

[0081] In the response bias scenario (matching degree < 0.7), the weights of various modalities are as follows: , , , , (Increase the weight of the intention-environment interaction modality).

[0082] Specifically, in scenarios where teleoperation response deviation is detected (i.e., the matching degree between the robot's actual actions and the operator's intention is lower than a preset threshold), the weight ratio of EEG signals and tactile signals in the fusion features is dynamically increased. For example, in normal scenarios, the weight of EEG signals is 0.2 and that of tactile signals is 0.1; in response deviation scenarios, their weights are increased to 0.35 and 0.15, respectively. This is because EEG signals directly reflect the operator's original intention (unaffected by action execution errors), while tactile signals provide real-time feedback on the operator's interaction with the environment (such as whether grip strength is sufficient). Increasing their weights enhances the fusion features' ability to represent "true intention" and "environmental interaction," thereby more efficiently correcting deviations.

[0083] (3) Fusion output layer: Fusion features The calculation formula is as follows:

[0084] Step 4, Control Command Generation: The controller generates control commands for the humanoid robot based on the fusion features and sends the control commands to the humanoid robot via 5G communication to realize remote operation of the humanoid robot.

[0085] In the specific implementation, firstly, the fused features are input into a preset instruction generation model, and the joint angle, velocity and acceleration parameters are output. The output parameters satisfy the joint motion range constraints of the humanoid robot. Then, control instructions are generated based on the joint angle, velocity and acceleration parameters. Among them, the velocity parameter is smoothed and corrected based on the motion trend of the visual signal in the multimodal fused features.

[0086] Specifically, firstly, the fusion features Input instructions to generate a model (such as an LSTM network), and output the angles of each joint of the robot. The parameters are velocity v and acceleration a, where the velocity parameter is smoothed based on the motion trend of the visual signal (the upper limit of velocity is reduced when the trend changes abruptly). Then, standardized, multi-system compatible control commands are generated based on these parameters, such as the shoulder joint angle when grasping an object. =30°, elbow joint speed .

[0087] Step 5, Response Deviation Correction: When the deviation detection module of the controller detects a deviation in the teleoperation response of the humanoid robot, the data processing unit of the controller dynamically corrects the control command generation strategy based on the fusion features, optimizes the operation trajectory and generates differentiated operation guidance information, and sends the relevant correction commands to the humanoid robot via 5G communication.

[0088] Please see Figure 5 , Figure 5 This is a visual interface diagram of the response deviation scenario provided in the embodiments of this application, such as... Figure 5 As shown in the diagram, the visualization interface for the response deviation scenario displays a comparison between the green dashed line (corrected trajectory) and the gray solid line (original trajectory), along with a prompt box indicating the cause of the deviation.

[0089] In practical implementation, a multimodal fusion model outputs the current operation intention features and the robot's actual action features in real time, calculating the matching degree between the intention features and action features. If the matching degree is lower than a preset threshold, it is determined that there is a response deviation in the teleoperation. Specifically, when a response deviation is detected (matching degree < 0.7): based on the fused features... Generate a correction trajectory (including a continuous adjustment sequence of joint angles and speeds) and generate differentiated operation guidance information: on the display, the correction trajectory is marked with a green dashed line, and the original trajectory is marked with a gray solid line, and the deviation reason prompt is displayed simultaneously (such as "weak tactile signal, please increase grip strength").

[0090] Step 6, Robot Motion Model Training: Feedback data from the humanoid robot's execution of control commands is sent to the controller via 5G communication. The controller's data processing unit then combines the previously received multimodal data with this feedback data to train the robot's motion model.

[0091] In the specific implementation, firstly, a reinforcement learning environment is constructed, with fused features as state input, robot action output as behavior, and completion degree and operation smoothness as reward; then, the robot action model is trained through reinforcement learning algorithm until the model converges; when the robot's autonomous operation action smoothness score is greater than the preset threshold, it switches to model control; when a response deviation is detected or the smoothness score is less than or equal to the preset threshold, it switches back to teleoperation control based on multimodal fusion.

[0092] Specifically, first, construct the reinforcement learning environment: state ,Behavior The reward R is as follows: +10 points for completing the grasping task, +5 points for smoothness of movement (based on the rate of change of acceleration) > 0.8, -5 points for colliding with objects, and -2 points for timeout. Then, a deep Q-network (DQN) is used to train the model: input state S, output Q-value (action value) of each behavior, and update the model parameters by maximizing the cumulative reward. During training, the model is evaluated once every 1000 tasks. When the average reward of 10 consecutive evaluations is > 8 points, the model converges. Finally, the control switches: after the model converges, when the smoothness score of autonomous operation is > 0.8, it switches to model control; when a response deviation is detected or the smoothness score is ≤ 0.8, it switches back to multimodal teleoperation control.

[0093] It should be noted that this embodiment is only a brief illustrative description of the overall process of the humanoid robot teleoperation method based on multimodal data fusion. Detailed descriptions of each step can be found in the relevant content of the foregoing embodiments, and will not be repeated here. It is understood that the present invention does not limit this.

[0094] This application embodiment collects the operator's raw multimodal data using a multimodal acquisition device and sends the raw multimodal data to a controller. The raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals. The controller preprocesses the raw multimodal data to obtain candidate multimodal data, and then performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into a multimodal fusion model and outputs fusion features. The controller generates control commands for the humanoid robot based on the fusion features and sends the control commands to the humanoid robot to enable the humanoid robot to perform teleoperation according to the control commands. If the controller detects a response deviation in the humanoid robot's teleoperation, the controller generates correction commands based on the fusion features and sends the correction commands to the humanoid robot. The correction commands include correction trajectory information and differentiated operation guidance information to enable the humanoid robot to execute the correction commands to correct the teleoperation trajectory. This application's embodiments, by fusing multimodal data such as brain modality, electromyography, tactile sensation, vision, and motion, fully leverage the advantages of each modality to improve the accuracy and real-time performance of teleoperation. Spatiotemporal alignment processing resolves the temporal asynchrony issue of multimodal data, ensuring data alignment in the time dimension. When a response deviation is detected in the humanoid robot's teleoperation, generating corrected trajectory information and differentiated operation guidance information helps the operator understand the robot's action status in a timely manner, thereby further improving the stability of teleoperation. Furthermore, by employing a dynamic weighting mechanism to adjust the weights of each modality (especially in response deviation scenarios), the effectiveness of the fused features is enhanced. By combining multimodal data with execution feedback to train the robot model, the robot's autonomous learning ability and environmental adaptability are improved.

[0095] In summary, the overall implementation flow of the humanoid robot teleoperation method based on multimodal data fusion provided in this application embodiment is as follows: First, collect and preprocess the operator's multimodal data; then, perform spatiotemporal alignment processing on the preprocessed multimodal data to obtain time-synchronized multimodal data; subsequently, input the time-synchronized multimodal data into a preset multimodal fusion model and output fusion features; next, generate control commands for the humanoid robot based on the fusion features to realize teleoperation of the humanoid robot. The method for detecting teleoperation response deviation of the humanoid robot includes: extracting operation intention features and robot actual action features through the multimodal fusion model, calculating the matching degree between the two and comparing it with a preset threshold. The multimodal fusion model adopts a dynamic weighting mechanism, dynamically adjusting the weights according to the task relevance and signal-to-noise ratio of each modality data (such as increasing the weight of EEG signals and tactile signals in response deviation scenarios); when a teleoperation response deviation of the humanoid robot is detected, adjust the control command generation strategy based on the fusion features, optimize the operation trajectory, and generate differentiated operation guidance information; then, obtain the feedback data after the humanoid robot executes the control commands, and train the robot action model by combining the multimodal data and the feedback data.

[0096] This application's embodiments improve the accuracy and real-time performance of teleoperation by fusing multimodal data such as brain modality, electromyography (EMG), tactile feedback, vision, and motion, fully leveraging the advantages of each modality (e.g., EEG reflecting intention, EMG reflecting movement trends, and tactile feedback providing environmental interaction). It addresses the time asynchrony issue of multimodal data through spatiotemporal alignment processing; it enhances the effectiveness of fused features by employing a dynamic weighting mechanism to adjust the weights of each modality (especially in response deviation scenarios); and it improves the robot's autonomous learning ability and environmental adaptability by combining multimodal data with execution feedback to train the robot model. Simultaneously, by generating differentiated operation guidance information, it helps operators understand the robot's action status in a timely manner, further improving operational stability.

[0097] Please see Figure 6 This application also provides a humanoid robot teleoperation device 600 based on multimodal data fusion, which can implement the above-mentioned method. The device includes the following modules: The data acquisition module 601 is used to acquire the operator's raw multimodal data through a multimodal acquisition device and send the raw multimodal data to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals and motion signals; The spatiotemporal alignment module 602 is used to preprocess the original multimodal data through the controller to obtain candidate multimodal data, and to perform spatiotemporal alignment processing on the candidate multimodal data through the controller to obtain time-synchronized target multimodal data. The feature fusion module 603 is used to input the target multimodal data into the multimodal fusion model through the controller and output fused features; The instruction generation module 604 is used to generate control instructions for the humanoid robot based on the fusion features through the controller, and send the control instructions to the humanoid robot so that the humanoid robot can perform teleoperation according to the control instructions; The deviation correction module 605 is used to generate a correction instruction based on the fusion features if the controller detects a response deviation in the teleoperation of the humanoid robot, and send the correction instruction to the humanoid robot. The correction instruction includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction instruction to correct the teleoperation trajectory.

[0098] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0099] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0100] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0101] Please see Figure 7 , Figure 7 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the methods described in the embodiments of this application. The input / output interface 703 is used to implement information input and output; The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704); The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.

[0102] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0103] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0104] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0105] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0106] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0107] The humanoid robot teleoperation method, apparatus, electronic device, storage medium, and program product based on multimodal data fusion provided in this application embodiment acquire the operator's raw multimodal data through a multimodal acquisition device and send the raw multimodal data to a controller. The raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals. The controller preprocesses the raw multimodal data to obtain candidate multimodal data, and performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into a multimodal fusion model and outputs fusion features. The controller generates control commands for the humanoid robot based on the fusion features and sends the control commands to the humanoid robot to enable the humanoid robot to perform teleoperation according to the control commands. If the controller detects a response deviation in the humanoid robot's teleoperation, the controller generates correction commands based on the fusion features and sends the correction commands to the humanoid robot. The correction commands include correction trajectory information and differentiated operation guidance information to enable the humanoid robot to execute the correction commands to correct the teleoperation trajectory. This application's embodiments, by fusing multimodal data such as brain modality, electromyography, tactile sensation, vision, and motion, fully leverage the advantages of each modality to improve the accuracy and real-time performance of teleoperation. Spatiotemporal alignment processing resolves the temporal asynchrony issue of multimodal data, ensuring data alignment in the time dimension. When a response deviation is detected in the humanoid robot's teleoperation, generating corrected trajectory information and differentiated operation guidance information helps the operator understand the robot's action status in a timely manner, thereby further improving the stability of teleoperation. Furthermore, by employing a dynamic weighting mechanism to adjust the weights of each modality (especially in response deviation scenarios), the effectiveness of the fused features is enhanced. By combining multimodal data with execution feedback to train the robot model, the robot's autonomous learning ability and environmental adaptability are improved.

[0108] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0109] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0112] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0113] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0114] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0115] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0116] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0118] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A teleoperation method for humanoid robots based on multimodal data fusion, characterized in that, The method includes the following steps: The operator's raw multimodal data is collected by a multimodal acquisition device and sent to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals and motion signals; The controller performs data preprocessing on the original multimodal data to obtain candidate multimodal data, and then performs spatiotemporal alignment processing on the candidate multimodal data to obtain time-synchronized target multimodal data. The controller inputs the target multimodal data into the multimodal fusion model and outputs fused features. The controller generates control commands for the humanoid robot based on the fused features and sends the control commands to the humanoid robot, so that the humanoid robot can perform teleoperation according to the control commands. If the controller detects a response deviation in the teleoperation of the humanoid robot, it generates a correction instruction based on the fusion features and sends the correction instruction to the humanoid robot. The correction instruction includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction instruction to correct the teleoperation trajectory. When the controller detects a response deviation in the teleoperation of the humanoid robot, it dynamically increases the weight ratio of the brain modality signal and the tactile signal in the fusion features. The multimodal fusion model includes a feature extraction layer, a dynamic weighting layer, and a fusion output layer. The step of inputting the target multimodal data into the multimodal fusion model via the controller and outputting fused features includes: The controller inputs the target multimodal data into the multimodal fusion model, and the feature extraction layer extracts the target modal features corresponding to each target modal data in the target multimodal data; wherein, the target modal features include operation intention features and environmental interaction features; Through the dynamic weighting layer, corresponding weights are assigned to each target modal feature based on the signal-to-noise ratio of each target modal feature and the correlation between each target modal feature and the current operation task; The fusion output layer performs weighted fusion of the target modal features based on the weights of each target modal feature, and outputs the fused feature.

2. The method according to claim 1, characterized in that, The method further includes: The controller obtains feedback data corresponding to the humanoid robot executing the correction command; The humanoid robot's motion model is trained by the controller in combination with the target multimodal data and the feedback data. When the motion model meets the model convergence condition, the trained motion model is obtained. When the smoothness score of the autonomous operation of the trained action model is greater than the preset score threshold, the system switches to model control. When the smoothness score or autonomous operation response deviation of the trained action model is less than or equal to the preset score threshold, the system switches to multimodal teleoperation control.

3. The method according to claim 1, characterized in that, The step of performing spatiotemporal alignment processing on the candidate multimodal data through the controller to obtain time-synchronized target multimodal data includes: The controller obtains the acquisition timestamp corresponding to each candidate modality data in the candidate multimodal data; The controller calibrates the acquisition timestamps corresponding to different types of candidate modal data based on a preset high-precision time reference. The controller performs interpolation processing on the calibrated candidate multimodal data to obtain the time-synchronized target multimodal data.

4. The method according to claim 1, characterized in that, The step of generating control commands for the humanoid robot based on the fused features through the controller includes: The controller inputs the fused features into a preset command to generate a model and outputs joint motion control parameters. The joint motion control parameters satisfy the joint motion range constraints of the humanoid robot. The joint motion control parameters include joint angle, joint velocity, and joint acceleration. The joint velocity is obtained by smoothing and correcting the motion trend of the visual signal in the fused features. The preset instruction generation model generates control instructions for the humanoid robot based on the joint motion control parameters.

5. The method according to claim 1, characterized in that, If the controller detects a response deviation in the teleoperation of the humanoid robot, the controller generates a correction command based on the fused features, including: The controller dynamically outputs the current operational intent features and the robot's actual action features through the multimodal fusion model. The controller calculates the degree of matching between the current operational intent features and the actual action features of the robot. If the matching degree is lower than the preset matching threshold, it is determined that the teleoperation of the humanoid robot has a response deviation; When the teleoperation of the humanoid robot results in a response deviation, the controller generates the correction command based on the fused features.

6. A humanoid robot teleoperation device based on multimodal data fusion, characterized in that, The device includes the following modules: The data acquisition module is used to acquire the operator's raw multimodal data through a multimodal acquisition device and send the raw multimodal data to the controller; the raw multimodal data includes brain modality signals, electromyographic signals, tactile signals, visual signals, and motion signals; The spatiotemporal alignment module is used to preprocess the original multimodal data through the controller to obtain candidate multimodal data, and to perform spatiotemporal alignment processing on the candidate multimodal data through the controller to obtain time-synchronized target multimodal data. The feature fusion module is used to input the target multimodal data into the multimodal fusion model through the controller and output fused features; The instruction generation module is used to generate control instructions for the humanoid robot based on the fusion features through the controller, and send the control instructions to the humanoid robot so that the humanoid robot can perform teleoperation according to the control instructions; A deviation correction module is used to generate a correction command based on the fusion features if the controller detects a response deviation in the teleoperation of the humanoid robot, and send the correction command to the humanoid robot. The correction command includes correction trajectory information and differentiated operation guidance information, so that the humanoid robot executes the correction command to correct the teleoperation trajectory. Specifically, when the controller detects a response deviation in the teleoperation of the humanoid robot, it dynamically increases the weight ratio of the brain modality signal and the tactile signal in the fusion features. The multimodal fusion model includes a feature extraction layer, a dynamic weighting layer, and a fusion output layer. The feature fusion module is specifically used for: The controller inputs the target multimodal data into the multimodal fusion model, and the feature extraction layer extracts the target modal features corresponding to each target modal data in the target multimodal data; wherein, the target modal features include operation intention features and environmental interaction features; Through the dynamic weighting layer, corresponding weights are assigned to each target modal feature based on the signal-to-noise ratio of each target modal feature and the correlation between each target modal feature and the current operation task; The fusion output layer performs weighted fusion of the target modal features based on the weights of each target modal feature, and outputs the fused feature.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Natural interaction method based on multi-modal information fusion

    CN108399427A

  • Robot wearing control system and application method thereof

    CN120461417A