Intelligent wearable robot system based on multi-modal large model and human-in-the-loop control and implementation method
The intelligent wearable robot system, which integrates a multimodal large model with human-in-the-loop control, solves the problems of multimodal perception and dynamic control in existing technologies. It achieves high-precision user intent recognition, environmental perception, and adaptive human-computer interaction, improving the system's intelligence level and security, and is suitable for various application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-05-05
AI Technical Summary
Existing wearable robot systems suffer from problems such as reliance on a single sensor and lack of multimodal data fusion and dynamic control in terms of intent recognition, environmental perception, and human-computer interaction, leading to misjudgment, slow response, insufficient adaptability, and poor user experience.
An intelligent wearable robot system employing a multimodal large model and human-in-the-loop control constructs an edge-cloud collaborative architecture. It rapidly responds to user requests through edge nodes, utilizes a large model in the cloud for highly complex inference and global optimization, combines federated learning and differential privacy to protect data privacy, employs a multi-head attention mechanism for multimodal feature fusion, and achieves adaptive control through algorithms such as impedance control and model predictive control.
It significantly improves the accuracy and robustness of user intent recognition, enhances environmental awareness and adaptive human-computer interaction, achieves high-precision action prediction and behavior planning, ensures data security and real-time performance, and is applicable to various scenarios such as medical rehabilitation, industrial assistance, and military operations.
Smart Images

Figure CN121973156A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wearable robot technology and intelligent control, specifically relating to an intelligent wearable robot system and its implementation method based on a multimodal large model and human-in-the-loop control. Background Technology
[0002] Traditional wearable robots have made initial progress in applications such as medical rehabilitation, industrial assistance, and daily life assistance. However, current technologies still generally rely on single sensors or simple algorithms for motion control, and their overall performance and intelligence level need improvement. Firstly, in terms of intent recognition, traditional wearable robots mainly rely on inertial measurement units (IMUs), single force sensors, or simple motion capture methods to acquire motion data. This low-dimensional data input cannot fully analyze the user's true action intent and psychological state. Especially in dynamic environments, complex motion patterns, or in the presence of noise interference, it is prone to misjudgment, slow response, or control deviation, making it difficult to achieve precise and personalized motion assistance. Secondly, in terms of environmental perception, most existing systems can only capture mechanical state information, lacking the ability to comprehensively perceive and fuse multi-source data such as vision, hearing, touch, force, and user physiological signals. This prevents the construction of a complete environmental understanding and scene model, resulting in insufficient adaptability and unnatural interaction when facing complex task scenarios. Secondly, regarding human-computer interaction, traditional control systems generally adopt a one-way command-driven mode, lacking the perception and adaptive adjustment mechanism for real-time user feedback (such as fatigue level, comfort, and movement deviation). This makes it impossible to flexibly optimize control strategies for different individuals and task requirements, compromising user experience and safety. Finally, in terms of intelligence, traditional wearable robots lack global reasoning capabilities based on large-scale models and multimodal data understanding capabilities. Their overall algorithms are mostly pre-set control logic, failing to achieve high-precision recognition of complex intentions and intelligent collaborative control. This limits the application depth of wearable robots in medical rehabilitation, complex industrial scenarios, and special operational tasks.
[0003] The emergence of Multi-Modal Large Models (MM-LM) offers a new technical approach to solving the aforementioned problems. These models can uniformly receive, fuse, and process multimodal information such as visual, speech, inertial signals, bioelectrical signals, and tactile / force sensations. Through deep neural networks and cross-modal feature association learning, they possess higher-dimensional environmental understanding and user state reasoning capabilities. This not only enables more accurate identification of user action intentions but also allows for prediction of task requirements and adaptation to different scenarios.
[0004] At the same time, the "Human-in-the-loop control" concept allows users' subjective preferences, real-time feedback, and physiological states to directly enter the control closed loop, enabling the system to dynamically adjust and personalize according to users' immediate needs, thereby significantly improving the naturalness and safety of human-machine collaboration.
[0005] Therefore, how to achieve multimodal perception, deep fusion algorithms, and user-in-the-loop optimization control in wearable robot design has become a core technical problem and research focus that needs to be addressed in current technologies. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide an intelligent wearable robot system and implementation method based on multimodal large model and human-in-the-loop control, which solves the problems of multimodal perception, deep fusion algorithm and user-in-the-loop optimization control in the prior art.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0008] A smart wearable robot system based on a multimodal large model and human-in-the-loop control constructs an edge-cloud collaborative architecture. Edge nodes rapidly respond to user requests, while the cloud-based large model performs complex inference and global optimization. Edge devices exchange model parameters with the cloud via secure communication protocols, and federated learning and differential privacy are combined to ensure data privacy protection.
[0009] End-side devices include the wearable mechanism and actuator of the smart wearable robot, as well as the multimodal perception module deployed on the body of the smart wearable robot;
[0010] The cloud-based large model, including a multimodal large model processing module and a human-in-the-loop control module, is used to incorporate the user's physiological state and subjective feedback into the control closed loop, enabling cross-terminal data security collaborative training.
[0011] Edge nodes, including feedback modules, are used to detect the status of actuators in real time and feed it back to the multimodal large model processing module and the human-in-the-loop control module to form closed-loop control.
[0012] The multimodal large model processing module, based on a deep neural network architecture, performs spatiotemporal feature fusion, user intent recognition, and environmental state perception on multimodal data, and achieves dynamic updates and personalized decision optimization through large model fine-tuning and federated learning.
[0013] The human-in-the-loop control module constructs a hybrid control strategy that includes impedance control, model predictive control, iterative learning control, and online parameter tuning using reinforcement learning. It dynamically adjusts control parameters based on user physiological feedback and operational needs, enabling human-computer interaction and motion assistance synchronized with the user.
[0014] The multimodal large model processing module adopts a lightweight Transformer or a multimodal fusion neural network, and uses a multi-head attention mechanism to achieve deep fusion and correlation mining of features of different modalities; it achieves low power consumption and high real-time deployment of local wearable devices through methods including parameter pruning, model quantization and distillation.
[0015] The human-in-the-loop control module further includes:
[0016] The real-time physiological feedback unit is used to dynamically adjust the assist force or control strategy based on the user's heart rate, respiratory rate, and muscle fatigue level indicators.
[0017] User-initiated input interfaces, including voice, gesture recognition, or touch panels, are used for users to independently adjust motion modes, speed, intensity, or assistance levels.
[0018] The adaptive control engine can update control parameters online based on reinforcement learning strategies, optimizing user experience and system energy efficiency.
[0019] The multimodal sensing module includes:
[0020] High-definition RGB / depth camera for recognizing the position, shape, and motion trajectory of objects in the environment;
[0021] An array microphone is used to receive user voice commands and detect ambient sound source characteristics;
[0022] sEMG sensors are used to collect signals of a user's muscle electrical activity to predict movement intentions;
[0023] Physiological sensors are used to monitor a user's physiological state in real time, including heart rate, brain waves, and blood flow characteristics.
[0024] The feedback module includes a force sensor, a position sensor, an angle encoder, and an inertial unit.
[0025] The actuator module employs a high-power-density brushless motor or a flexible actuator, and reduces the overall weight through carbon fiber composite materials and lightweight structural design.
[0026] Each module has a pluggable hardware interface for hardware adaptation and expansion in different application scenarios.
[0027] The implementation method of an intelligent wearable robot system based on a multimodal large model and human-in-the-loop control includes the following steps:
[0028] Step S1: Multimodal data acquisition and preprocessing. Physiological signals including vision, speech, IMU, sEMG, EEG, and ECG are acquired through the multimodal perception module and preprocessed including time alignment, noise filtering, feature standardization, and artifact removal.
[0029] Step S2, Multimodal Deep Fusion and Inference: Input the preprocessed data from Step S1 into the multimodal large model processing module to complete cross-modal feature extraction and semantic fusion, and output the prediction results of user intent, action pattern and environmental state.
[0030] Step S3: Human-in-the-loop adaptive control, combining the user's physiological feedback and subjective preference information, uses impedance control, MPC, ILC and RL online parameter tuning algorithms to generate the final control strategy;
[0031] Step S4: Execution of actions and safety control. The actuator module drives the robot joints according to the control strategy to achieve precise action output, while performing overload and anomaly detection.
[0032] Step S5: Feedback Acquisition and Closed-Loop Optimization. The feedback module transmits force, position, velocity, and fatigue index information in real time, which is used by the multimodal large model and the human-in-the-loop control module for strategy optimization and online correction.
[0033] The application of intelligent wearable robot systems based on multimodal large models and human-in-the-loop control is carried out, with functional adaptations made according to different application scenarios, including:
[0034] Rehabilitation exoskeleton for gait training and rehabilitation assistance;
[0035] Intelligent bionic prostheses are used to precisely simulate natural limb movements and grip control;
[0036] Industrial-assisted exoskeletons are used to reduce muscle fatigue during high-intensity work and improve work efficiency.
[0037] Military load-bearing exoskeleton for high-load marching and complex terrain operations.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. Significantly improved intent recognition accuracy: By utilizing multimodal signal fusion and large model inference mechanisms, it integrates physiological and environmental multimodal signals such as vision, speech, IMU, sEMG, EEG, and ECG, significantly improving the robustness and accuracy of user intent recognition; and compensating for the misjudgment problem caused by traditional single sensors.
[0040] 2. Comprehensive environmental perception: Visual, voice, and tactile data are analyzed together, enabling the system to respond flexibly to complex environments.
[0041] 3. Adaptive Human-Computer Interaction: By introducing user feedback and physiological state through "human-in-the-loop" control, an adaptive control framework is constructed that includes impedance control, model predictive control (MPC), iterative learning control (ILC), and reinforcement learning (RL) for online parameter tuning. This enables real-time, human-computer interaction and achieves dynamic optimization and personalized control.
[0042] 4. Improved intelligence level: By leveraging the global reasoning capabilities of large models, high-precision action prediction and behavior planning can be achieved.
[0043] 5. Balancing data security and real-time performance: Employing edge-cloud collaborative computing, federated learning, and differential privacy mechanisms, it adopts a three-tier architecture of real-time inference on the edge, lightweight collaborative computing at the edge, and multimodal large model training in the cloud. By combining federated learning and differential privacy technologies, it balances real-time performance and data security, ensuring user privacy while supporting low-latency online inference and control.
[0044] 6. Wide range of applications: Applicable to various scenarios such as medical rehabilitation, industrial assistance, military operations and daily life assistance, achieving a comprehensive improvement in intent recognition accuracy, control precision, interaction naturalness and security, with high scalability and practical value. Attached Figure Description
[0045] Figure 1 This is a schematic diagram of the physical form and human-machine coupling of a system based on digital twins, provided by the present invention.
[0046] Figure 2 The system module architecture diagram provided by this invention.
[0047] Figure 3 The algorithm logic and control strategy diagram provided for this invention.
[0048] Figure 4 The diagram illustrates the computing architecture and application scenarios provided by this invention. Detailed Implementation
[0049] The structure and working process of the present invention will be further described below with reference to the accompanying drawings.
[0050] The purpose of this invention is to address the shortcomings of existing wearable robot systems in terms of intent recognition, environmental perception, human-computer interaction, and intelligence level, such as reliance on a single sensor or simple algorithm, inability to effectively understand user intent in complex scenarios, and lack of multimodal data fusion and dynamic control. This invention proposes an intelligent wearable robot system and its implementation method based on a multimodal large model and human-in-the-loop control.
[0051] By organically integrating multimodal large models with human-in-the-loop control, a next-generation intelligent wearable robot system with high perception, intelligent decision-making, and adaptive interaction capabilities can be constructed, providing more intelligent and personalized solutions for scenarios such as medical rehabilitation, industrial assistance, military load-bearing, and daily life assistance.
[0052] A smart wearable robot system based on a multimodal large model and human-in-the-loop control constructs an edge-cloud collaborative architecture. Edge nodes rapidly respond to user requests, while the cloud-based large model performs complex inference and global optimization. Edge devices exchange model parameters with the cloud via secure communication protocols, and federated learning and differential privacy are combined to protect data privacy. Sensitive data is processed locally and transmitted securely. Furthermore, cloud-based model updates improve the overall intelligence level.
[0053] The edge device includes the wearable mechanism and actuator of the intelligent wearable robot, as well as a multimodal perception module deployed on the robot body. The multimodal perception module is used to collect multi-source heterogeneous signals, including visual camera data, voice commands, inertial measurement unit (IMU) data, surface electromyography (sEMG), electroencephalography (EEG), electrocardiography (ECG), tactile, photoplethysmography (PPG), and force sensor data. This module has data synchronization, noise suppression, and feature extraction functions, providing high-quality input for subsequent multimodal fusion. The actuator module, composed of a servo motor, hydraulic actuator, and compliant drive components, is responsible for executing control commands to achieve precise motion output. This module has built-in overload protection and limit mechanisms to ensure motion safety and response speed.
[0054] The cloud-based large model, including a multimodal large model processing module and a human-in-the-loop control module, is used to incorporate the user's physiological state and subjective feedback into the control closed loop, enabling cross-terminal data security collaborative training.
[0055] The Multimodal Large Model Processing Module (MM-LM) uses deep learning and multimodal fusion algorithms (including cross-modal attention mechanisms, spatiotemporal feature alignment, neural networks, etc.) to uniformly encode and extract features from perceived data, outputting user intent, action patterns, and environmental states. This module supports online transfer learning and personalized model fine-tuning, and achieves lightweight inference on the client-side through model distillation, quantization, and pruning; combined with federated learning and differential privacy technologies, it enables secure collaborative training across terminals.
[0056] The human-in-the-loop control module incorporates the user's physiological state and subjective feedback into the control closed loop. It employs a hybrid control algorithm with multiple strategies, including impedance control, MPC, ILC, and RL online parameter tuning, to dynamically optimize control parameters and control modes based on real-time sensing results, generating personalized and highly adaptable control commands.
[0057] Edge nodes, including feedback modules, are used to detect the status of actuators in real time and feed it back to the multimodal large model processing module and the human-in-the-loop control module to form a closed-loop control. The feedback module transmits information such as force, position, speed, and fatigue index in real time, which is used by the multimodal large model and control module for strategy optimization and online correction.
[0058] The human-in-the-loop control module further includes:
[0059] The real-time physiological feedback unit is used to dynamically adjust the assist force or control strategy based on the user's heart rate, respiratory rate, and muscle fatigue level indicators.
[0060] User-initiated input interfaces, including voice, gesture recognition, or touch panels, are used for users to independently adjust motion modes, speed, intensity, or assistance levels.
[0061] The adaptive control engine can update control parameters online based on reinforcement learning strategies, optimizing user experience and system energy efficiency.
[0062] The multimodal sensing module includes:
[0063] High-definition RGB / depth camera for recognizing the position, shape, and motion trajectory of objects in the environment;
[0064] An array microphone is used to receive user voice commands and detect ambient sound source characteristics;
[0065] sEMG sensors are used to collect signals of a user's muscle electrical activity to predict movement intentions;
[0066] Physiological sensors such as EEG / ECG / PPG are used to monitor the user's physiological state in real time, including heart rate, brain waves, and blood flow characteristics.
[0067] The feedback module includes a force sensor, a position sensor, an angle encoder, and an inertial unit, which are used to provide high-precision closed-loop feedback to support 1ms-level response control loops.
[0068] The actuator module employs a high-power-density brushless motor or flexible actuator, and reduces overall weight through carbon fiber composite materials and lightweight structural design to improve wearing comfort and movement flexibility.
[0069] Each module has a pluggable hardware interface for hardware adaptation and expansion in different application scenarios, enabling rapid replacement and upgrade of sensors, actuators and control units.
[0070] The control strategy can automatically switch between impedance control, MPC, ILC and other modes to adapt to different task objectives, user status and external environmental interference.
[0071] Specific embodiments, such as Figures 1 to 4 As shown:
[0072] I. This embodiment relates to a lower limb exoskeleton system applied in the field of medical rehabilitation. This system, by integrating multiple sensing technologies and intelligent control algorithms, can achieve high-precision recognition of the patient's movement intentions and environmental perception, thereby assisting the patient in completing rehabilitation training and improving the recovery of lower limb function. The system is equipped with multiple inertial measurement units (IMUs) and surface electromyography (sEMG) sensors to collect electromyographic signals and movement states of the patient's lower limbs in real time, accurately capturing the patient's movement intentions and action needs. Simultaneously, a vision camera is integrated to identify ground environment information in real time, including road surface smoothness, obstacles, and complex terrain such as stairs, thereby ensuring safety and environmental adaptability during movement. Based on the above multimodal data, the system's embedded multimodal large model, through deep learning and fusion algorithms, combined with the patient's rehabilitation progress and individual characteristics, dynamically adjusts the auxiliary torques of each joint of the exoskeleton to achieve precise and personalized auxiliary support. The human-loop control module continuously monitors the patient's physiological indicators such as heart rate, blood pressure, and fatigue index, incorporating this information into the control closed loop. Combining impedance control, admittance control, and model predictive control (MPC) strategies, it optimizes control parameters in real time to prevent muscle fatigue or injury due to overload. The actuators drive the joint motors according to control commands, assisting the patient in performing various daily living activities, including walking, sitting up, and climbing stairs, ensuring smooth and ergonomic movements. The feedback module continuously collects joint torque, positional error, and motion status data, performing real-time closed-loop adjustments and safety monitoring to ensure the safety and effectiveness of rehabilitation training. This system not only improves the accuracy of user intent recognition but also enhances environmental adaptability and natural interaction, enabling patients to complete high-quality rehabilitation training in a personalized and safe environment, significantly improving rehabilitation outcomes and quality of life.
[0073] II. This embodiment discloses an upper limb assistive exoskeleton system for industrial scenarios, aiming to reduce muscle fatigue in workers performing repetitive, heavy-duty tasks and improve operational efficiency and safety through multimodal sensor fusion and intelligent control strategies. The system is equipped with a high-resolution camera to accurately identify the position, shape, and dynamic changes of workpieces, acquiring real-time information about the work environment and providing accurate spatial positioning for grasping and handling actions. Simultaneously, the system utilizes distributed surface electromyography (sEMG) sensors and an inertial measurement unit (IMU) to collaboratively monitor the worker's electromyographic activity and arm movement status, capturing the worker's operational intentions and movement requirements in real time. Through deep fusion and inference of visual, electromyographic, and motion data using a multimodal large-scale model, the system automatically generates optimized grasping trajectories and output force distribution strategies for the actuators, ensuring precise and efficient assisted movements and significantly reducing the burden on the worker's upper limbs. The human-in-the-loop control module continuously monitors muscle fatigue and physiological indicators, dynamically adjusting the assist force based on feedback to avoid excessive or insufficient assistance, ensuring comfort and safety during operation. The actuators drive multi-degree-of-freedom joints according to control commands, enabling assistance with complex movements such as grasping, handling, and rotation, with rapid response and smooth motion. The feedback module monitors joint torque, positional deviation, and gripping stability in real time to ensure precise and stable gripping movements, preventing workpiece slippage or damage. This system integrates multimodal sensing, intelligent control, and real-time feedback to adapt to complex and changing industrial environments, improving operational safety and efficiency, and providing workers with reliable upper limb strength support and precise motion assistance.
[0074] The implementation method of an intelligent wearable robot system based on a multimodal large model and human-in-the-loop control includes the following steps:
[0075] Step S1: Multimodal data acquisition and preprocessing. Physiological signals including vision, speech, IMU, sEMG, EEG, and ECG are acquired through the multimodal perception module and preprocessed including time alignment, noise filtering, feature standardization, and artifact removal.
[0076] Step S2, Multimodal Deep Fusion and Inference: Input the preprocessed data from Step S1 into the multimodal large model processing module to complete cross-modal feature extraction and semantic fusion, and output the prediction results of user intent, action pattern and environmental state.
[0077] Step S3: Human-in-the-loop adaptive control, combining the user's physiological feedback and subjective preference information, uses impedance control, MPC, ILC and RL online parameter tuning algorithms to generate the final control strategy;
[0078] Step S4: Execution of actions and safety control. The actuator module drives the robot joints according to the control strategy to achieve precise action output, while performing overload and anomaly detection.
[0079] Step S5: Feedback Acquisition and Closed-Loop Optimization. The feedback module transmits force, position, velocity, and fatigue index information in real time, which is used by the multimodal large model and the human-in-the-loop control module for strategy optimization and online correction.
[0080] The application of intelligent wearable robot systems based on multimodal large models and human-in-the-loop control is carried out, with functional adaptations made according to different application scenarios, including:
[0081] Rehabilitation exoskeleton for gait training and rehabilitation assistance;
[0082] Intelligent bionic prostheses are used to precisely simulate natural limb movements and grip control;
[0083] Industrial-assisted exoskeletons are used to reduce muscle fatigue during high-intensity work and improve work efficiency.
[0084] Military load-bearing exoskeleton for high-load marching and complex terrain operations.
[0085] In summary, the system comprises a physical entity and cognitive control. The physical entity is fixed to key parts of the wearer's body using an organic combination of rigid support and flexible actuation. It integrates multimodal sensing components such as visual sensors, microphone arrays, inertial measurement units (IMUs), surface electromyography (sEMG) electrodes, and tactile / force sensors to collect environmental states, motion signals, and physiological information in real time. Cognitive control performs deep fusion and semantic understanding of the aforementioned high-dimensional sensing data through a multimodal large-scale model, accurately interpreting user intentions and action patterns, and combining individual anatomical parameters and environmental features to achieve dynamic modeling and virtual-real synchronization. The human-in-the-loop control module within the system integrates subjective user feedback (such as pain, fatigue, and comfort) and objective physiological states. This invention incorporates closed-loop control and utilizes advanced algorithms such as impedance control, model predictive control (MPC), and reinforcement learning (RL) to achieve online adaptive adjustment of control parameters, thereby generating high-precision, compliant, and safe motion control commands to drive the actuator. Simultaneously, the feedback module continuously transmits joint force, position, speed, fatigue index, and safety interaction indicators to iteratively optimize the control strategy. Furthermore, it employs techniques such as model distillation, quantization, federated learning, and differential privacy to achieve real-time inference and data security protection at the edge. This invention can significantly improve the accuracy of intent recognition, control precision, interaction naturalness, and human-machine collaboration safety in various application scenarios, including medical rehabilitation, industrial assistance, military operations, and daily life assistance, demonstrating broad application prospects and practical value.
[0086] To enable those skilled in the art to implement this invention clearly and unambiguously, and to avoid insufficient disclosure due to overly generalized core concepts in the claims, the specific implementation parameters and algorithm logic of the "multimodal large model processing module" and the "human-in-the-loop control module" in this system are further described in detail:
[0087] 1. Specific implementation parameters of Multimodal Large Model (MM-LM) and edge inference
[0088] In practical implementation, to achieve accurate feature extraction and low-latency operation, the 60fps video stream output by the visual sensor and the 1000Hz electromyography signal acquired by the sEMG sensor are used to extract single-modal features via lightweight residual networks (such as ResNet-18) and one-dimensional convolutional neural networks (1D-CNN), respectively. Subsequently, a cross-modal multi-head attention mechanism is used to align and fuse spatial environmental features with human temporal force exertion features. To meet the 1ms-level control loop requirement at the system's underlying level, before deploying the large model to the edge, INT8 quantization and knowledge distillation techniques are used. While ensuring that the recognition accuracy decreases by no more than 1.5%, the number of model parameters is compressed to 20% of the original, so that the overall inference latency at the edge is strictly controlled within 15ms.
[0089] 2. Specific Quantitative Logic of Human-in-the-Loop (HITL) Control and Adaptive Parameter Tuning
[0090] The "fatigue index" in the human-loop control module is quantitatively evaluated by analyzing the median frequency (MDF) decline rate of the sEMG signal in real time. The fatigue index normalization interval is set to [0, 1]. When the system detects that the user's real-time fatigue index exceeds a preset safety threshold (e.g., 0.6), the adaptive control engine is immediately triggered. At this time, the reward function of the reinforcement learning (RL) algorithm automatically increases the weight allocation to "minimizing the human body's active torque," thereby instructing the impedance controller to dynamically adjust the system's virtual stiffness parameter (K) and damping parameter (B). This online adjustment of quantitative parameters enables the actuator to instantly output a larger auxiliary torque to take over the load, thus achieving true "human-loop" real-time physiological closed-loop intervention.
[0091] Summary of Experimental Results and Performance Comparison Table
[0092] To further verify the implementation effect and technical advantages of the present invention, a 30-minute comparative experiment was conducted between the system proposed in this invention (multimodal large model + human-in-the-loop control) and a traditional wearable robot system (relying solely on IMU sensing + traditional PID control) in simulated complex terrain rehabilitation training and high-load industrial handling scenarios. Specific performance comparison data are shown in Table 1.
[0093]
[0094] The above comparative data fully demonstrates that this invention, through deep fusion of multimodal data such as vision, electromyography, and inertia, completely solves the problem of low intent recognition rate (accuracy rate of 98.2%) in complex scenarios using traditional single sensors. Simultaneously, by specifying and quantifying user physiological indicators (fatigue index) and model predictive control, a closed-loop "human-in-the-loop" dynamic parameter adjustment mechanism is formed, not only reducing trajectory tracking error to 1.1°, but also lowering the peak collision force of human-computer interaction by 68.8%. The implementation scheme provided by this invention possesses a clear algorithmic model architecture and quantitative control logic, completely overcoming the technical bottlenecks of poor generalization and inability to dynamically adjust to individual differences in existing technologies, demonstrating practical engineering feasibility and outstanding application value.
[0095] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. An intelligent wearable robot system based on a multimodal large model and human-in-the-loop control, characterized in that: An edge-cloud collaborative architecture is constructed, enabling rapid response to user requests through edge nodes, leveraging large cloud models for complex inference and global optimization, and facilitating model parameter exchange between edge devices and the cloud via secure communication protocols. Data privacy is protected through federated learning and differential privacy principles. End-side devices include the wearable mechanism and actuator of the smart wearable robot, as well as the multimodal perception module deployed on the body of the smart wearable robot; The cloud-based large model, including a multimodal large model processing module and a human-in-the-loop control module, is used to incorporate the user's physiological state and subjective feedback into the control closed loop, enabling cross-terminal data security collaborative training. Edge nodes, including feedback modules, are used to detect the status of actuators in real time and feed it back to the multimodal large model processing module and the human-in-the-loop control module to form closed-loop control.
2. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 1, characterized in that: The multimodal large model processing module, based on a deep neural network architecture, performs spatiotemporal feature fusion, user intent recognition, and environmental state perception on multimodal data, and achieves dynamic updates and personalized decision optimization through large model fine-tuning and federated learning. The human-in-the-loop control module constructs a hybrid control strategy that includes impedance control, model predictive control, iterative learning control, and online parameter tuning using reinforcement learning. It dynamically adjusts control parameters based on user physiological feedback and operational needs, enabling human-computer interaction and motion assistance synchronized with the user.
3. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 2, characterized in that: The multimodal large model processing module adopts a lightweight Transformer or a multimodal fusion neural network, and uses a multi-head attention mechanism to achieve deep fusion and correlation mining of features of different modalities; it achieves low power consumption and high real-time deployment of local wearable devices through methods including parameter pruning, model quantization and distillation.
4. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 2, characterized in that: The human-in-the-loop control module further includes: The real-time physiological feedback unit is used to dynamically adjust the assist force or control strategy based on the user's heart rate, respiratory rate, and muscle fatigue level indicators. User-initiated input interfaces, including voice, gesture recognition, or touch panels, are used for users to independently adjust motion modes, speed, intensity, or assistance levels. The adaptive control engine can update control parameters online based on reinforcement learning strategies, optimizing user experience and system energy efficiency.
5. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 1, characterized in that: The multimodal sensing module includes: High-definition RGB / depth camera for recognizing the position, shape, and motion trajectory of objects in the environment; An array microphone is used to receive user voice commands and detect ambient sound source characteristics; sEMG sensors are used to collect signals of a user's muscle electrical activity to predict movement intentions; Physiological sensors are used to monitor a user's physiological state in real time, including heart rate, brain waves, and blood flow characteristics.
6. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 1, characterized in that: The feedback module includes a force sensor, a position sensor, an angle encoder, and an inertial unit.
7. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 1, characterized in that: The actuator module employs a high-power-density brushless motor or a flexible actuator, and reduces the overall weight through carbon fiber composite materials and lightweight structural design.
8. The intelligent wearable robot system based on multimodal large model and human-in-the-loop control according to claim 1, characterized in that: Each module has a pluggable hardware interface for hardware adaptation and expansion in different application scenarios.
9. A method for implementing an intelligent wearable robot system based on a multimodal large model and human-in-the-loop control, characterized in that: Includes the following steps: Step S1: Multimodal data acquisition and preprocessing. Physiological signals including vision, speech, IMU, sEMG, EEG, and ECG are acquired through the multimodal perception module and preprocessed including time alignment, noise filtering, feature standardization, and artifact removal. Step S2, Multimodal Deep Fusion and Inference: Input the preprocessed data from Step S1 into the multimodal large model processing module to complete cross-modal feature extraction and semantic fusion, and output the prediction results of user intent, action pattern and environmental state. Step S3: Human-in-the-loop adaptive control, combining the user's physiological feedback and subjective preference information, uses impedance control, MPC, ILC and RL online parameter tuning algorithms to generate the final control strategy; Step S4: Execution of actions and safety control. The actuator module drives the robot joints according to the control strategy to achieve precise action output, while performing overload and anomaly detection. Step S5: Feedback Acquisition and Closed-Loop Optimization. The feedback module transmits force, position, velocity, and fatigue index information in real time, which is used by the multimodal large model and the human-in-the-loop control module for strategy optimization and online correction.
10. The application of an intelligent wearable robot system based on a multimodal large model and human-in-the-loop control, characterized by: Functionality is adapted to different application scenarios, including: Rehabilitation exoskeleton for gait training and rehabilitation assistance; Intelligent bionic prostheses are used to precisely simulate natural limb movements and grip control; Industrial-assisted exoskeletons are used to reduce muscle fatigue during high-intensity work and improve work efficiency. Military load-bearing exoskeleton for high-load marching and complex terrain operations.