Human-robot interaction control method for rehabilitation robot based on multi-modal and diffusion strategy
By combining multimodal signal fusion and diffusion probability models with rolling predictive control, the accuracy and safety issues in the human-computer interaction control of rehabilitation robots are solved, enabling personalized, smooth, and safe generation of auxiliary movements, thereby improving the effectiveness and safety of rehabilitation training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing human-computer interaction control methods for rehabilitation robots are insufficient to meet the precision, safety, and comfort requirements of clinical rehabilitation. Traditional methods cannot fully capture multi-dimensional complementary information such as the patient's muscle activation state, changes in movement posture, and human-computer interaction force. The accuracy of motion prediction is insufficient, and complex control algorithms are difficult to achieve an effective balance between precision and real-time performance.
By employing multimodal signal fusion technology, combined with diffusion probability model and rolling predictive control, a multimodal temporal fusion network is constructed by collecting force/torque, surface electromyography and skeletal joint posture signals to generate action sequences. Temporal consistency and dynamic constraints are introduced to achieve random action generation and closed-loop control.
It improves the comprehensive perception of the patient's movement status, enhances the flexibility and safety of movement generation, meets the personalized needs of rehabilitation training, achieves system stability and real-time performance, and reduces patient discomfort and safety risks.
Smart Images

Figure CN122425666A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical rehabilitation equipment and intelligent control technology, specifically to a human-computer interaction control method for rehabilitation robots based on multimodal and diffusion strategies. Background Technology
[0002] In the field of medical rehabilitation, rehabilitation robots, as key devices to assist patients in restoring motor function, provide personalized training support through human-computer interaction and have become an important means of rehabilitation treatment for motor dysfunction. However, the current mainstream human-computer interaction control methods for rehabilitation robots still face many technical bottlenecks, making it difficult to fully meet the precision, safety, and comfort requirements of clinical rehabilitation.
[0003] Most existing rehabilitation robot controls rely on single-modal sensing signals (such as skeletal posture signals or electromyography signals) for action intent recognition and assisted action generation. This one-sided signal perception method cannot comprehensively capture multi-dimensional complementary information such as the patient's muscle activation state, changes in movement posture, and human-computer interaction forces, resulting in insufficient accuracy in action prediction and difficulty in accurately matching the patient's actual movement needs. At the same time, the multi-source signals involved in rehabilitation training, such as force / torque, surface electromyography (sEMG), and skeletal joint posture, exhibit significant heterogeneity in terms of time scale (sampling frequency covering 30Hz-1kHz), physical meaning, and noise characteristics. Traditional data fusion methods cannot effectively mine the temporal dependencies within each modality and the complementary value across modalities, making multimodal information integration difficult and further affecting the system's comprehensive perception of the patient's movement state.
[0004] In the motion generation stage, traditional deterministic control strategies (such as LSTM regression and reinforcement learning DDPG methods) often struggle to characterize multiple reasonable assistive movements corresponding to the same motor intention. They are also poorly adapted to noise interference in the perceived signal and individual differences in movement among different patients, easily leading to problems such as large motion tracking errors and sudden fluctuations. This not only results in a large deviation between the assistive movements and the patient's actual motion trajectory, but also causes fatigue or discomfort during patient training due to insufficient motion smoothness and significant acceleration fluctuations. It can even pose safety hazards due to sudden peaks in interactive force, seriously affecting the effectiveness and safety of rehabilitation training.
[0005] Furthermore, the conflict between complex control algorithms and the need for real-time interaction also restricts the practical application of rehabilitation robots. While highly complex algorithms can improve control accuracy to some extent, they are often accompanied by significant computational delays, making it difficult to meet the real-time control requirements of rehabilitation robots above 50Hz. On the other hand, while low-complexity algorithms can guarantee real-time performance, they sacrifice the accuracy and robustness of motion generation, making it difficult to achieve an effective balance between accuracy and real-time performance.
[0006] To overcome the aforementioned technical limitations, there is an urgent need for a human-computer interaction control method that can integrate the advantages of multi-source heterogeneous signals, accurately characterize complex motion distributions, and balance safety and real-time performance. The development of multimodal data fusion technology and diffusion probability models offers new insights into this problem: effective fusion of multimodal signals enables comprehensive perception of the patient's motion state, while diffusion strategies can improve the system's robustness to noise and individual differences through random motion generation modeling, while ensuring the smoothness and physical feasibility of the motion.
[0007] From the perspective of existing technologies, representing robot strategies as a conditional denoising diffusion process, modeling complex and multi-peaked action distributions through action diffusion, and combining time-series diffusion Transformers with receding horizon control (rolling time-domain control) to achieve closed-loop action generation and execution can effectively handle high-dimensional action spaces and multimodal action output problems (Diffusion Policy: Visuomotor Policy Learning via Action Diffusion). Diffusion strategies are suitable for robot action sequence generation, especially for modeling multimodal action distributions, and improve execution stability through rolling prediction. However, for general robot operation or visual motion control tasks, the conditional inputs are mainly based on visual observations, and a standardized acquisition, temporal fusion, and safety constraint control framework for multimodal heterogeneous signals centered on force / torque, surface electromyography, and skeletal joint posture has not been established specifically for the human-computer interaction characteristics in rehabilitation robot scenarios. Furthermore, it has not addressed the unique challenges of individual patient differences, physiological noise sensitivity, interaction safety, and smooth auxiliary action generation in rehabilitation training. Therefore, it still cannot directly meet the comprehensive needs of rehabilitation robot human-computer interaction control for multimodal perception, safety, and personalized assistance.
[0008] An existing multi-sensor fusion sEMG control system for upper limb exoskeleton rehabilitation robots fuses sEMG data with sensor data such as angle, pressure, inertia, and torque to enhance motion intent recognition and control robustness. The system design employs a technical approach of signal acquisition, preprocessing, feature extraction, fusion, and machine learning classification (Multi-Sensor Fusion-Based Surface EMG Control System for Upper LimbExoskeleton Rehabilitation Robots). This system already involves upper limb rehabilitation robots, multi-sensor fusion, and sEMG control, which is quite similar to the application scenario of this invention. However, the core of this system still focuses on the mapping of control commands after intent recognition and classification, without further modeling the control problem as a short-term motion sequence probability distribution generation problem under multimodal constraints, nor does it use a diffusion probability model to model the distribution of complex auxiliary movements. Furthermore, although the system involves multi-sensor fusion, there is no unified fusion modeling for time-series heterogeneous signals such as skeletal joint three-dimensional pose, electromyography, and interactive forces, nor is a complete scheme combining diffusion denoising generation, motion smoothing constraints, dynamic constraints, and closed-loop rolling control disclosed. Therefore, there are still significant shortcomings in terms of complex motion distribution modeling capabilities, auxiliary motion diversity, motion smoothness, and unified implementation of physical safety constraints.
[0009] As can be seen from the existing technologies described above, one type of existing solution focuses on diffusion strategy action generation, which can solve the problems of multimodal action distribution and rolling prediction, but lacks multimodal perception and safety control design for human-computer interaction in rehabilitation scenarios; the other type focuses on multi-sensor fusion control of rehabilitation robots, which can improve the robustness of intent recognition, but lacks the ability to probabilistically model the distribution of complex auxiliary actions, and also lacks a collaborative control mechanism of diffusion generation and closed-loop constraints. This invention addresses these shortcomings by proposing a human-computer interaction control method for rehabilitation robots based on multimodal and diffusion strategies. By fusing force / torque, surface electromyography, and skeletal posture information, and combining a diffusion strategy to generate short-term action sequences, it introduces temporal consistency constraints, dynamic constraints, and a rolling predictive closed-loop control mechanism, thereby achieving more accurate, smooth, safe, and personalized human-computer interaction auxiliary control. Summary of the Invention
[0010] From the perspective of rehabilitation training development, the human motor control process is highly complex and dynamically evolving. Patients' motor intentions, muscle activation states, and human-computer interaction behaviors change continuously with training stages and individual differences. Traditional human-computer interaction analysis methods based on single modalities or static features have significant limitations and biases. Meanwhile, multimodal interaction signals commonly suffer from inconsistent sampling frequencies, large differences in physical dimensions, and strong noise interference during acquisition. Furthermore, in actual rehabilitation scenarios, it is difficult to accurately characterize the complex mapping relationship between patient motor intentions and assistive movements using a single deterministic model. This results in a lack of personalization and adaptability in the generated rehabilitation movements, affecting training safety and effectiveness. Stochastic modeling methods can characterize the uncertainty and ambiguity in the movement generation process, while multimodal fusion technology can collaboratively represent patient motor states and human-computer interaction characteristics through complementary information, combined with a closed-loop control mechanism to achieve continuous correction and optimization of the dynamic interaction process.
[0011] The present invention is achieved by at least one of the following technical solutions.
[0012] A human-computer interaction control method for rehabilitation robots based on multimodal and diffusion strategies includes the following steps: S1. Collect force signals, surface electromyography signals, and skeletal joint posture signals, and perform standardized preprocessing to obtain multimodal features; S2. Input the multimodal features into the multimodal temporal fusion network model to generate the fusion conditional features at the current time. S3. Model the distribution of complex actions using a diffusion probability model, and generate action sequences using fusion conditional features as constraints. S4. A rolling predictive control strategy is adopted, which selects only the first action in the generated action sequence as the current control input and collects data for the next control cycle.
[0013] Further, step S1 includes: (1) Collect the interaction force between the patient and the training device. The original force signal at each sampling time is represented as a six-dimensional vector. The original force signal is zero-point calibrated and gravity compensation is performed. The compensated signal is filtered by moving average and the statistical features of mean and standard deviation are extracted within a fixed time window. Then, the features of each channel are normalized by the Z-score normalization method to meet the uniform scale distribution, thereby realizing the standardized representation of the mechanical interaction features. (2) Collect electromyographic signals of the target muscle group, perform bandpass filtering on the electromyographic signals, and perform full-wave rectification and smoothing on the filtered signals to extract the envelope signal reflecting the change in muscle contraction intensity, and calculate the time domain characteristics of the root mean square value and the mean absolute value from the envelope signal. (3) Real-time acquisition of three-dimensional spatial coordinate information of key joints of the human body, calculation of joint motion velocity and joint acceleration through the difference between adjacent frames to characterize the trend of motion state; then, splicing the joint position, velocity and acceleration features to form a skeletal posture feature vector that comprehensively reflects the patient's motion state.
[0014] Further, in step S2, the multimodal temporal fusion network model generates the fusion condition features at the current moment, including the following steps: Temporal modeling is performed on the feature sequences of different modalities to extract the dynamic evolution characteristics within each modality. An attention mechanism is introduced in the fusion stage to adaptively evaluate the importance of different modalities at different times. The fusion condition features are obtained by weighted summation of the temporal features of each modality.
[0015] Furthermore, in step S3, the distribution of complex actions is modeled using a diffusion probability model, including the following steps: First, using fused conditional features as constraints, an action sequence is generated. The control problem is then modeled as a conditional probability distribution of the random variable of the action sequence under the fused conditional features: Secondly, a diffusion probability model is introduced to construct the forward noise addition process. The noise intensity is set by a linear scheduling strategy so that the action sequence approximately follows a simple Gaussian distribution at the end of the diffusion, thereby reducing the difficulty of reverse modeling. Then, in the reverse generation stage, a Transformer-based denoising network is constructed to perform a step-by-step inversion of the diffusion process under the guidance of fusion condition features.
[0016] Furthermore, a constraint reinforcement mechanism is introduced during the action generation process to impose temporal consistency constraints on the rate of change between adjacent actions:
[0017] in, for The single-step control action corresponding to each control cycle. The preset action change threshold is used to limit the change range between adjacent control actions to ensure the smoothness and safety of the generated action sequence.
[0018] Furthermore, in step S4, constraints are imposed on the rate of change between adjacent control inputs:
[0019] in, A preset threshold is used to limit mutations in control instructions, avoiding discomfort or potential risks to patients. This indicates the current control input.
[0020] Furthermore, in joint space control scenarios, the generated control inputs must also satisfy robot dynamics constraints:
[0021] in, For the robot at the joint position The inertia matrix below, Let the joint position vector be... The joint acceleration vector, This is a matrix of Coriolis force and centrifugal force terms related to joint position and joint velocity. For joint velocity vectors, The vector of the gravity term. This is the joint driving torque vector output by the closed-loop control strategy.
[0022] The system for implementing the human-computer interaction control method for the rehabilitation robot based on multimodal and diffusion strategies includes: The multimodal sensing module is used to acquire force signals, surface electromyography signals, and skeletal joint posture signals, and to perform preprocessing. The multimodal temporal fusion module is used to generate the fusion condition features at the current time. The action sequence generation module is used to generate action sequences under multimodal constraints. The closed-loop human-machine interaction module is used to take the action in the generated action sequence as the current control input through a rolling predictive control strategy.
[0023] A computer device according to the present invention includes a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, which, when executed by the processor, causes the processor to implement the method described herein.
[0024] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the method described herein.
[0025] Compared with existing technologies, this invention offers the following advantages: Human motion control is characterized by significant dynamism and individual variability. A patient's movement intentions, muscle activation states, and human-computer interaction behaviors continuously evolve during rehabilitation training. Traditional rehabilitation robot systems often rely on single-modal signals or deterministic control strategies, making it difficult to comprehensively and accurately represent the patient's true movement state. Furthermore, their adaptability to individual differences and perceptual noise is limited, easily leading to monotonous auxiliary movements, unstable interactions, and even safety hazards. To address these issues, this invention introduces multimodal fusion, random action generation, and closed-loop control mechanisms to construct a human-computer interaction control method system for rehabilitation training. This system effectively improves the comprehensiveness of movement intention perception, the flexibility of action generation, and the stability and safety of system operation, demonstrating significant technical advantages and application value.
[0026] 1. Compared with existing rehabilitation systems based on single sensor information or simple feature fusion, the multimodal heterogeneous interactive signal standardization acquisition and temporal fusion scheme proposed in this invention can collaboratively perceive and uniformly express force / torque signals, surface electromyography signals, and skeletal joint posture signals. It effectively solves the inconsistency problem of multimodal data in terms of sampling frequency, physical dimensions, and temporal characteristics, thereby more comprehensively and accurately representing the patient's movement intention and interaction state, and providing high-quality input for subsequent control decisions.
[0027] 2. The feature modeling method based on multimodal temporal fusion network proposed in this invention improves the sensitivity and robustness of perceiving changes in the patient's motion state by fully mining the temporal dependencies within each modal signal and realizing the dynamic fusion of cross-modal complementary information. It overcomes the shortcomings of traditional methods that rely heavily on historical static features and are insufficient in characterizing dynamic evolution processes.
[0028] 3. Compared with traditional deterministic control or rule-driven action generation methods, this invention models the human-computer interaction control problem as a random action sequence generation problem under multimodal constraints, and introduces a diffusion probability model to characterize the distribution of complex auxiliary actions. This can generate diverse and reasonable auxiliary action schemes while ensuring safety constraints, and improve the adaptability of rehabilitation action generation to different patients, different training stages and different interaction states.
[0029] 4. This invention introduces temporal consistency constraints and physical feasibility constraints during the action generation process, which effectively limits the abrupt changes and unexecutable situations of generated actions, ensures the smoothness and safety of auxiliary actions, reduces the risk of discomfort or secondary injury to patients during rehabilitation training, and enhances the reliability of the system in practical applications.
[0030] 5. The closed-loop human-computer interaction control mechanism based on rolling prediction proposed in this invention organically combines multimodal perception, random action generation and robot execution process. By periodically updating control decisions, it effectively suppresses the accumulation of prediction errors in the time dimension, improves the system's real-time response capability and interaction stability to changes in patient status, and meets the actual needs of rehabilitation training for continuity and real-time performance.
[0031] 6. The overall technical solution proposed in this invention ensures the accuracy of motion intention recognition and control safety while taking into account real-time computing performance. It is suitable for the online operation needs of actual rehabilitation robot systems and can provide patients with a more personalized, intelligent, safe and efficient rehabilitation training experience. It has good engineering feasibility and prospects for promotion and application. Attached Figure Description
[0032] Figure 1 is a flowchart of the multimodal perception-diffusion generation-closed-loop human-machine interaction control method of the embodiment.
[0033] Figure 2 is a flowchart of the standardized acquisition and timing fusion method for multimodal heterogeneous interactive signals in the embodiment.
[0034] Figure 3 is a flowchart of the random action generation modeling method based on diffusion strategy in the embodiment.
[0035] Figure 4 is a flowchart of the closed-loop human-computer interaction control implementation method of the embodiment. Detailed Implementation
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] This embodiment of a rehabilitation robot human-computer interaction control system based on a multimodal temporal fusion and diffusion strategy includes: The multimodal sensing module is used to acquire force / torque signals, surface electromyography signals, and skeletal joint posture signals, and to process and preprocess them within a time window of length T.
[0038] The multimodal temporal fusion module is used to generate fusion condition features for the current time step.
[0039] The action sequence generation module is used to generate action sequences under multimodal constraints.
[0040] The closed-loop human-machine interaction module is used to select only the first action in the generated action sequence as the current control input through a rolling predictive control strategy.
[0041] like Figure 1 As shown in this embodiment, a human-computer interaction control method for a rehabilitation robot based on a multimodal and diffusion strategy includes the following steps: 1. Collect multimodal heterogeneous interaction signals, standardize and preprocess them, and then generate fusion condition features at the current moment through a multimodal temporal fusion network model.
[0042] During rehabilitation training, different patients exhibit significant differences in their expression of motor intention, muscle activation patterns, and human-computer interaction behaviors. A single modal signal is insufficient to comprehensively and accurately reflect the patient's true motor state. This invention achieves comprehensive acquisition and refined modeling of patient motor behavior through the collaborative perception and fusion processing of multi-source heterogeneous interactive signals. This embodiment targets common mechanical, physiological, and kinematic information in rehabilitation training, simultaneously acquiring and processing human-computer interaction force information, surface electromyography signals, and skeletal joint posture signals to construct a unified multimodal data representation. The acquired multi-source signals contain rich temporal variations and state correlation information, providing a reliable data foundation for subsequent motion state analysis, action generation, and control strategy formulation. Due to the significant heterogeneity of different modal signals in terms of sampling frequency, physical dimensions, and temporal characteristics, standardized preprocessing and temporal fusion methods are required to achieve consistent expression and collaborative representation of multimodal information in time and feature space.
[0043] like Figure 2 As shown, step 1 is as follows: First, to address the need for acquiring human-computer interaction force information during rehabilitation training, a six-dimensional force / torque sensor installed at the end of the robotic arm is used to collect the interaction forces between the patient and the training equipment in real time. In one embodiment, the raw force / torque signal at each sampling moment can be represented as a six-dimensional vector:
[0044] in, Indicates time The acquired six-dimensional force / torque signal vector; Indicates along Components of the interaction force along the axial direction; Indicates along Components of the interaction force along the axial direction; Indicates along Components of the interaction force along the axial direction; Indicates circling The interaction torque components of the shaft; Indicates circling The interaction torque components of the shaft; Indicates circling The interaction torque components of the shaft; superscript This indicates transpose.
[0045] The acquired raw force signal contains interference components introduced by sensor zero-point drift and the weight of the end effector. Therefore, zero-point calibration and gravity compensation are performed to obtain the true human-machine interaction force signal. To reduce the impact of measurement noise on subsequent modeling, the compensated signal is processed by moving average filtering, and statistical features such as mean and standard deviation are extracted within a fixed-length time window. Subsequently, the features of each channel are normalized using the Z-score normalization method to ensure a uniform scale distribution, thereby achieving a standardized representation of the mechanical interaction features.
[0046] Secondly, to address the need for sensing the patient's muscle activation state, a multi-channel surface electromyography (EMG) acquisition device was used to simultaneously acquire EMG signals from the target muscle group. The raw EMG signal is a time-varying voltage sequence, containing useful neuromuscular activation information but also susceptible to power frequency interference and low-frequency drift. Therefore, the raw EMG signal was bandpass filtered, and the filtered signal underwent full-wave rectification and smoothing to extract the envelope signal reflecting changes in muscle contraction intensity. Based on this, time-domain features such as the root mean square (RMS) value and mean absolute value were calculated from the envelope signal, with the RMS feature being particularly important. It can be represented as:
[0047] in, This indicates the number of sampling points used to calculate the root mean square value; Indicates the first sampling window within the specified sampling window. The amplitude of the electromyographic signal corresponding to each sampling point; This is the sampling point number.
[0048] The aforementioned features can effectively characterize the intensity and duration of muscle activation. To mitigate the impact of differences in channel gain and individual variations on model training, the electromyographic temporal features of the extracted envelope signal are standardized to ensure good comparability and stability.
[0049] Then, to meet the need for acquiring the patient's overall movement posture and joint movement characteristics, depth vision sensing devices are used to collect real-time three-dimensional spatial coordinate information of key human joints. In one embodiment, a single frame of skeletal data can be represented as a set of coordinates of multiple joint points. :
[0050] in, Indicates time No. Three-dimensional position vectors of each joint; Indicates time No. Each key point is Coordinate components along the axis; Indicates time No. Each key point is Coordinate components along the axis; Indicates time No. Each key point is Coordinate components along the axis.
[0051] To fully extract dynamic information during the motion process, joint motion velocities are calculated by differencing adjacent frames based on the original position data. :
[0052] in, Indicates time No. The velocity vector of each joint; This represents the time interval between two adjacent frames of bone data, i.e., the sampling period of the bone pose signal.
[0053] Furthermore, joint acceleration is calculated to characterize the changing trend of motion state. Subsequently, the joint position, velocity, and acceleration features are concatenated to form a skeletal posture feature vector that comprehensively reflects the patient's motion state. By normalizing the above features, the scale effects caused by differences in body size and acquisition distance are eliminated, achieving a unified expression of skeletal motion features.
[0054] Finally, after constructing standardized features for force / torque, electromyography, and skeletal posture modalities, the multimodal temporal features were fused to construct a multimodal temporal fusion network model. This model performs temporal modeling on the feature sequences of different modalities to extract the dynamic evolution characteristics within each modality. An attention mechanism is introduced during the fusion stage to adaptively evaluate the importance of different modalities at different times. By weighted summation of the temporal features of each modality, a unified fusion condition feature is obtained. express:
[0055] in, Indicates the first Temporal feature representation of class modalities These are the corresponding adaptive weights. The fusion conditional features After nonlinear mapping and dimensional transformation, a fixed-dimensional conditional input vector is formed, which is used for subsequent rehabilitation action generation and human-machine collaborative control, thereby improving the system's ability to understand the patient's movement intentions and the level of intelligence in rehabilitation training.
[0056] 2. Random action sequence generation based on diffusion strategy.
[0057] In the human-computer interaction control process of rehabilitation training, patients may respond to multiple reasonable and safe auxiliary action schemes under the same motor intention and state conditions. The action generation process has obvious randomness and multiple solutions. Traditional deterministic control methods are difficult to effectively characterize the above uncertainties, which can easily lead to monotonous action patterns or insufficient adaptability to individual differences. To this end, this invention proposes a conditional action sequence probability model, which transforms the human-computer interaction control problem into a random action sequence generation problem under multimodal conditional constraints. By modeling the complex action distribution through a diffusion probability model, the robustness of the system to patient differences and perceptual noise is improved while ensuring the smoothness and physical feasibility of the actions.
[0058] like Figure 3 As shown, step 2 is as follows: First, a formal model is constructed for the human-computer interaction control problem. At each control moment... The system has a length of Multimodal features are acquired and fused within a time window to form a fused conditional feature representation. The control target no longer directly generates a single action, but instead integrates conditional features. As a conditional input, the initial noisy action sequence is subjected to multi-step reverse denoising to generate a sequence of lengths that match the patient's current movement intention, muscle activation state, and human-computer interaction state. short action sequence :
[0059] in, Indicates the current control time. The corresponding short-term action sequence; This represents the time interval between two adjacent control actions, i.e., the control cycle; This indicates the number of action steps included in the prediction time domain; Indicates the current time Single-step control actions.
[0060] This allows the control problem to be modeled as a condition under given fusion features. The following model is used to model the conditional probability distribution of the random variable of the action sequence:
[0061] This modeling approach can explicitly characterize the diversity among different action schemes, providing a theoretical basis for the subsequent introduction of random generation strategies.
[0062] Secondly, to characterize complex action distributions, a diffusion probability model is introduced to construct the forward noise addition process. This is achieved by defining... The forward diffusion process, consisting of steps, gradually injects Gaussian noise into the real action sequence, causing its distribution to gradually approximate a standard Gaussian distribution. The forward diffusion process can be represented as:
[0063] in, Indicates that in the known number of... Step action sequence Conditions, No. Step action sequence The conditional probability distribution; Indicates the diffusion process The noisy action sequence corresponding to each step; Indicates the diffusion process The action sequence corresponding to each step; Indicates a Gaussian distribution; Represents the identity matrix; Indicates the first The noise scheduling coefficients in the diffusion process are used to control the proportion of original action information retained and the noise injection intensity in this step. By setting the noise intensity through a linear scheduling strategy, the action sequence approximately follows a simple Gaussian distribution at the end of the diffusion, thereby reducing the difficulty of reverse modeling.
[0064] Then, in the reverse generation stage, a Transformer-based denoising network is constructed to fuse conditional features. Guided by this, the diffusion process is progressively inverted. To simplify the modeling process, a noise prediction paradigm is adopted, transforming the inverse denoising process into a regression problem on the noise term:
[0065] in, This represents the noise estimate obtained from the denoising network prediction. The parameter is The denoising network is a reverse denoising model within a diffusion strategy network, used to denoise based on the current denoising action sequence. , characteristics of fusion conditions and the current diffusion step The model predicts the noise components included in the current step. During the training of the inverse denoising model, effective learning of the action distribution is achieved by minimizing the mean square error between the real noise and the predicted noise.
[0066] in, This represents the training loss function of the denoising network; This represents the expectation operation, used to perform a statistical average of the training samples and noise samples during the diffusion process; This represents the real Gaussian noise injected into the real action sequence during the forward diffusion process; This represents the noise estimate obtained from the denoising network prediction. This represents the square of the L2 norm.
[0067] Multimodal conditional features continuously participate in the modeling process throughout the inverse denoising process, enabling the generated actions to gradually converge in a statistical sense to an action space that conforms to the patient's intentions and interaction state.
[0068] Finally, to ensure the executability of the generated actions in a practical rehabilitation robot system, a constraint reinforcement mechanism is further introduced during the action generation process. On one hand, a temporal consistency constraint is applied to the rate of change between adjacent actions:
[0069] in, For the first The single-step control action corresponding to each control cycle. This refers to the single-step control action corresponding to the previous control cycle. This is a preset threshold for action change.
[0070] To limit motion jitter and ensure the smoothness of generated motion; on the other hand, in joint space or Cartesian space control scenarios, robot dynamics constraints are introduced to make the generated motion meet the physical feasibility requirements of the system, thereby improving the safety and reliability of motion sequences in real systems.
[0071] 3. Closed-loop human-machine interaction control.
[0072] In rehabilitation training applications, human-computer interaction control systems not only need to accurately understand the patient's movement intentions but also meet engineering requirements such as real-time response, continuous interaction, and operational safety. Because the patient's movement state changes dynamically over time, and there is a time delay and uncertainty between perceived information and action generation, open-loop control methods are difficult to adapt to actual training needs. Therefore, this invention proposes a closed-loop human-computer interaction control method based on rolling prediction. By periodically updating perceived information and control decisions, adaptive and stable control is achieved during the human-computer interaction process.
[0073] like Figure 3 As shown, step 3 is as follows: First, during the online operation phase, the system operates with a fixed control cycle Δt. Within each control cycle, force / torque signals, surface electromyography signals, and skeletal joint posture signals are synchronously acquired through a multimodal sensing module and processed and pre-processed within a time window of length T. After processing by the multimodal temporal fusion network model, a fusion conditional feature representation for the current moment is generated:
[0074] in, This represents the set of multimodal signals acquired within a time window. This is a multimodal temporal fusion mapping function. The fusion condition features... It comprehensively reflects the patient's current movement intention, muscle activation state, and human-computer interaction force level.
[0075] Secondly, in the action generation stage, conditional features are fused. As a constraint, an initial noisy action sequence is sampled from a standard Gaussian distribution:
[0076] Guided by the diffusion strategy network, the action sequence is gradually recovered through a multi-step reverse denoising process:
[0077] in, This represents the initial noise action sequence of the diffusion reverse generation process; This represents the total number of steps in the diffusion process; Indicates the first The intermediate action sequence corresponding to the reverse denoising process; The parameter is A denoising mapping function based on multimodal conditions. After... After the reverse denoising step, the final recovered short-time action sequence is obtained. This process statistically approximates the movement distribution that aligns with the patient's current condition and rehabilitation goals.
[0078] Then, during the real-time execution phase, the system employs a rolling predictive control strategy, selecting only the short-term action sequence that is ultimately recovered. The first action in the input is used as the current control input:
[0079] This information is then sent to the robot's actuator. In the next control cycle t+Δt, the system re-performs multimodal perception, feature fusion, and motion generation, thus forming a continuous closed-loop control process.
[0080] in, This indicates the system feedback status after the robot interacts with the patient. This rolling update mechanism can effectively suppress the accumulation of prediction errors over time and improve the stability of the system during dynamic interactions.
[0081] Furthermore, to further ensure the safety and smoothness of the closed-loop control process, constraints are imposed on the rate of change between adjacent control inputs:
[0082] in, A preset threshold is used to limit abrupt changes in control commands and avoid causing discomfort or potential risks to the patient. In joint space control scenarios, the generated control inputs must also satisfy robot dynamics constraints:
[0083] in, For the robot at the joint position The inertia matrix below, Let the joint position vector be... The joint acceleration vector, This is a matrix of Coriolis force and centrifugal force terms related to joint position and joint velocity. For joint velocity vectors, The vector of the gravity term. This is the joint driving torque vector output by the closed-loop control strategy. By implicitly satisfying the above constraints during the control process, the physical feasibility and execution safety of the generated actions in the actual robot system are guaranteed.
[0084] Finally, by jointly optimizing the parameters of the multimodal fusion network and the diffusion strategy network, the computation time for perception, fusion, inference, and control within a single control cycle meets the requirements for real-time operation, thereby achieving safe, stable, and continuous closed-loop human-machine interaction control.
[0085] The purpose of this invention is to improve the accuracy of patient movement intention perception, the flexibility of movement generation, and the safety and real-time performance of interactive control in rehabilitation training human-computer interaction systems. By constructing a standardized method for the acquisition, processing, and temporal fusion of multimodal heterogeneous interactive signals, it achieves collaborative perception and unified expression of force / torque signals, surface electromyography signals, and skeletal joint posture signals, solving the heterogeneity and inconsistency issues in sampling frequency, physical dimensions, and temporal characteristics of multimodal interactive data. Based on this, human-computer interaction control is modeled as a random action sequence generation problem under multimodal constraints. A diffusion probability model is introduced to characterize the distribution of complex auxiliary movements, improving the adaptability of movement generation to individual patient differences and diverse interaction needs. Furthermore, by constructing a closed-loop human-computer interaction control mechanism based on rolling prediction, real-time closed-loop collaboration between multimodal perception, movement generation, and robot execution is achieved, ensuring the smoothness, physical feasibility, and stability of generated movements and system operation. This provides a more precise, personalized, safe, and efficient human-computer collaborative control solution for rehabilitation training.
[0086] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A human-computer interaction control method for rehabilitation robots based on multimodal and diffusion strategies, characterized in that, Includes the following steps: S1. Collect force signals, surface electromyography signals, and skeletal joint posture signals, and perform standardized preprocessing to obtain multimodal features; S2. Input the multimodal features into the multimodal temporal fusion network model to generate the fusion conditional features at the current time. S3. Model the distribution of complex actions using a diffusion probability model, and generate action sequences using fusion conditional features as constraints. S4. A rolling predictive control strategy is adopted, which selects only the first action in the generated action sequence as the current control input and collects data for the next control cycle.
2. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 1, characterized in that, Step S1 includes: (1) Collect the interaction force between the patient and the training device. The original force signal at each sampling time is represented as a six-dimensional vector. The original force signal is zero-point calibrated and gravity compensation is performed. The compensated signal is filtered by moving average and the statistical features of mean and standard deviation are extracted within a fixed time window. Then, the features of each channel are normalized by the Z-score normalization method to meet the uniform scale distribution, thereby realizing the standardized representation of the mechanical interaction features. (2) Collect electromyographic signals of the target muscle group, perform bandpass filtering on the electromyographic signals, and perform full-wave rectification and smoothing on the filtered signals to extract the envelope signal reflecting the change in muscle contraction intensity, and calculate the time domain characteristics of the root mean square value and the mean absolute value from the envelope signal. (3) Real-time acquisition of three-dimensional spatial coordinate information of key joints of the human body, calculation of joint motion velocity and joint acceleration through the difference between adjacent frames to characterize the trend of motion state; then, splicing the joint position, velocity and acceleration features to form a skeletal posture feature vector that comprehensively reflects the patient's motion state.
3. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 1, characterized in that, In step S2, the multimodal temporal fusion network model generates the fusion condition features at the current time, including the following steps: Temporal modeling is performed on the feature sequences of different modalities to extract the dynamic evolution characteristics within each modality. An attention mechanism is introduced in the fusion stage to adaptively evaluate the importance of different modalities at different times. The fusion condition features are obtained by weighted summation of the temporal features of each modality.
4. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 1, characterized in that, In step S3, the distribution of complex actions is modeled using a diffusion probability model, including the following steps: First, using fused conditional features as constraints, an action sequence is generated. The control problem is then modeled as a conditional probability distribution of the random variable of the action sequence under the fused conditional features: Secondly, a diffusion probability model is introduced to construct the forward noise addition process. The noise intensity is set by a linear scheduling strategy so that the action sequence approximately follows a simple Gaussian distribution at the end of the diffusion, thereby reducing the difficulty of reverse modeling. Then, in the reverse generation stage, a Transformer-based denoising network is constructed to perform a step-by-step inversion of the diffusion process under the guidance of fusion condition features.
5. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 4, characterized in that, A constraint reinforcement mechanism is further introduced during the action generation process to impose temporal consistency constraints on the rate of change between adjacent actions: in, for The single-step control action corresponding to each control cycle. The preset action change threshold is used to limit the change range between adjacent control actions to ensure the smoothness and safety of the generated action sequence.
6. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 1, characterized in that, In step S4, constraints are applied to the rate of change between adjacent control inputs: in, A preset threshold is used to limit mutations in control instructions, avoiding discomfort or potential risks to patients. This indicates the current control input.
7. The human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies according to claim 1, characterized in that, In joint space control scenarios, the generated control inputs must also satisfy robot dynamics constraints: in, For the robot at the joint position The inertia matrix below, Let the joint position vector be... The joint acceleration vector, This is a matrix of Coriolis force and centrifugal force terms related to joint position and joint velocity. For joint velocity vectors, The vector of the gravity term. This is the joint driving torque vector output by the closed-loop control strategy.
8. A system for implementing the human-computer interaction control method for a rehabilitation robot based on multimodal and diffusion strategies as described in claim 1, characterized in that, include: The multimodal sensing module is used to acquire force signals, surface electromyography signals, and skeletal joint posture signals, and to perform preprocessing. The multimodal temporal fusion module is used to generate the fusion condition features at the current time. The action sequence generation module is used to generate action sequences under multimodal constraints. The closed-loop human-machine interaction module is used to take the action in the generated action sequence as the current control input through a rolling predictive control strategy.
9. A computer device comprising a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, characterized in that: When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor implements the method as described in any one of claims 1 to 8.