Brain-Computer Interface-Based Control Method and System for Exoskeleton Robots
By performing three-level preprocessing on brain-computer interface data and using an integrated learning model to identify movement intentions, combined with dynamic adjustment algorithms to optimize control, the problems of response delay and poor adaptability in exoskeleton robot control have been solved, achieving efficient human-machine collaboration.
Patent Information
- Application Number
- CN202511323969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing exoskeleton robot control methods suffer from high response latency, difficulty in accurately capturing user intentions, susceptibility of EEG signals to environmental noise interference, limited and weak generalization ability of motion intention recognition models, and a lack of effective feedback regulation mechanisms.
By acquiring brain-computer interface data and performing three-level dynamic preprocessing, motion features are extracted using adaptive filtering, wavelet transform, and co-space pattern algorithms. Motion intention recognition is then performed by combining a heterogeneous ensemble learning model, and motion control is optimized through a dual-closed-loop dynamic adjustment algorithm.
It improves signal processing accuracy and intent recognition accuracy, realizes efficient collaboration between the exoskeleton robot and the user's movement intent, solves the problems of response delay and poor adaptability, and significantly improves human-machine collaboration efficiency and control safety.
Smart Images

Figure CN120816499B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, and in particular to a control method and system for exoskeleton robots based on brain-computer interfaces. Background Technology
[0002] Exoskeleton robots, as important devices for assisting human movement, typically consist of mechanical structures, drive systems, sensors, control systems, and power supplies. They can move in tandem with the human body, enhancing or assisting human motor abilities and physiological functions, leading to their widespread application in fields such as medical rehabilitation and assistive devices for people with disabilities. However, traditional control methods often rely on mechanical sensors or manual operation, resulting in problems such as high response latency and difficulty in accurately capturing user intentions.
[0003] Meanwhile, in existing brain-computer interface-based control technologies, EEG signals are easily affected by environmental noise and physiological artifacts, resulting in insufficient feature extraction accuracy; the motion intention recognition model is singular, has weak generalization ability, and is difficult to adapt to the signal characteristics of different users; and there is a lack of effective feedback adjustment mechanism, making it easy for control commands to become disconnected from actual motion needs.
[0004] In summary, there is an urgent need for a control method and system for exoskeleton robots that can improve signal processing robustness, intent recognition accuracy, and dynamic adjustment capabilities.
[0005] Therefore, how to provide control methods and systems for exoskeleton robots based on brain-computer interfaces is an urgent problem to be solved. Summary of the Invention
[0006] This invention provides a brain-computer interface-based exoskeleton robot control method and system to solve the aforementioned technical problems in the prior art.
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or to describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0008] According to a first aspect of the present invention, a brain-computer interface-based exoskeleton robot control method is provided.
[0009] In one embodiment, a brain-computer interface-based exoskeleton robot control method includes:
[0010] Acquire brain-computer interface data, preprocess the brain-computer interface data to obtain preprocessed brain-computer interface data, and extract motion features using the preprocessed brain-computer interface data to obtain motion feature vectors;
[0011] An ensemble learning model is used to identify motion intent from motion feature vectors, motion control commands are generated based on the motion intent identification results, and motion control is performed according to the motion control commands.
[0012] Based on the acquired motion state feedback data, a dynamic adjustment algorithm is used to adjust the motion state, and the motion control commands are dynamically adjusted according to the motion state adjustment results. The motion control is then optimized using the adjusted motion control commands.
[0013] According to a second aspect of the present invention, a brain-computer interface-based exoskeleton robot control system is provided.
[0014] In one embodiment, the brain-computer interface-based exoskeleton robot control system includes: a motion feature extraction module, a motion control module, and a motion control optimization module;
[0015] The motion feature extraction module is used to acquire brain-computer interface data, preprocess the brain-computer interface data to obtain preprocessed brain-computer interface data, and extract motion features using the preprocessed brain-computer interface data to obtain motion feature vectors.
[0016] The motion control module is used to identify motion intentions from motion feature vectors using an ensemble learning model, generate motion control commands based on the motion intention recognition results, and perform motion control according to the motion control commands.
[0017] The motion control optimization module is used to adjust the motion state based on the acquired motion state feedback data using a dynamic adjustment algorithm, and dynamically adjust the motion control commands according to the motion state adjustment results, thereby optimizing motion control through the adjusted motion control commands.
[0018] According to a third aspect of the present invention, a computer device is provided.
[0019] In some embodiments, the computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the brain-computer interface-based exoskeleton robot control method described above.
[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided.
[0021] In one embodiment, a computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the brain-computer interface-based exoskeleton robot control method described above.
[0022] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0023] 1. This invention extracts motion features to avoid interference from environmental noise and physiological artifacts in the extracted EEG signals, ensuring the accuracy of feature extraction. Furthermore, by using an integrated learning model for motion intention recognition, it avoids the problems of single and weak generalization ability of motion intention recognition models, thus adapting to the signal features of different users. Finally, by using a dynamic adjustment algorithm to adjust the motion state, it achieves effective feedback regulation, avoiding the disconnect between control commands and actual motion needs.
[0024] 2. This invention uses motion feature vectors to recognize motion intentions, solving problems such as low signal processing accuracy, insufficient accuracy of intention recognition, and rigid control strategies in existing technologies. It achieves efficient coordination between the exoskeleton robot and the user's motion intentions, avoiding problems such as high response delay and difficulty in accurately capturing user intentions in traditional control methods.
[0025] 3. The data preprocessing and integrated learning model of this invention improves signal processing accuracy and intent recognition accuracy. At the same time, it combines motion state feedback data for dynamic adjustment, which solves the problems of response delay and poor adaptability in traditional exoskeleton robot control. It significantly improves human-machine collaboration efficiency and control safety, and is applicable to various scenarios such as medical rehabilitation and assistive movement.
[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0028] Figure 1 This is a flowchart illustrating a brain-computer interface-based exoskeleton robot control method according to an exemplary embodiment.
[0029] Figure 2 This is a schematic diagram of the system structure according to an exemplary embodiment;
[0030] Figure 3 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment;
[0031] Figure 4 This is an overall framework diagram of a brain-computer interface-based exoskeleton robot control method according to an exemplary embodiment;
[0032] Figure 5 This is a schematic diagram illustrating the structure of an integrated learning model in a brain-computer interface-based exoskeleton robot control method according to an exemplary embodiment.
[0033] Figure label:
[0034] 201. Motion Feature Extraction Module; 202. Motion Control Module; 203. Motion Control Optimization Module. Detailed Implementation
[0035] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some embodiments may include or substitute parts and features of other embodiments. The scope of the embodiments herein encompasses the entire scope of the claims and all available equivalents thereof. Throughout this document, the terms “first,” “second,” etc., are used only to distinguish one element from another without requiring or implying any actual relationship or order between the elements. Indeed, a first element can also be referred to as a second element, and vice versa. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a structure, apparatus, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a structure, apparatus, or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the structure, apparatus, or device that includes said element. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0036] The terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" used in this document to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings. They are used solely for the convenience of describing the document and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description herein, unless otherwise specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two elements; they can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0037] In this document, unless otherwise stated, the term "multiple" means two or more.
[0038] In this text, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0039] In this article, the term "and / or" describes the relationship between objects, indicating that there can be three relationships. For example, A and / or B means: A or B, or A and B.
[0040] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0041] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0042] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0043] Figure 1 An embodiment of the brain-computer interface-based exoskeleton robot control method of the present invention is shown.
[0044] In this optional embodiment, the brain-computer interface-based exoskeleton robot control method includes:
[0045] Step S101: Obtain brain-computer interface data, preprocess the brain-computer interface data to obtain preprocessed brain-computer interface data, and extract motion features using the preprocessed brain-computer interface data to obtain motion feature vectors.
[0046] Step S102: Use an ensemble learning model to identify motion intentions from motion feature vectors, generate motion control commands based on the motion intention identification results, and perform motion control according to the motion control commands.
[0047] Step S103: Based on the acquired motion state feedback data, the motion state is adjusted using a dynamic adjustment algorithm, and the motion control command is dynamically adjusted according to the motion state adjustment result. The motion control is then optimized using the adjusted motion control command.
[0048] It should be further explained that the brain-computer interface data is acquired and then subjected to a three-level dynamic preprocessing process, namely, firstly through adaptive filtering, where the step size factor... q The data is dynamically adjusted based on the proportion of power frequency interference, with a value range of 0.01-0.1. Power frequency interference and electromyographic noise are removed, and then wavelet transform is applied. The number of decomposition layers is dynamically adjusted based on the proportion of electrooculogram artifacts, with a value range of 5-7 layers. Data decomposition and frequency band filtering are then performed. Finally, the common spatial pattern (CSP) algorithm is used to maximize the difference in motion intent to extract motion features, resulting in a motion feature vector. A heterogeneous ensemble learning model (i.e., an ensemble learning model) is used to identify motion intent from the motion feature vector. This heterogeneous ensemble learning model includes an optimized convolutional neural network model, an optimized support vector machine model, and an optimized random forest model. The optimized convolutional neural network model... The network model contains 3 convolutional layers and 2 residual connections. The optimized support vector machine model uses a radial basis function kernel, with the kernel parameter γ1 dynamically adjusted from 0.1 to 1. The optimized random forest model contains 100 decision trees and employs feature random subspace sampling. Dynamic weighted voting is used, where the weights are allocated in real-time based on the sub-model confidence, with an initial weight ratio of 0.35:0.3:0.35. The model outputs motion intent recognition results, generates motion control commands based on these results, and executes the motion control. Based on the acquired motion state feedback data, a dual-loop dynamic adjustment algorithm is used to adjust the motion state; specifically, the inner loop uses a reinforcement learning algorithm to update the state value function formula. V '( s )= V ( s )+ α [ r + γV ( s ')- V ( s To adjust the joint angle deviation, the outer ring uses an impedance control algorithm, the formula of which is: This allows for adjustments to the human-computer interaction force, and dynamic optimization of motion control commands based on the adjustment results.
[0049] In this optional embodiment, brain-computer interface (BCI) data is acquired, preprocessed to obtain preprocessed BCI data, and motion features are extracted using the preprocessed BCI data to obtain a motion feature vector, including:
[0050] Brain-computer interface data is acquired, and adaptive filtering is used to remove power frequency interference and electromyographic noise from the brain-computer interface data to obtain denoised brain-computer interface data.
[0051] Wavelet transform is used to decompose the denoised brain-computer interface data, and the decomposed brain-computer interface data is combined with preset frequency bands for data filtering to obtain preprocessed brain-computer interface data.
[0052] Based on the common space pattern algorithm, motion feature vectors are obtained by extracting motion features that maximize the differences in motion intentions from the preprocessed brain-computer interface data.
[0053] In this optional embodiment, motion feature extraction that maximizes the difference in motion intent is performed on the preprocessed brain-computer interface data based on the common space pattern algorithm, resulting in a motion feature vector including:
[0054] The covariance of motor intention is calculated based on the preprocessed brain-computer interface data to obtain the motor intention covariance matrix. The spatial filter matrix is then solved to obtain the motor intention spatial filter matrix.
[0055] The motion intent spatial filter matrix is optimized using a pre-constructed objective function, and the variance difference of different motion intent signals is maximized based on the optimization results, thus obtaining the variance difference results of different motion intent signals.
[0056] Motion features are extracted based on the variance differences of signals with different motion intentions, and motion feature vectors are generated from the extracted motion features.
[0057] It should be further explained that adaptive filtering is used to remove power frequency interference and electromyographic noise from the brain-computer interface data, that is, the power ratio of power frequency interference in the brain-computer interface data is calculated in real time. P 工频 ,like P 工频 If the percentage is greater than 30%, then adjust the step size factor of the adaptive filter. q =0.01+0.09×( P 工频 -30%) / 20%, and its value ranges from 0.01 to 0.1, and is expressed by the formula w ( n+1 ) =w ( n ) +2qe ( n ) x ( n Update the filter weights to make the power frequency interference suppression ratio ≥45dB.
[0058] It should be further explained that wavelet transform is used to perform data decomposition processing on the denoised brain-computer interface data, that is, to calculate the power ratio of electrooculography artifacts in the brain-computer interface data in real time. P 眼电 ,like P 眼电 If the percentage is greater than 20%, then the wavelet decomposition level will be increased from 5 levels to 5 + 2 × ( P 眼电-20%) / 10%, with a value range of 5-7 layers, and a modal entropy threshold is introduced. T entropy =0.8-0.2× P 眼电 / 50%, threshold filtering is performed on the electrooculogram (EOG) signals in the 3-8Hz frequency band to achieve an EOG artifact removal rate of ≥96%.
[0059] In this optional embodiment, the motion intention recognition is performed on the motion feature vector using an ensemble learning model, motion control commands are generated based on the motion intention recognition results, and motion control is performed according to the motion control commands, including:
[0060] An ensemble learning model is constructed and trained using pre-acquired brain-computer interface data to obtain a motion intention recognition model.
[0061] The motion intent recognition model is used to identify the motion intent of the motion feature vector by combining the preset motion intent categories, and the motion intent recognition result is obtained.
[0062] The motion intent recognition results are verified for confidence level. Based on the confidence level verification results, motion parameters are calculated by combining the preset motion parameter mapping rules and the pre-acquired human motion data. Motion control commands are generated according to the motion parameters and motion control is performed using the motion control commands.
[0063] In this optional embodiment, an ensemble learning model is constructed and trained using pre-acquired brain-computer interface data to obtain a motion intention recognition model, including:
[0064] Feature extraction is performed on the pre-acquired brain-computer interface data to obtain a brain-computer interface feature dataset, and the brain-computer interface feature dataset is divided into a training set, a validation set, and a test set based on a preset ratio;
[0065] The convolutional neural network model, support vector machine model, and random forest model were trained and optimized using the training set to obtain optimized convolutional neural network model, optimized support vector machine model, and optimized random forest model.
[0066] The optimized convolutional neural network model, the optimized support vector machine model, and the optimized random forest model are integrated in parallel to obtain the motion intention recognition model.
[0067] It should be further explained that the optimized convolutional neural network model, optimized support vector machine model, and optimized random forest model are integrated in parallel to obtain the motion intent recognition model. That is, the accuracy of each sub-model on the validation set is updated every 500 samples. Through formula ; Calculate the dynamic weights, where This makes the weights of the sub-models positively correlated with the recognition accuracy.
[0068] In this optional embodiment, the convolutional neural network model, support vector machine model, and random forest model are trained and optimized using the training set to obtain optimized convolutional neural network model, optimized support vector machine model, and optimized random forest model, including:
[0069] The convolutional neural network model is initialized using a normal distribution to obtain an initial convolutional neural network model. The initial convolutional neural network model is then iteratively trained using the training set, and the convolutional neural network model is updated using the cross-entropy loss function and backpropagation to obtain an optimized convolutional neural network model.
[0070] A support vector machine model with motion intent classification is constructed by kernel function selection and multi-classification strategy. The support vector machine model is trained based on the training set, and the objective function of the model is optimized by solving a convex quadratic programming problem to obtain an optimized support vector machine model.
[0071] The training set is sampled with replacement to obtain a training subset, and a decision tree is constructed using classification and regression tree algorithms. The splitting features and thresholds of the optimized decision tree are selected based on the Gini impurity. The optimized decision tree is then combined with the training subset to train a random forest model, resulting in an optimized random forest model.
[0072] In this optional embodiment, motion intent recognition is performed on the motion feature vector by combining a motion intent recognition model with a preset motion intent category, and the resulting motion intent recognition results include:
[0073] The convolutional structure based on the motion intent recognition model, combined with the preset motion intent category, performs temporally correlated motion intent recognition on the motion feature vector, and obtains the temporally correlated motion intent classification probability.
[0074] By using the radial basis function of the motion intent recognition model in combination with the preset motion intent categories, the motion feature vector is statistically identified to obtain the motion intent classification probability with statistical regularity.
[0075] By combining the decision tree of the motion intent recognition model with the preset motion intent categories, motion intent with boundary features is recognized from the motion feature vector, and the motion intent classification probability with boundary features is obtained.
[0076] The motion intent classification probabilities with temporal correlation, statistical regularity, and boundary features are integrated using a weighted voting method to obtain the motion intent recognition result.
[0077] It should be further explained that the weighted voting method for integrating classification probabilities also includes conflict resolution of sub-model probabilities. If the highest probability category of each of the three sub-models is different, the difference between the highest and second-highest probabilities of each model is calculated. Δp = p 1- p 2. Prioritize adoption Δp The largest model result, of which Δp A result with a confidence level ≥ 0.5 is considered a high-confidence result and is output directly. Δp If the value is less than 0.5, a secondary identification is triggered, and the sample collection time is extended to three seconds before recalculation.
[0078] In this optional embodiment, based on the acquired motion state feedback data, a dynamic adjustment algorithm is used to adjust the motion state, and the motion control commands are dynamically adjusted according to the motion state adjustment results. Motion control optimization is performed using the adjusted motion control commands, including:
[0079] The motion deviation is calculated from the acquired motion state feedback data, and the motion adjustment judgment is made by combining the motion deviation calculation result with a preset threshold to obtain the motion adjustment judgment result.
[0080] Based on the motion adjustment judgment results, a dynamic adjustment algorithm is used to adjust the motion state parameters of the acquired motion state feedback data to obtain the motion state adjustment results.
[0081] The motion control commands are dynamically adjusted based on the results of the motion state adjustment, and the motion control is optimized using the adjusted motion control commands.
[0082] It should be further explained that before adjusting the motion state based on the acquired motion state feedback data, user adaptation is also included. This involves collecting three minutes of motion imagery signals from new users and aligning the new user's feature sequence with a preset standardized template using the Dynamic Time Warping (DTW) algorithm. This means aligning it with an offline database containing more than 200 user features and calculating sequence similarity. Sim ,like Sim If the value is ≥0.75, then weighted transfer learning is used, making the weights and... Sim A positive correlation is established, and the parameters of the ensemble learning model are initialized, with an adaptation time of ≤10 minutes.
[0083] In this optional embodiment, based on the motion adjustment judgment result, a dynamic adjustment algorithm is used to adjust the motion state parameters of the acquired motion state feedback data to obtain the motion state adjustment result, including:
[0084] Based on the motion adjustment judgment results, the state value function and parameters of the dynamic adjustment algorithm are initialized to obtain the initialized dynamic adjustment algorithm;
[0085] The initial dynamic adjustment algorithm is combined with the acquired motion state feedback data to select actions based on an exploration probability greedy strategy, and the action to be executed is obtained.
[0086] The motion state adjustment parameters are updated by performing actions, and motion state feedback data of the performed actions is obtained. Real-time rewards are calculated based on the motion state feedback data of the performed actions, and the state value function is updated using the real-time reward calculation results.
[0087] The motion state is adjusted by using the updated state value function and the value iteration optimization strategy, and the motion state adjustment result is obtained.
[0088] It should be noted that the dynamic adjustment algorithm for motion state parameter adjustment also includes scene adaptation logic. If an outdoor scene is detected, i.e., the ambient noise is ≥65dB, the exploration probability of the algorithm will be dynamically adjusted. e The value decreased from 0.2 to 0.05 using the formula. Correct the joint linkage delay time to make the motion synchronization error ≤220ms.
[0089] Figure 2 An embodiment of the brain-computer interface-based exoskeleton robot control system of the present invention is shown.
[0090] In this optional embodiment, the brain-computer interface-based exoskeleton robot control system includes: a motion feature extraction module 201, a motion control module 202, and a motion control optimization module 203.
[0091] The motion feature extraction module 201 is used to acquire brain-computer interface data, preprocess the brain-computer interface data to obtain preprocessed brain-computer interface data, and extract motion features using the preprocessed brain-computer interface data to obtain motion feature vectors.
[0092] The motion control module 202 is used to perform motion intention recognition on motion feature vectors using an ensemble learning model, generate motion control commands based on the motion intention recognition results, and perform motion control according to the motion control commands.
[0093] The motion control optimization module 203 is used to adjust the motion state based on the acquired motion state feedback data using a dynamic adjustment algorithm, and dynamically adjust the motion control commands according to the motion state adjustment results, and optimize the motion control through the adjusted motion control commands.
[0094] It should be further explained that the motion feature extraction module is used to acquire brain-computer interface data and perform adaptive filtering based on dynamic adjustment of step size factor, as well as wavelet transform and common spatial pattern (CSP) feature extraction based on dynamic adjustment of layer number through a three-level dynamic preprocessing unit, outputting motion feature vectors; the motion control module includes a heterogeneous integrated learning unit and an instruction generation unit. The heterogeneous integrated learning unit contains a convolutional neural network submodule with 3 layers of convolution and 2 layers of residuals, a support vector machine submodule with radial basis function kernel function γ1=0.1-1, and a random forest submodule based on 100 decision trees. Through a dynamic weighted voting unit, where the weights are allocated according to the confidence of the sub-models, the module outputs the motion intention recognition result; the instruction generation unit is based on the recognition result and a preset motion parameter mapping rule, where the mapping rule includes user height and weight correction formulas. The system generates motion control commands; the motion control optimization module includes a dual closed-loop adjustment unit, namely, the inner loop adjusts the joint angle deviation through the state value function update logic of the reinforcement learning sub-unit, and the outer loop adjusts the human-machine interaction force through the impedance control sub-unit, and dynamically optimizes the motion control commands based on the adjustment results.
[0095] It should be further explained that brain-computer interface data includes raw EEG signals. The user's raw EEG signals are acquired through EEG acquisition equipment. These signals undergo preprocessing to remove noise and artifacts. From the preprocessed EEG signals, feature vectors related to motor intention are extracted to obtain motor feature vectors. EEG signal acquisition uses dry electrode EEG acquisition equipment, with electrodes laid out according to the international 10-20 system. The focus is on acquiring EEG signals from the C3 and C4 areas of the motor cortex and the Fz area of the prefrontal cortex. The sampling frequency is set to 250Hz to ensure signal coverage of motor-related areas. m Waves in the 8-13Hz frequency band and β The wave operates in the 13-30Hz frequency band. Signal preprocessing, i.e., through adaptive filtering formulas: y ( n ) =w T ( n ) x ( n ); e ( n ) =d ( n ) -y ( n ); w ( n+1 ) =w ( n ) +2qe ( n ) x ( n In the formula, y ( n () indicates the filtered output;w ( n ) represents the weight vector; x ( n ) represents the input signal vector; e ( n ) represents the error signal; d ( n ) represents the desired signal; q This represents the step size factor, with a value range of 0.01-0.1; w ( n+1 The ) represents the updated filter weights; the filter weights are updated in real time using an adaptive filtering formula to remove 50Hz power frequency interference and EMG noise, improving the signal-to-noise ratio to over 30dB after filtering; wavelet transform can use the db4 wavelet basis function to decompose the signal into 5 layers, decomposing the EEG signal into sub-bands in the 8-30Hz frequency band, obtaining the decomposed brain-computer interface data, retaining the sub-band signal in the 3-30Hz frequency band (i.e., the preset frequency band), further suppressing EEG and ECG artifacts, with an artifact removal rate of 92%, obtaining the preprocessed brain-computer interface data, this frequency band covers motor imagery related to... m wave and β Waves are used to effectively extract motion-related features.
[0096] Motion feature extraction employs the Common Spatial Pattern (CSP) algorithm, which optimizes the spatial filter through an objective function (i.e., a pre-constructed objective function) to maximize the variance difference between signals of different motion intentions. For example, for the two types of intentions, "left hand extended" and "right hand clenched fist," the feature vectors extracted by the CSP algorithm can improve the distinguishability between the two types of signals by 40%. The formula for extracting motion feature vectors using the Common Spatial Pattern algorithm is as follows:
[0097] ;
[0098] In the formula, J ( W () represents the objective function; W ∑1 represents the spatial filter matrix; ∑2 represents the signal covariance matrix corresponding to the two types of motion intentions; ∑1 represents the spatial filter matrix; ∑2 represents the spatial filter matrix; ∑1 represents the signal covariance matrix corresponding to the two types of motion intentions; tr This represents the matrix trace operation.
[0099] It should be further explained that an ensemble learning model is used to classify feature vectors, identify the user's motion intentions, generate control commands for the exoskeleton robot based on the motion intentions, and send them to the actuators. An ensemble learning model is constructed, comprising a Convolutional Neural Network (CNN), a Support Vector Machine (SVM), and a Random Forest (RF) model. The CNN uses a 3-layer convolutional structure and extracts local signal features using 3×3 convolutional kernels; the SVM uses radial basis function kernels to classify high-dimensional features; and the RF contains 100 decision trees. The motion intention recognition result is determined by integrating the results through a voting method, as shown in the formula:
[0100] ;
[0101] In the formula, Γ represents the integrated voting result of the final motion intention recognition, and the output is the index of the specific motion intention category, such as "elbow extension" and "elbow flexion"; v i ( k ) indicates the first i The classifier for the th classifier k Voting value for motion intent; argmax k This indicates the index corresponding to the maximum value.
[0102] A multi-source EEG signal dataset (i.e., pre-acquired brain-computer interface data) was collected, comprising at least 200 individuals of different ages (18-65 years old) and varying motor abilities (healthy individuals, patients with mild to moderate motor impairments, and patients with moderate motor impairments). For each individual, EEG signals were collected for six basic motor intentions, such as raising an arm, raising a leg, bending over, clenching a fist, extending, and remaining still. 100 valid samples were collected for each intention, with each sample lasting 2 seconds and a sampling frequency of 250Hz. After preprocessing, a 128-dimensional feature vector (i.e., the brain-computer interface feature dataset) was extracted. The dataset was divided into training, validation, and test sets in a 7:2:1 ratio. The training set was used for model parameter learning, the validation set for hyperparameter tuning, and the test set for evaluating the final model performance. The training set was augmented with a ±10ms time axis shift and a ±5% amplitude scaling to increase the sample size to 1.5 times the original size, thus preventing overfitting. A parallel ensemble architecture is adopted, with three sub-models—CNN, SVM, and RF—processing the input 128-dimensional feature vector and outputting their respective motion intent classification probabilities, i.e., 6-dimensional vectors corresponding to the probabilities of 6 types of intents. Finally, a weighted voting method is used to fuse these probabilities to obtain the final classification result (i.e., the motion intent recognition result). Weight allocation is dynamically determined based on the accuracy of each sub-model on the validation set. Initially, all weights are set to 1, and during training, they are adjusted every 5 epochs based on the validation set accuracy, using the following formula:
[0103] ;
[0104] In the formula, w i Indicates the first i The weights of each sub-model; acc i Indicates the first i The accuracy of each sub-model on the validation set; acc j Indicates the first j The accuracy of each sub-model on the validation set is measured. The total training epochs are set to 50 epochs, with each epoch containing a complete traversal of the training set. If the validation set accuracy does not improve for 10 consecutive epochs and the fluctuation is less than 0.5%, training is terminated early, and the current optimal model parameters are saved. The CNN model construction includes an input layer that receives 128-dimensional EEG feature vectors, which are transformed into a 16×8 two-dimensional feature matrix through a reshape operation to simulate spatial distribution; Convolutional Layer 1 uses 3×3 convolutional kernels (32 kernels), stride 1, "same" padding, ReLU activation function, and output feature map size 16×8×32; Pooling Layer 1 uses 2×2 max pooling, stride 2, and output feature map size 8×4×32; Convolutional Layer 2 uses 3×3 convolutional kernels (64 kernels), stride 1, "same" padding, ReLU activation function, and output feature map size 8×4×32. The first layer has a size of 8×4×64; the second pooling layer uses 2×2 max pooling with a stride of 2, resulting in an output feature map size of 4×2×64; the first fully connected layer flattens the features output from the pooling layer into a 512-dimensional vector, connects it to a 512-dimensional fully connected layer, uses ReLU activation, and sets the dropout rate to 0.3; the output layer is a 6-dimensional fully connected layer with softmax activation, outputting the probability distribution of 6 types of motion intentions; training parameters include using the Adam optimizer with an initial learning rate of 0.001, which decays to 0.5 times every 10 epochs; the cross-entropy loss function is:
[0105] ;
[0106] In the formula, L This represents the cross-entropy loss function value, which measures the difference between the model's prediction and the true label. The smaller the value, the more accurate the prediction. y i The first one representing the true label i Each component uses one-hot encoding, such as "1" to indicate that the category is a true value and "0" to indicate that it is not a true value; The model represents the first iThe output probabilities for each category range from [0,1], reflecting the model's confidence in predicting that category. The batch size is 32, meaning 32 samples are input for parameter updates each time. Model parameters are initialized using a He normal distribution. Iterative training is performed on the training set, calculating the loss value in each iteration and updating the model parameters through backpropagation until a preset number of training rounds is reached or an early stopping mechanism is triggered. The model is saved, specifically the model parameters with the highest accuracy on the validation set, as the optimized convolutional neural network model.
[0107] The kernel function chosen for constructing the SVM model is the radial basis function (RBF), and the formula is as follows:
[0108] ;
[0109] In the formula, K ( u , y The value represents the output of the radial basis function kernel, which measures the similarity between two samples. The larger the value, the higher the similarity. u This represents the current input sample feature vector, such as EEG signal features, EMG signal features, etc., with the same dimension as the sample feature space; y This represents the feature vector of the reference sample in the support vector set, i.e., the support vector selected during the training of the SVM model; g This represents the kernel function parameter, also known as the "bandwidth parameter," which controls the range of the radial basis function. The optimal value is determined through grid search. u - y || 2 This represents the squared Euclidean distance between two sample feature vectors, measuring the degree of difference between the samples in the original feature space. The multi-classification strategy employs a "one-vs-one" approach, constructing 15 binary SVM sub-models for 6 classes of intent, with the final multi-class result determined by voting. Parameter selection includes the grid search range and the penalty coefficient. C The search range for 1 is [0.1, 1, 10, 100]. g The search scope is
[0110] [0.001,0.01,0.1,1], the optimal parameter combination was determined on the validation set using 5-fold cross-validation (experimentally, the optimal combination was found to be [0.001,0.01,0.1,1]). C 1 = 10 g =0.1). The 128-dimensional feature vector of the input is standardized so that the mean of each feature is 0 and the standard deviation is 1. The formula is:
[0111] ;
[0112] In the formula, x' represents the standardized feature vector, which has the same 128 dimensions as the original feature vector, and has a mean of 0 and a standard deviation of 1; f The feature vector representing the original input is specifically 128-dimensional, such as multimodal features extracted from EEG or EMG signals; m The mean of the original feature vector is calculated using the training set. ,in, N For the sample size, f i Indicates the first i One original feature vector; s Let represent the feature standard deviation. The SVM model is trained using the training set, and the objective function is optimized by solving a convex quadratic programming problem. The objective function formula is:
[0113] ;
[0114] In the formula, N This represents the total number of training samples, i.e., the number of samples used in model training. w The weight vector represents the classification hyperplane, which determines the direction and tilt of the hyperplane, and its dimension is consistent with that of the feature space; x i Indicates the first i The non-negative slack variables for each sample are used to allow some samples to not meet the strict classification constraints, i.e., soft-margin SVM. The larger the value, the further the sample deviates from the ideal classification boundary. C 1 represents the penalty coefficient, and the optimal value has been determined to be 10 through grid search, which is used to balance "maximizing the classification margin" and "misclassification penalty". C The larger the value of 1, the heavier the penalty for misclassified samples, meaning the model tends to classify more strictly and may overfit. C A smaller value of 1 allows for more misclassifications, meaning the model focuses more on maximizing the margin and may underfit. Evaluate the model's performance on the validation set, adjust the parameters, and save the optimized support vector machine model.
[0115] The constraints are:
[0116] ;
[0117] In the formula, h i Indicates the first i The label of a sample is usually +1 or -1 in binary classification, representing the two categories respectively; x i Indicates the first i The feature vectors of each training sample, with standardized 128-dimensional features; bThe bias term, representing the classification hyperplane, determines the hyperplane's position in the feature space; it defines the hyperplane together with the weight vector. The core of this objective function is to find a balance between maximizing the classification margin and minimizing the classification error. This represents maximizing the classification margin, where the margin size is equal to || w || 2 Inversely proportional; This represents minimizing the total slack variable, i.e., minimizing the penalty for classification errors. The constraints ensure that samples are classified correctly as much as possible, while allowing a small number of samples to be tolerated through slack variables. By solving this convex quadratic programming problem, the optimal weight vector can be obtained. w and bias b This allows us to determine the optimal classification hyperplane.
[0118] The construction of the RF model involves determining the number of decision trees, specifically 100, through a validation set. This demonstrates that exceeding 100 decision trees does not significantly improve model performance. Parameters for a single decision tree include a maximum depth of 10 to avoid overfitting, and a random number of features selected during each node split. The minimum number of samples per leaf node is 5. Bootstrap sampling is used to sample the training set with replacement, generating an independent training subset for each decision tree, with the same sample size as the original training set. Each decision tree is constructed using the Classification and Regression Tree (CART) algorithm, selecting the optimal splitting feature and splitting threshold based on Gini impurity. The Gini impurity formula is:
[0119] ;
[0120] In the formula, G Gini impurity quantifies the degree of disorder in the sample classes within a node. The highest purity is achieved when all samples in a node belong to the same class. G =0; if the samples in the node are uniformly distributed in K The category with the highest disorder is... G =1-1 / K For example, the maximum value in binary classification is 0.5. r i Indicates the first node i The proportion of samples belonging to class 1 in this node. i The ratio of the number of samples in a class to the total number of samples in the nodes satisfies ,in TThe total number of sample categories. 100 decision trees classify the input samples, and the output of the RF model is determined by majority voting. The contribution of each feature to all decision tree splits is calculated, i.e., the sum of the reductions in Gini impurity by that feature. Feature importance is evaluated, and features with a contribution of less than 0.1% are removed to simplify the model and improve generalization ability, thus optimizing the model. Post-pruning is performed on each decision tree to remove branches that do not improve the accuracy on the validation set, reducing model complexity and obtaining an optimized random forest model. Through the above specific steps, the ensemble learning model and its sub-models can be completely constructed and trained, ensuring model reproducibility and avoiding the problem of insufficient patent disclosure. Experimental verification shows that the model trained according to the above steps achieves an accuracy of 93.5% on the test set, meeting the actual needs of exoskeleton robot control.
[0121] Testing showed that the model achieved a 93.5% accuracy rate in recognizing six basic motor intentions, such as raising hands, raising legs, and bending over, with a response time of ≤200ms. The specific testing hardware environment included an EEG acquisition device (a 32-channel dry electrode EEG machine with a sampling frequency of 250Hz, input impedance ≥100MΩ, and noise level ≤1μV); computing equipment included a workstation equipped with an Intel Xeon E5-2690 processor and an NVIDIA Tesla V100 graphics card with 32GB of video memory, ensuring a model inference latency of ≤200ms; and the exoskeleton robot's lower limb exoskeleton had a hip and knee joint range of motion of ±90° and a torque sensor accuracy of ±0.1N·m. The test dataset contains 200 participants: 120 healthy individuals, 50 individuals with mild movement disorders, and 30 individuals with moderate movement disorders. The male-to-female ratio is 1:1, and the participants' ages range from 18 to 65 years, with an average age of 38.5 years. For each participant, six types of movement intention samples were collected: raising the hand (shoulder flexion), raising the leg (hip flexion), bending over (lumbar flexion), clenching the fist (finger flexion), extending the arm (full body extension), and remaining still. 100 valid samples were collected for each intention type, excluding samples severely affected by blinking or electromyography (EMG). The sample features were preprocessed into 128-dimensional feature vectors, including temporal features such as mean, variance, and peak value (32 dimensions), frequency domain features of the 8-30Hz power spectrum (64 dimensions), and spatial domain features such as C3 / C4 electrode coherence (32 dimensions). Load the CNN, SVM, and RF weight files and the mean and standard deviation matrices from the standardized parameters in the trained ensemble model; initialize the exoskeleton control interface, set the communication baud rate to 115200bps, and ensure that the command transmission latency is ≤50ms. The single-sample testing process includes the test subject wearing an EEG device, performing a specified movement intention such as "raising a leg" according to instructions, and collecting 2 seconds of EEG signals; then the signal is processed by a step length factor. qAn adaptive filter with a resolution of 0.05 and a db4 wavelet transform are used to decompose the signal to 8-30Hz. A 128-dimensional feature vector is then extracted, standardized, and input into the ensemble model. Finally, CNN, SVM, and RF output classification probabilities, which are weighted according to a weighted voting method (e.g., CNN weight 0.35, SVM weight 0.3, RF weight 0.35) and fused to output the final intent (i.e., motion intent recognition result). The matching between the model's recognition result and the actual intent is recorded, and the single-sample recognition accuracy is calculated. Batch testing involves 200 people × 6 classes × 10 groups = 12000 samples. Accuracy is calculated separately for each intent class and each tester. The overall accuracy is calculated as the number of correctly identified samples / total number of samples; the category accuracy is calculated as the number of correctly identified intents / total number of samples in that class; and a 6×6 confusion matrix is used, with elements (…). i , j ) indicates the first i The class was misclassified as the first j The proportion of classes was as follows. The test results showed an overall accuracy of 93.5%, meaning 11,220 out of 12,000 samples were correctly identified; the category accuracy was 95.2% for raising the hand, 94.8% for raising the leg, 92.3% for bending over, 96.1% for clenching the fist, 91.5% for extending the arm, and 92.6% for remaining still; the confusion matrix key values included a misclassification rate of 3.2% for "raising the leg" and "bending over," as both involve trunk movements, and the misclassification rate among other categories was ≤2%; the average inference latency was 185ms, meaning that from signal acquisition to intent output, the real-time control requirement was ≤250ms. See Table 1 for details.
[0122] Table 1. Examples of comparison with single models
[0123] ;
[0124] Table 1's comparative analysis shows that the ensemble model improves accuracy by 6.2%-7.9% through complementarity, especially for patients with moderate movement disorders and poor signal quality, where accuracy is improved by 10.4%-16.6%. This is because SVM is sensitive to noise, while CNN and RF can compensate for this deficiency. Compared with existing technologies, a single CNN achieves 86.2% accuracy, but does not consider patients with movement disorders; the SVM+RF ensemble achieves 89.7% accuracy with an inference latency of 280ms, exceeding real-time requirements; this invention improves accuracy by 3.8%-7.3%, reduces latency by 34%-95ms, and for the first time achieves adaptation to patients with different degrees of movement disorders. The key technical reason for the improved accuracy is the multi-model complementarity mechanism. CNN excels at capturing local temporal correlations of features, such as... mThe instantaneous changes in wave attenuation demonstrate a clear signal characteristic recognition rate of 96.1% for the fine movement of "clenching a fist." RF excels at handling global statistical regularities of high-dimensional features, achieving a stable recognition rate of 92.6% for "stationary" signals. SVM performs exceptionally well when feature boundaries are clear; when combined with the former two, it reduces the misclassification rate from 8.7% to 3.2% for the fuzzy classification problem of "raising a leg" and "bending over." Data augmentation and generalization design simulates the reaction time differences of different test subjects through time axis shifting and amplitude scaling to simulate EEG signal intensity fluctuations, ensuring that the overfitting coefficient of the model on the test set (i.e., training accuracy - test accuracy) is ≤2.3%. For patients with movement disorders, weighted sampling increases the weight of patient samples in the training set by 1.5 times, enhancing the model's ability to recognize weak signals. Dynamic weight optimization is achieved by optimizing the CNN on the validation set at 8-13Hz. m When the wave characteristics are significant, the weight increases to 0.4, and the SVM at 30Hz... β When wave characteristics are clear, the weight is increased to 0.35 to ensure that each model plays its maximum role in its advantageous scenarios. Stability testing results were obtained by testing continuously for 10 days, with 1200 samples repeated daily. The accuracy fluctuation range was 93.5% ± 0.8%, with no significant drift. After replacing three different batches of exoskeleton devices, the accuracy remained at 93.2% due to the decoupling of the model from the hardware interface, verifying robustness. Through the above testing process and comparative analysis, it is clear that the 93.5% accuracy of the integrated model is reproducible and superior to existing technologies, fully meeting the control requirements of exoskeleton robots.
[0125] Based on the identified motion intent, it is mapped to parameters such as joint angles and movement speed of the exoskeleton robot. The mapping rules from motion intent to exoskeleton joint parameters are as follows: for 6 types of motion intent (i.e., preset motion intent categories), a kinematic parameter mapping table (i.e., preset motion parameter mapping rules) is pre-established, which clarifies the target joint corresponding to each type of intent, namely the hip joint, knee joint, shoulder joint, etc., the angle range, the speed range, and the movement sequence. The specific rules are shown in Table 2.
[0126] Table 2 Basic Mapping Table
[0127] ;
[0128] It is important to note that in the table, angles are defined as 0° in a neutral human position, with flexion being positive and extension negative; velocity refers to the angular velocity of joint movement. Dynamic parameter calculations include intent confidence verification, which involves extracting the probability distribution of motion intent output by the ensemble model, and determining the highest probability... P max When ≥0.85, the base parameter corresponding to this intent is used directly; when 0.7≤ P max When the value is less than 0.85, fuzzy correction is initiated, and a weighted calculation is performed using parameters from the second-highest probability intention. The formula is as follows: ,in, P max The confidence level, or probability value, represents the highest probability intention. It ranges from [0,1] and indicates the level of confidence the model has in the "most likely motion intention," such as the highest probability output by an SVM or decision tree. i target This represents the target control parameters after fuzzy correction, such as the joint angles and movement speed of the exoskeleton robot, which are used as the final output control commands. i 1 represents the angle of the highest probability intention, and the corresponding control parameter is the joint angle target value for the "elbow extension" intention. i 2 represents the angle of the second-highest probability intent. The control parameters corresponding to this second-highest probability intent, such as the joint angle target value corresponding to the "elbow flexion" intent, have a probability second only to the highest probability; that is, weights are allocated according to probability. Dynamic joint angle calculation includes adjustments based on user height and weight from pre-acquired human motion data. i base The range is corrected, and the angle is corrected. i adj The corrected formula is: ; ( H Height in cm W (Based on weight in kg, with 170cm and 65kg as baseline values); the upper limit of the angle is adjusted based on the user's exercise ability level (healthy / mild / moderate impairment) from pre-acquired human motion data. That is, healthy people use 100% of the base angle; those with mild impairment use 80% of the base angle (e.g., reducing the leg lift angle from 60° to 48°); and those with moderate impairment use 60% of the base angle (e.g., reducing the leg lift angle to 36°); the final target angle (i.e., the exercise parameter) is then determined. ( k ability (This represents the ability coefficient, 1 / 0.8 / 0.6).
[0129] The dynamic calculation of motion speed is based on the positive correlation between speed and EEG signal intensity. The calculation formula is as follows: ;in, v final The final velocity; v base Basic motion speed; S eeg The normalized intensity of the EEG signal (0-1, calculated using power in the 8-30Hz frequency band), for example, when S eeg =0.8 (signal strength), leg lift speed ;when S eeg =0.3 (weak signal), leg lift speed The upper limit of velocity shall not exceed the maximum value of the basic range, and the lower limit shall not be lower than the minimum value. The determination of motion timing parameters includes correction of joint linkage delay time based on spinal nerve conduction velocity; the correction formula is as follows: ; v nerve The spinal nerve conduction velocity is measured in m / s, with 80 m / s as the baseline. The normal range is 60-120 m / s. t delay This refers to the corrected joint linkage delay time; t base For the initial joint linkage delay time, for example, when v nerve =100m / s (fast conduction), knee joint hysteresis time when raising leg t =50ms × (1 - 0.01 × 20) = 40ms; when v nerve =60m / s (slow conduction), hysteresis time t =50ms × (1 - 0.01 × (-20)) = 60ms. The parameters are then formatted for output, i.e., the calculated values are displayed. i final 、v final 、t delay (i.e., motion parameters) are converted into a command format (i.e., motion control commands) that the exoskeleton control system can recognize. For example, angle commands are 16-bit binary codes (range -128° to +128°, accuracy 0.5°); speed commands are 8-bit binary codes (range 0° / s to 100° / s, accuracy 1° / s); and timing commands are 10-bit binary codes (range 0ms to 1023ms, accuracy 1ms).
[0130] The exoskeleton robot acquires motion state information (i.e., motion state feedback data) through sensors, constructs a feedback mechanism (i.e., a dynamic adjustment algorithm), and dynamically adjusts control commands. The feedback mechanism is implemented using a reinforcement learning algorithm, and the state value function update formula is:
[0131] ;
[0132] In the formula, V '( s ) represents the updated state value function; V ( s () represents the value of the current state; α Indicates the learning rate; r Indicates the reward value; c Indicates the discount factor; s 'Indicates the next state. Execution feedback verification, that is, after the exoskeleton executes the command, the actual angle is collected in real time through the joint encoder. i real Calculate the deviation: e θ =| i final - i real |。When e θ When the temperature exceeds 3°, a secondary correction of the trigger parameter is performed, with the correction amount being: "sign" is the sign function, applied until the deviation is ≤3°. For example, in the parameter mapping process for the intention of "raising leg", the intention is identified as "raising leg". P max =0.92 (≥0.85), call the basic parameters (hip joint 0°-60°, speed 20° / s-40° / s); based on the user's height of 180cm and weight of 70kg, perform angle correction as follows:
[0133] i adj =60°×(1+0.005×10)×(1+0.003×5)=60°×1.05×1.015≈63.9°; The user has a mild disability. k ability =0.8, i final =63.9° × 0.8 ≈ 51.1°; EEG signal intensity S eeg =0.6, speed v final =40° / s × (0.5 + 0.5 × 0.6) = 32° / s; Spinal cord conduction velocity v nerve =90m / s, knee joint hysteresis time t =50ms × (1 - 0.01 × 10) = 45ms; After the exoskeleton executes the command, the actual angle is 50.8°, with a deviation of 0.3° (≤3°), indicating the parameters are valid. Through the above steps, a precise mapping from motion intention to exoskeleton joint parameters is achieved, ensuring a high degree of match between the exoskeleton's movement and the user's intention under different user conditions and movement states. The "leg lift" intention corresponds to a hip joint target angle of 45° and a movement speed of 0.3m / s. The control command is transmitted to the actuator via the CAN bus, with a transmission delay of ≤50ms. Testing shows that the average error of parameter mapping is ≤2.5° (angle) and ≤2° / s (speed), meeting the accuracy requirements of exoskeleton control.
[0134] The exoskeleton acquires information such as joint torques and motion trajectories (i.e., motion state feedback data) through force sensors and an inertial measurement unit (IMU). Reinforcement learning algorithms (i.e., dynamic adjustment algorithms) are then used to dynamically adjust control parameters. If the actual motion angle deviates from the target angle by more than 5°, the control command is updated through a state value function, causing the deviation to quickly converge to within ±2°. The core elements of the reinforcement learning algorithm include the state space (S), which is defined as the feature vector of the exoskeleton's current motion state, containing joint angle deviations. e θ =| i final - i real | (unit: °); Joint angular velocity deviation e v =| v target - v real | (unit: ° / s); e v Indicates the deviation of joint angular velocity, subscript v This represents "Velocity," a non-negative value that reflects the degree of difference between the target angular velocity and the actual angular velocity. The smaller the value, the higher the control precision. v target The target angular velocity of the joint, i.e., the desired angular velocity, is calculated and generated by the control system based on the motion intention, such as based on the fuzzy-corrected target parameters. i target It is derived that; v real This represents the actual angular velocity of the joint, measured in real time by sensors such as encoders and IMUs to determine the robot joint's motion speed; joint torque. t real (Unit: N·m); EEG signal intensity S eeg (Standardized to 0-1); the state vector is represented as s = e θ , e v , t real , S eeg Each dimension is normalized (mapped to the [-1,1] interval); the motion space (A) is defined as the adjustable control parameter increment, including angle correction. Dth (Range -5° to +5°, step size 0.5°); Speed correction amount Δv (Range -5° / s to +5° / s, step size 0.5° / s); Torque correction factor kτ (Range 0.8-1.2, step size 0.05); the action vector is represented as follows a =[ Dth , Δv , k τ The reward function (R), designed based on motion bias and safety, is given by the following formula: ;in, R ( s , a () represents the reward value, which is the state. s and actions a The larger the value of the function, the more actions will be performed in state s. a The better the effect, the more it reflects the overall motion precision, speed control, and safety; s It represents the current state, such as multi-dimensional state vectors including robot joint angles, angular velocities, and patient biosignals; a This refers to the actions performed by the intelligent agent, such as control commands for adjusting robot joint torque and correcting motion trajectory. e θ This indicates the joint angle deviation, which is the difference between the target angle and the actual angle. The smaller the value, the more precise the angle control. e v This indicates the deviation of joint angular velocity; the smaller the value, the smoother the speed control. t real This represents the actual output torque of the robot joint, which is measured by a torque sensor. t max This represents the maximum safe torque of the joint at the current angle, referring to the dynamic boundary in the "Safety Constraint Working Domain," such as... , q 1 indicates the current angle; I ( t real > t max () represents an indicator function, also known as a "characteristic function". t real > t max If the actual torque exceeds the safety threshold, then ;otherwise It is used to punish unsafe behavior; This is an indicator function that indicates when the torque exceeds a safety threshold. t max The value is 1 when the condition is met, and 0 otherwise. e θ ≤2° and e v An additional 5 points are awarded for values ≤2° / s (to encourage convergence to the target range); State value function ( V( s ), which means in the state s The expected long-term cumulative reward obtained by following the current strategy is given by the formula: V ( s )=E[ G t | s t = s ]=E[ R t+1 +γ R t+2 +γ 2 R t+3 +...| s t = s ];in, V ( s ) represents the state value function, or simply "state value", which is the value of a function in the current state. s Under these conditions, the expected long-term cumulative reward that can be obtained by following the optimal strategy is the average value. The larger the value, the more "favorable" the state is. The mean operator, which represents the expected value of a random variable, is used to quantify the average value of long-term cumulative rewards due to the randomness of the environment and actions. G t Indicates from time t The initial accumulated reward, also known as "reward," is the weighted sum of all future rewards. G t = R t+1 +γ R t+2 +γ 2 R t+3 +...; s t = s This represents a conditional expression, namely, "at time t, the system is in state s". s t For a moment t state, s For a specific state; R t+z Indicates time t + z The immediate reward obtained, such as through the reward function mentioned above. R ( s , a The calculated value, z =1,2,… represents the future nth... zThe reward for each step; γ represents the discount factor, set to 0.9, ranging from [0,1], used to "decay" the weight of future rewards. z Follow z As γ increases, it decreases; when γ=0.9, it places greater emphasis on short-term rewards, but also considers long-term benefits, serving as a balance between immediate and future rewards. G t From time t The initial cumulative reward. The initialization phase of reinforcement learning, which dynamically adjusts control parameters, uses a state-value function. V ( s Initialize as a 0 matrix (covering all possible states); learning rate α =0.1, exploration probability (20% probability of randomly selecting an action, 80% probability of selecting the optimal action); Set the maximum number of iterations. T =100 (maximum number of steps in a single adjustment). Real-time adjustment includes state awareness, i.e., data is collected every 50ms via force sensors and IMU to calculate the current state. s t =[ e θ , e v , t real , S eeg ],when e θ When the angle is >5°, the reinforcement learning adjustment mechanism is triggered; action selection is based on... A greedy strategy selects actions based on probability. Randomly select actions a t ; in terms of probability Choose the action that maximizes the immediate reward. a t =argmax a R( s t ,a). a t Indicates at time t The specific action chosen maximizes the immediate reward R( s t a) Determine; for example, when e θ =6° (deviation exceeds standard), prioritize selection Dth =+1° (positive correction) action; execute actions and state transitions, and perform actions through the exoskeleton. a t Update control parameters:
[0135] ;in, i target 'Indicates the updated joint target angle, i.e., the corrected angle command; i target This indicates the joint target angle before the update, i.e., the original angle command; Dth This indicates the angle correction amount, which is the angle adjustment value calculated based on the control strategy or feedback. v target 'Indicates the updated joint target angular velocity, i.e., the corrected velocity command; v target This indicates the joint target angular velocity before the update, i.e., the original velocity command; Δv This indicates the speed correction amount, which is the speed adjustment value calculated based on the control strategy or feedback; t target 'Indicates the updated joint target torque, i.e., the corrected torque command; t target This indicates the target torque of the joint before the update, i.e., the original torque command; k τ This represents the torque correction coefficient, used to proportionally adjust the target torque, such as when calculated based on safety constraints or load changes; it also collects the new state after execution (i.e., motion state feedback data of the executed action), and the new state is... s t+1 =[ e θ ', e v ', t real ', S eeg '];in, s t+1 Indicates time t The new state vector is incremented by 1, which represents the system state after the action is performed. e θ 'Indicates the updated value of the angle deviation, that is, the deviation between the corrected target angle and the actual angle; e v 'Indicates the updated value of the angular velocity deviation, that is, the deviation between the corrected target angular velocity and the actual angular velocity; t real 'This indicates that after the update is performed, the actual output torque of the joint can be measured by the sensor; S eeg 'Indicates the updated EEG signal characteristics, i.e., the processing results of EEG signals collected after the action was performed; immediate reward is calculated based on the new state. R t = R ( s t , at And update the state value function, specifically by substituting the following formula: V ( s t )+ α [ R t +γ V ( s t+1 )- V ( s t )];in, V ( s t () indicates time t state s t The value function, i.e., the estimated value before the update; α The learning rate, ranging from [0,1], controls the update magnitude of the value function, such as... α =0.1 represents a 10% deviation for each update; R t Indicates time t The immediate reward, i.e., based on the reward function R ( s t , a t ) Calculation, reflected in the state s t Execute action a t The effect; V ( s t+1 () indicates time t +1 New Status s t+1 The value function (estimated value); R t +γ V ( s t+1 ) represents the target value, also known as the TD target, and also represents an ideal estimate of the value of the current state; R t +γ V ( s t+1 )- V ( s t The ) represents the time-series difference error, or TD error, which measures the deviation between the current value estimate and the ideal estimate and is used to drive the update of the value function; for example, if V ( s t )=5, R t =8,V ( s t+1 If )=6, then the updated result is 5+0.1×(8+0.9×6-5)=5+0.1×(8+5.4-5)=5.84; if e θ If the temperature is ≤2° and stable for 3 consecutive cycles, the adjustment is terminated; otherwise, return to state awareness until the maximum number of iterations is reached. T The strategy optimization involves refining the strategy through value iteration after every 100 adjustment cycles, making action selection more inclined towards higher-value states. in, Representing state s The optimal strategy under the following conditions, i.e. The asterisk in the code indicates the optimal state, and the output is the state in which the condition is met. s The optimal action to choose is argmax. a Indicates the action a The maximum value operation selects the action that maximizes the value of the subsequent expression from all possible actions. The mathematical expectation operator represents the average value of a random variable, indicating that randomness exists due to environmental feedback. R Indicates the execution of an action a The immediate reward obtained afterward, i.e., through the reward function R ( s , a )calculate; V ( s ') indicates the execution of an action. a The new state after transition s The state value function also represents the long-term value of a new state; s 'Indicates the execution of an action' a The new state after transition, i.e., the symbol of the original state. s Distinguish by using an apostrophe to indicate subsequent states; | s , a This represents a conditional expression, and also indicates the current state. s Next action a Under these conditions, gradually reduce the probability of exploration. (Decrease by 0.02 every 100 iterations, down to a minimum of 0.05) to improve policy stability. State-value function V ( s () is the core basis for action selection, in state s Next, prioritize those that can make V ( s Maximize the action a For example, when V ( s 1) = 10 (corresponding actions) a1), V ( s 2) = 7 (corresponding actions) a 2) When, among which, V ( s 1) Indicates the execution of an action a The new state after 1 s 1. State value function; V ( s 2) Indicates the execution of an action. a 2. The new state after which it was transferred s 2. State value function; the algorithm will choose a 1. To pursue higher long-term rewards. Value update-driven parameter convergence is achieved through continuous updates. V ( s The algorithm can learn from historical adjustment experience, and when a certain action... a It can rapidly reduce deviation ( e θ (from 6° to 2°), corresponding to R t Increase V ( s t The higher the torque, the higher the probability that the action will be selected in similar future states; when the action causes the torque to exceed the limit ( t real > t max ), R t It becomes a negative value. V ( s t If the rate of change of the state-value function decreases, the algorithm will avoid selecting that action again. ΔV = V ( s t+1 )- V ( s t It is positively correlated with the convergence speed of the angle deviation, when ΔV When the deviation is greater than 0, the convergence speed of the bias increases (an average reduction of 0.8° per step); when ΔV When the value is ≤0, the algorithm triggers a policy correction, increasing the exploration probability to find a better action. The deviation adjustment process for the "lifting leg" intention includes an initial state of... e θ =7° (exceeding the standard) e v =3° / s, t real =12 N·m (not exceeding the limit) s =[7,3,12,0.6], V ( s=3; Action selection is selected a =[+1°,+0.5° / s,1.0] (Increases the target angle and increases the speed); the new state is... e θ =5.5° e v =2.2° / s, R t =10 - 2 × 5.5 - 1 × 2.2 = 10 - 11 - 2.2 = -3.2; Value updated to
[0136] The iteration result is after 8 adjustments. e θ =1.8° (converging to within ±2°), at this point V ( s =12 (high-value state). Through the above steps, the reinforcement learning algorithm uses the state value function to achieve dynamic optimization of control parameters, ensuring that the exoskeleton converges quickly when the deviation exceeds the limit. The test results show that the average adjustment time is ≤0.5s, which is significantly better than traditional PID control (1.2s).
[0137] In summary, the multi-level signal preprocessing algorithm significantly improves the quality of EEG signals, providing a reliable data foundation for subsequent recognition; the integrated learning model enhances the accuracy and generalization ability of motor intention recognition, adapting to the signal characteristics of different users; the feedback adjustment mechanism enables dynamic optimization of control commands, reducing motor deviation and improving the safety of human-machine collaboration; and the control response delay is ≤250ms, meeting real-time control requirements and suitable for high-precision scenarios such as rehabilitation training.
[0138] Taking lower limb rehabilitation training for patients with spinal cord injuries as an example, the patient wears an EEG acquisition device with electrodes attached to the C3, C4, and Fz areas. After system initialization, 5 minutes of resting-state EEG signals are collected as a baseline. The patient imagines "raising the left leg," and the device collects the raw EEG signals. The signals are then adaptively filtered to remove 50Hz power frequency interference and wavelet transform is used to remove electrooculogram artifacts. The CSP algorithm extracts feature vectors, which are input into the ensemble learning model. The CNN layer outputs a probability of 0.91 for "raising the left leg," the SVM outputs 0.88, and the RF outputs 0.93. After voting, the intention is confirmed. Control commands are generated, namely, a target hip joint angle of 30° and a target knee joint angle of 20°, driving the exoskeleton to raise the left leg at a speed of 0.2 m / s. The force sensor detects an abnormal hip joint torque (exceeding the preset threshold of 15 N·m) and feeds it back to the control system. Through reinforcement learning, the commands are adjusted to reduce the torque to 12 N·m, and the angle deviation is controlled within 3°. During training, the system updates the model parameters every 10 minutes to adapt to changes in the patient's signal characteristics and ensure long-term control stability. When applied to different users, the Dynamic Time Warping (DTW) algorithm aligns feature sequences, enabling the model to quickly adapt to new users, reducing the adaptation time to less than 10 minutes. The DTW algorithm aligns feature sequences by defining the sequence of features, where the feature sequences to be aligned are a template sequence and a test sequence. T The standard motion intent feature sequence from the training set has a length of [length missing]. n , represented as T =[ t 1, t 2,..., t n ],in For the first i 128-dimensional feature vectors at each time step (corresponding to 8-30Hz EEG features); test sequence X The feature sequence to be identified is collected in real time, with a length of [length missing]. m ( m and n (may not be equal), represented as X =[ x 1, x 2,..., x m ],in For the first j The feature vector at each moment. The difference in sequence length stems from the difference in the speed of motion execution among different users (e.g., a healthy person needs 1.2 seconds to complete a "leg raise," while a person with a movement disorder needs 1.8 seconds), resulting in different feature sequence lengths. n =300 (1.2s × 250Hz) m =450 (1.8s × 250Hz). Time-axis normalization is performed on the features of each dimension in the sequence to eliminate amplitude differences: In the formula, Indicates the first element in the template feature sequence k Under the feature dimension, the first i The original feature values of each time step. The template feature sequence here can be understood as a predefined feature sequence used as a reference, such as the standard feature sequence when a healthy person completes the "leg raise" action. This indicates that after the test feature sequence is standardized along the time axis, the first... k Under the feature dimension, the first j The feature values of each time step are defined as follows: the test feature sequence is the feature sequence to be analyzed and compared with the template sequence, such as the feature sequence of a patient with a movement disorder when performing the "leg raise" action. The purpose of standardization is to eliminate the amplitude differences caused by factors such as differences in the speed of movement execution among different users, so that it can be compared with the template sequence on the same scale in the future. Indicates the first element in the test feature sequence k Under the feature dimension, the first j The raw feature values at each time step are the unstandardized raw data. k =1,2,...,128 are the indexes of the feature dimensions. m k , s k For the training set k The mean and standard deviation of the features are calculated to ensure that the template and test sequences are compared on the same scale. The construction of the DTW distance matrix includes local distance calculation, specifically calculating the Euclidean distance between any two feature vectors in the template and test sequences, which is used as elements of the local distance matrix D. The calculation formula is: In the formula, Indicates the first position in the template sequence i The standardized feature vectors of each time step. In addition, the template sequence is a pre-determined feature sequence used as a reference standard, such as the brain-computer interface feature sequence when a healthy person completes a certain action, which has been processed by time axis standardization, that is, the feature vector after eliminating amplitude differences. Indicates the first test sequence j The standardized feature vectors of each time step, where the test sequence is the feature sequence to be analyzed and compared with the template sequence, such as the brain-computer interface feature sequence when a patient with a movement disorder completes the same action, which has also undergone time axis standardization. Indicates the first position in the template sequence i In the feature vector of the nth time step, the th k Standardized feature values for each feature dimension; Indicates the first test sequence j In the feature vector of the nth time step, the th kThe standardized feature values of each feature dimension. i =1,2,..., n , j =1,2,..., m ,matrix D The size is n × m ; Optimization processing is as follows D ( i , j )> i dist (When the distance threshold is set to 15, force) D ( i , j ) = +∞ (avoiding irrelevant point matching); the cumulative distance matrix is constructed by defining the cumulative distance matrix. C ,in C ( i , j ) indicates that t i and x j The minimum cumulative distance during alignment is calculated according to the following rules: C ( i , j )= D ( i , j )+min{ C ( i -1, j ), C ( i , j -1), C ( i -1, j -1)};Its boundary conditions are: C (1,1)= D (1,1); C ( i ,1)= D ( i ,1)+ C ( i -1,1) (The first column can only accumulate from top to bottom); C (1, j )= D (1, j )+ C (1, j -1) (The first row can only accumulate from left to right). The search for the optimal aligned path includes setting path constraints to avoid excessive path distortion and setting slope constraints (Sakoe-Chiba band): in α=0.1 is the constraint coefficient, i.e., the alignment point ( i , j The path should fall within 10% of the bandwidth on both sides of the main diagonal to reduce computation while ensuring proper alignment. A backtracking method is used to find the optimal path, starting from the bottom right corner of the matrix. C ( n , m ) Begin backtracking; among which n yes i The maximum value, m yes j To find the maximum value, each time select the direction with the smallest accumulated distance from the previous step (up, left, top left), and record the path point. i , j );like C ( i , j )- D ( i , j )= C ( i -1, j If ), then the previous step is ( i -1, j (Vertical movement, stretching template sequence); if C ( i , j )- D ( i , j )= C ( i , j -1), then the previous step is ( i , j -1) (Horizontal movement, tensile test sequence); if C ( i , j )- D ( i , j )=C( i -1, j -1), then the previous step is ( i -1, j -1) (Diagonal movement, synchronized alignment). Termination condition is backtracking to... C Stop at (1,1) to obtain the optimal alignment path. P =[( i 1, j 1),( i 2, j 2),...,( i L , j L )],in Lpath length ( L ≥max( n , m )). Path smoothing is a process for smoothing paths. P Perform moving average filtering to eliminate local jitter: The first and last points are not processed. This represents the optimal alignment path after moving average filtering. P The Middle k The template sequence time step index corresponding to each path point. The purpose of the moving average filter is to eliminate local jitter in the path and make the aligned path smoother. This is achieved by averaging the current and adjacent template sequence time step indices. This represents the optimal alignment path after moving average filtering. P The Middle k The time step index of the test sequence corresponding to each path point is also used to eliminate local jitter in the path. A smoother result is obtained by averaging the time step indices of the current and adjacent test sequences, ensuring the continuity of the aligned path. The time warping of the aligned sequence is implemented based on the optimal path. P , with a length of m test sequence X Regularized into template sequence T Equal length ( n ) sequence For the template sequence of the first... i Points t i Find all in the path i k = i Corresponding test points If there are multiple test points, take the average: ;in, K 1 represents the number of matching points; if it corresponds to a single test point, assign the value directly. Alignment effectiveness evaluation includes calculating the similarity of the aligned sequences: ;in, C ( n , m The total cumulative distance is also shown in the bottom right corner of the matrix. The alignment is determined to be valid at that time. For example, the alignment of the sequence with the intention of "raising the leg", the template sequence T ( n =300), which is the EEG characteristic sequence of a healthy person "raising their leg", key time points i =150 corresponds to a hip joint angle of 30°; Test sequence X ( m =450); the characteristic sequence of the patient's "leg raising" movement, due to the slow movement, corresponds to a 30° angle. j=250; the alignment result is the path. P The key matching point is (150, 250). After normalization... The length is 300, and T similarity Sim =0.82, ensuring temporal consistency during subsequent intent recognition. Through the above steps, the DTW algorithm can effectively solve the problem of feature sequence length mismatch caused by differences in movement rhythm among different users. Testing showed that the recognition accuracy of the integrated model improved by 4.2% after alignment (from 89.3% to 93.5%), especially significantly improving the intent recognition effect for patients with movement disorders. Dynamic cascaded data preprocessing breaks through the signal-to-noise ratio bottleneck of EEG signals. Compared to traditional preprocessing methods where adaptive filtering and wavelet transform are mostly executed independently and sequentially, without considering the dynamic changes in noise features, the preprocessing mechanism of this invention analyzes the noise types in the signal in real time (such as the proportion of power frequency interference, electromyographic noise, and electrooculogram artifacts) through a noise feature monitoring module, automatically adjusting the weight allocation of the two-stage processing. When the proportion of power frequency interference exceeds 30%, the step size factor of the adaptive filtering is enhanced. q (Value range 0.01-0.1, dynamically adjusted). An adaptive filtering formula accelerates filter convergence, increasing the 50Hz interference suppression ratio to over 45dB. When the proportion of electrooculogram (EOG) artifacts exceeds 20%, the number of wavelet transform decomposition layers is optimized (dynamically increased from 5 to 7). A modal entropy threshold is introduced based on the db4 wavelet basis function to specifically suppress EOG signals in the 3-8Hz frequency band, increasing the artifact removal rate from 92% to 96%. Dynamic matching of noise type and processing intensity is achieved, addressing the insufficient adaptability of a single preprocessing algorithm in complex scenarios. This provides high-purity signals for subsequent feature extraction (signal-to-noise ratio stabilized at 35dB±2dB), making it particularly suitable for noisy environments such as hospitals and homes. A heterogeneous ensemble learning model enables super-resolution recognition of motion intent. While existing ensemble learning methods often employ homogeneous model fusion, this invention constructs a CNN-SVM-RF heterogeneous ensemble learning model, improving recognition accuracy through a feature complementarity mechanism.
[0139] The CNN-SVM-RF heterogeneous ensemble learning model and voting mechanism adopt a three-layer architecture of "parallel input-feature splitting-result fusion" (e.g., Figure 5As shown), the input layer receives a 128-dimensional feature vector (including 32-dimensional time domain, 64-dimensional frequency domain, and 32-dimensional spatial domain), and transmits it to three sub-models: CNN, SVM, and RF. The processing layer computes independently for each sub-model: CNN focuses on capturing the local spatiotemporal correlation of features, SVM excels at high-dimensional feature boundary delineation, and RF focuses on global statistical regularities. The fusion layer integrates the output probabilities of the three sub-models through a weighted voting mechanism to generate the final motion intent. The data interaction interface uses a standardized format for the feature vector (float32 type, range [-1,1]), and data transfer between sub-models is achieved through a shared memory buffer, with a transmission latency ≤5ms. The output of each sub-model is uniformly a probability distribution vector for 6 types of motion intents (…). p 1, p 2,..., p 6), of which and The CNN model execution includes feature reconstruction, which reshapes a 128-dimensional vector into a 16×8×1 two-dimensional feature matrix (simulating the spatial distribution of EEG signals), which is then used as the CNN input; feature extraction, where convolutional layer 1 (3×3 kernels × 32) performs sliding convolution on the input matrix, outputting a 16×8×32 feature map with the ReLU activation function; pooling layer 1 (2×2 max pooling) downsamples to an 8×4×32 feature map; convolutional layer 2 (3×3 kernels × 64) outputs an 8×4×64 feature map with the ReLU activation function; pooling layer 2 (2×2 max pooling) downsamples to a 4×2×64 feature map; classification output, which is flattened into a 512-dimensional vector, compressed through a fully connected layer (512-64-6); and outputting the probability distribution through the softmax function. The inference time is approximately 150ms.
[0140] The SVM model execution includes kernel function mapping, which uses the RBF kernel function to map the 128-dimensional features to a high-dimensional space; multi-classification computation, which involves constructing 15 binary SVM sub-classifiers (combining 6 intentions in pairs); each sub-classifier outputs a class judgment (+1 / -1), assigning the highest probability to the class with the highest cumulative votes; and probability calibration, which uses Platt scaling to transform the decision values into a probability distribution. Inference takes approximately 220ms. The RF model execution includes feature sampling, which involves randomly sampling 128-dimensional features (each sampling...). The process involves generating independent feature sets for each decision tree; decision tree voting, where 100 CART trees are computed in parallel, with each tree splitting according to the principle of minimizing Gini impurity; single tree outputting class predictions (1-6), and calculating the class vote rate of the 100 trees; and probability generation, which involves normalizing the vote rate and generating a probability distribution. The inference time is approximately 200ms.
[0141] The weighted voting mechanism includes dynamic weight calculation, which involves evaluating the performance of sub-models and updating the accuracy of each model on the validation set every 500 samples. acc CNN , acc SVM , acc RF Then, weight normalization is performed, and the weight coefficients are calculated: , In actual testing, it remained stable at... w CNN =0.35, w SVM =0.3, w RF =0.35; where, w CNN The weight coefficients of the CNN sub-models reflect the importance of the model in the fusion decision. The larger the weight, the stronger the impact on the final result. acc CNN This represents the accuracy of the convolutional neural network sub-model on the validation set, reflecting the model's accuracy in recognizing motion intentions, and its range is [0,1]. acc SVM This represents the accuracy of the support vector machine sub-model on the validation set; acc RF This represents the accuracy of the random forest submodel on the validation set; w SVM Represents the weight coefficients of the SVM sub-model; w RF This represents the weight coefficients of the RF sub-model; probability fusion is performed through weighted summation, with the probabilities of the six intent categories being weighted and accumulated separately. The weighted accumulation expression is as follows: ;in, p i final Indicates the first i The final fusion probability of the class of motion intentions serves as the basis for multi-class decision-making; the larger the value, the higher the credibility of that class of intentions. p i CNN This indicates that the CNN sub-model predicts the first... i The probability of a certain type of motion intention, ranging from [0,1]; p i SVM This indicates that the SVM sub-model predicts the first... i The probability of a certain type of motion intention, ranging from [0,1]; p i RF This indicates that the RF sub-model predicts the first... i The probability of a certain type of motion intention, ranging from [0,1]; wherei =1,2,...,6 (corresponding to 6 types of motion intentions); the result is determined by selecting the category with the highest probability as the final output; intention = argmax i { p i final If the maximum probability This triggers secondary recognition (increasing sample collection time to 3 seconds); where argmax i Indicates the index of motion intention categories i The maximum value operation involves iterating through all six intent categories to find the final fusion probability and selecting the category index with the highest probability. i The intent corresponding to that index will be used as the final output. The conflict resolution mechanism involves initiating confidence backtracking when the highest probability classes of the three sub-models are all different (e.g., CNN selects "leg raise", SVM selects "bend over", and RF selects "still"). This backtracking calculates the difference between the second highest probability class and the highest probability class for each model. Δp = p 1- p 2; Prioritize adoption Δp The largest model result (highest confidence), for example, CNN. Δp =0.82-0.1=0.72; SVM's Δp =0.6-0.3=0.3; RF Δp =0.55-0.2=0.35 The final result of the CNN's "leg raise" is adopted. For example, the sub-model output of the ensemble recognition of the "fist clenching" intention is the CNN's... P CNN =[0.02,0.01,0.01,0.92,0.03,0.01] (Category 4, "clenched fist," has the highest probability); SVM's P SVM =[0.03,0.02,0.02,0.85,0.05,0.03] (Category 4 "clenched fist" probability is the highest); RF P RF =[0.01,0.01,0.02,0.88,0.06,0.02] (Category 4, "clenched fist," has the highest probability). Weighted fusion is... p 4 final =0.35×0.92+0.3×0.85+0.35×0.88=0.322+0.255+0.308=0.885 Other categories p i final<0.1, ultimately outputting the intention to clench a fist. Performance comparison: The accuracy of single models in recognizing a clenched fist is as follows: CNN accuracy is 92.1%, SVM accuracy is 89.3%, and RF accuracy is 91.5%; the ensemble model accuracy is 96.1%, with an error rate reduction of 43%-57%, demonstrating the advantages of heterogeneous ensemble. Through the above steps, the CNN-SVM-RF heterogeneous ensemble learning model achieves complementary advantages of different learning mechanisms. The weighted voting mechanism ensures the reliability of the results. Validated by 12,000 test samples, the average recognition accuracy of the ensemble model reaches 93.5%, and its robustness to noise interference is significantly better than that of the single model. The CNN branch adopts a 3-layer convolution + 2-layer residual connection structure, with alternating use of 3×3 convolution kernels and 1×1 convolution kernels, specifically capturing local temporal features of EEG signals (such as...). m The output feature map is processed by global average pooling to obtain a 256-dimensional feature vector (wave attenuation slope). The SVM branch, based on the radial basis function kernel (γ1 dynamically adjusted to 0.1-1), performs nonlinear classification on the spatial features extracted by CSP to solve the boundary ambiguity problem under small sample conditions. The RF branch's 100 decision trees use a feature random subspace (each node randomly selects 20% of the features) to focus on learning the statistical features of the signal (such as variance and peak frequency) and suppress outlier interference. A voting mechanism introduces confidence weighting, with the formula as follows: ( c i Set the classifier confidence level (values range from 0.8 to 1.0) to avoid low confidence results affecting the final decision.
[0142] The three models cover temporal, spatial, and statistical feature dimensions, respectively, improving the recognition accuracy of six basic motor intentions to 95.8%. Even for spinal cord injury patients with weak signal strength (EEG amplitude <50μV), the recognition accuracy remains above 90%, breaking through the feature coverage limitations of traditional single models. A dual closed-loop feedback system integrating reinforcement learning and impedance control enables compliant human-machine collaborative control.
[0143] This invention can construct a dual closed-loop feedback mechanism including an inner loop (angle closed loop), that is, acquiring joint angles through an IMU. i real , and the target angle i cmd deviation Dth = i cmd - i real Input reinforcement learning module, using state value function, when Dth When the temperature is >5°, the reward value will be dynamically adjusted. r (Positive rewards are inversely proportional to the square of the deviation), which improves the angle convergence speed by 40%; the outer loop (force closed loop) is the human-computer interaction force collected by the force sensor. FWith Expectation F 0 (Dynamically set according to user weight, such as when weight is 50kg) F 0 Deviation of =10N) ΔF Through impedance control formula ; ( K 2 is the stiffness coefficient. B (where is the damping coefficient), where This represents the target joint angular velocity, i.e., the target angle. i cmd The rate of change over time, i.e., the desired angular velocity of joint rotation, where, i Represents joint angle; This represents the actual joint angular velocity, which is the actual rotational angular velocity of the joint collected by sensors such as an IMU (Inertial Measurement Unit). It is a measurement of the true speed of joint movement; joint torque is adjusted in real time. t This ensures that the fluctuation of the interaction force is controlled within ±2N. It resolves the contradiction of "precise angle but large force impact" in traditional control, ensuring both the accuracy of the movement trajectory (deviation ≤2°) and avoiding secondary injury caused by excessive interaction force in the rehabilitation training of spinal cord injury patients, achieving the dual control goal of "precision + safety". A user-adaptive feature alignment network enables rapid cross-user adaptation. Addressing the differences in EEG signal characteristics among different users (e.g., a 1-3cm shift in the activation position of the motor cortex), an adaptive network based on Dynamic Time Warping (DTW) and feature transfer learning is designed. In the offline stage, a multi-user feature library (containing EEG features of 200+ users of different ages and injury levels) is constructed. The DTW algorithm is used to align the feature sequences of similar movement intentions along the time axis, generating standardized feature templates. The standardized feature templates for similar movement intentions are generated using the DTW algorithm. The screening and preprocessing of similar movement intention samples include sample set construction. Specifically, for a single movement intention (e.g., "leg lift"), feature sequences from N=50 healthy subjects are collected, with 3 sets of valid samples collected from each subject at 24-hour intervals, resulting in a total of 150 original sequences (denoted as N=50). X 1, X 2,..., X 150 ); each sequence X k =[ x k,1 ,x k,2 ,..., x k,mk ],in m k The sequence length (due to individual differences in action speed) (corresponding to a sampling duration of 0.8-2.0s). It is a 128-dimensional feature vector.
[0144] Outlier removal involves calculating the kinematic characteristics of each sequence, specifically the mean angular velocity. ( Δt =4ms is the sampling interval), characteristic variance ;in, Indicates the first k In the feature sequence, the first t The time step, the first d The original feature values of each feature dimension; i max Indicates the maximum angle; i min Indicates the minimum angle; excludes Exceeding [20° / s, 50° / s] or s k 2 Samples exceeding the mean ± 2 standard deviations were retained, with 120 valid sequences retained. Feature standardization involved global standardization of all retained sequences. ;in, Indicates the first k In the feature sequence, the first t The time step, the first d The feature values of each feature dimension after global standardization; m d , s d For the first d The total mean and standard deviation of the dimensional feature across 120 sequences; that is m d Indicates the first d The total mean of the dimensional feature across all 120 retained valid sequences is used for feature standardization to eliminate scale differences in this dimensional feature between different sequences. s d Indicates the first d The total standard deviation of the dimensional feature across all 120 retained valid sequences, and the total mean. m d In conjunction with this, feature standardization is performed, mapping features to a scale with a mean of 0 and a standard deviation of 1, ensuring consistent feature scales across samples. In DTW-based multi-sequence hierarchical alignment, reference sequence selection involves calculating the average DTW distance of each sequence to all other sequences. ;in, l Represents a sequence index, used for traversing except the first... k All other sequences besides the first sequence, to calculate the first... k The average DTW distance between this sequence and other sequences; Indicates the first lThe first feature sequence is one of the comparison objects when calculating the average DTW distance, and is compared with the second feature sequence. k bar sequence X k Perform DTW distance calculation; select The smallest sequence is used as the initial reference sequence. X ref (That is, it has the highest overall similarity with sequences of the same type), and its length is m ref The first-level alignment involves aligning all sequences with the reference sequence, and for each sequence... X k Execute the DTW algorithm and X ref Alignment (refer to the aforementioned DTW implementation steps) is achieved by calculating the local distance matrix. ;in, Indicates the first k In the sequence to be aligned, the first i The standardized feature vector at each time step contains multiple feature dimensions and is a feature representation after previous standardization. Represents the initial reference sequence X ref In the middle, the first j The standardized feature vectors at each time step are used to calculate the local distance with the feature vectors of the sequence to be aligned, thereby achieving dynamic time warping; a cumulative distance matrix is constructed. C Configure Sakoe-Chiba bandwidth α =0.15; Backtracking yields the optimal path. P k =[( i 1, j 1),( i 2, j 2),...,( i L , j L )];Will X k Organize into X ref Equal-length sequences (length m ref ),in ;in, Indicates the first k The sequence undergoes primary alignment, i.e., alignment with the initial reference sequence. X ref After alignment and normalization, in the first j The feature mean at each time step, and in addition, the feature mean mapped to the same time step. j All iThe corresponding eigenvalues are averaged to normalize the sequence to be aligned to the same length as the reference sequence. Secondary alignment includes consistency optimization of the aligned sequence, which involves calculating the pointwise mean of 120 normalized sequences. ,in ;by For the new reference sequence, repeat the first-level alignment step, applying it to all... Reorganized into Eliminate the accumulated error of the first alignment; calculate the sequence consistency after the second alignment:
[0145] ;in, Indicates the first k The sequence undergoes two-level alignment, i.e., using a pointwise mean sequence. After re-aligning and normalizing the new reference sequence, at the... j The time step, the first d The feature values of each feature dimension; when Consistency Iteration stops when the value is ≥0.85; otherwise, the second-level alignment is repeated (in actual testing, two iterations are sufficient). The standardized feature template is generated from the 120 sequences after the second-level alignment. Calculate the point-by-point mean as the standardization template. ,in, m T = m ref (Uniform length, e.g., 300 frames correspond to 1.2 seconds); Calculate the pointwise standard deviation of the template. ,in, Indicates the first k In the sequence after secondary alignment, the first... j The feature vectors at each time step are the feature representations of the sequence at the corresponding time step after the second-level alignment operation is completed; Represents a standardized feature template T In the middle, the first j The point-by-point mean at the nth time step is the result of analyzing 120 double-aligned sequences at the nth time step. j The average of the eigenvalues at each time step is used as the typical feature representation of the standardized template at that time step; the confidence interval (±2) of the template is used as the template. s j Key time points are marked by combining kinematic data (joint angles) and marking three keyframes in template T, specifically the start frame. j s That is, the difference between the feature vector and the static state exceeds the threshold for the first time. );in, Indicates the start frame j sThe corresponding feature vector is used to determine whether the difference between the feature vector and the static state exceeds a threshold for the first time. That is, by calculating the difference between this feature vector and the static state feature vector... t rest The norm is used to measure and determine the start time of the action; This represents the feature vector in a static state, used to determine whether the feature vector has regressed to the vicinity of the static state, i.e., by calculating the feature vector and... t rest The difference is measured by the norm; peak frame j p That is, the feature dimension most relevant to the movement intention (such as...) β Wave power reaches its maximum value; frame terminates. j e That is, the eigenvectors regress to the vicinity of the resting state. ;in, Indicates the termination frame j e The corresponding feature vector, that is, the feature vector at the time of termination when the feature vector regresses to near the static state. For example, in the "leg raise" template, j s =50 (Action initiated) j p =150 (maximum hip angle) j e =250 (action completed). The test set validation involved performing DTW matching between the template and similar intent sequences from 20 new subjects, calculating the average similarity. Among them, Sim≥0.75 is required; cross-category discrimination, that is, the average distance between the template and the other 5 intention templates is calculated, which must be greater than twice the distance between the template and the same type of sequence (to ensure category uniqueness).
[0146] The process of generating the standardized template for "leg lift" intent includes an initial dataset of 150 "leg lift" feature sequences, ranging in length from 220 to 480 frames, from which 120 sequences are selected and retained. The reference sequence is the one with the smallest average DTW distance. X ref (Length 300 frames); First-level alignment involves normalizing 120 sequences to 300 frames, consistency index Consistency =0.79; Secondary alignment uses the mean sequence as a new reference and then aligns again. Consistency =0.88; Template generated as T The keyframes are the mean sequence of 300 frames. j s =48、 j p =145、 j e=242; the verification result is the average similarity between the new sample and the template. Sim =0.81, and the distance to the "bending over" template is 2.3 times that of similar templates, meeting the validity requirements. The standardized feature template generated through the above steps solves the problem of sequence inconsistency caused by individual differences in similar motor intentions, providing a reliable benchmark for subsequent DTW real-time matching. Testing showed that the intention recognition accuracy based on this template improved by 5.3%, especially significantly improving the recognition stability for patients with movement disorders (fluctuation range reduced from ±8% to ±3%). In the online adaptation phase, after a new user wears the device, 3 minutes of motor imagery signals are collected. Using a feature distance metric (calculating the Euclidean distance between the new feature and the template), the three most similar templates are automatically matched from the library. Weighted transfer learning (weights are inversely proportional to the distance) is used to initialize the ensemble learning model parameters. During the adaptation process, 10% of the model parameters are updated every 5 minutes through incremental learning to avoid catastrophic forgetting.
[0147] The Transformer-based multimodal fusion attention mechanism builds upon existing heterogeneous ensemble learning models by introducing the Transformer architecture to construct a multimodal fusion attention mechanism, achieving deep correlation between EEG signals and auxiliary modality data. The data input and output of the Transformer-based multimodal fusion attention mechanism include preprocessing of the multimodal data input layer, namely, data modality segmentation and feature dimension unification. The input data contains three modalities, specifically EEG feature sequences. E 128 dimensions / frame, duration 2s (500 frames), derived from preprocessed 8-30Hz frequency band signals; electromyography (EMG) characteristic sequence. M 64 dimensions / frame, duration 2 seconds (500 frames), surface electromyography signals extracted from 6 muscles including the quadriceps; motion posture sequence P 32 dimensions / frame, duration 2 seconds (500 frames), composed of acceleration and angular velocity data acquired by IMU; electromyography and posture sequences were standardized (similar to EEG characteristics). m , s (Parameters), ensuring consistent feature scale across the three modality classes. Modality embedding and location encoding include adding a dedicated embedding vector for each modality class: E embed ( i )= E ( i )+ e eeg , M embed ( i )= M ( i )+ e emg , P embed (i )= P ( i )+ e pose ;in, E embed ( i ) indicates that after modal embedding, the first EEG feature sequence is... i The feature vector at each time step is derived from the original EEG feature vector. E ( i Based on this, EEG-specific embedding vectors were added. e eeg The results obtained later are used for subsequent fusion and processing of multimodal features; E ( i ) represents the first EEG feature sequence. i The original feature vectors of each time step, i.e., EEG features without modality embedding operations; M embed ( i ) indicates that after modal embedding, the electromyographic feature sequence is... i The feature vector at each time step is a feature vector from the original electromyography feature vector. M ( i Based on this, add electromyography-specific embedding vectors. e emg Obtained; M ( i ) represents the first in the electromyographic feature sequence. i The original feature vectors of each time step, i.e., the electromyographic features without modality embedding operation; P embed ( i ) represents the motion posture sequence after modal embedding, the th i The feature vector at each time step is a feature vector of the original motion posture. P ( i Based on this, add motion pose-specific embedding vectors. e pose Obtained; P ( i ) represents the first position in the motion posture sequence. i The original feature vectors at each time step, i.e., motion pose features without modality embedding operations; e eeg Represents a learnable embedding vector specific to an EEG modality, with dimension R. 128 This is used to add modality-specific embedding information to EEG features, which facilitates the subsequent differentiation and fusion of multimodal features; e emg Represents a learnable embedding vector specific to an electromyographic modality, with dimension R. 128Its function is to add electromyographic modality-specific embedding information to electromyographic features; e pose The learnable embedding vector representing the motion pose modality has dimension R. 128 This is used to add motion posture modality-specific embedding information to motion posture features; among which, For learnable modal embeddings (initially randomly generated, optimized through training); sinusoidal position encoding is added to capture temporal information: , ,in, d =128 is the feature dimension. pos For frame index, k =0,1,...,63, finally yielding the input sequences for the three modalities. The cross-modal attention computation of the Transformer multimodal fusion attention layer includes constructing a query (Q), key (K), and value (V) matrix; using EEG sequences as the query source and EMG and posture sequences as the key-value sources. ;in ) is the projection matrix, which compresses the 128-dimensional features to 64 dimensions. X m ; X p [] indicates that the electromyography and pose sequences are concatenated in the frame dimension (1000×128). Attention weights are calculated. ;in M The mask matrix (to prevent information leakage in future frames) has 0s on the diagonal and lower triangle, and -∞s on the upper triangle; the output is EEG-multimodal fusion features. F eeg =Linear( Attention ( Q , K , V ))+ X e (Residual connectivity + LayerNorm); Dimensionality maintained at 500×128. Intramodal self-attention enhancement, i.e., enhancement of the fused EEG features. F eeg Perform self-attention calculation: , F eeg =Linear( Attention ( Q ', K ', V '))+ F eeg ;in, Q ' represents the query matrix in intramodal self-attention computation, which is composed of fused EEG features. F eeg With projection matrixW q The product obtained by multiplication is used to query relevant information in the self-attention mechanism; K 'Represents the key matrix in intramodal self-attention computation, which is also composed of fused EEG features. F eeg With projection matrix W q Multiplying them together, we get the result here. K '= Q ', is a common setting in self-attention mechanisms where the key and query are calculated using the same method, used to match with the query matrix to calculate attention weights; V 'Represents the value matrix in intramodal self-attention computation, which is also composed of fused EEG features. F eeg With projection matrix W q The product obtained by multiplication is the feature matrix ultimately used for weighted summation in the attention mechanism; W q ' represents the projection matrix in intramodal self-attention computation, used to project the fused EEG features. F eeg Projected onto a specific dimension to generate a query, key, and value matrix, which participates in the self-attention calculation process; F eeg 'This represents the EEG characteristics enhanced by intramodal self-attention. It is achieved by performing a linear transformation (Linear operation) on the self-attention calculation results, combining it with residual connections, and adding the original...' F eeg The result, obtained after layer normalization (LayerNorm), enhances the EEG feature representation after attentional association within the features; it also enhances the temporal correlation within the EEG sequence and highlights feature peaks related to motor intention (such as...). m (Wave suppression period). Interaction with the original heterogeneous ensemble model includes feature splitting and sub-model adaptation, i.e., the fused features output by the Transformer. F eeg The processing involves extracting the features from the last frame (128-dimensional) and inputting them into the SVM and RF models (maintaining compatibility with the original architecture); preserving the complete sequence (500×128) and inputting it into the improved CNN model (adding a temporal convolutional layer); temporal adaptation of the CNN model, i.e., adding a 1D convolutional layer (kernel size 3, output channels 32) before the original structure to convert the 500×128 sequence to 500×32, and then obtaining a 128-dimensional feature vector through global average pooling. Attention weight-guided sub-model weighting, i.e., extracting the weight matrix of the Transformer cross-modal attention layer. (Attention level of EEG frames to EMG / pose frames); Calculation of modal confidence: , c emg + pose )=1- c eeg ,in, pose This represents the pose modality, one of the three modalities of input data: EEG, EMG, and pose. It describes motion-related posture information and consists of acceleration, angular velocity, and other data collected by the IMU. The sub-model weights are dynamically adjusted. ;in, w CNN ' represents the dynamically adjusted weights of the Convolutional Neural Network (CNN) sub-model, which are the weights of the original CNN sub-model. w CNN Based on this, combined with EEG modal confidence c eeg The weighted average of the CNN sub-model outputs is obtained after dynamic adjustment and is used for subsequent multi-model probability fusion. w SVM ' represents the dynamically adjusted weights of the Support Vector Machine (SVM) sub-model, which are the original weights of the SVM sub-model. w SVM Based on this, confidence levels of electromyography and motor posture modalities were combined. c emg + pose as well as w RF The weighted average of the SVM sub-model outputs is obtained by dynamically adjusting the relevant balance terms and is used for multi-model probabilistic fusion. w RF ' represents the weights of the Random Forest (RF) sub-models. During the dynamic adjustment of these sub-model weights... w RF 'Keep the weights constant as a balancing term to ensure the sum of all sub-model weights is 1, thus reasonably fusing the outputs of each sub-model; keep the weights constant as a balancing term to ensure the sum of the weights is 1.' The multi-model probability fusion of the output layer includes the sub-model output probability distribution: Weighted fusion formula: ; Decision result output and verification; The ultimate intention is argmax( P final Simultaneously output the modal consistency score: ;when Score If the value is ≥0.85, the result is output directly; otherwise, multimodal reacquisition is triggered (extended to 3s) to ensure decision reliability.
[0148] The multimodal fusion process of the "leg-raising" intention includes input data, namely EEG sequences. E 500 frames, 128 dimensions (including) m Wave suppression characteristics); electromyographic sequences M 500 frames, 64 dimensions (significant quadriceps activation features); pose sequence P 500 frames, 32 dimensions (increasing hip angle). Transformer processing for cross-modal attention display of EEG frames 150-250 (corresponding to the mid-stage of the leg raise movement) to EMG frames 180-280. α =0.8 (highly correlated); fusion feature F eeg Highlight the feature peaks between frames 150-250. The sub-model output is... P CNN =[0.02,0.93,0.01,0.01,0.02,0.01] (Category 2 "Leg Raise"); P SVM =[0.03,0.89,0.02,0.02,0.03,0.01]; P RF =[0.01,0.91,0.01,0.02,0.03,0.02]. Weights and fusion are... c eeg =0.7, therefore w CNN =0.35×(0.5+0.35)=0.35×0.85=0.2975, w SVM =0.3×(0.5+0.15)=0.195, w RF =0.35; P final(2) = 0.2975 × 0.93 + 0.195 × 0.89 + 0.35 × 0.91 ≈ 0.276 + 0.174 + 0.319 = 0.769, outputting the intention to "raise leg". By introducing the Transformer's multimodal fusion attention mechanism, the model's adaptability to complex scenarios (such as noise interference and action variations) is significantly improved. Compared with the original heterogeneous ensemble model, the average recognition accuracy increased from 93.5% to 96.2%, and the cross-user generalization error decreased by 40%, fully demonstrating the technical advantages of multimodal information complementarity. Modal input, namely the 256-dimensional feature vector extracted from the EEG signal by CSP, incorporates two types of auxiliary modal data: 16-channel muscle activity signals collected by the electromyography sensor, which are decomposed into 64-dimensional features by wavelet packet decomposition; and 3-axis acceleration and 3-axis angular velocity recorded by the inertial measurement unit (IMU), which are statistically analyzed by sliding window to obtain 32-dimensional kinematic features. The three modalities of data were mapped to a 128-dimensional vector space through linear projection layers, forming modality-specific feature sequences (EEG sequence length L1=30, EMG sequence length L2=20, kinematic sequence length L3=15). Intramodal self-attention mechanisms were implemented, with each modality branch equipped with an independent self-attention layer to calculate dependencies within the feature sequences; the EEG branch, specifically targeting… m Wave, β The temporal correlation of the waves is addressed using causal masking self-attention to ensure that only historical signals are considered at the current moment. The attention weight calculation formula is as follows: ;in M It is a causal mask matrix (the lower triangle is 0, and the upper triangle is -∞). d k =128 is the feature dimension. The electromyography and kinesiology branch employs maskless self-attention to capture global feature correlations, such as the synchronicity between muscle activity intensity and joint movement speed. The cross-modal cross-attention mechanism specifically constructs three sets of cross-attention layers to achieve intermodal information interaction; using EEG features as the query ( Q ), electromyographic characteristics are key ( K ) and value ( V ), learning the mapping relationship between "brain electrical intention and muscle response"; using kinematic characteristics as Q EEG characteristics are K / V Strengthen the temporal alignment of "motor state - EEG command"; design bidirectional cross-attention to make electromyography and kinematic features mutually reinforcing. Q / K / VThis compensates for the latency limitations of single-modal processing. The cross-attention output is normalized by LayerNorm and connected to the residual, preserving the original features while incorporating cross-modal information. A global fusion attention layer aggregates the cross-attention outputs from the three modalities, forming a 65-bit (30+20+15) hybrid feature sequence, which is then input into the global self-attention layer. A multi-head attention mechanism is employed (number of heads...). h =8), each head independently calculates the attention distribution for different subspaces, using the following formula: MultiHead ( Q , K , V )= Concat ( head 1,..., head h ) W O in head i = Attention ( QW i Q , KW i K , VW i V ), W i Q , W i K , W i V , W OThe model uses a learnable parameter matrix. Attention weights are visualized to dynamically adjust the contribution of each modality; for example, EEG weights account for 60% when the motor intention is clear, while EMG weights increase to 50% during the motor execution phase. A classification head design is used, where the [CLS] label vector of the global attention output is mapped to a 6-class motor intention space via a feedforward neural network (FFN). The FFN contains two linear layers (intermediate dimension 512, activation function GELU), and a dropout layer (probability 0.3) is introduced to suppress overfitting. Comparative experiments show that after adding multimodal fusion, the recognition accuracy for complex intentions such as "fine grasping" increases from 89.2% to 94.7%, and the tolerance for signal interruption (<200ms) increases by 30%. The specific experimental test subjects for multimodal fusion included 20 healthy adults (10 men and 10 women, aged 22-45 years, with no history of motor disorders) and 10 patients with mild hand dysfunction (post-stroke sequelae, able to perform simple grasping but with weak fine control). The experimental equipment included: electroencephalography (EEG) acquisition, specifically a 64-channel wet electrode EEG (500Hz sampling rate, 0.1-500Hz bandwidth); electromyography (EMG) acquisition, specifically an 8-channel surface EMG sensor (1000Hz sampling rate, attached to the flexor and extensor muscles of the forearm); posture tracking, specifically an inertial measurement unit (IMU, 200Hz sampling rate, measuring three-dimensional acceleration and angular velocity of the hand); and motion capture, specifically an optical motion capture system (0.1mm accuracy, recording finger joint angles). The test intent categories included complex intents such as fine grasping (pinching a 5mm diameter steel ball), multi-finger coordination (twisting a bottle cap), and force-controlled grasping (picking up an egg); and simple intents such as clenching a fist, extending a palm, and remaining still (as a control). The specific data collection for the experiment involved three sets of experiments for each subject, with each set containing 10 repetitions of each of the six types of intentions, spaced 1 minute apart. The data collection duration was 3 seconds for each type of intention, with the first 2 seconds being the execution phase and the last 1 second being the resting recovery phase. Signal interruption simulation involved randomly inserting a 200ms EEG signal interruption (replaced with baseline noise) into 10% of the samples to simulate signal loss in actual use. The model testing process included a control group using only the original heterogeneous ensemble model (CNN+SVM+RF) based on EEG features; and an experimental group incorporating a Transformer multimodal fusion model (EEG + EMG + posture). The test set was divided according to "external grouping of subjects" (80% training, 20% testing, to ensure model generalization to new users). Performance evaluation metrics included accuracy (number of correctly identified samples / total number of samples); fault tolerance (accuracy during signal interruption / accuracy during normal signal × 100%); and average recognition latency (time from signal input to intention output, required to be ≤300ms). The comparative experimental results are shown in Tables 3 and 4.
[0149] Table 3 Comparison of Accuracy Rates for Complex Intent Recognition
[0150] ;
[0151] As shown in Table 3, the recognition of complex intentions is more significantly improved because multimodal fusion compensates for the ambiguity of EEG in the representation of fine motor skills (e.g., electromyography can provide details of finger force exertion, and posture can reflect the spatial trajectory of the hand).
[0152] Table 4 Comparison of Signal Interruption Tolerance Rate
[0153] ;
[0154] As shown in Table 4, in the experimental group, when the signal was interrupted, electromyography (EMG) and postural features could temporarily replace electroencephalography (EEG) for decision-making (e.g., the peak EMG pattern during grasping has strong recognizability), while the control group's accuracy plummeted due to its sole reliance on EEG. In a typical sample comparison example, the control group had 5 incorrect cases due to EEG errors. β The wave (13-30Hz) characteristics were not obvious, and were misidentified as "force-controlled grip"; the experimental group's correction mechanism correctly identified all 5 samples through the "little finger abductor-polus pollicis brevis synergistic pattern" of electromyography (feature vector cosine similarity greater than 0.85). In the signal interruption samples, the control group had 7 cases of "multi-finger coordination" misidentified as "fist clenching" when the signal was interrupted for 200ms; the experimental group retained the correct identification of 6 of these cases through the "wrist pronation angle greater than 30°" feature of the posture sensor.
[0155] In the latency performance comparison, the average latency of the control group was 245ms; the average latency of the experimental group was 280ms (due to the increased computational load from multimodal fusion), which still met the real-time control requirements (≤300ms).
[0156] In patient testing, the accuracy rate of "fine grasping" recognition among patients with mild disabilities was 76.4% in the control group (poor EEG signal quality) and 89.7% in the experimental group (EMG and posture compensated for EEG deficiencies), an improvement of 13.3%, significantly better than the improvement in healthy individuals. These experiments demonstrate that multimodal fusion significantly improves performance in complex intentions and signal-unstable scenarios (p < 0.01), validating the effectiveness of the Transformer architecture in cross-modal information complementarity and avoiding the limitations of relying on single EEG features. Experimental data show that this improvement significantly enhances the robustness of exoskeleton robots in recognizing complex movement intentions while maintaining real-time performance. Breaking through the limitations of traditional single-modal or simple splicing and fusion methods, this technology automatically uncovers implicit correlations between EEG, EMG, and kinematic data through the Transformer's attention mechanism. Specifically, even when EEG signal quality is poor (e.g., signal-to-noise ratio <25dB), it maintains over 90% recognition accuracy by leveraging complementary information from EMG and kinematic modalities. Furthermore, it utilizes cross-attention to address the latency differences in multimodal data acquisition (EEG leads EMG by approximately 150ms), reducing the intention-action response synchronization error to ±50ms. This provides richer decision-making support for the "intention-execution-feedback" closed loop in rehabilitation training; for example, it corrects the execution intensity of EEG intentions based on muscle activity intensity, preventing excessively large movements.
[0157] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above-described embodiment of the brain-computer interface-based exoskeleton robot control method.
[0158] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0159] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described embodiments of the brain-computer interface-based exoskeleton robot control method.
[0160] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described embodiments of the brain-computer interface-based exoskeleton robot control method.
[0161] Those skilled in the art will understand that implementing all or part of the processes in the brain-computer interface-based exoskeleton robot control methods of the above embodiments can be accomplished by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the brain-computer interface-based exoskeleton robot control methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0162] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. A brain-computer interface-based exoskeleton robot control method, characterized in that, include: Brain-computer interface data is acquired, and adaptive filtering is used to remove power frequency interference and electromyographic noise from the brain-computer interface data to obtain denoised brain-computer interface data. Wavelet transform is used to decompose the denoised brain-computer interface data, and the decomposed brain-computer interface data is combined with preset frequency bands for data filtering to obtain preprocessed brain-computer interface data. Based on the common space pattern algorithm, motion feature vectors are obtained by extracting motion features from preprocessed brain-computer interface data that maximize the differences in motion intentions. An ensemble learning model is constructed and trained using pre-acquired brain-computer interface data to obtain a motion intention recognition model. The motion intent recognition model is used to identify the motion intent of the motion feature vector by combining the preset motion intent categories, and the motion intent recognition result is obtained. The motion intent recognition results are verified for confidence. Based on the confidence verification results, motion parameters are calculated by combining the preset motion parameter mapping rules and the pre-acquired human motion data. Motion control commands are generated according to the motion parameters, and motion control is performed using the motion control commands. The motion deviation is calculated from the acquired motion state feedback data, and the motion adjustment judgment is made by combining the motion deviation calculation result with a preset threshold to obtain the motion adjustment judgment result. Based on the motion adjustment judgment results, a dynamic adjustment algorithm is used to adjust the motion state parameters of the acquired motion state feedback data to obtain the motion state adjustment results. The motion control commands are dynamically adjusted based on the results of the motion state adjustment, and the motion control is optimized using the adjusted motion control commands.
2. The brain-computer interface-based exoskeleton robot control method according to claim 1, characterized in that, The co-space pattern algorithm is used to extract motion features from the preprocessed brain-computer interface data to maximize the difference in motion intent, resulting in motion feature vectors including: The covariance of motor intention is calculated based on the preprocessed brain-computer interface data to obtain the motor intention covariance matrix. The spatial filter matrix is then solved to obtain the motor intention spatial filter matrix. The motion intent spatial filter matrix is optimized using a pre-constructed objective function, and the variance difference of different motion intent signals is maximized based on the optimization results, thus obtaining the variance difference results of different motion intent signals. Motion features are extracted based on the variance differences of signals with different motion intentions, and motion feature vectors are generated from the extracted motion features.
3. The brain-computer interface-based exoskeleton robot control method according to claim 1, characterized in that, The construction of the ensemble learning model and the training of the ensemble learning model using pre-acquired brain-computer interface data to obtain the motion intention recognition model includes: Feature extraction is performed on the pre-acquired brain-computer interface data to obtain a brain-computer interface feature dataset, and the brain-computer interface feature dataset is divided into a training set, a validation set, and a test set based on a preset ratio; The convolutional neural network model, support vector machine model, and random forest model were trained and optimized using the training set to obtain optimized convolutional neural network model, optimized support vector machine model, and optimized random forest model. The optimized convolutional neural network model, the optimized support vector machine model, and the optimized random forest model are integrated in parallel to obtain the motion intention recognition model.
4. The brain-computer interface-based exoskeleton robot control method according to claim 3, characterized in that, The process of training and optimizing the convolutional neural network model, support vector machine model, and random forest model using the training set to obtain optimized convolutional neural network models, optimized support vector machine models, and optimized random forest models includes: The convolutional neural network model is initialized using a normal distribution to obtain an initial convolutional neural network model. The initial convolutional neural network model is then iteratively trained using the training set, and the convolutional neural network model is updated using the cross-entropy loss function and backpropagation to obtain an optimized convolutional neural network model. A support vector machine model with motion intent classification is constructed by kernel function selection and multi-classification strategy. The support vector machine model is trained based on the training set, and the objective function of the model is optimized by solving a convex quadratic programming problem to obtain an optimized support vector machine model. The training set is sampled with replacement to obtain a training subset, and a decision tree is constructed using classification and regression tree algorithms. The splitting features and thresholds of the optimized decision tree are selected based on the Gini impurity. The optimized decision tree is then combined with the training subset to train a random forest model, resulting in an optimized random forest model.
5. The brain-computer interface-based exoskeleton robot control method according to claim 1, characterized in that, The process of recognizing motion intent by combining a motion intent recognition model with a preset motion intent category to identify motion feature vectors and obtaining motion intent recognition results includes: The convolutional structure based on the motion intent recognition model, combined with the preset motion intent category, performs temporally correlated motion intent recognition on the motion feature vector, and obtains the temporally correlated motion intent classification probability. By using the radial basis function of the motion intent recognition model in combination with the preset motion intent categories, the motion feature vector is statistically identified to obtain the motion intent classification probability with statistical regularity. By combining the decision tree of the motion intent recognition model with the preset motion intent categories, motion intent with boundary features is recognized from the motion feature vector, and the motion intent classification probability with boundary features is obtained. The motion intent classification probabilities with temporal correlation, statistical regularity, and boundary features are integrated using a weighted voting method to obtain the motion intent recognition result.
6. The brain-computer interface-based exoskeleton robot control method according to claim 1, characterized in that, The process of adjusting motion state parameters based on motion adjustment judgment results using a dynamic adjustment algorithm on the acquired motion state feedback data to obtain motion state adjustment results includes: Based on the motion adjustment judgment results, the state value function and parameters of the dynamic adjustment algorithm are initialized to obtain the initialized dynamic adjustment algorithm; The initial dynamic adjustment algorithm is combined with the acquired motion state feedback data to select actions based on an exploration probability greedy strategy, and the action to be executed is obtained. The motion state adjustment parameters are updated by performing actions, and motion state feedback data of the performed actions is obtained. Real-time rewards are calculated based on the motion state feedback data of the performed actions, and the state value function is updated using the real-time reward calculation results. The motion state is adjusted by using the updated state value function and the value iteration optimization strategy, and the motion state adjustment result is obtained.
7. A brain-computer interface-based exoskeleton robot control system, characterized in that, The brain-computer interface-based exoskeleton robot control system includes: a motion feature extraction module, a motion control module, and a motion control optimization module; The motion feature extraction module is used to acquire brain-computer interface (BCI) data, remove power frequency interference and electromyographic noise from the BCI data using adaptive filtering to obtain denoised BCI data; decompose the denoised BCI data using wavelet transform, and filter the data based on the decomposed BCI data and preset frequency bands to obtain preprocessed BCI data; and extract motion features from the preprocessed BCI data based on a common spatial pattern algorithm to maximize the difference in motion intent, thereby obtaining motion feature vectors. The motion control module is used to construct an integrated learning model and train the integrated learning model using pre-acquired brain-computer interface data to obtain a motion intention recognition model; the motion intention recognition model is used to identify motion intentions from motion feature vectors by combining the motion intention recognition model with preset motion intention categories to obtain motion intention recognition results; the motion intention recognition results are used to perform motion intention confidence verification; based on the motion intention confidence verification results, motion parameters are calculated by combining the motion intention confidence verification results with preset motion parameter mapping rules and pre-acquired human motion data; motion control commands are generated according to the motion parameters; and motion control is performed using the motion control commands. The motion control optimization module is used to calculate motion deviation from the acquired motion state feedback data, and to make motion adjustment judgments by combining the motion deviation calculation results with preset thresholds to obtain motion adjustment judgment results; based on the motion adjustment judgment results, to adjust the motion state parameters of the acquired motion state feedback data using a dynamic adjustment algorithm to obtain motion state adjustment results; and to dynamically adjust the motion control commands according to the motion state adjustment results, and to optimize motion control through the adjusted motion control commands.
Citation Information
Patent Citations
Mechanical exoskeleton rehabilitation training system and method based on brain-computer interface
CN120514396A
Exoskeleton system assistance control method and device based on man-machine co-fusion strategy
CN120560495A