Virtual reality navigation rehabilitation training system and method
By combining EEG and EMG signals through a multimodal feature coupling network, a virtual reality navigation rehabilitation training system was realized, which solved the problem of low accuracy in motor intention recognition in stroke cognitive rehabilitation systems, improved recognition accuracy and system response speed, and promoted the rehabilitation training effect of patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF WENZHOU ZHEJIANG UNIV
- Filing Date
- 2026-04-30
- Publication Date
- 2026-05-29
AI Technical Summary
Existing stroke cognitive rehabilitation systems based on neurophysiological signals suffer from severe feature overlap when distinguishing between similar actions such as "going straight" and "stopping," resulting in limited recognition accuracy and making it difficult to achieve effective motor intention-driven rehabilitation training without relying on exoskeleton robot platforms.
A multimodal feature coupling network is used, which combines EEG and EMG signals. The signals are collected through virtual reality devices, and the multimodal feature coupling network is used for preprocessing and feature extraction. By fusing EEG rhythm features and muscle force features, the virtual character drives the movement task in the virtual movement task scenario, so as to achieve accurate recognition of movement intention.
It improves the accuracy of motor intention recognition and system robustness, better assists rehabilitation training, shortens system warm-up and response delays, ensures rapid and accurate response to patient control commands, and enhances stroke rehabilitation outcomes.
Smart Images

Figure CN122117239A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of rehabilitation training technology, and in particular to a virtual reality navigation rehabilitation training system and method. Background Technology
[0002] With the increasing number of patients suffering from lower limb paralysis due to stroke, spinal cord injury, and other causes, how to conduct cognitive-motor rehabilitation training in a safe and controllable environment to promote the recovery of cognitive control function in lower limb movement has become an important research question in the field of rehabilitation medicine. Compared with traditional rehabilitation methods that mainly rely on passive stretching and passive joint movement, active brain-computer interfaces (BCI / MCI) that combine multimodal neurophysiological signals can achieve "motor intention-driven" cognitive training by decoding the patient's motor intention patterns without relying on complex robotic platforms such as exoskeletons. This is considered a current research hotspot in the intersection of cognitive rehabilitation and artificial intelligence.
[0003] However, stroke cognitive rehabilitation systems based on neurophysiological signals still face serious challenges in practical applications. In terms of neurophysiological signal decoding and modeling, related technologies mainly rely on time-domain or simple frequency-domain feature extraction methods, leading to severe feature overlap when distinguishing between similar actions such as "moving straight" and "stopping," thus limiting recognition accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a virtual reality navigation rehabilitation training system and method that can improve the accuracy of motion intention recognition and better assist subjects in rehabilitation training.
[0005] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a virtual reality navigation rehabilitation training system, comprising: The multimodal signal acquisition module is used to: acquire target electroencephalogram (EEG) signals and target electromyogram (EMG) signals of the subject under the visual cues of a virtual reality device; The preprocessing module is used to preprocess the target EEG signal and the target EMG signal respectively to obtain the preprocessed EEG signal and the preprocessed EMG signal. The motor intention prediction module is used to: input preprocessed EEG signals and preprocessed EMG signals into a trained multimodal feature coupling network to obtain the subject's target motor intention prediction result; wherein, the multimodal feature coupling network includes an EEG network, an EMG network, a feature fusion layer, and a classifier; the EEG network is used to extract the power spectral density features of the preprocessed EEG signals in multiple rhythm frequency bands, and extract the EEG rhythm features after logarithmic transformation and batch normalization; the EMG network is used to extract features from the preprocessed EMG signals to obtain muscle exertion features; the feature fusion layer is used to: generate gating signals by mutual interaction of EEG rhythm features and muscle exertion features, perform cross-modal weighted modulation on each other, and enhance the cross-modal coupling of the modulated EEG features and modulated EMG features to obtain fused features; the classifier is used to predict motor intention based on the fused features to obtain the subject's target motor intention prediction result; The virtual reality navigation rehabilitation module is used to drive a virtual character to perform movement tasks in a virtual movement task scenario based on the subject's predicted movement intention.
[0006] Secondly, this application provides a rehabilitation training method based on the aforementioned virtual reality navigation rehabilitation training system, comprising: After the subjects wore the virtual reality device, the target EEG and target EMG signals were collected under the visual cues of the virtual reality device. The target EEG signal and target EMG signal are preprocessed separately to obtain preprocessed EEG signal and preprocessed EMG signal; The preprocessed EEG and EMG signals are input into a trained multimodal feature coupling network to obtain the target movement intention prediction results of the subject. Virtual characters are driven to perform motion tasks in virtual motion task scenarios based on the predicted motion intentions of the subjects.
[0007] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a virtual reality navigation rehabilitation training system and method. The multimodal feature coupling network, using signals from both electroencephalogram (EEG) and electromyography (EMG) levels as input, can extract features that more comprehensively express the subject's motor intentions, achieving complementary representation of central nervous system commands and peripheral motor performance, thus improving the accuracy of motor intention recognition. Compared to single-modality systems, this application can capture motor intentions more comprehensively, significantly improving decoding accuracy and system robustness under complex navigation tasks. It solves the problem of low recognition accuracy in related stroke cognitive rehabilitation systems based on neurophysiological signals, and can better assist subjects in rehabilitation training. Furthermore, this application uses a multimodal feature coupling network for motor intention prediction, which improves computational efficiency. It only requires collecting a very short period (seconds) of the patient's multimodal neurophysiological signals to accurately identify motor intentions. This "short-window start" mechanism greatly shortens system warm-up and response delays, ensuring rapid and accurate response to the patient's control commands and improving stroke rehabilitation outcomes. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort: Figure 1 This is a schematic diagram of the structure of a virtual reality navigation rehabilitation training system provided in one embodiment of this application; Figure 2 A schematic diagram of a virtual reality navigation rehabilitation training process for multimodal neurophysiological signal fusion decoding provided in an embodiment of this application; Figure 3 This is a schematic diagram of a data acquisition paradigm for the training phase provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the characteristic distribution of electroencephalogram (EEG) and electromyogram (EMG) signals during straight-line motion, provided in an embodiment of this application; wherein, Figure 4 (a) in the figure is a distribution map of EEG signal characteristics during straight-line movement; Figure 4 (b) in the figure shows the characteristic distribution of electromyographic signals during straight-line movement; Figure 5 A schematic diagram illustrating the principle of a multimodal feature coupling network provided in an embodiment of this application; Figure 6 This is a schematic diagram of a dual-band spectral domain modeling network provided in an embodiment of this application; Figure 7 This is a schematic diagram of the texture-energy dual-flow electromyography branching process provided in an embodiment of this application; Figure 8A graph showing the recognition accuracy in a four-category navigation task provided in an embodiment of this application; Figure 9 This is a classification confusion matrix diagram on a constructed dataset provided in one embodiment of this application. Detailed Implementation
[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0011] In one exemplary embodiment, such as Figure 1 As shown, a virtual reality navigation rehabilitation training system is provided, including a multimodal signal acquisition module, a preprocessing module, a motion intention prediction module, and a virtual reality navigation rehabilitation module.
[0012] The multimodal signal acquisition module is used to acquire target electroencephalogram (EEG) signals and target electromyogram (EMG) signals of subjects under the visual cues and guidance of virtual reality (VR) devices.
[0013] The preprocessing module is used to preprocess the target EEG signal and the target EMG signal respectively to obtain the preprocessed EEG signal and the preprocessed EMG signal.
[0014] The motor intention prediction module is used to: input preprocessed EEG signals and preprocessed EMG signals into a trained multimodal feature coupling network to obtain the subject's target motor intention prediction result; wherein, the multimodal feature coupling network includes an EEG network, an EMG network, a feature fusion layer, and a classifier; the EEG network is used to extract the power spectral density features of the preprocessed EEG signals in multiple rhythm frequency bands, and extract the EEG rhythm features after logarithmic transformation and batch normalization; the EMG network is used to extract features from the preprocessed EMG signals to obtain muscle exertion features; the feature fusion layer is used to: generate gating signals by mutual generation of EEG rhythm features and muscle exertion features, and perform cross-modal weighted modulation on each other, and perform cross-modal coupling enhancement on the modulated EEG features and modulated EMG features to obtain fused features; the classifier is used to predict motor intention based on the fused features to obtain the subject's target motor intention prediction result.
[0015] The motion intention prediction results include the predicted probability of each action category, and the action category corresponding to the highest predicted probability is taken as the final predicted action category; the action categories include stop, go straight, turn left and turn right.
[0016] The virtual reality navigation rehabilitation module is used to drive a virtual character to perform movement tasks in a virtual movement task scenario based on the subject's predicted movement intention. The module constructs a virtual hospital corridor scene using Unity, displays directional prompts via dynamic arrows at the top of the screen, and overlays the waveforms of the target's electroencephalogram (EEG) and electromyogram (EMG) signals, along with the predicted movement intention, in real time using the LSL protocol.
[0017] The system first guides the subject to generate corresponding multimodal neurophysiological signals, including target EEG and target EMG signals, through dynamic arrow prompts at the top of the screen. Then, it receives decoded motion commands (the decoded motion commands refer to the final predicted action category) via WiFi UDP protocol, driving the virtual character in real-time to perform path tasks such as walking straight, turning, and stopping in a hospital corridor scene. Simultaneously, it overlays and displays the multimodal neurophysiological signal waveforms and action categories in real-time using the LSL protocol. The screen displays the real-time straight-line distance of the virtual character from the target location, and updates visual cues and instructs the virtual character to stop and wait for further motion commands when it reaches the target location. The virtual hospital corridor scene is as follows: Figure 2 As shown.
[0018] The decoding model construction and virtual scene adaptive adjustment specifically include training and inference phases. The virtual reality navigation rehabilitation training process based on multimodal neurophysiological signal fusion decoding is as follows: Figure 2 As shown.
[0019] Training Phase: The multimodal feature coupling network extracts EEG and EMG features through independent branch networks and employs a feature-level fusion strategy for feature fusion and classification. Specifically, the EEG network captures movement intent through frequency domain energy analysis, while the EMG network uses a two-stream architecture to consider both motion texture and energy intensity. After extraction, a bidirectional gated modulation module achieves nonlinear coupling of multimodal signals. The fused features are then input into the classifier, and weighted cross-entropy loss and label smoothing techniques are introduced to address class imbalance issues (such as "going straight" versus "stopping") and improve the model's generalization ability.
[0020] The system also includes a model training module; the model training module is used to: acquire a dataset; the dataset includes sample EEG signals and sample EMG signals from several subjects, as well as the corresponding movement category label for each subject; input the sample EEG signals and sample EMG signals into a multimodal feature coupling network to obtain the sample movement intention prediction results; calculate the loss function value based on the sample movement intention prediction results and movement category labels; update the model parameters of the multimodal feature coupling network based on the loss function value to obtain the trained multimodal feature coupling network. Specifically, this is performed in the following steps S1 to S6.
[0021] Step S1: Constructing a multimodal neurophysiological signal dataset. Subjects wear appropriate acquisition devices to collect multimodal neurophysiological signals. Then, synchronized test segment data is extracted based on labels. Unique preprocessing procedures are applied to EEG and EMG signals respectively. Subsequently, a sliding window technique is used to divide the test segments into non-overlapping small data segments. Finally, training and validation sets are divided to obtain a dataset containing four types of movement commands: "left turn, right turn, straight ahead, and stop."
[0022] Step S1.1: Subjects wore a 64-channel EEG acquisition device and a 4-channel EMG acquisition device, and used an exoskeleton robot for motion assistance to simulate a natural standing state. Each motion acquisition cycle consisted of three phases: visual guidance, actual execution, and rest. The system cyclically presented these states until the end of the experiment. The specific process included: a 4-second visual guidance phase, where a central fixation point was displayed on the screen to guide the subject into an immersive walking preparation state; subsequently, arrow prompts for "stop," "go straight," "turn left," and "turn right" appeared randomly on the screen, guiding the patient to generate specific lower limb walking drive or turning intentions; a 16-second execution phase, where the subject performed a 16-second long-duration deep motor imagery based on the visual prompts, constructing a lower limb walking or turning motion program in the cerebral cortex, while simultaneously performing weak muscle contractions in place to activate specific rhythms in the sensorimotor cortex. The specific movement paradigm was as follows: Straight walking: The subject's dominant supporting leg is straight (simulating the supporting side), while the other leg exerts a slight force (simulating the lifting side), and the force is maintained for 16 seconds; Stop: Keep both legs straight and exert force continuously for 16 seconds; Turn left: Straighten your right leg (simulating the supporting side), exert a slight force with your left leg (simulating the lifting side), and continue exerting force for 16 seconds; Turn right: Straighten your left leg (simulating the supporting side), exert a slight force with your right leg (simulating the lifting side), and continue exerting force for 16 seconds.
[0023] During the 4-second rest period, the screen dims, and the subjects stop imagining and relax their brains and leg muscles to eliminate the neural inertia of the previous task and prepare for the next round of the experiment.
[0024] The experiment was divided into three blocks, each containing 48 action tasks covering all action categories. To avoid fatigue effects, the order of the actions was randomly shuffled, and a rest period of at least 5 minutes was set between blocks. The total experiment duration was 40 minutes. This workflow design ensured standardized action execution and consistent data collection timing, contributing to improved data quality and experimental controllability.
[0025] Based on the above process, a dataset simulating a real stroke rehabilitation training environment based on multimodal neurophysiological signals was constructed. This dataset includes 64-channel EEG signal sequences (denoted as raw EEG signals) and 4-channel EMG signal sequences (denoted as raw EMG signals) from 30 subjects, along with corresponding action labels. The data covers four types of lower limb movements: stopping, walking straight, turning left, and turning right. The data acquisition paradigm during the training phase is as follows: Figure 3 As shown, the characteristic distribution of EEG and EMG signals during straight-line walking is as follows: Figure 4 As shown.
[0026] Step S2, Differentiated preprocessing of multimodal signals. Differentiated preprocessing procedures are designed to maximize the signal-to-noise ratio based on the spectral characteristics of different physiological signals.
[0027] EEG: The raw EEG signal was bandpass filtered from 1.0 Hz to 40.0 Hz using an FIR / IIR filter to remove low-frequency drift and high-frequency noise; a notch filter was set at 50 Hz to eliminate power line interference; Independent Component Analysis (ICA): The effective EEG channel data was extracted and decomposed using Fast Independent Component Analysis (FICA) to automatically identify and remove the electrooculogram (EOG) component; Common Average Reference (CAR): An average rereference was applied to the effective EEG channels to eliminate common-mode interference and obtain the sample EEG signal.
[0028] This application uses an EEG recording device to acquire 64-channel EEG signals and a surface electromyography (EMG) device to acquire 4-channel EMG signals from the calves and thighs. The signal spectrum acquired by these devices is mainly distributed in the following ranges: EEG (1Hz-40Hz) and EMG (20Hz-300Hz). A 16-second synchronized segment is extracted based on action category labels and numbered. EEG (59 channels, effective channel count) and EMG (4 channels) data matrices with corresponding indices are extracted to construct a standardized dataset. The 16-second segment is then divided into eight 2-second segments, each segment consisting of one raw EEG or raw EMG signal. All data segments are randomly divided into training and validation sets in an 8:2 ratio. To prevent data leakage, data segments from the same segment will not appear simultaneously in both the training and validation sets. The validated multimodal feature coupling network (the best trained model) is obtained by using the validation set and meeting the validation requirements.
[0029] The raw EEG signal was sequentially bandpass filtered (1-40Hz) and then notched at 50Hz. To eliminate common-mode interference, an averaged rereference was used to obtain... Sample EEG signals at time points : (1); in, and The first and EEG signals from the channel; The number of effective EEG channels.
[0030] Electromyography (EMG): Considering the high-frequency characteristics of muscle activation signals, the raw EMG signal was subjected to a broadband pass filter from 20.0 Hz to 300.0 Hz to retain rich high-frequency texture and transient burst features, while filtering out motion artifacts, thus obtaining the sample EMG signal. The raw EMG signal was then bandpass filtered from 20-300 Hz to preserve the texture details of muscle contraction, resulting in the sample EMG signal. .
[0031] (2); in, Indicates bandpass filtering. For frequency.
[0032] Step S3: Model the multimodal feature coupling network architecture, such as... Figure 5 As shown.
[0033] Step S3.1: Extraction of EEG and EMG signal features.
[0034] The EEG network includes a Fast Fourier Transform unit, a power spectral density calculation unit, a logarithmic transform unit, a first splicing unit, a Batch Normalization (BN) layer, and Fully Connected Layers (FC).
[0035] The Fast Fourier Transform (FFT) unit maps the preprocessed EEG signal from the time domain to the frequency domain to obtain frequency domain features. The power spectral density (PSD) calculation unit extracts PSD features for different rhythm frequency bands based on the PSD features. The logarithmic transform unit performs a logarithmic transform on the PSD features of different rhythm frequency bands to obtain the frequency band features of each rhythm frequency band. The dynamic range is compressed by the logarithmic transform to Gaussianize the feature distribution. The first concatenation unit concatenates the frequency band features of different rhythm frequency bands to obtain the original spectral domain feature vector. The batch normalization (BN) layer performs batch normalization on the original spectral domain feature vector to obtain the batch normalized feature vector. The fully connected (FC) network inputs the batch normalized feature vector into the fully connected coding network to extract EEG rhythm features.
[0036] The EEG network employs a dual-band spectral network (DBSN), with rhythm frequency bands including... Main frequency bands and The main frequency band of the wave, The main frequency band of the wave is 8Hz to 14Hz. The main frequency band of the wave is 15Hz to 25Hz.
[0037] like Figure 6 As shown, firstly, preprocessed EEG signals are acquired. The frequency resolution is dynamically calculated based on the current sampling frequency and sample length. The formula is as follows: (3); in, The system sampling rate, This represents the number of time axis sampling points for the current input data segment. It is the set of real numbers.
[0038] Subsequently, the system can call the Real-Number Fast Fourier Transform (rFFT) to map the preprocessed EEG signal, which is currently a time-domain signal, to the positive frequency range, and calculate the amplitude spectrum to obtain the energy intensity at each frequency point, as shown in the following formula: (4); in, Frequency index; These are time-domain sampled values; The imaginary unit; Pi is the mathematical constant of a circle.
[0039] Furthermore, the system uses preset functional frequency bands ( Main frequency band: 8Hz-14Hz The main frequency band of the wave is 15Hz-25Hz, and the calculated frequency resolution is as follows. Dynamically determine the frequency band index range: (5); in, and These are the lower and upper limits of the frequency band index range, respectively. The target physical frequency; This represents the floor operator.
[0040] Subsequently, the system, based on the rhythmic features related to motion imagery, indexed the frequency band range [ Summing the amplitude spectrum (power spectral density) within the range of ] is used to compress the dynamic range of signal features and reduce non-stationarity caused by individual differences, thus extracting the signal. The dimensional features are subjected to logarithmic transformation, and a nonlinear logarithmic operator is introduced for feature compression and normalization to obtain the frequency band features of a single channel. : (6); in, Indicates the channel index; For a minimum bias constant (if possible) ), used to ensure numerical stability.
[0041] Next, the channels obtained from the aforementioned calculations will be... Main frequency band characteristics and Main frequency band characteristics of the wave Vector concatenation is performed to construct the original spectral domain feature vector. The formula is as follows: (7); in, This represents the total number of effective EEG acquisition channels (59 in this embodiment). It has a specific dimensional space to accommodate multi-channel dual-rhythm features.
[0042] To eliminate feature scale differences across samples and accelerate model convergence, the BN layer modifies the original spectral domain feature vectors. Batch normalization processing: (8); in, and These are the mean vector and variance vector of the EEG features within the current training batch, respectively. and These are the learnable affine transformation gain and bias parameters, respectively. is the numerical stability constant.
[0043] Finally, the feature vectors after batch normalization. The data is fed into a fully connected coding network (i.e., a fully connected layer). Through nonlinear mapping of hidden layer neurons (reducing the dimensionality from 118 dimensions to 64 dimensions, and then further to 32 dimensions), a nonlinear topological transformation is performed to extract deep neural representations. This achieves dimensionality reduction extraction from the original physical frequency domain features to high-dimensional neural dynamic features, ultimately outputting 32-dimensional EEG rhythm features. The formula is as follows: (9); in, and In the fully connected layer Layers and Layer weight matrix; and In the fully connected layer Layers and The layer's dedicated bias vector; It is a non-linear activation function; It is a random deactivation operator.
[0044] Texture-Energy Network (TEN): This architecture constructs a parallel processing architecture for both texture and energy pathways, targeting the transient characteristics and spatial distribution of electromyographic signals. For example... Figure 5 , Figure 7 As shown, the electromyographic signals after the aforementioned preprocessing It will enter the texture pathway and the energy pathway in parallel.
[0045] Electromyographic networks include texture pathways, energy pathways, second splicing units, and dimensionality reduction units.
[0046] 1) The texture pathway includes an instance normalization layer, a first convolutional layer, a first ReLU layer, a second convolutional layer, a second ReLU layer, a squeeze-and-excitation (SE) attention layer, and a global average pooling layer connected in sequence, which are used to extract muscle texture vectors based on preprocessed electromyographic signals.
[0047] First, the instance normalization layer independently normalizes and maps the electromyographic channels within each preprocessed electromyographic signal, ensuring that the characteristics of different subjects are mapped to a uniform distribution space to eliminate absolute amplitude deviations caused by individual skin impedance and electrode tension. The formula is as follows: (10); in, For the normalized first Each electromyographic channel in The sampled value at time; For the first Each electromyographic channel in The original sampled value at time; and The first Mean and variance of electromyographic channels during the current 2-second test segment; This is a minimal bias term specific to this branch to prevent the denominator from being zero; and These are learnable scaling and translation parameters.
[0048] Subsequently, the normalized electromyographic signals The data is fed into a feature extractor consisting of multiple one-dimensional convolutions, through a process with a size of [missing information]. The convolutional kernel captures the microscopic waveform of muscle contraction. The feature extractor includes a first convolutional layer, a first ReLU layer, a second convolutional layer, and a second ReLU layer. The formula for the first convolutional layer is as follows: (11); in, For the texture feature map, the first Each channel is in The response value at any given time; This is the dedicated weight matrix for the first convolutional layer; This is the dedicated bias term corresponding to the first convolutional layer; The number of input channels is set to 4; The kernel size is set to 7. The step size is set to 2; The nonlinear activation function selected for this branch.
[0049] After two convolutional layers, the importance of feature channels is adaptively recalibrated through a channel attention mechanism to achieve deep feature extraction.
[0050] First, compression is performed. Global average pooling is then applied to each channel over time to generate channel-level statistics. (12); in, For the first Compression characteristics of each channel, ; This is the original feature map output by the feature extractor.
[0051] A subsequent excitation operation is performed. A squeeze-and-excitation (SE) attention layer is introduced to compute channel attention weights, thereby adaptively enhancing muscle texture features highly correlated with left / right turning movements. Channel attention weights are generated through two fully connected layers: (13); in: This is the weight matrix of the first fully connected layer (dimension reduced, reduction=4). This is the weight matrix for the second fully connected layer (increased dimension). These are the bias terms for the two fully connected layers (bias=False in this embodiment); The Sigmoid activation function has an output range of... This is the channel attention weight vector.
[0052] Finally, a scaling operation is performed, combining the generated channel attention weight vector with the original feature map. Perform channel-by-channel multiplication: (14); in, This is the channel attention weight vector corresponding to the channel.
[0053] Subsequently, convolutional feature maps After global average pooling, the time dimension is compressed to 1, thus outputting a 64-dimensional muscle texture vector that reflects the muscle exertion pattern. : (15); in, The time axis length of the convolutional feature map.
[0054] 2) Energy pathway, used for: calculating the channel time-averaged energy, inter-channel standard deviation, and left-right leg energy ratio based on preprocessed electromyography (EMG) signals; where the left-right leg energy ratio is the ratio of the average energy of the left leg to the average energy of the right leg; concatenating the channel time-averaged energy, inter-channel standard deviation, and left-right leg energy ratio to obtain spatial features; dividing the preprocessed EMG signals into several sub-windows in the time dimension; calculating the absolute average amplitude of each sub-window to obtain the energy level of each sub-window; and calculating the variance based on the energy levels of all sub-windows to obtain the force exertion dynamics. The system calculates the temporal fluctuation characteristics of left-right differences, channel range, temporal slope, and peak ratio based on the preprocessed electromyography (EMG) signals. The temporal characteristics are composed of the dynamic fluctuation characteristics of force exertion, the temporal fluctuation characteristics of left-right differences, the temporal fluctuation characteristics of channel range, temporal slope, and peak ratio. The system calculates the root mean square logarithm of each channel based on the preprocessed EMG signals to obtain the global instantaneous activation energy characteristics. The spatial characteristics, temporal characteristics, and global instantaneous activation energy characteristics are concatenated to obtain the spatiotemporal energy characteristics. The spatiotemporal energy characteristics are then processed through a linear mapping layer to obtain the muscle energy vector.
[0055] 2.1) Spatial Feature Extraction in the Energy Path: For the steady-state energy difference between "stopping" and "straight-through" states, 14-dimensional spatiotemporal features were explicitly calculated. The spatial features include six dimensions such as channel average / standard deviation and left-right energy ratio. In the temporal dimension, the signal was divided into eight equal-length sub-windows, each 0.25s long, and eight temporal features, including energy variance (Temporal Var) and peak-to-average ratio (PAR), were extracted. Subsequently, the spatial and temporal features were concatenated to form a 14-dimensional spatiotemporal energy feature.
[0056] Within the energy pathway, spatial statistical features are first extracted over the entire time period to characterize the overall activation level of each electromyographic channel and the asymmetry of force exertion between the left and right limbs. Let the length of a single test segment be... One sampling point, Indicates the first Each electromyographic channel at time The preprocessed electromyographic signal, then the first The time-averaged energy of each channel within this test section Defined as: (16).
[0057] To measure the dispersion of the overall activation distribution across the four channels, the standard deviation of the channel average energy is further calculated. : (17); in, This is the arithmetic mean of the time-averaged energies of the four channels. It reflects the degree of spatial imbalance in the force exertion intensity of different muscle channels.
[0058] Considering that this application focuses on the coordinated force exertion pattern of the left and right lower limbs, channels 1 and 3 correspond to the left leg muscle group, and channels 2 and 4 correspond to the right leg muscle group, the average energy of the left leg is defined as follows: Average energy of the right leg for: (18).
[0059] Based on this, in order to quantify the asymmetry of force exerted on the left and right sides, the energy ratio of the left and right legs is introduced. index: (19); in, To prevent the minimum bias constant with a denominator of zero (taken in this embodiment) ).
[0060] Ultimately, the spatial characteristic vector of the energy pathway is determined by the time-averaged energy of each channel. Standard deviation between channels and the energy ratio of the left and right legs The concatenation yields a 6-dimensional feature vector. : (20).
[0061] 2.2) Temporal Feature Extraction in the Energy Pathway: First, the system divides the 2-second raw electromyographic sequence into eight non-overlapping, equal-length sub-windows using a time-domain slicing operator. For the first... Each sub-window is used to calculate its absolute average amplitude to extract the energy level of that time-domain segment. : (twenty one); in, for The amplitude of the electromyographic signal after time-phase preprocessing; The length of a single sub-window sampling point; This is the index number of the current child window.
[0062] Subsequently, the system captures the dynamic fluctuation trend of force over time by calculating the time variance of the energy of all sub-windows, as shown in the following formula: (twenty two); in, For channel The corresponding dynamic fluctuation characteristics of force exertion; This is the arithmetic mean of the energy of all sub-windows; The total number of sub-windows divided within the 2s test segment (set to 8 in this embodiment).
[0063] Furthermore, the system calculates the temporal fluctuation characteristics of left-right differences, channel ranges, temporal slopes, and peak-to-peak ratios to comprehensively characterize the temporal dynamics of electromyographic signals. Simultaneously, it calculates the root mean square logarithm of each channel to represent the global instantaneous activation energy. (twenty three); in, For channel The global instantaneous activation energy characteristics; Indicates the channel index; It is the minimum bias constant.
[0064] Finally, the system concatenates the spatial feature vector and the temporal feature vector to output a complete 14-dimensional spatiotemporal feature vector. : (twenty four); (25); in, This represents a splicing operation; It has 6-dimensional spatial features; It has 8-dimensional time features; The time fluctuation characteristics of the left-right difference; The time fluctuation characteristics of the channel range; The time slope; This represents the peak ratio.
[0065] Spatiotemporal feature vectors A 16-dimensional muscle energy vector is obtained after a linear mapping layer. .
[0066] 2.3) The second stitching unit is used to: stitch the 64-dimensional muscle texture vector... With 16-dimensional muscle energy vector The concatenation yields an 80-dimensional muscle concatenation feature vector; a dimensionality reduction unit is used to reduce the dimensionality of the muscle concatenation feature vector through a linear layer to obtain a 16-dimensional muscle exertion feature. Dimensionality reduction units can be composed of fully connected layers.
[0067] Step S3.2: Feature-level cross-modal fusion layer and four-class classification.
[0068] The feature fusion layer comprises a bidirectional gating modulation module, a third splicing unit, and a cross-enhancement sub-network. The bidirectional gating modulation module is used to: generate gating signals for EEG rhythm features and muscle exertion features using the gating module; perform cross-modal weighted modulation of the EEG rhythm features based on the gating signals for muscle exertion features to obtain modulated EEG features; and perform cross-modal weighted modulation of the muscle exertion features based on the gating signals for EEG rhythm features to obtain modulated electromyographic features. The third splicing unit is used to splice the modulated EEG features and the modulated electromyographic features to obtain joint features. The cross-enhancement sub-network (i.e., Figure 5 The feature enhancement module (in the code) is used to: sequentially enhance the joint features through a linear layer, a ReLU layer, and a Dropout layer to obtain enhanced features; and fuse the enhanced features and joint features based on the enhancement coefficients to obtain fused features. The gating module includes a linear layer, a ReLU layer, a linear layer, and a Sigmoid layer.
[0069] Step S3.21, Generate intermodulation gating signals: Utilizing EEG rhythm characteristics Generate features for modulating muscle exertion. The gating vector, i.e., the gating signal of the brainwave rhythm characteristics. The formula is as follows: (26); in, This represents the Sigmoid activation function, which maps the output value to... interval; and These are the weight matrices for the hidden layer (i.e., the first linear layer of the gated module) and the output layer (i.e., the second linear layer of the gated module), respectively.
[0070] Step S3.22: Implement weighted residual modulation and adaptively weight the original features using the obtained gated signal. To ensure training stability, a modulation intensity coefficient is introduced. And residual connections, to obtain modulated electromyographic characteristics. : (27); in, This represents element-wise product. Similarly, it utilizes the characteristics of muscle exertion. Gating signals for generating muscle exertion characteristics The brainwave rhythm characteristics were modulated to obtain modulated brainwave characteristics. .
[0071] Step S3.23, splicing and feature-level enhancement: The modulated bimodal features (including modulated EEG features and modulated EMG features) are spliced together to obtain joint features. .
[0072] Finally, in order to capture deep cross-modal coupling information, The data is fed into the enhancement subnetwork. The enhancement subnetwork consists of enhancement layers and fusion layers. The enhancement layers possess cross-modal capabilities with linear transformation, nonlinear activation, and Dropout regularization (dropout rate of 0.2). Residual connections are used to alleviate the gradient vanishing problem in deep networks. Enhancement features are then calculated. : (28); in, This is the mapping matrix for the enhancement layer. Then, the enhancement coefficients are used... The enhanced feature is fed back to each modality to obtain the fused feature. : (29).
[0073] Step S3.24, fuse the features The data is fed into the classifier. The classifier consists of an encoding layer (Embedding Layer), a linear layer, and a classification head. Features are then fused. The input is fed into the encoding layer to obtain a stable encoded vector. This encoded vector is then fed into the linear layer to obtain a linear vector. Finally, the linear vector is fed into the classification head to obtain a four-class motion intent probability vector, which is the sample motion intent prediction result. This sample motion intent prediction result includes the predicted probability of each action category. The encoding layer consists of a linear mapping layer, a non-linear activation (ReLU) layer, and a dropout layer: the linear layer is used to project the fused features onto a fixed-dimensional embedding space.
[0074] The embedding vector is obtained after passing through a linear layer. : (30); in, and This represents the weights and biases of the linear layer. This embedding vector... The data is then fed into the output layer and mapped to a 4-category classification probability distribution, which represents the predicted motion intent of the sample. : (31); in, The classification head weights are used. The output of the classification head is fed into the Softmax activation function, which outputs a probability distribution vector for the four action classes. During training, the system calculates a weighted focal loss based on the sample motion intent prediction results and embeds the vectors... Calculate the center loss to ensure that similar features are highly aggregated in Euclidean space.
[0075] Step S4: Construct the joint objective function of weighted discriminant constraints and central clustering constraints. Next, weighted discriminative constraints and end-to-end training are constructed. To address the classification bias caused by the extreme similarity between "go straight" and "stop" intentions in physical space, this application constructs a class-weighted cross-entropy loss function, as shown in the following formula: (32); in, This represents the total number of samples in the current batch. For the first Each sample in its true category The predicted probability; Specific weighting factors are set for different action categories. Based on the prior distribution of the categories, this application sets the weighting factor for straight-moving actions to 0.8, the weighting factor for stopping actions to 2.0, and the weighting factor for other actions to 1.0, so as to impose stronger discriminative constraints on easily confused categories such as "stop / straight-moving". By applying a weighting factor of 0.8 to the prediction error of the "straight-moving" category and a weighting factor of 2.0 to the prediction error of the "stopping" category during backpropagation, the model is forced to learn the subtle differences in the temporal dynamic features of the two types of actions. At the same time, the label smoothing technique with a coefficient of 0.1 is used to transform the one-hot encoded label into a soft distribution, suppressing model overfitting caused by multimodal signal noise and enhancing the discriminative robustness of each action cluster in the feature space.
[0076] Furthermore, to suppress overfitting of the model to specific training samples and mitigate the risk of overfitting after multimodal fusion, a label smoothing strategy is introduced into the weighted cross-entropy loss to calibrate the target distribution, as shown in the following formula: (33); in, The smoothed target probability distribution; This is the smoothing factor for the fusion stage (set to 0.1 in this embodiment). The total number of action categories; The vector is composed entirely of 1s. After label smoothing, the classification loss can be written as the distribution form of the weighted focus loss function: (34); in, This is the weighted focus loss function; Action category The corresponding weighting factor; Action category The corresponding smoothed target probability distribution; For the first Each sample in the action category The predicted probability is that a set of sample EEG signals and sample EMG signals correspond to one sample.
[0077] Weighted Focal Loss assigns different weights to each class, with higher weights indicating that the class is "more important / harder to learn." During training, errors in this class are amplified, thus mitigating the impact of class imbalance. Center Loss measures the distance of each sample to its class center vector. During training, it aims to bring features of samples belonging to the same class as close as possible to their respective class centers, making features within the same class more compact and reducing intra-class differences; and making feature distributions more separated between different classes, resulting in clearer inter-class boundaries.
[0078] Furthermore, to make similar samples more compact in the embedding space and the inter-class boundaries clearer, this application introduces a center loss to constrain the intra-class aggregation of the embedding vectors, in addition to the classification loss. The center loss function is defined as follows: (35); in, For center loss function; For the first The embedding vector obtained by linear mapping of each sample before the classifier head; Action category The corresponding center vector.
[0079] The loss function value is calculated from the total loss function. Finally, in this embodiment, the weighted focus loss and the center loss are jointly optimized to obtain the total loss function: (36); in, The value of the loss function; The center loss is a tradeoff coefficient used to control the strength of the center clustering constraint relative to the discriminative learning.
[0080] Step S5: End-to-End Joint Optimization Based on Backpropagation: This embodiment employs an end-to-end training strategy based on backpropagation. In each batch iteration, the system first inputs EEG and EMG features into the cross-modal fusion network to obtain the four-class prediction distribution. and the corresponding embedding vector Then, the weighted focal loss and the center loss are calculated separately, and the total loss function is used as the objective function for gradient backpropagation.
[0081] Regarding updating the model parameters of the multimodal feature coupling network, the model training module is used to: collaboratively update the model parameters of the EEG network, EMG network, feature fusion layer, and classifier; and individually update the center vector corresponding to each action category. In terms of parameter update strategy, this embodiment separates and optimizes the "main network parameters" and "center vector parameters."
[0082] 1) The Adam optimizer is used to collaboratively update all learnable main network parameters, including EEG network, EMG network, feature fusion and enhancement layer, classification layer, etc.: The Adam optimizer is used to perform end-to-end joint training on all network parameters by minimizing the joint loss function composed of class-weighted focus loss and center loss, and save the best offline training model, i.e. the trained multimodal feature coupling network.
[0083] 2) Regarding the center vector parameters An independent optimizer (e.g., SGD) is used for updates to stabilize the convergence of the center vector and avoid interference with the adaptive momentum estimation of the main network parameters. (37); in, The learning rate is the value of the center vector. The gradient of the center vector. Through the end-to-end training mechanism of "joint backpropagation and separate update" described above, the classification loss improves the discriminative power between classes, and the center loss enhances the compactness within classes, thereby significantly reducing the misclassification rate and improving the overall robustness in easily confused intentions such as "stop / go straight".
[0084] To ensure stable convergence of the model when processing physiological signals with strong non-stationarity, the learning rate is dynamically adjusted, initially set to [value missing]. (Effective adjustment range is) An adaptive gradient scaling mechanism is used to balance gradient differences between different modes; at the same time, the batch size is set to... This aims to improve the utilization of the computing device's memory while ensuring the accuracy of gradient estimation; furthermore, by setting a fully connected layer... The random deactivation operator, combined with the weight decay coefficient ( This effectively suppressed the overfitting of the model to the physiological noise of specific subjects and enhanced the generalization performance across samples.
[0085] Step S6: As Figure 8 As shown, the trained multimodal feature coupling network exhibits excellent convergence characteristics. In the early stages of training, the loss function value decreases rapidly, while the recognition accuracy steadily increases with each iteration, reaching a performance plateau after the 43rd iteration. Experimental data indicates that the system achieves a four-class intent recognition accuracy of 87.93% on the target test set, with end-to-end inference latency controlled within a certain range. Within this range, the physiological real-time requirements for real-time interactive control of lower limb exoskeleton robots are met.
[0086] like Figure 9 As shown, through large-scale data testing on 30 subjects, the normalized confusion matrix generated by the system shows that samples of each category are highly clustered on the main diagonal.
[0087] (II) Inference Phase: The trained multimodal feature coupling network receives preprocessed multimodal neurophysiological signal sequences guided by the VR device in real time. It extracts key representations from the trained network branches and performs feature fusion. Finally, the classifier outputs the predicted probabilities and movement intentions for four navigation actions: "turn left, turn right, go straight, and stop." The virtual hospital corridor scene adaptively adjusts scene parameters (such as specific scene, viewpoint, and route) based on the predicted movement intentions.
[0088] In the online phase (i.e., the inference phase), this embodiment presents a closed-loop interactive process of "VR task guidance - multimodal decoding - VR adaptive adjustment" to achieve real-time guidance and generation of the subject's movement intentions (stop / forward / left turn / right turn) in a virtual hospital corridor scene, and to accurately identify and control them.
[0089] Subjects wore VR glasses, a 64-channel EEG acquisition device, and a 4-channel EMG acquisition device, standing still while waiting for the VR device to start. The VR glasses then loaded a virtual hospital corridor scene constructed using Unity, displaying a visual cue with a directional arrow centered at the top of the screen to guide the subject to generate enhanced EEG and EMG responses. While the subject viewed the visual cue and generated movement intentions, target EEG and EMG signals were simultaneously acquired and transmitted in real-time via WiFi to the online inference terminal. Simultaneously, the Unity terminal received the signals via the LSL protocol and dynamically drew a transparently overlaid EEG / EMG waveform in the upper left corner of the screen.
[0090] The target EEG and EMG signal data are subjected to a preprocessing procedure consistent with offline training to obtain preprocessed EEG and EMG signals. The preprocessed EEG and EMG signals are then input into the trained multimodal feature coupling network (the best model after training) to output the target motion intention prediction result.
[0091] The predicted motion intent of the target is sent to Unity via UDP commands. Unity receives the data in real time and drives the virtual character to perform the corresponding motion, which in turn changes the virtual scene parameters. This simultaneously updates the UI prompts and the motion status display in the upper right corner of the screen, achieving a closed-loop interaction of "VR task guidance - multimodal decoding - VR adaptive adjustment".
[0092] First, regarding the command dimension, traditional methods mostly focus on binary classification recognition at the left and right hand limb level, making it difficult to distinguish the cognitive intentions of complex lower limb rehabilitation movements such as "turn left, turn right, go straight, and stop," especially since "turning" and "going straight" have serious overlap and confusion in the feature space. Second, in terms of neurophysiological signal decoding and modeling, related technologies are mainly based on time-domain or simple frequency-domain feature extraction methods, often ignoring the dynamic energy evolution of different rhythms in specific functional frequency bands. Furthermore, the fusion of multimodal neurophysiological signals often adopts simple splicing or cascade fusion, failing to fully consider the complementary relationship between the two types of signals in terms of time scale, spatial distribution, and semantic level. There is a lack of targeted construction of cross-modal alignment and co-enhancement mechanisms, resulting in serious feature overlap when the model distinguishes similar movements such as "going straight" and "stopping," thus limiting the recognition accuracy. Finally, in terms of application paradigm, most studies are limited to gait simulation on treadmills, generally adopting non-continuous gait execution paradigms (such as the stride-to-touch pattern), and spatial guidance mostly relies on static external display devices (such as fixed displays), making it difficult to achieve continuous gait spatial guidance capabilities in real-world environments.
[0093] To overcome the aforementioned limitations, this application constructs a controllable and interactive motion task scenario using VR devices. Visual task prompts guide and enhance the subject's EEG and EMG responses. The collected multimodal signals are preprocessed and input into a fusion decoding model, which outputs a classification result of the motion intention. The decoding result further drives the scene parameters such as scene and perspective in the virtual environment to adaptively update in real time, so as to realize the simulation and feedback training of motion behavior, thereby forming a closed-loop human-computer interaction navigation rehabilitation process of "VR task guidance - multimodal decoding - VR adaptive adjustment".
[0094] The beneficial effects of this application are mainly reflected in: 1. Multimodal Heterogeneous Fusion Perception: By constructing a multimodal feature coupling network including EEG and EMG, complementary representations of central nervous system commands and peripheral motor performance are achieved. Compared with single-modality, this application can capture motor intentions more comprehensively, significantly improving decoding accuracy and system robustness under complex navigation tasks; 2. Highly efficient real-time response and rapid startup: This application boasts extremely high computational efficiency and real-time performance, accurately identifying motor intentions by collecting only a very short period (seconds) of multimodal neurophysiological signals from the subject. This "short window startup" mechanism significantly shortens system warm-up and response delays, ensuring rapid and accurate response to the subject's control commands and improving stroke rehabilitation outcomes. 3. Enhanced Immersive Vision-Motor Coupling: Employing a virtual reality-guided rehabilitation mode, a vision-motor closed loop is constructed, enhancing sensorimotor cortex activation through real-time visual cues. The system replaces the physical environment with a highly realistic virtual scene, achieving "zero-risk" simulation of real rehabilitation training, balancing the training safety and rehabilitation effectiveness for stroke subjects.
[0095] Based on the same inventive concept, this application also provides a rehabilitation training method based on the virtual reality navigation rehabilitation training system described above. The solution provided by this method is similar to the solution described in the system above; therefore, the specific limitations in one or more rehabilitation training method embodiments provided below can be found in the limitations of the virtual reality navigation rehabilitation training system described above, and will not be repeated here.
[0096] In one exemplary embodiment, a rehabilitation training method based on the above-described virtual reality navigation rehabilitation training system is provided, comprising the following steps: After the subjects wore the virtual reality device, the target EEG and target EMG signals were collected under the visual cues of the virtual reality device. The target EEG signal and target EMG signal are preprocessed separately to obtain preprocessed EEG signal and preprocessed EMG signal; The preprocessed EEG and EMG signals are input into a trained multimodal feature coupling network to obtain the target movement intention prediction results of the subject. Virtual characters are driven to perform motion tasks in virtual motion task scenarios based on the predicted motion intentions of the subjects.
[0097] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of the relevant data are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the corresponding device.
[0098] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0099] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A virtual reality navigation rehabilitation training system, characterized in that, The system includes: The multimodal signal acquisition module is used to: acquire target electroencephalogram (EEG) signals and target electromyogram (EMG) signals of the subject under the visual cues of a virtual reality device; The preprocessing module is used to preprocess the target EEG signal and the target EMG signal respectively to obtain the preprocessed EEG signal and the preprocessed EMG signal. The motor intention prediction module is used to: input preprocessed EEG signals and preprocessed EMG signals into a trained multimodal feature coupling network to obtain the subject's target motor intention prediction result; wherein, the multimodal feature coupling network includes an EEG network, an EMG network, a feature fusion layer, and a classifier; the EEG network is used to extract the power spectral density features of the preprocessed EEG signals in multiple rhythm frequency bands, and extract the EEG rhythm features after logarithmic transformation and batch normalization; the EMG network is used to extract features from the preprocessed EMG signals to obtain muscle exertion features; the feature fusion layer is used to: generate gating signals by mutual interaction of EEG rhythm features and muscle exertion features, perform cross-modal weighted modulation on each other, and enhance the cross-modal coupling of the modulated EEG features and modulated EMG features to obtain fused features; the classifier is used to predict motor intention based on the fused features to obtain the subject's target motor intention prediction result; The virtual reality navigation rehabilitation module is used to drive a virtual character to perform movement tasks in a virtual movement task scenario based on the prediction results of the subject's movement intention.
2. The virtual reality navigation rehabilitation training system according to claim 1, characterized in that, The brainwave network is used for: The preprocessed EEG signal was mapped from the time domain to the frequency domain using Fast Fourier Transform to obtain frequency domain features; Power spectral density features of different rhythm frequency bands are extracted based on frequency domain features; Logarithmic transformations were performed on the power spectral density characteristics of different rhythm frequency bands to obtain the frequency band characteristics of each rhythm frequency band; The frequency band features of different rhythm frequency bands are spliced together to obtain the original spectral domain feature vector; Batch normalization is performed on the original spectral domain eigenvectors to obtain batch normalized eigenvectors; The batch-normalized feature vectors are input into a fully connected encoding network to extract EEG rhythm features.
3. The virtual reality navigation rehabilitation training system according to claim 1, characterized in that, The electromyographic network includes a texture pathway, an energy pathway, a second splicing unit, and a dimensionality reduction unit; The texture pathway includes an instance normalization layer, a first convolutional layer, a first ReLU layer, a second convolutional layer, a second ReLU layer, an SE attention layer, and a global average pooling layer connected in sequence, used to extract muscle texture vectors based on preprocessed electromyographic signals; Energy pathways are used to calculate the channel time-averaged energy, inter-channel standard deviation, and energy ratio of the left and right legs based on preprocessed electromyographic signals. The left-right leg energy ratio is the ratio of the average energy of the left leg to the average energy of the right leg; the spatial features are obtained by splicing the channel time-averaged energy, the standard deviation between channels, and the left-right leg energy ratio. In the time dimension, the preprocessed electromyography signal is divided into several sub-windows; the absolute average of the amplitude of each sub-window is calculated to obtain the energy level of each sub-window; the variance is calculated based on the energy levels of all sub-windows to obtain the dynamic fluctuation characteristics of force exertion. Based on the preprocessed electromyography (EMG) signals, the temporal fluctuation characteristics of left-right differences, channel ranges, temporal slopes, and peak ratios are calculated. The dynamic fluctuation characteristics of force exertion, the temporal fluctuation characteristics of left-right differences, the temporal fluctuation characteristics of channel ranges, temporal slopes, and peak ratios constitute the temporal features. Based on the preprocessed EMG signals, the root mean square logarithm of each channel is calculated to obtain the global instantaneous activation energy features. The spatial features, temporal features, and global instantaneous activation energy features are concatenated to obtain the spatiotemporal energy features. The spatiotemporal energy features are then processed through a linear mapping layer to obtain the muscle energy vector. The second splicing unit is used to splice the muscle texture vector and the muscle energy vector to obtain the muscle splicing feature vector. The dimensionality reduction unit is used to reduce the dimensionality of the muscle splicing feature vector through a linear layer to obtain the muscle force characteristics.
4. The virtual reality navigation rehabilitation training system according to claim 1, characterized in that, The feature fusion layer includes a bidirectional gated modulation module, a third splicing unit, and an enhancement sub-network; The bidirectional gating modulation module is used to: generate gating signals for EEG rhythm features and gating signals for muscle exertion features using the gating module; perform cross-modal weighted modulation on the EEG rhythm features based on the gating signals for muscle exertion features to obtain modulated EEG features; and perform cross-modal weighted modulation on the muscle exertion features based on the gating signals for EEG rhythm features to obtain modulated electromyographic features. The third splicing unit is used to splice the modulated EEG features and the modulated EMG features to obtain joint features; The enhancement subnetwork is used to enhance the joint features sequentially through a linear layer, a ReLU layer, and a Dropout layer to obtain enhanced features. The enhanced features and joint features are fused based on the enhancement coefficient to obtain the fused features.
5. The virtual reality navigation rehabilitation training system according to claim 1, characterized in that, The rhythm frequency band includes Main frequency bands and The main frequency band of the wave, The main frequency band of the wave is 8Hz to 14Hz. The main frequency band of the wave is 15Hz to 25Hz.
6. The virtual reality navigation rehabilitation training system according to claim 5, characterized in that, The system further includes a model training module; the model training module is used for: Obtain the dataset; the dataset includes sample EEG and sample EMG signals from several subjects, as well as the corresponding movement category label for each subject; The sample EEG and EMG signals are input into a multimodal feature coupling network to obtain the sample motion intention prediction results; The loss function value is calculated based on the sample motion intention prediction results and motion category labels; The model parameters of the multimodal feature coupling network are updated based on the loss function value to obtain the trained multimodal feature coupling network.
7. The virtual reality navigation rehabilitation training system according to claim 6, characterized in that, The loss function value is calculated from the total loss function, which is expressed as follows: ; ; ; in, The value of the loss function; This is the weighted focus loss function; This represents the total number of samples in the current batch. Action category The corresponding weighting factor; The total number of action categories; Action category The corresponding smoothed target probability distribution; For the first Each sample in the action category The predicted probability is that a set of sample EEG signals and sample EMG signals corresponds to one sample. For center loss function; For the first The embedding vector obtained by linear mapping of each sample before the classifier head; Action category The corresponding center vector; The tradeoff coefficient for the center loss.
8. The virtual reality navigation rehabilitation training system according to claim 7, characterized in that, In updating the model parameters of a multimodal feature-coupled network, the model training module is used for: The model parameters of the EEG network, EMG network, feature fusion layer, and classifier are updated collaboratively. The center vector corresponding to each action category is updated separately.
9. The virtual reality navigation rehabilitation training system according to claim 1, characterized in that, The virtual reality navigation and rehabilitation module constructs a virtual hospital corridor scene using Unity, displays directional prompts via dynamic arrows at the top of the screen, and overlays the waveforms of the target's electroencephalogram (EEG) and electromyogram (EMG) signals, as well as the prediction results of the target's movement intentions, in real time using the LSL protocol.
10. A rehabilitation training method based on the virtual reality navigation rehabilitation training system according to any one of claims 1-9, characterized in that, The method includes: After the subjects wore the virtual reality device, the target EEG and target EMG signals were collected under the visual cues of the virtual reality device. The target EEG signal and target EMG signal are preprocessed separately to obtain preprocessed EEG signal and preprocessed EMG signal; The preprocessed EEG and EMG signals are input into a trained multimodal feature coupling network to obtain the target movement intention prediction results of the subject. Virtual characters are driven to perform motion tasks in virtual motion task scenarios based on the predicted motion intentions of the subjects.