Voice-based motor imagery training system for visually impaired patient

Through a motor imagination training system based on multimodal feature extraction and environmental adaptive adjustment based on speech cues, the problem of insufficient applicability of the motor imagination system in visually impaired patients is solved, and the accuracy and training effect of motor imagination are improved.

CN120408340AActive Publication Date: 2025-08-01JILI INNOVATION (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510557075.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing motor imagination brain-computer interface system is not suitable for visually impaired patients. How to design effective sound cues strategies and improve the accuracy of motor imagination is still a difficult point.

Method used

A motor imagination training system based on speech cues is designed, using multimodal feature extraction and environmental adaptive adjustment, combining speech stimulation module, EEG signal acquisition module, data processing and feature extraction module and motion imagination recognition module, the motor imagination classification model completed is used for identification, and the speech stimulation output strategy is adjusted through environmental noise data.

Benefits of technology

It improves the accuracy and training effect of visually impaired patients, ensures that speech stimulation is clearly heard under different ambient noise conditions, and improves the comfort and experience of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408340A_ABST
    Figure CN120408340A_ABST
Patent Text Reader

Abstract

The invention discloses a voice-based motor imagery training system for visually impaired patients, and the system comprises a voice stimulation module which is configured to output voice stimulation after a motor imagery training task starts; the environment adaptive adjustment module collects environment noise data and adjusts an output strategy of voice stimulation according to the noise data. The electroencephalogram signal acquisition module acquires electroencephalogram signals; the data processing and feature extraction module preprocesses the electroencephalogram signals and extracts frequency domain features and time domain features; and the motor imagery recognition module performs motor imagery classification recognition by utilizing a trained motor imagery classification model based on the preprocessed electroencephalogram signals, the frequency domain features and the time domain features to obtain a recognition result. The environment noise data is collected, the voice stimulation output strategy is adjusted according to the noise data, noise interference is avoided, and the training effect is improved; frequency domain features and time domain features are extracted, and motor imagery recognition is performed in combination with the electroencephalogram signals, so that the accuracy of motor imagery is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brain-computer interfaces (BCIs) and multimodal interaction, and particularly to a motor imagery training system for visually impaired patients based on voice. Background Art

[0002] Motor imagery brain-computer interface (MI-BCI) is an important brain-computer interface technology. It controls external devices or computer systems by analyzing electroencephalogram signals generated based on motor imagery when the user does not actually move. Motor imagery refers to the way an individual simulates body movements through internal representation without actual movement. Compared with physical movements, the neural activation patterns of motor imagery are highly similar, so it has become a core task in BCI systems.

[0003] MI-BCI has broad application prospects in the fields of rehabilitation medicine, intelligent assistive devices, and gaming and entertainment. In the field of rehabilitation, MI-BCI systems can be used for the motor function recovery training of stroke or spinal cord injury patients. By repeatedly performing motor imagery, patients can promote neuroplasticity and thus improve the damaged motor function. In addition, MI-BCI can also be used to control intelligent assistive devices such as prosthetics and wheelchairs, enabling disabled people to live more independently. In games and virtual reality, MI-BCI can also provide a more immersive interaction experience and enhance the user's entertainment experience.

[0004] Existing motor imagery brain-computer interface (MI-BCI) systems mostly rely on a single modality (such as vision or touch), and are not suitable enough for visually impaired patients. Research shows that appropriate sound cues can enhance the user's motor imagery ability and thus improve the recognition accuracy of the system. Although traditional voice cue systems can alleviate the problem of visual dependence, however, how to design effective sound cue strategies and how to improve the accuracy of motor imagery are still one of the current research hotspots and difficulties. Summary of the Invention

[0005] The technical objective of the present invention is to design a motor imagery training system based on voice cues, adopt effective sound cue strategies, and extract multimodal features for motor imagery training, so as to improve the accuracy of motor imagery and gradually enhance the actual motor performance.

[0006] To achieve the above technical objective, the present application adopts the following technical solutions.

[0007] The embodiment of the present application provides a motor imagery training system for visually impaired patients based on voice, including:

[0008] A voice stimulation module, configured to output voice stimulation after the start of the motor imagery training task;

[0009] An environment adaptive adjustment module, configured to collect ambient noise data during the output of voice stimuli by the voice stimulus module, and adjust the output strategy of the voice stimuli according to the noise data;

[0010] An electroencephalogram (EEG) signal acquisition module, configured to acquire EEG signals of the patient performing a motor imagery training task under the voice stimuli, and transmit the EEG signals to a data processing and feature extraction module;

[0011] A data processing and feature extraction module, configured to preprocess the EEG signals to obtain preprocessed EEG signals, and extract frequency domain features and time domain features;

[0012] The motor imagery recognition module is configured to perform motor imagery classification and recognition based on the preprocessed EEG signals, frequency domain features, and time domain features, using a trained motor imagery classification model to obtain a recognition result.

[0013] Further, the system further includes: an emotional state monitoring module, configured to collect the patient's galvanic skin response data and heart rate data before the start of the motor imagery training task;

[0014] Extract features based on the galvanic skin response data and heart rate data, and perform recognition of the patient's current emotional state based on the extracted features to determine whether the patient's current emotional state is suitable for starting the motor imagery training.

[0015] Still further, determining whether the patient's current state is suitable for starting the motor imagery training based on the features includes:

[0016] Still further, input the extracted features into a support vector machine (SVM), and use the following formula to perform recognition of the patient's current emotional state:

[0017]

[0018] where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, ξ i is the slack variable, α i and β i are the Lagrange multipliers, x i is the feature vector of the i-th training sample, y i is the class label of the i-th training sample, and N is the number of samples.

[0019] Further, the environment adaptive adjustment module specifically performs:

[0020] Based on the ambient noise data, calculate the noise energy of the time window to determine the current noise level;

[0021] Determine the current output volume level according to the current noise level and the difference between the preset output volume level of the stimulation audio and the background noise level, and determine the gain adjustment factor according to the output volume level and the default volume level;

[0022] Adjust the volume of the voice stimulation according to the gain adjustment factor.

[0023] Furthermore, the formula for determining the gain adjustment factor is as follows:

[0024]

[0025] Where L output = L noise + ΔL; L output is the output volume level, L noise is the noise level, ΔL is the difference between the preset output volume level of the stimulation audio and the background noise level, and L default is the default volume level.

[0026] Furthermore, the system further includes a central control module, which is configured to map the recognition result to a specific control instruction and send the control instruction to the device to be controlled, so that the device to be controlled performs corresponding operations.

[0027] Furthermore, the voice stimulation module is configured to include four voice output devices, and the four voice output devices are arranged around the patient in a uniform angular layout or a scattered non-uniform distance layout;

[0028] In the uniform angular layout, the four voice output devices are centered on the patient and are respectively arranged in the front, back, left, and right directions, and the angular intervals between the devices are equal;

[0029] In the scattered non-uniform distance layout, the four voice output devices are scattered around the patient, and the distances from the patient are not equal.

[0030] Furthermore, the data processing and feature extraction module performs:

[0031] Divide the EEG signal using a preset time window, and calculate the P300 peak amplitude and the corresponding time points within each time window;

[0032] Smooth the obtained P300 peak amplitude sequence to obtain a smoothed signal;

[0033] And calculate the power spectrum of the preselected key frequency bands;

[0034] Input the preprocessed EEG signals, the calculated P300 peak amplitudes and corresponding time points of each time window, the smoothed signals, and the power spectrum into the motor imagery recognition module.

[0035] Further, the motor imagery classification model includes a time-domain convolutional layer, a multi-head attention layer, a spatial convolutional layer, a frequency-domain fusion branch, and a decision layer connected in sequence;

[0036] The time-domain convolutional layer inputs the preprocessed EEG signals, performs one-dimensional convolution, and extracts local time-domain features;

[0037] The multi-head attention layer maps the local time-domain features output by the time-domain convolutional layer into queries, keys, and values, calculates the correlations of each subspace, and splices them after parallel calculation by multiple attention heads to obtain a global feature representation that fuses global information;

[0038] The spatial convolutional layer is used to combine the global feature representation with the cross-channel information extracted by the spatial convolutional layer to obtain depth features in the spatio-temporal dimension;

[0039] The frequency-domain fusion branch fuses the depth features with the calculated P300 peak amplitudes and corresponding time points of each time window, the smoothed signals, and the power spectrum to form multi-modal fusion features;

[0040] The decision layer is used to input the multi-modal fusion features into a fully-connected classification layer, output a probability distribution through the Softmax function, and obtain a classification result through a threshold determination mechanism.

[0041] Further, the motor imagery training includes initial training and enhanced training; in the initial training stage, the voice stimulation module outputs voice stimulations in a fixed order to guide the patient to establish an association between the voice and the movement direction; in the enhanced training stage, the voice stimulation module outputs voice stimulations randomly to verify the training effect.

[0042] Compared with the prior art, the voice-based motor imagery training system for visually impaired patients provided by the embodiments of the present application has the following beneficial technical effects: realizing multi-feature motor imagery recognition, the motor imagery recognition module extracts frequency-domain features and time-domain features from EEG signals, and combines the EEG signals themselves, and uses the trained motor imagery classification model to perform motor imagery classification and recognition; this multi-feature fusion method can capture the motor imagery information in EEG signals more comprehensively, improve the accuracy of recognition, help to understand the patient's motor imagery state more accurately, and then optimize the training plan. The environment adaptive adjustment module can collect environmental noise data during voice stimulation and adjust the output strategy of voice stimulation according to the noise data. This can ensure that the patient can clearly hear the voice stimulation under different environmental noise conditions, avoid noise interference, and improve the comfort and experience of training. Brief Description of the Drawings

[0043] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure of the present application in any way. Additionally, the shapes and proportional dimensions of the various components in the figures are merely schematic and are used to assist in understanding the present application, rather than specifically defining the shapes and proportional dimensions of the various components of the present application. Those skilled in the art can, under the teachings of the present application, select various possible shapes and proportional dimensions according to specific circumstances to implement the present application. In the drawings:

[0044] Figure 1 Schematic diagram of the structure of the speech-based motor imagery training system for visually impaired patients provided for the embodiment;

[0045] Figure 2 Schematic diagram of the workflow of the speech-based motor imagery training system for visually impaired patients provided for the embodiment;

[0046] Figure 3 Schematic diagram of the workflow of electroencephalogram signals in the embodiment;

[0047] Figure 4 Schematic diagram of a layout mode of the speech stimulation module in the embodiment;

[0048] Figure 5 Schematic diagram of the signal processing flow in the embodiment;

[0049] Figure 6 Schematic diagram of the one-dimensional time-domain convolution of the original data adding a multi-head attention spatial depthwise separable convolution decision layer in the embodiment;

[0050] Figure 7 Schematic diagram of the convolutional neural network EEGNet combined with the multi-head attention classification model in the embodiment. Detailed Description of the Embodiments

[0051] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0052] When visually impaired patients perform motor imagery training, they have difficulty perceiving the direction of movement and achieving poor training results due to the lack of visual information. In MI-BCI systems, auditory cues are widely used to guide users in performing specific motor imagery tasks. This multi-sensory cueing method can help users better focus their attention and improve the concentration and effectiveness of training. However, how to design effective auditory cueing strategies and how to improve the accuracy of motor imagery remain one of the current research hotspots and difficulties.

[0053] This application designs a motor imagery training system based on voice prompts (hereinafter referred to as the "system") to help visually impaired patients establish an association between sounds and control commands through sound signals, thereby improving the accuracy of motor imagery and gradually enhancing actual motor performance.

[0054] The following further describes this application in conjunction with the accompanying drawings of the specification and specific embodiments.

[0055] Please refer to Figure 1 , the motor imagery training system for visually impaired patients based on voice includes a voice stimulation module, an environment adaptive adjustment module, an electroencephalogram (EEG) signal acquisition module, a data processing and feature extraction module, and a motor imagery recognition module.

[0056] The voice stimulation module is configured to output voice stimulation after the start of the motor imagery training task.

[0057] The environment adaptive adjustment module is configured to collect environmental noise data during the output of voice stimulation by the voice stimulation module and adjust the output strategy of voice stimulation according to the noise data;

[0058] The EEG signal acquisition module is configured to collect the EEG signals of patients performing the motor imagery training task under voice stimulation and transmit the EEG signals to the data processing and feature extraction module.

[0059] The data processing and feature extraction module is configured to preprocess the EEG signals to obtain the preprocessed EEG signals and extract frequency domain features and time domain features.

[0060] The motor imagery recognition module is configured to perform motor imagery classification recognition based on the preprocessed EEG signals, frequency domain features, and time domain features using the trained motor imagery classification model to obtain the recognition result.

[0061] In the embodiment, the voice stimulation module includes four voice output devices (such as high-fidelity audio devices), which can be arranged around the patient in either a uniform angular layout or a scattered non-uniform distance layout; in the uniform angular layout, such as Figure 4As shown, the four voice output devices are centered around the patient and are arranged in the front, back, left, and right directions respectively, with equal angular intervals between the voice output devices. Looking from the horizontal direction, the four audio devices divide the 360° space into four equal parts, each part having an angle of 90°. This layout enables the patient to receive sound with a relatively balanced probability from all directions during training, which helps the patient establish a clear sense of sound direction and better associate the sound source direction with the corresponding motor imagery. For example, in the initial training stage, the system sequentially plays the vowel sounds emitted by each device, and the patient performs motor imagery based on the sound source direction. The uniform angular layout allows the patient to more accurately perceive the sound direction and improve the training effect.

[0062] In the scattered non-uniform distance layout (not shown), the four voice output devices are scattered around the patient and have unequal distances from the patient. This layout breaks the conventional symmetric distribution pattern, resulting in differences in the distance and time for the sound to reach the patient. Since the sound travels different distances, the intensity and time when it reaches the patient's ears will also be different, thus forming a more spatially layered and variable sound environment. In the enhanced training stage, the system designs complex task scenarios, and the sounds corresponding to different direction motor imagery are randomly generated. The sound variations in the scattered non-uniform distance layout can better simulate the complex sound environment in real life and further test and improve the patient's reaction ability and adaptability to the association between sound and movement instructions.

[0063] In specific embodiments, the layout mode with the optimal effect can be selected through experiments.

[0064] In an embodiment, the voice stimulation module outputs voice stimulation through multiple voice output devices arranged around the patient according to a preset layout method, which can provide rich and spatially sense auditory information for visually impaired patients, help them better concentrate, and more effectively trigger motor imagery, thereby improving the training effect.

[0065] As an example, each voice output device plays specific vowel sounds of "a", "e", "o", "u", and these sounds correspond to specific movement control instructions for the subjects. Using only fixed vowels lacks diversity and is likely to cause user fatigue and decreased adaptability. In some embodiments, on the basis of the original vowels (a, e, o, u), new digital voice stimulations ("1", "2", "3", "4") are added, which are respectively mapped to different movement instructions corresponding to the original four different instructions to enhance the stimulation diversity and user perception.

[0066] The system uses a voice output device (audio device) and voice prompts, eliminating the need to rely on complex hardware and visual information, making the training process more straightforward and intuitive. It can adjust the frequency, volume, and content of the voice prompts according to the needs of different patients, providing personalized training programs to ensure the effectiveness of the training.

[0067] In some embodiments, each audio device has a built-in sound signal generation module that can play preset vowel sounds and the sounds of different digits. The frequency, volume, and clarity of the sound can be adjusted according to the specific needs of the patient to ensure that the voice prompts can be clearly and effectively conveyed to the patient. The audio device can use an embedded sensor to collect ambient noise data in real time. When the detected background noise exceeds a preset threshold (e.g., 50 dB), the system automatically activates the adjustment of the stimulation volume based on the LMS algorithm and simultaneously corrects the subsequent playback strategy through a feedback mechanism.

[0068] Combined Figure 2 、 Figure 3 and Figure 5 As shown in

[0069] The training method of the system provided in this application may include: 1. Initial training: The patient sits at the center of the training room, and four audio devices are evenly placed in the front, back, left, and right directions or scattered non-uniformly according to two different speaker distribution positions. The system sequentially plays the vowel sounds emitted by each device, and the patient performs corresponding motor imagery according to the direction of the sound source. When hearing "a", the patient should imagine the corresponding control instruction one, and so on. The control instructions can be used for various different designs, such as controlling the wheelchair to move forward, backward, left, and right or controlling four different instructions corresponding to the robotic arm. The initial training stage is the sequential arrangement of sounds from four directions. Through repeated practice, it helps the patient establish the association between the sound and the movement direction.

[0070] In some embodiments, the system maps the vowel sounds played by each audio device to specific movement instructions. When hearing "a", the system maps it to movement instruction one - the wheelchair moves forward; when hearing "o", it is mapped to movement instruction two - the wheelchair turns left; when hearing "e", it is mapped to movement instruction 3 - the wheelchair moves backward; when hearing "u", it is mapped to movement instruction four - the wheelchair turns right. The movement instructions correspondingly control the movement directions of the wheelchair in which the subject is sitting, and real - time feedback is given to the subject according to the movement of the wheelchair. These mapped instructions can also be used to operate smart home devices, enabling patients to directly control the home environment through motor imagery. These control instructions can be used for the operation of smart home devices, such as turning lights on and off, adjusting the temperature, etc. This not only improves the practicality of training but also enhances the patient's independence and convenience in life.

[0071] The system provided by this application can provide rich and spatially - sensed auditory information for visually - impaired patients, helping them better concentrate and more effectively trigger motor imagery, thus improving the training effect. The multi - feature fusion method can more comprehensively capture the motor imagery information in EEG signals, improve the accuracy of recognition, help to more precisely understand the patient's motor imagery state, and then optimize the training plan. Adjusting the output strategy of speech stimuli according to noise data can ensure that patients can clearly hear the speech stimuli under different environmental noise conditions, avoid noise interference, and enhance the comfort and experience of training.

[0072] The application scenarios of the system provided by the embodiments of this application may include:

[0073] Rehabilitation training: This system is widely used in the rehabilitation training of visually - impaired patients. By combining sound cues with motor imagery, it helps patients improve their motor coordination ability and enhance their self - care ability in daily life.

[0074] Smart home control: The system can be integrated with the smart home system, enabling patients to control various devices in the home through motor imagery, such as turning lights on and off, adjusting the temperature, etc., improving their quality of life and independence.

[0075] Virtual reality and navigation: The system can also be combined with virtual reality technology to provide an immersive training environment for visually - impaired patients. In addition, the system can be applied to a navigation assistance system to help patients perform safe navigation and obstacle avoidance training in complex environments.

[0076] In some embodiments, the voice output device integrates an embedded microphone array. Through an adaptive filtering algorithm, it real - time collects environmental noise and dynamically adjusts the gain parameter of the speech stimulus. Combining short - time energy analysis, it ensures that the signal - to - noise ratio of the speech signal is stable at 10 - 15 dB.

[0077] In some embodiments, the voice-based motor imagery training system for visually impaired patients further includes a central control unit, which is responsible for managing the operating states of each audio device, such as the duration, interval, and volume of sound playback. This module can also record the response data of the patient and provide feedback through analysis to help optimize the training plan.

[0078] In the embodiment, after the patient wears the electroencephalogram signal acquisition module (wears the electroencephalogram signal acquisition device) to acquire the electroencephalogram signal, the data processing and feature extraction module performs preprocessing through a filtering algorithm. Commonly used filters include band-pass filters, which are used to remove electromyogram noise and power interference to ensure the purity of the signal.

[0079] In the embodiment, in the initial training stage, the system induces the patient's motor imagery by playing vowel prompt sounds in different directions through the voice stimulation module. The mixed stimulation signal can be pre-generated by the MP3 voice processing module, and the signal library contains pre-recorded vowels and digital voices. In the initial training stage, the voice stimulation module plays in a fixed order in turn, and in the enhanced training stage, the voice stimulation module uses a random sequence generator to ensure that the order of the stimulation signals has no fixed pattern to avoid the user's memory adaptation.

[0080] The environment adaptive adjustment module collects environmental noise data during the voice stimulation output by the voice stimulation module, and adjusts the output strategy of the voice stimulation according to the noise data to automatically adjust the playback volume according to the environmental noise data, and can ensure a stable output above 70 dB. The environmental noise can be collected by a microphone and short-time energy analysis is performed in the time domain to calculate the current noise level L noise . Given the audio signal x(i), the sampling frequency is s, and the short-time window size is N0, the noise energy of a certain time window is calculated as follows:

[0081]

[0082] where E(k) represents the noise energy within this time window at time k, k is the start time identifier of the time window, N0 is the length of the time window, and x(i) is the audio signal value collected at time i.

[0083] After averaging multiple windows, calculate the noise level of the environmental noise, that is, the equivalent sound pressure level (SPL, Sound Pressure Level):

[0084]

[0085] where L noise is the noise level, E(j) is the short-time energy of the j-th time window; M is the number of windows for calculating the moving average, and C0 is a preset calibration constant for the voice output device (usually pre-calibrated under laboratory conditions).

[0086] To ensure the audibility of the stimulation audio, the environment adaptive adjustment module can set the output volume to be 10–15 dB higher than the background noise. The environment adaptive adjustment module can limit this calculated value between the maximum output volume and the minimum output volume. Assuming the default volume level of the current stimulation audio is Ldefault, the formula for the required gain adjustment factor is as follows:

[0087] where L output = L noise + ΔL; L output is the output volume level, L noise is the noise level, ΔL is the difference between the preset stimulation audio output volume level and the background noise level, and L default is the default volume level.

[0088] Frame-by-frame gradual smoothing transition is performed through the gain adjustment factor G parameter to avoid discomfort caused by sudden volume changes.

[0089] In the system, the audio stimulation signal is stored in numerical form (floating-point array), and the embodiment can achieve automatic volume adjustment by multiplying the audio signal by the gain adjustment factor G.

[0090] As an example, the specific steps are as follows:

[0091] 1. Load the audio signal: Read the audio stimulation file into a one-dimensional array, where each element in the array represents the amplitude value of a sampling point, ranging from [-1, 1].

[0092] 2. Apply gain adjustment: Multiply each sampling point x(i) by the gain adjustment factor G to obtain the adjusted signal:

[0093] xnew(i) = x(i) × G; This is equivalent to magnifying or reducing the amplitude of the entire audio waveform, achieving an increase or decrease in volume.

[0094] 3. Prevent signal overflow:

[0095] If there are values in the adjusted audio signal that exceed the range of [-1, 1][-1, 1][-1, 1], it may cause distortion or playback errors, and the entire audio signal needs to be normalized to scale it back to the safe range.

[0096] 4. Play the audio: Pass the adjusted audio signal to the playback device for playback. At this time, the audio has been automatically volume-adjusted according to the ambient noise.

[0097] Example: A certain sampling value in the original audio data is 0.5; the calculated gain adjustment factor G = 1.5, indicating that the volume should be increased; after adjustment, this sampling value becomes: xnew = 0.5 × 1.5 = 0.75; the same processing will be applied to the entire audio segment, thereby increasing the volume overall.

[0098] To avoid discomfort to patients caused by sudden volume changes, the system's environmental adaptive adjustment module will perform a gradual and smooth transition in frames through the gain adjustment factor G parameter. The audio signal is divided into multiple frames at a certain time interval. In each frame, the volume is gradually adjusted according to the G value, so that the volume change remains smooth between adjacent frames instead of changing suddenly. When the background noise is detected to increase at a certain moment, the calculated G value increases accordingly. The system will not instantly adjust the volume to the new level, but in each subsequent frame, gradually increase the volume, making it almost imperceptible to the patient the process of volume change and enhancing the training experience.

[0099] Through real-time noise detection and gain adjustment, this system can provide a stable stimulation volume (such as above the target 70 dB) under different environmental noise levels. At the same time, exponential smoothing and adaptive filtering are adopted to avoid discomfort caused by sudden volume changes and enhance the audibility of the signal. Through experimental verification, the system can maintain a signal-to-noise ratio of 10 - 15 dB within the background noise range of 40 - 80 dB, ensuring high recognizability of the stimulation signal.

[0100] In the embodiment, the electroencephalogram signal acquisition module uses a 16-channel high-precision electroencephalogram acquisition device to ensure good electrode contact during the acquisition process and reduce motion artifacts. The data is transmitted in real time to the data processing and feature extraction module at a fixed sampling rate of 250 Hz for preprocessing, such as applying band-pass filtering (0.5 - 40 Hz) to remove power frequency interference and low-frequency drift.

[0101] P300 is an important component in event-related potential (ERP) and usually appears within 250 - 400 ms after stimulation. To ensure the stability of feature extraction, in the embodiment, the data processing and feature extraction module uses the sliding window method to extract the key time-domain features of P300.

[0102] The selected window width is 400 ms, which can cover the typical time range of P300; the step size is 100 ms to balance the calculation efficiency and time resolution; since P300 is usually more significant in the parietal and central regions, the channels selected are Fz and Pz. Through these parameters, the peak amplitude of P300 and the corresponding time points within each sliding window are calculated for feature extraction. The specific method is as follows:

[0103]

[0104] x(t) is the EEG signal, and the peak value of P300 within the window is calculated. Among them, AP300 is the P300 peak amplitude, T P300 is the time point corresponding to the peak, and t0 is the starting time of the sliding window.

[0105] Since EEG is affected by environmental noise and physiological artifacts, the directly extracted P300 signal may have large fluctuations. To improve the feature stability, exponential moving average filtering (EMA) is used for time-domain smoothing. Let the original P300 peak signal sequence be An, and the smoothed signal be Sn. The EMA calculation is as follows:

[0106] S n = αA n +(1 - α)S n-1 ;

[0107] where α is the smoothing coefficient set to 0.3, and Sn-1 is the smoothing result of the previous time point.

[0108] To further enhance the classification performance, in addition to the P300 peak, the power spectral density (PSD) features in its corresponding time period are also extracted to obtain frequency-domain information. The PSD is calculated using the Welch method to calculate the power spectrum of the key frequency band (8 - 30 Hz):

[0109]

[0110] X n (f) is the result of the short-time Fourier transform (STFT) of the signal, and N is the number of windows.

[0111] Finally, the data processing and feature extraction module of the system extracts the time-domain P300 and frequency-domain PSD information to ensure complementary signal features and improve the classification accuracy; a motor imagery classification model is constructed based on the preprocessed EEG signals, P300 time-domain features, PSD frequency-domain features, and the smoothed signals to improve the recognition accuracy and system stability.

[0112] In the embodiment, the motor imagery classification model adopted by the motor imagery recognition module performs the following steps: dividing the EEG signal using a preset time window, calculating the P300 peak amplitude and the corresponding time point within each time window; performing smoothing processing on the obtained P300 peak amplitude sequence to obtain the smoothed signal; and calculating the power spectrum of the preselected key frequency band; inputting the preprocessed EEG signal, the calculated P300 peak amplitude and the corresponding time point of each time window, the smoothed signal, and the power spectrum into the trained motor imagery classification model for motor imagery classification recognition to obtain the recognition result.

[0113] Specifically, please refer to Figure 6 and Figure 7. To improve the signal classification performance of the motor imagery brain-computer interface (MI-BCI), this system is optimized based on EEGNet, introducing a multi-head attention mechanism and frequency-domain feature fusion, aiming to enhance the ability of the motor imagery classification model to capture spatio-temporal information and its adaptability to background noise. In addition, the motor imagery classification model adopts an adaptive training mechanism, allowing parameters to be continuously optimized during online operation to improve long-term stability.

[0114] 1. Time-domain convolutional layer

[0115] First, perform one-dimensional convolution (1D-CNN) on the preprocessed EEG signals to extract local temporal features. Since EEG signals have strong temporal dependence, this layer can effectively learn the waveform changes within a short time, thereby capturing preliminary feature representations. Let the input EEG signal be x ∈ R C×T , where C is the number of channels and T is the number of time steps. Then the time-domain convolution is calculated as follows:

[0116] H1 = σ(W1 * X + b1);

[0117] where W1 is the convolution kernel weight matrix, b1 is the bias term, * represents the convolution operation, and σ is the non-linear activation function (ReLU).

[0118] 2. Multi-head attention layer

[0119] Traditional EEGNet only uses convolutional neural networks (CNNs) to extract features and lacks attention to global features. To make up for this defect, this model introduces a multi-head attention mechanism (Multi-Head Self-Attention, MHSA) to model long-range dependencies. This layer maps the output of the convolutional layer into queries (Query), keys (Key), and values (Value), calculates the correlations within each subspace, thereby enhancing the global perception ability of the model. Let H1 be the input feature, then the attention calculation is as follows:

[0120] Q = W q H1K = W k H1V = W v H1;

[0121]

[0122] where W q , W k , W v are the projection matrices for queries, keys, and values, d k is the scaling factor, and Softmax is used to normalize the attention distribution. After multiple attention heads are calculated in parallel, they are concatenated:

[0123] H3 = Concat(head1, head2, …, head h )W o ;

[0124] W o is the final linear transformation matrix, and h is the number of attention heads.

[0125] 3. Spatial Convolution Layer

[0126] In EEG signal processing, there is spatial correlation between different channels. To effectively extract cross-channel information, the model uses depthwise separable convolution. This layer consists of two parts: depthwise convolution and pointwise convolution, which reduces the computational complexity while retaining sufficient feature information. The formula for spatial convolution is as follows:

[0127] H3 = σ(W 3,d *H2 + b3) + σ(W 3,p (H2 + b3)

[0128] where W 3,d is the depthwise convolution kernel (acting only within a single channel), W 3,p is the pointwise convolution kernel (used for inter-channel fusion), and b3 is the bias term.

[0129] 4. Frequency Domain Fusion Branch

[0130] In addition to time-domain features, the frequency features of EEG signals are also crucial. This model extracts the power spectral density (PSD) and wavelet transform coefficients (Wavelet Coefficients) in the front-end signal preprocessing, and reduces the dimension through a fully connected network (FCN). Let the original frequency domain feature be F, after dimensionality reduction, it is concatenated with the feature H3 obtained from the spatial convolution layer to form a multi-modal fusion feature:

[0131] F′ = σ(W f F + b f );

[0132] H4 = Concat(H3, F′);

[0133] 5. Decision Layer

[0134] The fused features enter the final fully-connected classification layer and output a probability distribution through the Softmax function. To improve the stability of classification, a threshold determination mechanism (P300 > 0.4 and relative beta wave power > 1.25 times the baseline) is added based on the Softmax result for secondary screening to ensure the reliability of the classification result. The final classification decision is as follows:

[0135] P = Softmax(W d H4 + b d );

[0136]

[0137] In the embodiment, it also includes motor imagery classification model training and online real-time parameter adjustment. The improved motor imagery classification model (EEGNET) is pre-trained using the offline collected dataset to ensure that the model has a high classification accuracy under various noise levels; during the online operation process, some model parameters (such as attention weights and fully-connected layer parameters) are continuously updated through a feedback mechanism, enabling the model to have the ability of adaptive online adjustment.

[0138] In the embodiment, the motor imagery classification model uses an improved EEGNET classification model: by introducing a multi-head attention mechanism to capture global dependencies and adopting a frequency-domain fusion module to integrate multi-scale features, the classification performance is improved by more than 15% compared with the traditional model.

[0139] In some embodiments, as Figure 1 shown, the system further includes: an emotional state monitoring module, which is configured to collect the electrodermal response data and heart rate data of the patient before the start of the motor imagery training task; extract features based on the electrodermal response data and heart rate data, and identify the current emotional state of the patient based on the extracted features to determine whether the current emotional state of the patient is suitable for starting the motor imagery training.

[0140] As an example, determining whether the current state of the patient is suitable for starting the motor imagery training based on the features includes:

[0141] Input the extracted features into the support vector machine SVM, and use the following formula to identify the current emotional state of the patient:

[0142]

[0143] where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, ξ i is the slack variable, α i and β i are the Lagrange multipliers, x i is the feature vector of the i-th training sample, y iis the class label of the i-th training sample, and N is the number of samples.

[0144] When patients are in a suitable emotional and physical state, they can focus more on the training tasks, cooperate more actively, and the brain's response to voice stimuli will also be more sensitive, which is more conducive to triggering motor imagination and making the training achieve better results.

[0145] In some embodiments, such as Figure 1 shown, the system further includes a central control module. After classifying the motor imagination of EEG signals through the EEGNet with multi-head attention, the system will use the central control module to map these classification results to specific control instructions. For example, when the motor imagination recognition module of the system recognizes that the patient imagines pronouncing 'a', it will generate control instruction 1, and the control instructions correspond one-to-one with the device operations.

[0146] After the control instruction is generated, the central control module of the system will send the instruction to the device to be controlled through a wireless communication module (Bluetooth), and the device will perform the corresponding operation. At the same time, in some embodiments, the central control module monitors the execution situation of the device in real time through sensors and feeds back the response (movement direction, state) of the device to the user. For example, if the device controlled by the central control module is a wheelchair, the central control module will feed back the actual movement of the wheelchair to the user to confirm the execution effect of the instruction through visual, audio, or tactile feedback modules.

[0147] In some embodiments, the central control module makes adaptive adjustments based on the feedback information. If it detects that the device does not respond as expected (for example, there is a deviation between the movement imagined by the patient and the actual control), it will fine-tune the instruction through the or PID (Proportional-Integral-Derivative) control algorithm to ensure that the device executes according to the user's intention. For example, if the wheelchair moves too fast or deviates in direction, the central control module will gradually adjust the parameters until the device reaches the optimal control effect.

[0148] In some embodiments, the central control module can adapt to different types of intelligent devices by designing a general interface protocol. For example, through a unified API and Bluetooth or WiFi communication protocols, the central control module can not only control wheelchairs but also interconnect with other smart home devices (such as smart lights, smart curtains, etc.). This design allows users to conveniently operate more devices through motor imagination.

[0149] The central control module provides a personalized configuration function, and users can customize control scenarios according to different needs. Users or doctors can set different motor imaginations to correspond to different device control instructions. Through a graphical interface, users can pre-define control schemes, mapping different motor imagination signals to specific operations to ensure that the system better meets the actual needs of users.

[0150] To achieve a wider range of applications, the central control module supports cloud data management functions. All training records, EEG signals, device control instructions, and feedback data can be uploaded to the cloud for storage and analysis. Doctors or technicians can monitor the user's rehabilitation progress through a remote platform and optimize system parameters based on big data analysis. In addition, the system can form an Internet of Things (IoT) ecosystem with medical devices, smart devices, etc. through the cloud platform to further expand its functions.

[0151] The voice-based motor imagery training system for visually impaired patients of the present invention, as a motor imagery brain-computer interface system, can accurately identify the motor imagery signals of patients, intelligently adjust the training difficulty, and achieve the control of smart home devices, ultimately providing effective rehabilitation training and life convenience for visually impaired patients.

[0152] The above has introduced in detail the voice-based motor imagery training system for visually impaired patients provided in this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the concept of this application and should not be construed as a limitation on the protection scope of this application.

Claims

1. A voice-based motor imagery training system for visually impaired patients, characterized in that, Including: A voice stimulation module, configured to output voice stimulation after the start of a motor imagery training task; An environment adaptive adjustment module, configured to collect ambient noise data during the output of voice stimulation by the voice stimulation module, and adjust the output strategy of voice stimulation according to the noise data; An electroencephalogram (EEG) signal acquisition module, configured to acquire the EEG signals of the patient performing the motor imagery training task under the voice stimulation, and transmit the EEG signals to a data processing and feature extraction module; A data processing and feature extraction module, configured to preprocess the EEG signals to obtain preprocessed EEG signals, and extract frequency domain features and time domain features; The motor imagery recognition module, configured to perform motor imagery classification recognition based on the preprocessed EEG signals, frequency domain features and time domain features, using a trained motor imagery classification model to obtain a recognition result.

2. The voice-based motor imagery training system for visually impaired patients according to claim 1, wherein The system further includes: an emotional state monitoring module, configured to collect the skin conductance response data and heart rate data of the patient before the start of the motor imagery training task; Extract features based on the skin conductance response data and heart rate data, and perform recognition of the patient's current emotional state based on the extracted features to determine whether the patient's current emotional state is suitable for starting the motor imagery training.

3. The voice-based motor imagery training system for visually impaired patients according to claim 2, wherein, Judging whether the patient's current state is suitable for starting the motor imagery training based on the features includes: Inputting the extracted features into a support vector machine (SVM), and using the following formula to perform recognition of the patient's current emotional state: where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, and ξ i is the slack variable, α i and β i are the Lagrange multipliers, x i is the feature vector of the i-th training sample, y i is the class label of the i-th training sample, and N is the number of samples.

4. The voice-based motor imagery training system for visually impaired patients according to claim 1, wherein, The environment adaptive adjustment module specifically executes: Based on the ambient noise data, calculate the noise energy of a time window to determine the current noise level; According to the current noise level and the difference between the preset output volume level of the stimulation audio and the background noise level, determine the current output volume level, and determine a gain adjustment factor according to the output volume level and the default volume level; Adjust the volume of the output voice stimulation according to the gain adjustment factor.

5. The voice-based motor imagery training system for visually impaired patients according to claim 4, characterized in that, The formula for determining the gain adjustment factor is as follows: Where L output = L noise + ΔL; L output is the output volume level, L noise is the noise level, and ΔL is the difference between the preset stimulating audio output volume level and the background noise level, L default is the default volume level.

6. The voice-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The system further includes a central control module, configured to map the recognition result to a specific control instruction, and send the control instruction to the device to be controlled, so that the device to be controlled performs corresponding operations.

7. The voice-based motor imagery training system for visually impaired patients according to claim 1, wherein The voice stimulation module is configured to include four voice output devices, and the four voice output devices are arranged around the patient in a uniform angle layout manner or a scattered non-uniform distance layout manner; In the uniform angle layout manner, the four voice output devices are centered on the patient and are respectively arranged in the front, back, left and right directions, and the angle intervals between the devices are equal; In the scattered non-uniform distance layout manner, the four voice output devices are scattered around the patient, and the distances from the patient are not equal.

8. The speech-based motor imagery training system for visually impaired patients according to claim 1, wherein The data processing and feature extraction module executes: Divide the EEG signals by using a preset time window, and calculate the P300 peak amplitude and the corresponding time points within each time window; Perform smoothing processing on the obtained P300 peak amplitude sequence to obtain a smoothed signal; And calculate the power spectrum of a preselected key frequency band. The preprocessed EEG signals, the calculated P300 peak amplitudes and corresponding time points of each of the time windows, the smoothed signals, and the power spectrum are input into the motor imagery recognition module.

9. The voice-based motor imagery training system for visually impaired patients according to claim 1, wherein, The motor imagery classification model includes a time-domain convolutional layer, a multi-head attention layer, a spatial convolutional layer, a frequency-domain fusion branch, and a decision layer connected in sequence; The time-domain convolutional layer inputs the preprocessed EEG signals, performs one-dimensional convolution, and extracts local time-domain features; The multi-head attention layer maps the local time-domain features output by the time-domain convolutional layer into queries, keys, and values, calculates the correlations of each subspace, and splices them after parallel calculation by multiple attention heads to obtain a global feature representation that fuses global information; The spatial convolutional layer is used to combine the global feature representation with the cross-channel information extracted by the spatial convolutional layer to obtain deep features in the spatio-temporal dimension; The frequency-domain fusion branch fuses the deep features with the calculated P300 peak amplitudes and corresponding time points of each of the time windows, the smoothed signals, and the power spectrum to form multi-modal fusion features; The decision layer is used to input the multi-modal fusion features into a fully connected classification layer, output a probability distribution through the Softmax function, and obtain a classification result through a threshold decision mechanism.

10. The voice-based motor imagery training system for visually impaired patients according to claim 1, wherein The motor imagery training includes initial training and enhanced training; in the initial training stage, the speech stimulation module outputs speech stimulations in a fixed order to guide the patient to establish an association between sounds and movement directions; In the enhanced training stage, the speech stimulation module randomly outputs speech stimulations to verify the training effect.

Citation Information

Patent Citations

  • Volume adjustment method and device and equipment

    CN107124149A

  • Hand rehabilitation training method based on motor imagery

    CN111110982A

  • Epilepsy prediction method based on domain adversarial multi-level deep convolutional feature fusion network

    CN115886840A

  • Motor imagery electroencephalogram signal decoding method and system, medium and equipment

    CN117609852A

  • Intelligent internal medicine nursing monitoring system

    CN117854739A