Voice-based motor imagery training system for visually impaired patients

CN120408340BActive Publication Date: 2026-09-01JILI INNOVATION (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510557075.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2026-09-01
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

传统语音提示系统虽能缓解视觉依赖问题,然而,如何设计有效的声音提示策略,以及如何提高运动想象的准确性,仍然是当前研究的热点和难点之一

Benefits of technology

[0042]与现有技术相比,本申请实施例提供的基于语音的视障患者运动想象训练系统具有以下有益技术效果:实现了多特征运动想象识别,运动想象识别模块从脑电信号中提取频域特征以及时域特征,并结合脑电信号本身,利用训练完成的运动想象分类模型进行运动想象分类识别;这种多特征融合的方式能够更全面地捕捉脑电信号中的运动想象信息,提高识别的准确性,有助于更精确地了解患者的运动想象状态,进而优化训练方案。环境自适应调节模块能够在语音刺激期间采集环境噪音数据,并根据噪音数据调节语音刺激的输出策略。这样可以确保患者在不同的环境噪音条件下都能清晰地听到语音刺激,避免噪音干扰,提升训练的舒适度和体验感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408340B_ABST
    Figure CN120408340B_ABST
Patent Text Reader

Abstract

This application discloses a speech-based motor imagery training system for visually impaired patients, including a speech stimulation module configured to output speech stimuli after the start of a motor imagery training task; an environment adaptive adjustment module that collects environmental noise data and adjusts the speech stimulation output strategy based on the noise data; an electroencephalogram (EEG) signal acquisition module that collects EEG signals; a data processing and feature extraction module that preprocesses the EEG signals and extracts frequency domain and time domain features; and a motor imagery recognition module that, based on the preprocessed EEG signals, frequency domain features, and time domain features, uses a trained motor imagery classification model to classify and recognize motor imagery, obtaining the recognition result. This application collects environmental noise data and adjusts the speech stimulation output strategy based on the noise data to avoid noise interference and improve training effectiveness; it extracts frequency domain and time domain features and combines them with the EEG signals themselves for motor imagery recognition, thereby improving the accuracy of motor imagery recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of brain-computer interface (BCI) and multimodal interaction technology, specifically to a voice-based motor imagery training system for visually impaired patients. Background Technology

[0002] Motor imagery brain-computer interface (MI-BCI) is an important brain-computer interface technology that controls external devices or computer systems by analyzing the brainwave signals generated by motor imagery without actual movement. Motor imagery refers to an individual simulating bodily movements through internal representation without actually moving. The neural activation patterns of motor imagery are highly similar to those of physical movement, thus making it a core task in BCI systems.

[0003] MI-BCI has broad application prospects in rehabilitation medicine, intelligent assistive devices, and gaming. In rehabilitation, MI-BCI systems can be used for motor function recovery training in patients with stroke or spinal cord injury. Through repeated motor imagery, patients can promote neuroplasticity, thereby improving impaired motor function. Furthermore, MI-BCI can be used to control intelligent assistive devices such as prostheses and wheelchairs, enabling people with disabilities to achieve a more independent life. In games and virtual reality, MI-BCI can also provide a more immersive interactive experience, enhancing the user's entertainment experience.

[0004] Existing motor imagery brain-computer interface (MI-BCI) systems mostly rely on a single modality (such as vision or touch), making them insufficiently applicable to visually impaired patients. Research indicates that appropriate audio cues can enhance users' motor imagery abilities, thereby improving the system's recognition accuracy. While traditional voice prompt systems can alleviate visual dependence, designing effective audio cue strategies and improving the accuracy of motor imagery remain key research areas and challenges. Summary of the Invention

[0005] The technical objective of this invention is to design a motion imagery training system based on voice prompts, employing effective voice prompt strategies and extracting multimodal features for motion imagery training, thereby improving the accuracy of motion imagery and gradually enhancing actual motion performance.

[0006] To achieve the above technical objectives, this application adopts the following technical solution.

[0007] This application provides a voice-based motor imagery training system for visually impaired patients, including:

[0008] The speech stimulation module is configured to output speech stimulation after the start of the motor imagery training task;

[0009] An environment adaptive adjustment module is configured to collect environmental noise data during the output of speech stimulation by the speech stimulation module, and adjust the output strategy of speech stimulation according to the noise data.

[0010] The EEG signal acquisition module is configured to acquire the EEG signals of the patient performing a motor imagery training task under the speech stimulation, and transmit the EEG signals to the data processing and feature extraction module.

[0011] The data processing and feature extraction module is configured to preprocess the EEG signal to obtain the preprocessed EEG signal and extract frequency domain features and time domain features.

[0012] The motor imagery recognition module is configured to classify and recognize motor images based on preprocessed EEG signals, frequency domain features, and time domain features, using a trained motor imagery classification model to obtain recognition results.

[0013] Furthermore, the system also includes an emotional state monitoring module, which is configured to collect the patient's skin conductance response data and heart rate data before the start of the motor imagery training task;

[0014] Based on the skin conductance data and heart rate data, features are extracted, and the patient's current emotional state is identified based on the extracted features to determine whether the patient's current emotional state is suitable for starting motor imagery training.

[0015] Furthermore, determining whether the patient's current state is suitable for starting motor imagery training based on the aforementioned characteristics includes:

[0016] Furthermore, the extracted features are input into a support vector machine (SVM), and the patient's current emotional state is identified using the following formula:

[0017]

[0018] Where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, and ξ is the weight vector. i As a slack variable, α i and β i For Lagrange multipliers, x i Let y be the feature vector of the i-th training sample. i Let N be the class label of the i-th training sample, and N be the number of samples.

[0019] Furthermore, the environmental adaptive adjustment module specifically performs the following:

[0020] Based on the environmental noise data, noise energy is calculated for a time window to determine the current noise level;

[0021] The current output volume level is determined based on the current noise level and the difference between the preset stimulus audio output volume level and the background noise level. The gain adjustment factor is determined based on the output volume level and the default volume level.

[0022] The volume of the speech stimulus is adjusted according to the gain adjustment factor.

[0023] Furthermore, the formula for determining the gain adjustment factor is as follows:

[0024]

[0025] Where L output =L noise +ΔL;L output For output volume level, L noise ΔL represents the noise level, where ΔL is the difference between the preset stimulus audio output volume level and the background noise level. default This is the default volume level.

[0026] Furthermore, the system also includes a central control module, which is configured to map the identification results to specific control commands and send the control commands to the device to be controlled so that the device to be controlled can perform corresponding operations.

[0027] Furthermore, the speech stimulation module is configured to include four speech output devices, which are arranged around the patient in a uniform angle layout or a non-uniform scattering distance layout.

[0028] In the uniform angle layout, the four voice output devices are arranged in front of, behind, to the left and to the right of the patient, with equal angular intervals between each device.

[0029] In the aforementioned non-uniform scattering distance layout, the four voice output devices are distributed in a scattering pattern around the patient, and the distances between them and the patient are unequal.

[0030] Furthermore, the data processing and feature extraction module performs the following:

[0031] The EEG signal is divided into preset time windows, and the P300 peak amplitude and corresponding time point are calculated in each time window.

[0032] The obtained P300 peak amplitude sequence is smoothed to obtain the smoothed signal.

[0033] And calculate the power spectrum of the pre-selected key frequency bands;

[0034] The preprocessed EEG signal, the calculated P300 peak amplitude and corresponding time point of each time window, the smoothed signal, and the power spectrum are input to the motor imagery recognition module.

[0035] Furthermore, the motion imagery classification model includes a temporal convolutional layer, a multi-head attention layer, a spatial convolutional layer, a frequency domain fusion branch, and a decision layer connected in sequence;

[0036] The temporal convolutional layer takes the preprocessed EEG signal as input, performs one-dimensional convolution, and extracts local temporal features.

[0037] The multi-head attention layer maps the local temporal features output by the temporal convolutional layer into queries, keys, and values, calculates the correlation between each subspace, and concatenates the parallel calculations of multiple attention heads to obtain a global feature representation that integrates global information.

[0038] The spatial convolutional layer is used to combine the global feature representation with the cross-channel information extracted by the spatial convolutional layer to obtain deep features in the spatiotemporal dimension.

[0039] The frequency domain fusion branch fuses the depth features, the calculated P300 peak amplitude and corresponding time point of each time window, the smoothed signal, and the power spectrum to form a multimodal fusion feature.

[0040] The decision layer is used to input the multimodal fusion features into the fully connected classification layer, output the probability distribution through the Softmax function, and obtain the classification result through a threshold determination mechanism.

[0041] Furthermore, the motor imagery training includes initial training and reinforcement training; in the initial training stage, the speech stimulation module outputs speech stimuli in a fixed order to guide the patient to establish an association between sound and movement direction; in the reinforcement training stage, the speech stimulation module outputs speech stimuli randomly to verify the training effect.

[0042] Compared with existing technologies, the speech-based motor imagery training system for visually impaired patients provided in this application has the following beneficial technical effects: It achieves multi-feature motor imagery recognition. The motor imagery recognition module extracts frequency domain features and time domain features from EEG signals and combines them with the EEG signals themselves, using a trained motor imagery classification model for classification and recognition. This multi-feature fusion approach can more comprehensively capture motor imagery information in EEG signals, improve recognition accuracy, and help to more accurately understand the patient's motor imagery state, thereby optimizing the training program. The environmental adaptive adjustment module can collect environmental noise data during speech stimulation and adjust the speech stimulation output strategy according to the noise data. This ensures that patients can clearly hear the speech stimulation under different environmental noise conditions, avoids noise interference, and improves the comfort and experience of training. Attached Figure Description

[0043] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this application in any way. Furthermore, the shapes and scales of the components in the drawings are merely illustrative to aid in understanding this application and do not specifically limit the shapes and scales of the components. Those skilled in the art, guided by the teachings of this application, can select various possible shapes and scales to implement this application according to specific circumstances. In the drawings:

[0044] Figure 1 A schematic diagram of the structure of a voice-based motor imagery training system for visually impaired patients provided in an embodiment;

[0045] Figure 2 A schematic diagram of the workflow of the speech-based motor imagery training system for visually impaired patients provided in this embodiment;

[0046] Figure 3 This is a schematic diagram of the EEG signal workflow in the embodiment;

[0047] Figure 4 This is a schematic diagram of one layout of the speech stimulation module in the embodiment;

[0048] Figure 5 This is a schematic diagram of the signal processing flow in the embodiment;

[0049] Figure 6 This is a schematic diagram of adding a multi-head attention spatial depth separable convolutional decision layer to the original data in the embodiment;

[0050] Figure 7 This is a schematic diagram of the convolutional neural network EEGNet combined with a multi-head attention classification model in the embodiment. Detailed Implementation

[0051] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0052] Visually impaired patients often experience difficulties perceiving direction and achieve poor training results during motor imagery training due to a lack of visual information. In the MI-BCI system, audio cues are widely used to guide users through specific motor imagery tasks. This multi-sensory approach helps users better concentrate, improving training focus and effectiveness. However, designing effective audio cues and improving the accuracy of motor imagery remain key research areas and challenges.

[0053] This application designs a motor imagery training system based on voice prompts (hereinafter referred to as "the system") to help visually impaired patients establish a connection between sound and control commands through sound signals, thereby improving the accuracy of motor imagery and gradually enhancing actual motor performance.

[0054] The present application will be further described below with reference to the accompanying drawings and specific embodiments.

[0055] Please see Figure 1 A speech-based motor imagery training system for visually impaired patients includes a speech stimulation module, an environment adaptive adjustment module, an electroencephalogram (EEG) signal acquisition module, a data processing and feature extraction module, and a motor imagery recognition module.

[0056] The speech stimulation module is configured to output speech stimulation after the start of the motor imagery training task.

[0057] The environmental adaptive adjustment module is configured to collect environmental noise data during the output of speech stimulation by the speech stimulation module, and adjust the output strategy of speech stimulation according to the noise data.

[0058] The EEG signal acquisition module is configured to acquire the EEG signals of the patient performing a motor imagery training task under voice stimulation, and transmit the EEG signals to the data processing and feature extraction module.

[0059] The data processing and feature extraction module is configured to preprocess the EEG signal, obtain the preprocessed EEG signal, and extract frequency domain features and time domain features.

[0060] The motor imagery recognition module is configured to classify and recognize motor images based on preprocessed EEG signals, frequency domain features, and time domain features, using a trained motor imagery classification model to obtain recognition results.

[0061] In this embodiment, the voice stimulation module includes four voice output devices (such as high-fidelity audio devices), which can be arranged around the patient using either a uniform angle layout or a non-uniform scattering distance layout; in a uniform angle layout, such as... Figure 4As shown, four voice output devices are positioned centered on the patient, positioned in front, behind, to the left, and to the right, with equal angular spacing between each device. Horizontally, the four devices divide the 360° space into four equal parts, each with a 90° angle. This arrangement ensures a relatively even probability of the patient receiving sound from all directions during training, helping them develop a clear sense of sound direction and better associate the sound's direction with corresponding motor imagery. For example, in the initial training phase, the system sequentially plays vowel sounds from each device, and the patient visualizes movement based on the sound's direction. The evenly spaced angles allow the patient to perceive sound direction more accurately, improving training effectiveness.

[0062] In a non-uniform scattering distance layout (not shown), four voice output devices are distributed in a scattering pattern around the patient at varying distances. This layout breaks the conventional symmetrical distribution pattern, causing differences in the distance and time it takes for sound to reach the patient. Due to the different distances the sound travels, the intensity and time it takes to reach the patient's ears also vary, creating a more spatially layered and varied sound environment. During the enhancement training phase, the system designs complex task scenarios, randomly generating sounds corresponding to imagined movements in different directions. The sound variations under the non-uniform scattering distance layout can better simulate complex sound environments in real life, further testing and improving the patient's responsiveness and adaptability to the association between sound and movement commands.

[0063] In a specific embodiment, the layout pattern with the best effect can be selected through experiments.

[0064] In this embodiment, the speech stimulation module outputs speech stimulation through multiple speech output devices arranged in a preset manner around the patient. This provides visually impaired patients with rich and spatial auditory information, helping them to better concentrate and more effectively evoke motor imagination, thereby improving training results.

[0065] As an example, each voice output device plays specific vowel sounds "a", "e", "o", and "u", which correspond to specific motor control commands performed by the subject. Using only fixed vowels lacks diversity and can easily lead to user fatigue and decreased adaptability. In some embodiments, in addition to the original vowels (a, e, o, u), new digital voice stimuli ("1", "2", "3", "4") are added, each mapped to a different motor command, corresponding to the original four different commands, to enhance stimulus diversity and user perception.

[0066] The system uses voice output devices (speaker equipment) and voice prompts, eliminating the need for complex hardware and visual information, making the training process simpler and more intuitive. It can adjust the frequency, volume, and content of the voice prompts according to the needs of different patients, providing personalized training plans to ensure the effectiveness of the training.

[0067] In some embodiments, each audio device has a built-in sound signal generation module capable of playing preset vowel sounds and different numbers. The frequency, volume, and clarity of the sound can be adjusted according to the patient's specific needs, ensuring that the sound cues are clearly and effectively conveyed to the patient. The audio device can use embedded sensors to collect ambient noise data in real time. When the background noise exceeds a preset threshold (e.g., 50 dB), the system automatically activates an LMS-based algorithm to adjust the stimulus volume, and simultaneously corrects the subsequent playback strategy through a feedback mechanism.

[0068] Combination Figure 2 , Figure 3 and Figure 5 As shown, the training method of the system provided in this application may include: 1. Initial training: The patient sits in the center of the training room. Four audio devices are arranged evenly in the front, back, left, and right directions, or in a non-uniformly scattered layout, according to two different speaker distribution positions. The system plays vowel sounds emitted by each device in sequence, and the patient imagines the corresponding movement according to the direction of the sound source. When hearing "a", the patient should imagine the corresponding control command one, and so on. The control commands can be used for various different designs, such as controlling the wheelchair to turn forward, backward, left, and right, or controlling four different commands corresponding to the robotic arm. The initial training stage is based on the sequential arrangement of sounds from four directions. Through repeated practice, it helps the patient establish the association between sound and movement direction.

[0069] 2. Enhanced Training: Once the patient is sufficiently familiar with the association between the initial auditory cues and movement directions, the system enters the advanced training phase. At this stage, the system designs various complex task scenarios, with movement imagery in different directions being randomly generated instead of the original sequential arrangement to verify the training effectiveness. The system can also further enhance the training effect by incorporating tactile cues from vibrators installed at different locations on the subject.

[0070] In some embodiments, the system maps the vowel sounds played by each audio device to specific motion commands. Hearing "a" corresponds to the motion command 1: wheelchair forward; hearing "o" corresponds to the motion command 2: wheelchair turn left; hearing "e" corresponds to the motion command 3: wheelchair backward; and hearing "u" corresponds to the motion command 4: wheelchair turn right. These motion commands control the forward, backward, left, and right movements of the wheelchair, providing real-time feedback to the subject based on the wheelchair's movements. These mapped commands can also be used to operate smart home devices, allowing patients to directly control their home environment through motion visualization. These control commands can be used to operate smart home devices, such as turning lights on and off, and adjusting the temperature. This not only improves the practicality of the training but also enhances the patient's independence and convenience in daily life.

[0071] The system provided in this application offers visually impaired patients rich and spatially-oriented auditory information, helping them to better concentrate and more effectively evoke motor imagery, thereby improving training effectiveness. The multi-feature fusion approach can more comprehensively capture motor imagery information from EEG signals, improving recognition accuracy and helping to more accurately understand the patient's motor imagery state, thus optimizing the training program. Adjusting the speech stimulus output strategy based on noise data ensures that patients can clearly hear the speech stimulus under different environmental noise conditions, avoiding noise interference and improving training comfort and experience.

[0072] The application scenarios of the system provided in this application embodiment may include:

[0073] Rehabilitation training: This system is widely used in the rehabilitation training of visually impaired patients. By combining sound cues with motor imagery, it helps patients improve their motor coordination and enhance their self-care abilities in daily life.

[0074] Smart home control: The system can be integrated with smart home systems, allowing patients to control various devices in their homes through motor imagery, such as turning lights on and off and adjusting the temperature, thereby improving their quality of life and independence.

[0075] Virtual Reality and Navigation: The system can also be integrated with virtual reality technology to provide an immersive training environment for visually impaired patients. Furthermore, the system can be applied to navigation assistance systems to help patients navigate safely and practice obstacle avoidance in complex environments.

[0076] In some embodiments, the voice output device integrates an embedded microphone array, which uses an adaptive filtering algorithm to collect ambient noise in real time and dynamically adjust the gain parameters of the voice stimulus. Combined with short-time energy analysis, this ensures that the signal-to-noise ratio of the voice signal remains stable at 10-15 dB.

[0077] In some embodiments, the voice-based motor imagery training system for visually impaired patients also includes a central control unit responsible for managing the operational status of various audio devices, such as the duration, interval, and volume of sound playback. This module can also record patient response data and provide feedback through analysis to help optimize the training program.

[0078] In this embodiment, after the patient wears the EEG signal acquisition module (wearing an EEG signal acquisition device) to acquire EEG signals, the data processing and feature extraction module performs preprocessing using filtering algorithms. Commonly used filters include bandpass filters, which are used to remove electromyographic noise and power supply interference to ensure signal purity.

[0079] In this embodiment, during the initial training phase, the system uses a speech stimulation module to play vowel cues from different directions to induce motor imagery in the patient. A mixed stimulation signal can be pre-generated using an MP3 speech processing module, with a signal library containing pre-recorded vowels and digital speech. During the initial training phase, the speech stimulation module plays the signals in a fixed sequence, while during the reinforcement training phase, the speech stimulation module uses a random sequence generator to ensure that the order of the stimulation signals is not fixed, thus preventing the user from developing memory-based adaptation.

[0080] The environmental adaptive adjustment module collects environmental noise data during the speech stimulation output by the speech stimulation module. Based on this noise data, it adjusts the speech stimulation output strategy to automatically adjust the playback volume, ensuring a stable output above 70dB. Environmental noise is collected via a microphone and short-time energy analysis is performed in the time domain to calculate the current noise level L. noise Let the audio signal be x(i), the sampling frequency be s, and the short time window size be N0. Then the noise energy of a certain time window is calculated as follows:

[0081]

[0082] Where E(k) represents the noise energy within the time window at time k, k is the starting time marker of the time window, N0 is the length of the time window, and x(i) is the audio signal value collected at time i.

[0083] After averaging the values ​​over multiple windows, the noise level of the environment, i.e., the equivalent sound pressure level (SPL), is calculated.

[0084]

[0085] Where L noise Where E(j) is the noise level, E(j) is the short-time energy of the j-th time window, M is the number of windows for calculating the moving average, and C0 is the preset voice output device calibration constant (usually pre-calibrated under laboratory conditions).

[0086] To ensure the audibility of the stimulus audio, the ambient adaptive adjustment module can be set to output volume 10–15 dB higher than the background noise. The module can limit this calculated value between the maximum and minimum output volume. Assuming the default volume level of the current stimulus audio is Ldefault, the formula for the required gain adjustment factor is as follows:

[0087] Where L output =L noise +ΔL;L output For output volume level, L noise ΔL represents the noise level, where ΔL is the difference between the preset stimulus audio output volume level and the background noise level. default This is the default volume level.

[0088] The frame transition is gradually smoothed by using the gain adjustment factor G parameter to avoid discomfort caused by sudden volume changes.

[0089] In the system, the audio stimulus signal is stored in numerical form (floating-point array), and the embodiment can achieve automatic volume adjustment by multiplying the audio signal by a gain adjustment factor G.

[0090] As an example, the specific steps are as follows:

[0091] 1. Load audio signal: Read the audio stimulus file into a one-dimensional array. Each element in the array represents the amplitude value of a sampling point, ranging from [-1, 1].

[0092] 2. Applying gain adjustment: For each sampling point x(i), multiply by the gain adjustment factor G to obtain the adjusted signal:

[0093] xnew(i) = x(i) × G; This is equivalent to amplifying or reducing the amplitude of the entire audio waveform, thereby increasing or decreasing the volume.

[0094] 3. Prevent signal overflow:

[0095] If any value in the adjusted audio signal exceeds the range of [-1,1][-1,1][-1,1], it may cause distortion or playback errors. In this case, the entire audio signal needs to be normalized to scale it back to a safe range.

[0096] 4. Play audio: The adjusted audio signal is transmitted to the playback device for playback. At this time, the audio volume has been automatically adjusted according to the ambient noise.

[0097] For example: In the original audio data, a certain sample value is 0.5; the calculated gain adjustment factor G = 1.5, indicating that the volume should be increased; after adjustment, the sample value becomes: xnew = 0.5 × 1.5 = 0.75; the same processing will be applied to the entire audio segment, thereby increasing the overall volume.

[0098] To avoid patient discomfort caused by sudden volume changes, the system's environmental adaptive adjustment module uses a gain adjustment factor G parameter to gradually and smoothly transition between frames. The audio signal is divided into multiple frames at regular time intervals. Within each frame, the volume is gradually adjusted based on the G value, ensuring smooth volume changes between adjacent frames rather than abrupt shifts. If increased background noise is detected at a certain moment, the calculated G value increases accordingly. The system does not instantly adjust the volume to a new level; instead, it gradually increases the volume in subsequent frames, making the volume change virtually imperceptible to the patient and improving the training experience.

[0099] Through real-time noise detection and gain adjustment, this system can provide a stable stimulus volume (e.g., above 70dB) under different ambient noise levels. It also employs exponential smoothing and adaptive filtering to avoid discomfort caused by sudden volume changes and enhance signal audibility. Experimental verification shows that the system maintains a signal-to-noise ratio of 10–15dB within a background noise range of 40–80dB, ensuring high intelligibility of the stimulus signal.

[0100] In this embodiment, the EEG signal acquisition module uses a 16-channel high-precision EEG acquisition device to ensure good electrode contact during acquisition and reduce motion artifacts. The data is transmitted in real time to the data processing and feature extraction module at a fixed sampling rate of 250Hz for preprocessing, such as applying bandpass filtering (0.5–40Hz) to remove power frequency interference and low-frequency drift.

[0101] P300 is an important component of event-related potentials (ERPs), typically appearing within 250-400 ms after stimulation. To ensure the stability of feature extraction, in this embodiment, the data processing and feature extraction module uses a sliding window method to extract key temporal features of P300.

[0102] The selected window width is 400ms, which covers the typical time range of P300; the step size is 100ms, balancing computational efficiency and temporal resolution; since P300 is usually more significant in the apical and central regions, the channels Fz and Pz are selected. These parameters are used to calculate the peak amplitude of P300 and the corresponding time point within each sliding window for feature extraction. The specific method is as follows:

[0103]

[0104] x(t) is the EEG signal, and the peak value of P300 within the window is calculated. Where AP300 For P300 peak amplitude, T P300 t0 represents the time point corresponding to the peak value, and t0 is the starting time of the sliding window.

[0105] Due to interference from environmental noise and physiological artifacts, the directly extracted P300 signal from EEG may exhibit significant fluctuations. To improve feature stability, exponential moving average (EMA) filtering is used for time-domain smoothing. Let the original P300 peak signal sequence be An, and the smoothed signal be Sn. The EMA is calculated as follows:

[0106] S n =αA n +(1-α)S n-1 ;

[0107] Where α is the smoothing coefficient set to 0.3, and Sn-1 is the smoothing result at the previous time point.

[0108] To further enhance classification performance, in addition to the P300 peak value, the power spectral density (PSD) features for the corresponding time period were extracted to obtain frequency domain information. The PSD calculation employed the Welch method to calculate the power spectrum of the key frequency band (8-30Hz).

[0109]

[0110] X n (f) represents the short-time Fourier transform (STFT) result of the signal, where N is the number of windows.

[0111] Finally, the system's data processing and feature extraction module extracts time-domain P300 and frequency-domain PSD information to ensure complementary signal features and improve classification accuracy. Based on the preprocessed EEG signal, P300 time-domain features, PSD frequency-domain features, and smoothed signal, a motor imagery classification model is constructed to improve recognition accuracy and system stability.

[0112] In this embodiment, the motor imagery recognition module uses a motor imagery classification model that performs the following steps: dividing the EEG signal into preset time windows, calculating the P300 peak amplitude and corresponding time point within each time window; smoothing the obtained P300 peak amplitude sequence to obtain a smoothed signal; calculating the power spectrum of a pre-selected key frequency band; and inputting the pre-processed EEG signal, the calculated P300 peak amplitude and corresponding time point for each time window, the smoothed signal, and the power spectrum into the trained motor imagery classification model to perform motor imagery classification and recognition, and obtain the recognition result.

[0113] Specifically, please see Figure 6 and Figure 7To improve the signal classification performance of the motor imagery brain-computer interface (MI-BCI), this system is optimized based on EEGNet, incorporating a multi-head attention mechanism and frequency domain feature fusion. This aims to enhance the motor imagery classification model's ability to capture spatiotemporal information and improve its adaptability to background noise. Furthermore, the motor imagery classification model employs an adaptive training mechanism, allowing for continuous parameter optimization during online operation to improve long-term stability.

[0114] 1. Temporal convolutional layer

[0115] First, a one-dimensional convolution (1D-CNN) is performed on the preprocessed EEG signal to extract local temporal features. Since EEG signals have strong temporal dependencies, this layer can effectively learn waveform changes over a short period, thereby capturing preliminary feature representations. Let the input EEG signal be x∈R. C×T Where C is the number of channels and T is the time step, the temporal convolution is calculated as follows:

[0116] H1 = σ(W1*X + b1);

[0117] Where W1 is the convolution kernel weight matrix, b1 is the bias term, * represents the convolution operation, and σ is the non-linear activation function (ReLU).

[0118] 2. Bullish Attention Layer

[0119] Traditional EEGNet only uses convolutional neural networks (CNNs) to extract features, lacking attention to global features. To compensate for this deficiency, this model introduces a multi-head self-attention (MHSA) mechanism to model long-distance dependencies. This layer maps the output of the convolutional layer to queries, keys, and values, calculating the relevance within each subspace, thereby enhancing the model's global perception capability. Let H1 be the input feature, then the attention is calculated as follows:

[0120] Q = W q H1K = W k H1V = W v H1;

[0121]

[0122] Among them W q W k W v For the projection matrix of query, key, and value, d k Softmax is used as a scaling factor to normalize the attention distribution. Multiple attention heads are computed in parallel and then concatenated.

[0123] H3=Concat(head1,head2,…,head h W o ;

[0124] W o Let h be the final linear transformation matrix, and h be the number of attention heads.

[0125] 3. Spatial Convolutional Layer

[0126] In EEG signal processing, spatial correlations exist between different channels. To effectively extract cross-channel information, the model employs depthwise separable convolution. This layer comprises both depthwise convolution and pointwise convolution, reducing computational complexity while preserving sufficient feature information. The spatial convolution calculation formula is as follows:

[0127] H3=σ(W 3,d *H2+b3)+σ(W 3,p (H2+b3)

[0128] Among them W 3,d For depthwise convolution kernels (acting only within a single channel), W 3,p b3 is the pointwise convolution kernel (used for inter-channel fusion), and b3 is the bias term.

[0129] 4. Frequency domain fusion branch

[0130] Besides time-domain features, the frequency characteristics of EEG signals are equally crucial. This model extracts the power spectral density (PSD) and wavelet transform coefficients during front-end signal preprocessing and then reduces their dimensionality using a fully connected network (FCN). Let the original frequency domain feature F be dimensionality-reduced and then concatenated with the feature H3 obtained from the spatial convolutional layer to form a multimodal fusion feature:

[0131] F′=σ(W f F+b f );

[0132] H4 = Concat(H3, F′);

[0133] 5. Decision-making level

[0134] The fused features are fed into the final fully connected classification layer, and the probability distribution is output through the Softmax function. To improve the stability of the classification, a threshold determination mechanism (P300 greater than 0.4 and relative β-wave power greater than 1.25 times the baseline) is added to the Softmax result for secondary screening to ensure the reliability of the classification results. The final classification decision is as follows:

[0135] P = Softmax(W d H4+b d );

[0136]

[0137] The embodiment also includes training the motion imagery classification model and online real-time parameter tuning. The improved motion imagery classification model (EEGNET) is pre-trained using an offline-collected dataset to ensure high classification accuracy under various noise levels. During online operation, some model parameters (such as attention weights and fully connected layer parameters) are continuously updated through a feedback mechanism, enabling the model to adaptively adjust online.

[0138] In this embodiment, the motion imagery classification model uses an improved EEGNET classification model: by introducing a multi-head attention mechanism to capture global dependencies and using a frequency domain fusion module to integrate multi-scale features, the classification performance is improved by more than 15% compared with the traditional model.

[0139] In some embodiments, such as Figure 1 As shown, the system also includes: a sensory state monitoring module, which is configured to collect the patient's skin conductance response data and heart rate data before the start of the motor imagery training task; extract features based on the skin conductance response data and heart rate data; identify the patient's current emotional state based on the extracted features; and determine whether the patient's current emotional state is suitable for starting motor imagery training.

[0140] As an example, determining whether a patient's current state is suitable for starting motor imagery training based on characteristics includes:

[0141] The extracted features are input into a support vector machine (SVM), and the patient's current emotional state is identified using the following formula:

[0142]

[0143] Where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, and ξ is the weight vector. i As a slack variable, α i and β i For Lagrange multipliers, x i Let y be the feature vector of the i-th training sample. iLet N be the class label of the i-th training sample, and N be the number of samples.

[0144] When patients are in a suitable emotional and physical state, they are able to focus more on training tasks, cooperate more actively, and their brains are more sensitive to verbal stimuli, which is more conducive to triggering motor imagery and making the training more effective.

[0145] In some embodiments, such as Figure 1 As shown, the system also includes a central control module. After classifying motor imagery of EEG signals using EEGNet with multi-head attention, the system uses the central control module to map these classification results to specific control commands. For example, when the system's motor imagery recognition module recognizes the patient's imagined pronunciation of 'a', it generates control command 1, with each control command corresponding to a specific device operation.

[0146] After a control command is generated, the system's central control module sends the command to the device to be controlled via a wireless communication module (Bluetooth), and the device executes the corresponding operation. Simultaneously, in some embodiments, the central control module monitors the device's execution in real time using sensors and provides feedback on the device's response (direction of movement, state) to the user. For example, if the central control module controls a wheelchair, it will provide feedback on the wheelchair's actual movement to the user, confirming the effectiveness of the command execution through visual, audio, or tactile feedback modules.

[0147] In some embodiments, the central control module adaptively adjusts based on feedback information. If it detects that the device is not responding as expected (e.g., a deviation between the patient's imagined movement and actual control), it fine-tunes the commands using a PID (proportional-integral-derivative) control algorithm to ensure the device executes according to the user's intent. For example, if the wheelchair moves too fast or deviates in direction, the central control module will gradually adjust the parameters until the device achieves optimal control.

[0148] In some embodiments, the central control module can be adapted to different types of smart devices by designing a universal interface protocol. For example, through a unified API and Bluetooth or WiFi communication protocol, the central control module can not only control wheelchairs but also interconnect with other smart home devices (such as smart lights, smart curtains, etc.). This design allows users to conveniently operate more devices through motion visualization.

[0149] The central control module offers personalized configuration capabilities, allowing users to customize control scenarios to meet diverse needs. Users or doctors can assign different control commands to different motor images. Through a graphical interface, users can predefine control schemes, mapping different motor image signals to specific operations to ensure the system better suits their actual requirements.

[0150] To enable broader applications, the central control module supports cloud-based data management. All training records, EEG signals, device control commands, and feedback data can be uploaded to the cloud for storage and analysis. Doctors or technicians can monitor the user's rehabilitation progress remotely and optimize system parameters based on big data analysis. Furthermore, the system can form an Internet of Things (IoT) ecosystem with medical equipment and smart devices through the cloud platform, further expanding its functionality.

[0151] The speech-based motor imagery training system for visually impaired patients of the present invention, as a motor imagery brain-computer interface system, can accurately identify the patient's motor imagery signals, intelligently adjust the training difficulty, and realize the control of smart home devices, ultimately providing effective rehabilitation training and convenience for visually impaired patients.

[0152] The above provides a detailed description of the speech-based motor imagery training system for visually impaired patients provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the concept of this application and should not be construed as limiting the scope of protection of this application.

Claims

1. A voice-based motor imagery training system for visually impaired patients, characterized in that, include: The speech stimulation module is configured to output speech stimulation after the start of the motor imagery training task; An environment adaptive adjustment module is configured to collect environmental noise data during the output of speech stimulation by the speech stimulation module, and adjust the output strategy of speech stimulation according to the noise data. The EEG signal acquisition module is configured to acquire the EEG signals of the patient performing a motor imagery training task under the speech stimulation, and transmit the EEG signals to the data processing and feature extraction module. The data processing and feature extraction module is configured to preprocess the EEG signal to obtain the preprocessed EEG signal and extract frequency domain features and time domain features. The motor imagery recognition module is configured to classify and recognize motor images based on preprocessed EEG signals, frequency domain features, and time domain features, using a trained motor imagery classification model to obtain recognition results. The environmental adaptive adjustment module specifically performs the following: Based on the environmental noise data, noise energy is calculated for a time window to determine the current noise level; The current output volume level is determined based on the current noise level and the difference between the preset stimulus audio output volume level and the background noise level. The gain adjustment factor is determined based on the output volume level and the default volume level. The volume of the output speech stimulus is adjusted according to the gain adjustment factor; The formula for determining the gain adjustment factor is as follows: ; in ; For output volume level, Noise level, The difference between the preset stimulus audio output volume level and the background noise level. This is the default volume level.

2. The speech-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The system also includes an emotional state monitoring module, which is configured to collect the patient's skin conductance response data and heart rate data before the start of the motor imagery training task. Based on the skin conductance data and heart rate data, features are extracted, and the patient's current emotional state is identified based on the extracted features to determine whether the patient's current emotional state is suitable for starting motor imagery training.

3. The speech-based motor imagery training system for visually impaired patients according to claim 2, characterized in that, Determining whether a patient's current state is suitable for starting motor imagery training based on the aforementioned characteristics includes: The extracted features are input into a support vector machine (SVM), and the patient's current emotional state is identified using the following formula: ; Where ω is the weight vector, b is the bias term, λ is the hyperparameter, C is the penalty parameter, and ξ is the weight vector. i As a slack variable, α i and β i For Lagrange multipliers, x i Let y be the feature vector of the i-th training sample. i Let N be the class label of the i-th training sample, and N be the number of samples.

4. The speech-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The system also includes a central control module, which is configured to map the identification results to specific control commands and send the control commands to the device to be controlled so that the device to be controlled performs the corresponding operation.

5. The speech-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The speech stimulation module is configured to include four speech output devices, which are arranged around the patient in a uniform angle layout or a non-uniform scattering distance layout. In the uniform angle layout, four voice output devices are arranged in front of, behind, to the left and to the right of the patient, with equal angular intervals between each device. In the aforementioned non-uniform scattering distance layout, the four voice output devices are distributed in a scattering pattern around the patient, and the distances between them and the patient are unequal.

6. The speech-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The data processing and feature extraction module performs the following: The EEG signal is divided into preset time windows, and the P300 peak amplitude and corresponding time point are calculated in each time window. The obtained P300 peak amplitude sequence is smoothed to obtain the smoothed signal. And calculate the power spectrum of the pre-selected key frequency bands; The preprocessed EEG signal, the calculated P300 peak amplitude and corresponding time point of each time window, the smoothed signal, and the power spectrum are input to the motor imagery recognition module.

7. The speech-based motor imagery training system for visually impaired patients according to claim 6, characterized in that, The motion imagery classification model includes a temporal convolutional layer, a multi-head attention layer, a spatial convolutional layer, a frequency domain fusion branch, and a decision layer connected in sequence. The temporal convolutional layer takes the preprocessed EEG signal as input, performs one-dimensional convolution, and extracts local temporal features. The multi-head attention layer maps the local temporal features output by the temporal convolutional layer into queries, keys, and values, calculates the correlation between each subspace, and concatenates the parallel calculations of multiple attention heads to obtain a global feature representation that integrates global information. The spatial convolutional layer is used to combine the global feature representation with the cross-channel information extracted by the spatial convolutional layer to obtain deep features in the spatiotemporal dimension. The frequency domain fusion branch fuses the depth features, the calculated P300 peak amplitude and corresponding time point of each time window, the smoothed signal, and the power spectrum to form a multimodal fusion feature. The decision layer is used to input the multimodal fusion features into the fully connected classification layer, output the probability distribution through the Softmax function, and obtain the classification result through a threshold determination mechanism.

8. The speech-based motor imagery training system for visually impaired patients according to claim 1, characterized in that, The motor imagery training includes initial training and enhancement training; in the initial training phase, the speech stimulation module outputs speech stimuli in a fixed order to guide the patient to establish an association between sound and movement direction; During the reinforcement training phase, the speech stimulation module randomly outputs speech stimuli to verify the training effect.

Citation Information

Patent Citations

  • Volume adjustment method and device and equipment

    CN107124149A

  • Hand rehabilitation training method based on motor imagery

    CN111110982A

  • Intelligent internal medicine nursing monitoring system

    CN117854739A