Noninvasive brain-computer interface system and method based on dynamic attention multi-modal fusion

By employing flexible graphene dry electrodes, dynamic adaptive noise reduction, and lightweight recognition technology, the problems of signal instability and rigid fusion strategies in non-invasive brain-computer interfaces have been solved, enabling efficient and real-time brain-computer interaction. This technology is applicable to fields such as rehabilitation medicine, daily intelligent interaction, and special environment operations.

CN122018685APending Publication Date: 2026-05-12ZHONGSHAN HOSPITAL FUDAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGSHAN HOSPITAL FUDAN UNIV
Filing Date
2026-01-26
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing non-invasive brain-computer interfaces suffer from poor signal acquisition stability, rigid multimodal feature fusion strategies, and insufficient real-time command recognition, making it difficult to meet the needs of large-scale applications in everyday scenarios.

Method used

The system employs a flexible wearable multimodal acquisition module, a dynamic adaptive noise reduction module, a spatiotemporal dynamic attention fusion module, and a lightweight command recognition module, combined with a cross-device adaptive control module, to form a full-link optimization solution. This solution includes flexible graphene dry electrode acquisition, dynamic adaptive noise reduction, dynamic attention fusion, and a lightweight recognition model, achieving high anti-interference capability, scene adaptation, and efficient real-time interaction of the signal.

Benefits of technology

It improves the stability and anti-interference capability of signal acquisition, enhances the flexibility of multimodal feature fusion and the real-time performance of recognition models, meets the real-time interaction needs of portable devices, and expands application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018685A_ABST
    Figure CN122018685A_ABST
Patent Text Reader

Abstract

The first aspect of the technical scheme of the invention is to disclose a noninvasive brain-computer interface system based on dynamic attention multi-modal fusion, which comprises a flexible wearable multi-modal acquisition module, a dynamic adaptive noise reduction module, a time-space domain dynamic attention fusion module, a lightweight instruction identification module and a cross-device adaptation control module, and a full-link optimization scheme of acquisition, noise reduction, fusion, identification and control is formed. The second aspect of the technical scheme of the invention is to disclose a non-invasive brain-computer interface method based on dynamic attention multi-modal fusion. According to the non-invasive brain-computer interface system and method, the environment change can be dynamically adapted, the signal anti-interference capability and recognition accuracy are improved, and the system and method have real-time interaction performance, so that the bottlenecks of unstable signals, rigid fusion strategies and poor real-time performance existing in a traditional scheme can be solved; and a safe, efficient and accurate brain-computer interaction solution can be provided for different people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a non-invasive brain-computer interface system and method based on dynamic attention multimodal fusion, belonging to the field of brain-computer interface technology. Background Technology

[0002] Brain-computer interface (BCI), as a core interactive technology that directly connects the brain to external devices, breaks through the physiological limitations of traditional neuromuscular conduction and provides revolutionary solutions for the reconstruction of motor function in patients with neurological diseases and human-computer intelligent interaction. It has become a research hotspot in the interdisciplinary fields of biomedical engineering, neuroscience and artificial intelligence.

[0003] From a technological perspective, brain-computer interfaces (BCIs) can be divided into invasive and non-invasive types. While invasive methods can acquire high-resolution EEG signals, they require surgical implantation of electrodes, posing medical risks such as infection and bleeding. Furthermore, the surgery is costly and post-operative maintenance is complex, limiting their application to clinical research on a small number of critically ill patients and hindering large-scale deployment. In contrast, non-invasive BCIs, which acquire EEG signals via scalp electrodes (such as electroencephalography, EEG), offer significant advantages including being non-invasive, easy to operate, cost-effective, and suitable for a wide range of patients. They have become the mainstream direction for the industrialization of BCIs and are widely used in rehabilitation medicine, smart wearables, industrial control, and other applications.

[0004] Despite some progress in non-invasive brain-computer interface technology, it still faces numerous technical bottlenecks in practical applications, severely restricting its interactive performance and large-scale deployment. These bottlenecks manifest in the following three core aspects: Firstly, signal acquisition stability is poor, and anti-interference capability is insufficient. Most existing non-invasive brain-computer interfaces rely on a single EEG modality (such as event-related potentials and motor imagery EEG) for signal acquisition. However, scalp EEG signals themselves have weak amplitudes (microvolts), making them highly susceptible to external environmental interference (such as 50Hz power frequency interference and electromagnetic radiation) and physiological artifacts (such as EEG, EMG, and ECG artifacts), resulting in an extremely low signal-to-noise ratio. Furthermore, traditional solutions often use wet electrodes for signal acquisition, requiring the use of conductive gel. This is not only cumbersome and uncomfortable to wear, but also suffers from signal quality degradation after the conductive gel dries, limiting the continuous usability of the device. In addition, existing noise reduction algorithms are mostly designed with fixed parameters, only able to process specific types of interference, and cannot dynamically adapt to complex and changing real-world application environments, further exacerbating signal instability.

[0005] Secondly, the multimodal feature fusion strategy is rigid and lacks scene adaptability. To improve signal reliability, some solutions attempt to introduce multimodal signals (such as combining EEG with EEG and EMG), but existing feature fusion methods still have obvious defects: most solutions use fixed-weight weighted fusion or simple feature splicing, without fully considering the inherent correlation and complementarity of different modal features in the spatiotemporal dimension; more importantly, the fixed-weight strategy cannot dynamically adjust the contribution of each modality feature according to the actual application scenario (such as quiet indoor, noisy outdoor, changes in user physiological state, etc.), which makes it impossible to give full play to the advantages of multimodal fusion in complex environments, and even leads to feature redundancy after fusion and a significant drop in recognition accuracy.

[0006] Third, the command recognition models are inefficient and fail to meet real-time interaction requirements. The core requirement of non-invasive brain-computer interfaces is to achieve real-time interaction between the brain and external devices, which places extremely high demands on the inference speed of the recognition model. However, existing command recognition models are mostly based on complex deep learning network structures (such as deep convolutional neural networks and recurrent neural networks), with a huge number of parameters. This not only requires powerful hardware computing power but also suffers from high inference latency, making it difficult to deploy on portable and embedded devices. At the same time, existing models are mostly customized for specific tasks and lack generalization ability. When the types of interaction commands increase or the application scenarios change, the recognition performance of the model will significantly decrease, making it unable to adapt to diverse daily application needs.

[0007] In summary, the technical shortcomings of existing non-invasive brain-computer interfaces in terms of signal acquisition stability, multimodal feature fusion flexibility, and the real-time performance and generalization of recognition models make it difficult for them to meet the needs of large-scale applications in everyday scenarios. Summary of the Invention

[0008] This invention aims to address the technical bottlenecks of existing non-invasive brain-computer interfaces, such as poor signal acquisition stability, rigid multimodal feature fusion strategies, and insufficient real-time command recognition, and provides a non-invasive brain-computer interface solution that combines high anti-interference capability, scene adaptability, and efficient real-time interaction performance.

[0009] To achieve the above objectives, the first aspect of the present invention discloses a non-invasive brain-computer interface system based on dynamic attention multimodal fusion, comprising a flexible wearable multimodal acquisition module, a dynamic adaptive noise reduction module, a spatiotemporal dynamic attention fusion module, a lightweight command recognition module, and a cross-device adaptive control module, forming a full-link optimization scheme of "acquisition-noise reduction-fusion-recognition-control", wherein: The flexible wearable multimodal acquisition module adopts flexible graphene dry electrode acquisition technology to achieve low impedance signal acquisition of less than 1kΩ without the need for conductive paste. It also acquires three-modal physiological signals through a built-in high-precision synchronous acquisition chip. The three-modal physiological signals include EEG containing event-related potentials and / or motor imagery EEG, EEG containing eye movement trajectory-related signals, and EMG used to help eliminate artifacts and supplement motor intentions. The dynamic adaptive noise reduction module integrates a variable step-size adaptive notch filter and a multi-scale residual wavelet denoising algorithm to form a two-layer noise reduction architecture of "targeted anti-interference + global artifact removal", outputting high-purity multimodal signals. Among them, the variable step-size adaptive notch filter dynamically tracks power frequency interference and harmonic signals in the 45-65Hz range, accurately filtering out power grid interference; the multi-scale residual wavelet denoising algorithm performs multi-scale decomposition and threshold denoising on the signal, and at the same time repairs the wrongly filtered EEG detail signals through the residual compensation mechanism. The spatiotemporal dynamic attention fusion module adopts a neural network architecture that combines spatial and temporal attention sublayers. Specifically, the spatial attention sublayer uses a channel attention mechanism, calculating the weight coefficients of features from each acquisition channel through fully connected layers and a sigmoid activation function. This focuses on the effective features of core channels and suppresses the interference features of redundant channels. Different acquisition channels correspond to different brain regions and modalities. The temporal attention sublayer is designed based on long short-term memory units to capture the dynamic correlation features of signal temporal dimensions (such as the rhythmic change trend of motor imagery EEG) and accurately identify key features within different time windows. The lightweight instruction recognition module constructs a "teacher-student" network training framework based on knowledge distillation technology. It uses a high-precision deep convolutional recurrent neural network as the teacher network and a lightweight convolutional recurrent neural network as the student network. By transferring the knowledge and recognition experience of the teacher network, the performance of the student network is improved. The trained student network can be directly deployed on embedded devices, including smartphones or portable EEG terminals, without the need for additional computing power. The cross-device adaptable control module has a built-in multi-protocol communication chip and adaptive switching unit, supporting real-time adaptive switching of multiple communication protocols such as Bluetooth 5.0, WiFi 6, and serial ports (RS232 / RS485). It can automatically select the optimal protocol according to the type of external device, communication distance, and data transmission requirements, and automatically identify the type of external device and complete the protocol matching through device fingerprint recognition technology. It can achieve precise control with plug and play without manual configuration.

[0010] Preferably, the flexible graphene dry electrode acquisition technology uses graphene / carbon fiber composite conductive material to prepare flexible dry electrodes for low impedance signal acquisition. The electrode surface is treated with micro-nano texture to form a self-adhesive elastic structure, which adheres tightly to the scalp through van der Waals forces. It is effectively adapted to various complex environments such as dry scalp, sweating, and short-term rain. It also has excellent biocompatibility and can be worn for more than 24 hours without obvious discomfort.

[0011] Preferably, the high-precision synchronous acquisition chip supports synchronous acquisition of 16 or more multimodal signals, and the sampling frequency can be adaptively adjusted within the range of 250-1000Hz according to the user's usage scenario to ensure that the signal accuracy requirements in different scenarios can be met. Specifically, in static rehabilitation training scenarios, a sampling rate of 250-500Hz is used to balance accuracy and power consumption; in high-precision scenarios, including dynamic virtual reality interaction, the sampling rate is automatically switched to 500-1000Hz.

[0012] Preferably, the variable step size adaptive notch filter is optimized based on the minimum mean square error algorithm, and the convergence speed of the filter is improved by more than 30% by dynamically adjusting the step size coefficient.

[0013] Preferably, the multi-scale residual wavelet denoising algorithm uses the db4 wavelet basis function to decompose the signal into 5-8 different scales. Based on threshold denoising, a residual compensation mechanism is introduced to supplement high-frequency detail components during reverse reconstruction, effectively solving the signal distortion problem caused by traditional wavelet denoising. The measured signal-to-noise ratio of the reconstructed signal is improved by more than 25%, and the fidelity preservation effect is particularly significant for weak EEG signals.

[0014] Preferably, the dynamic adaptive noise reduction module has a built-in interference identification unit that automatically determines the type of interference based on signal characteristics and adaptively selects the core noise reduction algorithm combination to further improve the targeting of noise reduction. If the main interference is power frequency, then a variable step size adaptive notch filter is used to process the signal. If the signal contains a large number of physiological artifacts such as those caused by eye movements or muscle twitching, a multi-scale residual wavelet denoising algorithm is used to process the signal. If the interference is mixed, a combination strategy of variable step size adaptive notch filter and multi-scale residual wavelet denoising algorithm is adopted to complete the noise reduction process in steps.

[0015] Preferably, the weight coefficients of the spatial-temporal dual attention sublayer are dynamically updated every 10ms based on the signal quality of the current scene and the user's physiological state. The contribution of each modality and channel feature is adjusted in real time according to different scenarios, including static rehabilitation, outdoor movement, and noisy environments, so as to significantly improve the feature recognition in complex environments. After testing, the inter-class distance of the features after fusion is improved by more than 40%. The signal quality includes signal-to-noise ratio and interference intensity, and the user's physiological state is judged in real time by auxiliary physiological signals.

[0016] Preferably, the student network core uses depthwise separable convolution to replace traditional convolution operations, splitting standard convolution into depthwise convolution and pointwise convolution. While maintaining feature extraction capabilities, this reduces the number of model parameters by more than 60%. At the same time, channel pruning technology is introduced to remove redundant channels and invalid parameters, further compressing the model size. After optimization, the model inference latency is less than 50ms, meeting the requirements of real-time interaction. Moreover, in a test set containing 20 common commands (such as body movement intentions and device control commands), the recognition accuracy remains above 95%.

[0017] Preferably, the cross-device adaptation control module has a built-in protocol library covering mainstream devices including rehabilitation robots, smart wheelchairs, VR devices, and home control systems.

[0018] Preferably, the cross-device adaptation control module has instruction encryption and verification functions. It encrypts the identified instructions using a hash algorithm to prevent tampering during transmission and ensure control security.

[0019] The second aspect of the technical solution of this invention discloses a non-invasive brain-computer interface method based on dynamic attention multimodal fusion, comprising the following steps: S1. Adaptive synchronous acquisition of multimodal signals: Determine the user's usage scenario and automatically match the corresponding sampling frequency parameters; The flexible graphene dry electrode array is activated to simultaneously collect the user's electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) physiological signals. During the collection process, a high-precision clock chip is used to achieve precise alignment of the timestamps of the three modal signals. S2, Dynamic Adaptive Noise Reduction Processing: After receiving the acquired raw signal, the dynamic adaptive noise reduction module uses an interference identification unit to determine the main types of interference in the signal. If the main interference is power frequency interference, then activate the variable step size adaptive notch filter to dynamically track and accurately filter out interference signals in the 45-65Hz range. If the signal contains a large number of physiological artifacts, the multi-scale residual wavelet denoising algorithm is activated to perform multi-scale decomposition and threshold denoising on the signal. At the same time, the residual compensation mechanism is used to repair the EEG detail signals that were mistakenly filtered out. If the interference is mixed, a combination strategy of variable step size adaptive notch filter and multi-scale residual wavelet denoising algorithm is adopted to complete the noise reduction process in steps. Ultimately, a high-purity multimodal signal is output; S3, Three-dimensional feature extraction: For the noise-reduced EEG, EEG, and EMG signals, a three-dimensional feature extraction system in the time domain, frequency domain, and spatial domain is constructed respectively. The three-dimensional features of the three modalities are integrated to construct a comprehensive high-dimensional feature set. S4. Spatiotemporal dynamic attention feature fusion: The high-dimensional feature set constructed in step S3 is input into the spatiotemporal dynamic attention fusion module: The weight coefficients of each channel and modality feature are calculated by spatial attention sublayer, focusing on strengthening the weight of EEG features in core brain regions and weakening the weight of features in electrooculography and electromyography that are unrelated to the command intention. By capturing the dynamic changes of features in the temporal dimension through the time attention sublayer, we can focus on key feature segments before and after the instruction is triggered. The dual attention sublayer updates the weight coefficients every 10ms and dynamically adjusts the fusion strategy according to the current signal quality and scene, ultimately completing the dynamic fusion of features and generating a highly recognizable one-dimensional fusion feature vector, effectively reducing feature redundancy. S5, lightweight command quick recognition: The fused feature vector output from step S4 is input into a lightweight convolutional recurrent neural network. The network first extracts the local correlation information of the features through a depthwise separable convolutional layer, then captures the temporal dependency of the features through an LSTM layer, and finally outputs the instruction recognition result through a fully connected layer and a Softmax activation function. S6, Cross-device adaptive control: The lightweight command recognition module outputs the recognition results to the cross-device adaptation control module. The cross-device adaptation control module automatically parses the meaning of the command and matches the corresponding control protocol. It adaptively switches between Bluetooth, WiFi, or serial communication protocols according to the type of external device. Through the built-in protocol library, it completes the fast connection and data transmission with the external device, realizing precise control of the external device. At the same time, the cross-device adaptation control module receives feedback signals from the external device in real time. If control failure or signal loss occurs, it automatically switches the communication protocol and resends the command to ensure the stability and reliability of control, truly realizing plug-and-play convenient interaction.

[0020] The non-invasive brain-computer interface system and method disclosed in this invention can dynamically adapt to environmental changes, improve signal anti-interference ability and recognition accuracy, and have real-time interactive performance. It can solve the bottlenecks of traditional solutions such as unstable signals, rigid fusion strategies and poor real-time performance. Its core applications cover rehabilitation medicine, daily intelligent interaction, special environment operation and scientific research support, and can provide safe, efficient and accurate brain-computer interaction solutions for different groups of people.

[0021] This invention improves the accuracy and stability of EEG signal recognition in complex environments by introducing a multimodal signal acquisition and dynamic attention fusion mechanism. Through a lightweight command recognition model design, it meets the deployment requirements of portable devices and real-time interaction needs. Ultimately, this enables the large-scale application of this technology in multiple fields such as rehabilitation medicine, daily intelligent interaction, and special scenario operations, providing safe, convenient, and accurate brain-computer interface services for different groups and promoting the industrialization of non-invasive brain-computer interface technology. Compared with existing technologies, this invention has the following beneficial effects: (1) Innovative acquisition technology: The flexible graphene dry electrode array is adopted, which does not require conductive paste, is comfortable to wear and has a long service life. The low impedance characteristics of the electrodes ensure signal stability in various environments, breaking through the limitations of traditional wet electrodes. (2) Noise reduction algorithm innovation: The dynamic adaptive noise reduction module combines variable step size notch filtering and residual wavelet denoising, which can track and filter out dynamic interference in real time, improve the signal-to-noise ratio of signal reconstruction by more than 25%, and significantly enhance the anti-interference ability. (3) Innovation of fusion mechanism: The spatiotemporal dynamic attention fusion mechanism realizes the real-time dynamic update of feature weights. Compared with the traditional fixed weight fusion, the recognition accuracy in complex scenarios is improved by 15%-20%, and the cross-scenario adaptability is greatly improved. (4) Model design innovation: The lightweight convolutional recurrent neural network achieves dimensionality reduction through knowledge distillation technology, reducing the number of parameters by more than 60% and the inference latency to less than 50ms, meeting the real-time interaction requirements of portable devices; (5) Innovative control scheme: The cross-device adaptable control module supports multi-protocol adaptive switching, realizes plug and play, and greatly expands the application scenarios of the system. Attached Figure Description

[0022] Figure 1 This is a diagram of the overall architecture of the present invention; Figure 2 This is a flowchart of the present invention. Detailed Implementation

[0023] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0024] The first aspect of this invention discloses a non-invasive brain-computer interface system based on dynamic attention multimodal fusion, primarily used for rehabilitation training of stroke patients and daily home control scenarios. Through the collaborative design of five functional modules, it achieves highly stable, highly accurate, and low-latency brain-computer interaction. Combined with... Figure 1 The system disclosed in this embodiment of the invention specifically includes a flexible wearable multimodal acquisition module, a dynamic adaptive noise reduction module, a spatiotemporal dynamic attention fusion module, a lightweight command recognition module, and a cross-device adaptation control module. Each module is connected through a high-speed data bus to form a closed-loop system of "signal acquisition - noise reduction processing - feature fusion - command recognition - device control".

[0025] The flexible wearable multimodal acquisition module serves as the core of the signal input, employing a 16-channel flexible graphene dry electrode array design. The electrode body is made of graphene / polydimethylsiloxane (PDMS) composite material, combining excellent conductivity and flexible fit. The measured electrode impedance is stable at 3.2kΩ (test conditions: frequency 10Hz-1kHz, dry scalp), achieving high-quality signal acquisition without the need for conductive paste. The deployment position of the 16-channel flexible graphene dry electrode array has been precisely optimized, with the specific fitting scheme as follows: Six channels were deployed in the prefrontal cortex region, corresponding to the Fp1, Fp2, F3, F4, Fz, and F7 sites of the international 10-20 system, to collect EEG signals related to visual evoked potentials and prefrontal cognition. Six channels were deployed in the parietal lobe region, corresponding to the Cz, Pz, C3, C4, P3, and P4 sites of the international 10-20 system, to focus on collecting EEG signals related to motor imagery, such as the μ rhythm changes corresponding to the intention to raise a hand or clench a fist. Four channels are deployed in the masseter muscle region of the face to simultaneously collect electromyographic signals. On the one hand, this helps to eliminate EEG artifacts caused by masseter muscle contraction, and on the other hand, it can serve as a supplementary control modality, such as triggering emergency commands by exerting force through the masseter muscle.

[0026] The flexible wearable multimodal acquisition module has a built-in adaptive sampling control unit that can automatically adjust the sampling frequency according to the user's usage scenario: in static rehabilitation training scenarios (such as limb movement imagination training in a seated position), the sampling frequency is set to 250Hz to reduce device power consumption while ensuring signal accuracy; in dynamic virtual reality interaction scenarios (such as limb virtual movement control in a VR environment), the sampling frequency is automatically increased to 1000Hz to ensure the capture of rapidly changing EEG signal details and meet the needs of high dynamic interaction.

[0027] The dynamic adaptive noise reduction module adopts a two-layer architecture of "targeted anti-power frequency interference + global removal of physiological artifacts." Its core parameters have been optimized through extensive experiments to balance noise reduction effect and signal fidelity. This includes a variable step-size adaptive notch filter and a multi-scale residual wavelet denoising algorithm. The step-size adjustment coefficient of the variable step-size adaptive notch filter is set to 0.02. This parameter is obtained through iterative optimization using the minimum mean square error (LMS) algorithm, which improves the filter's convergence speed by 35% compared to the fixed step-size scheme. Its center frequency dynamic tracking range is set to 45-65Hz, accurately covering the 50Hz power frequency and surrounding harmonic interference in my country. Even when power grid voltage fluctuations cause interference frequency shifts, it can quickly locate the interference source and efficiently filter it. The multi-scale residual wavelet denoising algorithm uses an improved db4 wavelet basis, which has good temporal localization characteristics and is suitable for processing non-stationary EEG signals. The signal decomposition layer is set to 5 layers, which can accurately decompose the original signal to different frequency scales, achieving effective separation of key EEG bands such as delta waves, theta waves, and alpha waves from physiological artifacts. At the same time, a residual compensation coefficient of 0.15 is introduced to supplement high-frequency detail components through the reverse reconstruction process, solving the signal smoothing distortion problem that easily occurs in traditional wavelet denoising. According to actual measurements, the signal-to-noise ratio of the reconstructed signal after processing by the dynamic adaptive denoising module can reach more than 38dB, which is 28% higher than that of traditional denoising schemes, and the fidelity effect is particularly significant for weak motor imagery EEG signals.

[0028] The spatiotemporal dynamic attention fusion module, as a core innovative unit, adopts a collaborative architecture of spatial-temporal dual attention sublayers. Its specific parameters and structural design are adapted to the fusion requirements of multimodal signals. The spatial attention sublayer uses a 3×3 two-dimensional convolutional layer, combined with batch normalization and ReLU activation functions, to efficiently extract spatial feature differences from different acquisition channels (different brain regions, different modalities). By calculating the weight coefficients of features from each channel, it emphasizes the contribution of parietal motor cortex EEG signals (weight ratio can reach 60%-70%), while weakening redundant features in facial electromyography unrelated to control commands. The temporal attention sublayer uses an LSTM layer with a stride of 1 and 128 hidden layer units, accurately capturing the dynamic correlation of signals in the temporal dimension (such as the rhythmic change of alpha wave suppression before motor imagery command triggering and β wave enhancement after triggering), avoiding feature misjudgment due to temporal misalignment. The weight update cycle of the spatial-temporal dual attention sublayer is strictly set to 10ms. This cycle was determined through a large number of real-time tests. It can ensure that the weight coefficients adapt to the dynamic changes of signals and scenarios in a timely manner, without increasing the system's computing power burden due to frequent updates. Ultimately, it achieves accurate fusion and redundancy suppression of multimodal features.

[0029] The lightweight instruction recognition module uses the Light-ConvLSTM model as its core recognition architecture. This model is obtained by knowledge distillation from a large ConvLSTM teacher network. During the distillation process, a Softmax function with a temperature coefficient of 2 and a mean squared error loss function are used to ensure that the high-precision recognition experience of the teacher network is efficiently transferred to the student network. The optimized Light-ConvLSTM model has only 1.2M parameters, which is 65% less than the traditional ConvLSTM model (3.4M parameters). At the same time, depthwise separable convolution replaces the traditional convolution operation, further compressing the model size. To verify the model's performance, a dataset containing 20 common commands (such as "raise hand", "clench fist", "control wheelchair forward", "turn on light") was used for testing. The test set included multimodal signal data from 100 subjects (including 50 stroke patients and 50 healthy individuals). The measured model recognition accuracy reached 96.2%, with the recognition accuracy for motor imagery commands reaching as high as 97.5%. The model's inference latency was only 38ms, far below the human perception latency threshold (100ms), fully meeting the requirements for real-time interaction and can be directly deployed on embedded devices such as smartphones and portable EEG terminals.

[0030] The cross-device adaptation control module serves as the interaction hub between the system and external devices. It incorporates a rich communication protocol library, covering mainstream protocols such as Bluetooth 5.0, WiFi 802.11n, and RS232 serial port. It also pre-stores communication protocols and control command sets for over 30 common devices, including rehabilitation robotic arms, smart home control systems, VR headsets, and smart wheelchairs. When an external device connects, the cross-device adaptation control module automatically identifies the device type using device fingerprint recognition technology (reading the device's MAC address and protocol identifier), completing the connection with the corresponding protocol within 500ms. For example, when connecting to a rehabilitation robotic arm, it automatically switches to the RS232 serial port protocol (ensuring accurate transmission of control commands); when connecting to a VR headset, it switches to the WiFi 802.11n protocol (meeting high-bandwidth data transmission requirements); and when connecting to a smart home control system, it switches to the Bluetooth 5.0 protocol (reducing power consumption). The cross-device adaptation control module supports simultaneous connection and switching control of multiple devices, and can control up to 8 external devices simultaneously. It also features command encryption and feedback verification functions. The control commands are encrypted using the AES encryption algorithm, and the module receives execution feedback signals from external devices in real time. If a command transmission failure or execution abnormality occurs, the module automatically retransmits the command and switches to a backup communication protocol to ensure the stability and security of the control, truly achieving a plug-and-play convenient operation experience.

[0031] In practical applications, the non-invasive brain-computer interface system of this embodiment can flexibly adjust its working mode according to the rehabilitation stage of stroke patients: in the early stage of rehabilitation, it is mainly used for motor imagery training, which controls the rehabilitation robotic arm to complete auxiliary movements by accurately recognizing the patient's motor intentions and promoting neural circuit remodeling; in the later stage of rehabilitation, it can be used for daily home control, where patients can control devices such as lights, TVs, and curtains through EEG commands, which greatly improves their self-care ability.

[0032] The use of the above-mentioned non-invasive brain-computer interface system includes the following steps: Step 1: Preparation before use: Step 1: Device Wearing: Stroke patients, with the assistance of medical staff or family members, wear the flexible wearable multimodal acquisition module (the core of which is a 16-channel graphene / PDMS dry electrode array). No conductive gel is needed during wear; simply attach the electrodes to the designated locations: align the prefrontal cortex region (6 channels including Fp1 and Fp2) and the parietal cortex region (6 channels including Cz and Pz) with their respective brain regions; attach the 4 channels in the masseter muscle region to the masseter muscle of the cheek. Ensure the electrodes are in close contact with the scalp / skin. Normal use is achieved when the measured impedance is stable at approximately 3.2kΩ.

[0033] Step 2, External Device Connection: Based on usage requirements, bring the target external device (connecting to a rehabilitation robotic arm in the early stages of rehabilitation, and to a smart home control system, smart wheelchair, etc. in the later stages) close to the cross-device adapter control module. The cross-device adapter control module automatically reads the MAC address and protocol identifier of the external device using device fingerprint recognition technology, completing protocol matching and connection within 500ms (e.g., RS232 serial protocol for connecting to a rehabilitation robotic arm, Bluetooth 5.0 protocol for connecting to a smart home system). No manual configuration is required; it is plug-and-play.

[0034] Step 3, Scene Mode Selection: Select the usage scene through the supporting terminal (such as a smartphone or portable EEG terminal). You can choose from modes such as "static rehabilitation training", "dynamic virtual reality interaction" and "home control". The system will automatically match the corresponding sampling frequency (static rehabilitation is set to 250Hz, and dynamic interaction is set to 1000Hz).

[0035] Step 2: System Startup and Initialization Step 1: Start the system: Click the "Start" button on the terminal. The system will automatically activate the five modules and complete the initialization. The dynamic adaptive noise reduction module will preload preset parameters (variable step size adaptive notch filter step size adjustment coefficient 0.02, multi-scale residual wavelet denoising algorithm set to 5-level decomposition + 0.15 residual compensation coefficient). The spatiotemporal dynamic attention fusion module and the lightweight instruction recognition module (Light-ConvLSTM model) will enter the ready state simultaneously.

[0036] Step 2, Signal Calibration: The patient remains still for 30 seconds. The system simultaneously acquires EEG, EEG, and EMG signals via a flexible acquisition module to complete baseline calibration. The dynamic adaptive noise reduction module tracks the current environmental power frequency interference (45-65Hz range) in real time, initially filtering out impurity signals to ensure that the reconstructed signal-to-noise ratio after calibration is above 38dB.

[0037] Step 3: Core usage process for each rehabilitation stage: Step 1, Early Stage of Rehabilitation: Motor Imagery Training (Adapted to Rehabilitation Prosthetic Arm): Step 101, Instruction Intent Trigger: The patient performs a specified motor imagery (such as "raising hand" or "clenching fist") under the guidance of medical staff. Electrodes in the parietal lobe region will accurately collect EEG signals related to the motor imagery (such as μ rhythm changes), and electrodes in the masseter muscle region of the face will simultaneously collect EMG signals (to help eliminate artifacts, and if necessary, emergency instructions can be triggered by the force exerted by the masseter muscle).

[0038] Step 102, Signal Processing and Command Recognition: The acquired multimodal raw signals are transmitted to the dynamic adaptive denoising module via a high-speed data bus. First, a variable step-size adaptive notch filter removes power frequency interference, and then a multi-scale residual wavelet denoising algorithm removes physiological artifacts such as electrooculograms and electromyograms. Simultaneously, residual compensation repairs signal distortion. The processed signals enter the spatiotemporal dynamic attention fusion module. The spatial attention sublayer enhances the parietal motor cortex signal (with a weighting of 60%-70%), and the temporal attention sublayer captures the EEG rhythm changes triggered by motor imagery (such as alpha wave suppression and beta wave enhancement). The weights are updated every 10ms to complete feature fusion. Finally, the fused features are input into the Light-ConvLSTM model, and command recognition can be completed within 38ms (accuracy of 97.5% for motor imagery commands).

[0039] Step 103, Device Execution and Feedback: After recognition, commands such as "raise hand" and "clench fist" are transmitted to the rehabilitation robotic arm via the cross-device adaptation control module. The robotic arm precisely executes the corresponding movements to assist the patient in completing rehabilitation training. Simultaneously, the module receives execution feedback signals from the robotic arm. If a deviation occurs, it automatically retransmits the command and switches to a backup protocol to ensure stable training.

[0040] Step 2, Post-Rehabilitation: Daily Home Control (Adapting to Smart Home Systems, Smart Wheelchairs, etc.): Step 201, Mode Switching: Switch the system mode to "Home Control" via the terminal. The system will automatically adjust the parameters (the sampling frequency will be kept at 250Hz to balance power consumption and accuracy), and the cross-device adaptation control module will synchronously update the protocol library to adapt to smart home devices such as lights, TVs, and curtains.

[0041] Step 202, Command Triggering and Execution: The patient sends a preset EEG command intention (such as "want to turn on the light" corresponding to a specific motor imagery or visual evoked potential). The flexible acquisition module collects the signal, and after noise reduction, fusion, and recognition processing, the command is quickly transmitted to the smart home control center. For example, after recognizing the "turn on the light" command, the module sends a control signal via Bluetooth 5.0 protocol, and the light immediately turns on. If it is necessary to control the movement of the wheelchair, it is only necessary to switch the external device connection, and the system automatically adapts to the corresponding protocol without reconfiguration.

[0042] Step 4: Post-use instructions: Step 1: System shutdown: Click the "Shut Down" button on the terminal. The system will shut down each module in sequence. The cross-device adapter control module will automatically disconnect from external devices to ensure data security.

[0043] Step 2, Device Removal and Storage: Gently tear off the flexible wearable data collection module. No special cleaning is required; simply store it. External devices can be moved away directly, and they will automatically reconnect when brought near the module for the next use.

[0044] The second aspect of this invention discloses a non-invasive brain-computer interface method based on dynamic attention multimodal fusion. This method relies on the aforementioned system to achieve end-to-end signal processing and interactive control, specifically including the following steps: S1. Adaptive synchronous acquisition of multimodal signals: The scene recognition unit determines the user's usage scenario (such as static rehabilitation training, dynamic virtual reality interaction, outdoor mobile control, etc.) and automatically matches the corresponding sampling frequency parameters. The flexible graphene dry electrode array is activated to simultaneously collect the user's electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) physiological signals. During the collection process, a high-precision clock chip is used to achieve precise alignment of the timestamps of the three modal signals (time deviation less than 1ms) to avoid feature distortion caused by timing misalignment. The collected data is stored in local cache in real time and transmitted synchronously to the next module.

[0045] S2, Dynamic Adaptive Noise Reduction Processing: After receiving the acquired raw signal, the dynamic adaptive noise reduction module uses an interference identification unit to determine the main type of interference in the signal (power frequency interference, electrooculogram artifacts, electromyogram artifacts, or mixed interference). If the main interference is power frequency interference, then activate the variable step size adaptive notch filter to dynamically track and accurately filter out interference signals in the 45-65Hz range. If the signal contains a large number of physiological artifacts (such as artifacts caused by eye movements and muscle twitching), the multi-scale residual wavelet denoising algorithm is activated to perform multi-scale decomposition and threshold denoising on the signal. At the same time, the residual compensation mechanism is used to repair the EEG detail signals that were mistakenly filtered out. If there is mixed interference, a combination strategy of "notch filtering + wavelet denoising" is adopted to complete the noise reduction process in steps. Ultimately, a high-purity multimodal signal is output.

[0046] S3, Three-dimensional feature extraction: For the denoised EEG, EEG, and EMG signals, three-dimensional feature extraction systems in the time domain, frequency domain, and spatial domain were constructed, respectively. Time-domain features include 12 types of statistical features such as peak value, mean, variance, kurtosis, and waveform factor; Frequency domain features are extracted using Fast Fourier Transform (FFT) and wavelet packet transform, covering key parameters such as power spectral density and center frequency for delta waves (0.5-4Hz), theta waves (4-8Hz), alpha waves (8-13Hz), beta waves (13-30Hz), and gamma waves (above 30Hz). Spatial domain features are used for the extraction of multi-channel EEG signals, including correlation coefficients and phase synchronization indicators between different channels; Finally, by integrating the three-dimensional features of the three modalities, a comprehensive high-dimensional feature set is constructed, providing a rich data foundation for subsequent fusion processing.

[0047] S4. Spatiotemporal dynamic attention feature fusion: The high-dimensional feature set constructed in step S3 is input into the spatiotemporal dynamic attention fusion module: The weight coefficients of each channel and modality feature are calculated by the spatial attention sublayer. The weight of EEG features in core brain regions (such as motor cortex and visual cortex) is strengthened, while the weight of features in electrooculography and electromyography that are not related to the intention of the command is weakened. By capturing the dynamic changes of features in the temporal dimension through the time attention sublayer, we can focus on key feature segments before and after the instruction is triggered. The dual attention sublayer updates the weight coefficients every 10ms, dynamically adjusting the fusion strategy based on the current signal quality and scene, ultimately completing the dynamic fusion of features and generating a highly recognizable one-dimensional fused feature vector, effectively reducing feature redundancy.

[0048] S5, lightweight command quick recognition: The fused feature vector output from step S4 is input into a lightweight convolutional recurrent neural network. The network first extracts local correlation information of features through depthwise separable convolutional layers, then captures the temporal dependencies of features through LSTM layers, and finally outputs the instruction recognition result through fully connected layers and a softmax activation function. The inference latency of the entire recognition process is less than 50ms, which can quickly respond to the user's brain commands, and the recognition accuracy remains above 95%. It can accurately recognize various types of commands such as limb movement intentions and device control commands.

[0049] S6, Cross-device adaptive control: The lightweight command recognition module outputs the recognition results, which are then transmitted to the cross-device adaptation control module. This module automatically parses the command meaning and matches the corresponding control protocol. Based on the type of external device (e.g., rehabilitation robot, smart wheelchair), it adaptively switches between Bluetooth, WiFi, or serial communication protocols. Through its built-in protocol library, it completes rapid connection and data transmission with the external device, enabling precise control. Simultaneously, the cross-device adaptation control module... It receives feedback signals from external devices in real time. If control failure or signal loss occurs, it automatically switches the communication protocol and resends the command to ensure the stability and reliability of control, and truly achieves convenient plug-and-play interaction.

Claims

1. A non-invasive brain-computer interface system based on dynamic attention multimodal fusion, characterized in that, This includes a flexible wearable multimodal acquisition module, a dynamic adaptive noise reduction module, a spatiotemporal dynamic attention fusion module, a lightweight command recognition module, and a cross-device adaptive control module, forming a full-link optimization solution of "acquisition-noise reduction-fusion-recognition-control", wherein: The flexible wearable multimodal acquisition module adopts flexible graphene dry electrode acquisition technology to achieve low impedance signal acquisition of less than 1kΩ without the need for conductive paste. It also acquires three-modal physiological signals through a built-in high-precision synchronous acquisition chip. The three-modal physiological signals include EEG containing event-related potentials and / or motor imagery EEG, EEG containing eye movement trajectory-related signals, and EMG used to help eliminate artifacts and supplement motor intentions. The dynamic adaptive noise reduction module integrates a variable step-size adaptive notch filter and a multi-scale residual wavelet denoising algorithm to form a two-layer noise reduction architecture of "targeted anti-interference + global artifact removal", outputting high-purity multimodal signals. Among them, the variable step-size adaptive notch filter dynamically tracks power frequency interference and harmonic signals in the 45-65Hz range, accurately filtering out power grid interference; the multi-scale residual wavelet denoising algorithm performs multi-scale decomposition and threshold denoising on the signal, and at the same time repairs the wrongly filtered EEG detail signals through the residual compensation mechanism. The spatiotemporal dynamic attention fusion module adopts a neural network architecture that combines spatial and temporal attention sublayers. Specifically, the spatial attention sublayer uses a channel attention mechanism, calculating the weight coefficients of features from each acquisition channel through fully connected layers and a sigmoid activation function. This focuses on the effective features of core channels and suppresses the interference features of redundant channels. Different acquisition channels correspond to different brain regions and modalities. The temporal attention sublayer is designed based on long short-term memory units to capture the dynamic correlation features of signal temporal dimensions and accurately identify key features within different time windows. The lightweight instruction recognition module constructs a "teacher-student" network training framework based on knowledge distillation technology. It uses a high-precision deep convolutional recurrent neural network as the teacher network and a lightweight convolutional recurrent neural network as the student network. By transferring the knowledge and recognition experience of the teacher network, the performance of the student network is improved. The trained student network can be directly deployed on embedded devices, including smartphones or portable EEG terminals, without the need for additional computing power. The cross-device adaptable control module has a built-in multi-protocol communication chip and adaptive switching unit, which supports real-time adaptive switching of multiple communication protocols. It can automatically select the optimal protocol according to the type of external device, communication distance and data transmission requirements, and automatically identify the type of external device and complete the protocol matching through device fingerprint recognition technology. It can achieve precise control with plug and play without manual configuration.

2. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The flexible graphene dry electrode acquisition technology uses graphene / carbon fiber composite conductive material to prepare flexible dry electrodes for low impedance signal acquisition. The electrode surface is processed with micro-nano texture to form a self-adhesive elastic structure, which adheres tightly to the scalp through van der Waals forces.

3. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The high-precision synchronous acquisition chip supports synchronous acquisition of 16 or more multimodal signals. The sampling frequency can be adaptively adjusted within the range of 250-1000Hz according to the user's usage scenario to ensure that the signal accuracy requirements in different scenarios can be met. Specifically, in static rehabilitation training scenarios, a sampling rate of 250-500Hz is used to balance accuracy and power consumption; in high-precision scenarios, including dynamic virtual reality interaction, the sampling rate is automatically switched to 500-1000Hz.

4. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The variable step size adaptive notch filter is optimized based on the minimum mean square error algorithm. By dynamically adjusting the step size coefficient, the convergence speed of the filter is improved by more than 30%.

5. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The multi-scale residual wavelet denoising algorithm uses the db4 wavelet basis function to decompose the signal into 5-8 different scales. Based on threshold denoising, a residual compensation mechanism is introduced to supplement high-frequency detail components during reverse reconstruction.

6. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The dynamic adaptive noise reduction module has a built-in interference identification unit that automatically determines the type of interference based on signal characteristics and adaptively selects the core noise reduction algorithm combination to further improve the targeting of noise reduction. If the main interference is power frequency, then a variable step size adaptive notch filter is used to process the signal. If the signal contains a large number of physiological artifacts such as those caused by eye movements or muscle twitching, a multi-scale residual wavelet denoising algorithm is used to process the signal. If the interference is mixed, a combination strategy of variable step size adaptive notch filter and multi-scale residual wavelet denoising algorithm is adopted to complete the noise reduction process in steps.

7. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The weight coefficients of the spatial-temporal dual attention sublayer are dynamically updated every 10ms based on the signal quality of the current scene and the user's physiological state. The contribution of each modality and channel feature is adjusted in real time according to different scenarios, including static rehabilitation, outdoor movement, and noisy environments, so as to significantly improve the feature recognition in complex environments. The signal quality includes signal-to-noise ratio and interference intensity, and the user's physiological state is determined in real time by auxiliary physiological signals.

8. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The core of the student network uses depthwise separable convolution to replace the traditional convolution operation, splitting the standard convolution into depthwise convolution and pointwise convolution. At the same time, channel pruning technology is introduced to remove redundant channels and invalid parameters, further compressing the model size.

9. The non-invasive brain-computer interface system based on dynamic attention multimodal fusion as described in claim 1, characterized in that, The cross-device adaptation control module has a built-in protocol library covering mainstream devices including rehabilitation robots, smart wheelchairs, VR devices, and home control systems. The cross-device adaptation control module has command encryption and verification functions. It uses a hash algorithm to encrypt the identified commands to prevent them from being tampered with during transmission and to ensure control security.

10. A non-invasive brain-computer interface method based on dynamic attention multimodal fusion, characterized in that, The non-invasive brain-computer interface system according to claim 1 includes the following steps: S1. Adaptive synchronous acquisition of multimodal signals: Determine the user's usage scenario and automatically match the corresponding sampling frequency parameters; The flexible graphene dry electrode array is activated to simultaneously collect the user's electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) physiological signals. During the collection process, a high-precision clock chip is used to achieve precise alignment of the timestamps of the three modal signals. S2, Dynamic Adaptive Noise Reduction Processing: After receiving the acquired raw signal, the dynamic adaptive noise reduction module uses an interference identification unit to determine the main types of interference in the signal. If the main interference is power frequency interference, then activate the variable step size adaptive notch filter to dynamically track and accurately filter out interference signals in the 45-65Hz range. If the signal contains a large number of physiological artifacts, the multi-scale residual wavelet denoising algorithm is activated to perform multi-scale decomposition and threshold denoising on the signal. At the same time, the residual compensation mechanism is used to repair the EEG detail signals that were mistakenly filtered out. If the interference is mixed, a combination strategy of variable step size adaptive notch filter and multi-scale residual wavelet denoising algorithm is adopted to complete the noise reduction process in steps. Ultimately, a high-purity multimodal signal is output; S3, Three-dimensional feature extraction: For the noise-reduced EEG, EEG, and EMG signals, a three-dimensional feature extraction system in the time domain, frequency domain, and spatial domain is constructed respectively. The three-dimensional features of the three modalities are integrated to construct a comprehensive high-dimensional feature set. S4. Spatiotemporal dynamic attention feature fusion: The high-dimensional feature set constructed in step S3 is input into the spatiotemporal dynamic attention fusion module: The weight coefficients of each channel and modality feature are calculated by spatial attention sublayer, focusing on strengthening the weight of EEG features in core brain regions and weakening the weight of features in electrooculography and electromyography that are unrelated to the command intention. By capturing the dynamic changes of features in the temporal dimension through the time attention sublayer, we can focus on key feature segments before and after the instruction is triggered. The dual attention sublayer updates the weight coefficients every 10ms and dynamically adjusts the fusion strategy according to the current signal quality and scene, ultimately completing the dynamic fusion of features and generating a highly recognizable one-dimensional fusion feature vector, effectively reducing feature redundancy. S5, lightweight command quick recognition: The fused feature vector output from step S4 is input into a lightweight convolutional recurrent neural network. The network first extracts the local correlation information of the features through a depthwise separable convolutional layer, then captures the temporal dependency of the features through an LSTM layer, and finally outputs the instruction recognition result through a fully connected layer and a Softmax activation function. S6, Cross-device adaptive control: The lightweight command recognition module outputs the recognition results to the cross-device adaptation control module. The cross-device adaptation control module automatically parses the meaning of the command and matches the corresponding control protocol. It adaptively switches between Bluetooth, WiFi, or serial communication protocols according to the type of external device. Through the built-in protocol library, it completes the fast connection and data transmission with the external device, realizing precise control of the external device. At the same time, the cross-device adaptation control module receives feedback signals from the external device in real time. If control failure or signal loss occurs, it automatically switches the communication protocol and resends the command to ensure the stability and reliability of control, truly realizing plug-and-play convenient interaction.