Joint preprocessing and fusion decoding method based on electroencephalogram and functional near-infrared signals

By employing synchronous acquisition and deep feature fusion, the problem of information utilization of EEG and fNIRS signals under noise interference was solved, achieving efficient decoding and adaptive enhancement of the multimodal brain-computer interface system.

CN122046262APending Publication Date: 2026-05-15SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-04-15
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing multimodal brain signal fusion methods struggle to effectively suppress noise interference and fully utilize the complementary information between EEG and fNIRS signals, resulting in limited decoding performance and adaptability of multimodal brain-computer interface systems in complex task scenarios.

Method used

By employing synchronous acquisition, modal adaptive preprocessing, joint noise modeling, and deep feature fusion, and combining cross-modal noise feature extraction and denoising with attention mechanism for feature interaction modeling, we can achieve the collaborative utilization of multimodal signals.

Benefits of technology

It significantly improves the quality of multimodal brain signals, enhances the robustness and adaptability of the decoding system, and improves decoding accuracy and performance in complex task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122046262A_ABST
    Figure CN122046262A_ABST
Patent Text Reader

Abstract

The invention discloses a combined preprocessing and fusion decoding method based on electroencephalogram signals and functional near-infrared signals, which comprises the following steps of: synchronously acquiring the electroencephalogram signals and the functional near-infrared signals, taking a task trigger event as a unified time reference, and carrying out basic preprocessing of time alignment and modal self-adaption on two modal signals to obtain a functional near-infrared signal; a pre-trained cross-modal noise joint modeling module is utilized to perform joint modeling on cross-modal joint noise features caused by a common noise source, collaborative denoising processing is performed on multi-modal signals based on the cross-modal joint noise features, multi-modal denoising representation is obtained, a cross-modal feature interactive modeling mode based on an attention mechanism is obtained, and the multi-modal noise is obtained. And dynamically modeling the correlation between different modal features, realizing deep fusion of multi-modal complementary information, carrying out decoding processing based on the fused features, and outputting a corresponding brain-computer interface control instruction or task identification result. According to the method, the decoding accuracy and robustness of the brain-computer interface system in a complex task scene are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of brain-computer interface and multimodal brain signal processing, specifically relating to a joint preprocessing and fusion decoding method based on EEG and functional near-infrared signals. Background Technology

[0002] Brain-computer interface (BCI) technology enables direct information interaction between the human brain and external devices through the acquisition and decoding of brain signals, and has broad application prospects in fields such as neurorehabilitation, assisted control, human-computer interaction, and cognitive state assessment. Among them, electroencephalogram (EEG) signals are widely used in BCI systems due to their high temporal resolution and relatively simple acquisition equipment; while functional near-infrared spectroscopy (fNIRS) signals can reflect changes in blood oxygenation in the cerebral cortex, and have advantages such as high spatial resolution and relatively low sensitivity to environmental interference.

[0003] However, EEG signals and fNIRS signals differ significantly in their physiological generation mechanisms, time response scales, and signal-to-noise ratio characteristics: EEG signals can reflect rapid changes in neural electrical activity but are susceptible to noise interference such as eye movement artifacts and electromyography artifacts; fNIRS signals can reflect relatively stable hemodynamic changes but exhibit response delays and are easily affected by motion artifacts and physiological rhythm noise. Therefore, how to effectively preprocess, denoise, and fuse EEG and fNIRS signals while ensuring the synchronization of multimodal signals has become a key technical challenge for multimodal brain-computer interface systems.

[0004] Existing multimodal brain signal fusion methods typically employ independent preprocessing and denoising for different modalities, making it difficult to fully account for correlation interference caused by common noise sources at the multimodal signal level. Furthermore, existing methods often use fixed weights or shallow fusion strategies during the fusion stage, making it difficult to dynamically model the correlations between features of different modalities. This results in the underutilization of complementary multimodal information, thus limiting the decoding performance and adaptability of the system in complex task scenarios.

[0005] Therefore, there is an urgent need for a brain signal processing method that can effectively suppress noise interference and improve the ability to collaboratively utilize information between EEG signals and fNIRS signals in multimodal signal processing, taking into account the differences in noise sources, time response characteristics and information expression methods between EEG signals and fNIRS signals. This would improve the overall decoding performance and application adaptability of multimodal brain-computer interface systems in complex task scenarios. Summary of the Invention

[0006] The main objective of this invention is to overcome the shortcomings and deficiencies of the prior art and provide a joint preprocessing and fusion decoding method based on EEG and functional near-infrared signals. By synchronously acquiring multimodal brain signals, performing modality-adaptive preprocessing, joint noise modeling, and deep feature fusion, the method effectively extracts and synergistically utilizes task-related information from different modal brain signals, thereby improving the decoding accuracy, robustness, and adaptability to complex task states of the brain-computer interface system.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a method for joint preprocessing and fusion decoding based on electroencephalography (EEG) and functional near-infrared (FIN) signals, comprising the following steps:

[0009] The system synchronously acquires raw sequence data during task execution, including EEG signals and functional near-infrared signals, records task triggering events, and performs time alignment on the raw sequence data based on the task triggering events to obtain aligned data.

[0010] The aligned data is subjected to modal adaptive preprocessing to obtain multimodal preprocessed data;

[0011] By utilizing a pre-trained noise joint modeling module, feature extraction is performed on multimodal preprocessed data to obtain cross-modal joint noise features;

[0012] Based on cross-modal joint noise features, the multimodal preprocessed data is denoised, and the denoised multimodal preprocessed data is mapped to a unified feature representation space to obtain the multimodal denoised representation.

[0013] Cross-modal feature interaction modeling is performed on the multimodal denoising representation based on an attention mechanism to obtain fused features;

[0014] Construct a corresponding decoding model based on the task type, decode the fused features, and output brain-computer interface control commands or task recognition results.

[0015] As a preferred technical solution, the step of performing modal adaptive preprocessing on the aligned data to obtain multimodal preprocessed data specifically includes:

[0016] Perform at least one of the following processing on the EEG signal: DC drift removal, bandpass filtering, notch filtering, rereference processing, eye movement artifact suppression, and electromyography artifact suppression;

[0017] Perform at least one of the following processing steps on the functional near-infrared signal: bandpass filtering, optical density conversion, hemoglobin concentration calculation, baseline correction, and motion artifact correction;

[0018] Based on the task triggering event within a preset time range, the short-term neural response time window length for the adapted EEG signal and the delayed blood flow response time window length for the adapted near-infrared signal are respectively extracted.

[0019] As a preferred technical solution, the short-time neural response time window length and the delayed blood flow response time window length are set with reference to the same task triggering event, and the short-time neural response time window length is less than the delayed blood flow response time window length.

[0020] As a preferred technical solution, the pre-training process of the noise joint modeling module includes:

[0021] Using a multimodal dataset including various known noise labels, a joint noise modeling module is constructed. This module includes a shared feature extraction structure and performs joint feature extraction on the multimodal signals of the i-th task, representing its cross-modal joint noise feature representation. As shown in the following formula: ,

[0022] in, This indicates the joint noise modeling module. This represents the temporal window characteristics of the EEG signal under the i-th task. This represents the time window characteristics of the functional near-infrared signal under the i-th task;

[0023] A cross-modal noise consistency constraint is introduced. By minimizing the differences between noise representations of different modes, cross-modal joint noise features are obtained. As shown in the following formula: ,

[0024] in, Representational paradigm, and These are the single-modal noise feature representations extracted from the EEG signal and functional near-infrared signal under the same task i, respectively.

[0025] As a preferred technical solution, the denoising of multimodal preprocessed data based on cross-modal joint noise features specifically includes:

[0026] For the i-th task, a multimodal denoising mapping function is defined for EEG signals and functional near-infrared signals, as follows: ,

[0027] in, This represents the preprocessed input signal data for the m-mode of the i-th task. This represents a nonlinear normalization function, and the output is a noise suppression mask with the same dimension as the input signal. This means mapping the joint noise features to the suppression weight space corresponding to the modal signals. This represents the single-modal noise feature representation of EEG signals or functional near-infrared signals extracted under the same task i. Indicates the bias term. This represents the denoised signal data of the i-th task in mode m. This indicates element-wise multiplication.

[0028] As a preferred technical solution, a unified feature representation space is defined as follows: , ,

[0029] in, Represents the feature mapping function of EEG signals. This represents the feature mapping function of the near-infrared signal. It is used to map EEG signals or functional near-infrared signals to a unified feature space with consistent feature dimensions and aligned temporal structure. This represents the normalization function.

[0030] As a preferred technical solution, the step of performing cross-modal feature interaction modeling on the multimodal denoising representation based on an attention mechanism to obtain fused features includes:

[0031] The multimodal denoising representation is modeled with intramodal self-attention to obtain the self-attention output corresponding to each modality;

[0032] We utilize a cross-attention mechanism to construct dependencies between cross-modal features and obtain cross-modal fusion feature representations.

[0033] As a preferred technical solution, the method of constructing the dependency relationship of cross-modal features using the cross-attention mechanism to obtain the cross-modal fusion feature representation is as follows: ,

[0034] in, This represents the cross-modal fusion feature representation of the i-th task. This represents the cross-modal query vector for the i-th task. This represents the cross-modal key vector of the i-th task. This represents the cross-modal value vector of the i-th task. Here, is the transpose symbol, and d is the feature dimension of the key vector. This represents the normalization function.

[0035] As a preferred technical solution, the step of constructing a corresponding decoding model based on the fusion features and the task type for decoding processing, and outputting brain-computer interface control commands or task recognition results, specifically includes:

[0036] Decoding models are constructed based on different experimental paradigms or tasks to decode the fused features and output the corresponding brain-computer interface control commands or task recognition results. The decoding model is as follows: ,

[0037] in, Indicates the decoding function. This represents the cross-modal fusion feature representation of the i-th task. This represents the decoding output result corresponding to the i-th task.

[0038] As a preferred technical solution, in the discrete task recognition scenario, the decoding output is a task category label: ,

[0039] in, This represents the index function for finding the maximum value. Represents the normalization function;

[0040] In continuous state estimation or brain-computer interface control scenarios, the decoding output is a continuous control parameter or state variable.

[0041] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0042] (1) This invention can achieve precise suppression of cross-modal noise. By constructing a noise joint modeling module to extract cross-modal joint noise features and performing collaborative denoising, it fully considers the cross-modal interference caused by common noise sources in EEG and fNIRS signals. Compared with independent denoising strategies, the noise suppression effect is more significant, effectively improving the quality of multimodal brain signals.

[0043] (2) This invention achieves deep fusion of multimodal information. Based on self-attention and cross-attention mechanisms, a cross-modal feature interaction model is constructed, which can not only enhance the feature dependence within a single modality, but also dynamically model the feature correlation between different modalities, realize the adaptive weighted fusion of multimodal features, and fully explore the complementary information of EEG and fNIRS.

[0044] (3) This invention adapts to the inherent time response characteristics of signals. Targeting the characteristics of fast neural response in EEG and slow blood flow response in fNIRS, a differentiated task-related time window extraction strategy is designed, achieving accurate matching of the temporal features of the two signals, laying a high-quality data foundation for subsequent fusion and decoding.

[0045] (4) This invention improves the robustness and adaptability of the decoding system. Differentiated decoding models are constructed according to different application scenarios such as discrete task recognition and continuous state estimation. At the same time, the denoised signal is mapped to a unified feature representation space, eliminating the scale difference between modalities. This enables the system to adapt to complex brain-computer interface task scenarios and significantly improves the accuracy and robustness of decoding. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart of the joint preprocessing and fusion decoding method based on EEG and functional near-infrared signals according to an embodiment of the present invention;

[0048] Figure 2 This is a flowchart illustrating the joint noise modeling, multimodal denoising, and unified representation described in an embodiment of the present invention. Detailed Implementation

[0049] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0050] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0051] Example 1.

[0052] Please see Figure 1 This embodiment provides a joint preprocessing and fusion decoding method based on EEG and functional near-infrared signals, including the following steps:

[0053] S1. Synchronously collect raw sequence data during task execution, including EEG signals and functional near-infrared signals, record task triggering events, and perform time alignment on the raw sequence data based on the task triggering events to obtain aligned data.

[0054] The data synchronization acquisition module is used to simultaneously acquire the user's electroencephalogram (EEG) signals and functional near-infrared spectroscopy (fNIRS) signals. The EEG signals are acquired through a multi-channel EEG acquisition device, preferably with a sampling frequency of 250Hz to 1000Hz; the fNIRS signals are acquired through a functional near-infrared spectrometer, preferably with a sampling frequency of 50Hz to 100Hz.

[0055] During the data acquisition process, the task triggering events corresponding to each task during the experiment are recorded simultaneously. Let the time sequence of the task triggering events recorded during the experiment be: ,

[0056] in, This indicates the time point of the Nth task or trial during the experiment.

[0057] The synchronously acquired EEG signal and fNIRS signal are represented as follows: ,

[0058] Among them, C e With C f T represents the number of channels for the EEG signal and the fNIRS signal, respectively. e With T f These represent the number of sampling points for the EEG signal and the fNIRS signal, respectively.

[0059] S2. Perform modal adaptive preprocessing on the aligned data to obtain multimodal preprocessed data.

[0060] like Figure 2 As shown, the EEG signal is processed by DC drift removal, bandpass filtering, and notch filtering; the fNIRS signal is processed by bandpass filtering, optical density conversion, and hemoglobin concentration calculation.

[0061] Based on the task triggering event, the processed signal extracts the short-term neural response time window length of EEG and the delayed blood flow response time window length of fNIRS. The short-term neural response time window length and the delayed blood flow response time window length are set with the same task triggering event as a reference, and the short-term neural response time window length is shorter than the delayed blood flow response time window length.

[0062] Next, for the i-th task, the task-related time window feature of EEG is represented as follows: ,

[0063] in, This represents the start time of the i-th task. This indicates the length of the short-term neural response time window in EEG. This indicates the time offset of the EEG task.

[0064] The task-related time window features of fNIRS are represented as follows: ,

[0065] in, This indicates the time offset of the fNIRS task. Indicates the blood flow response delay time. This indicates the length of the blood flow response time window for fNIRS.

[0066] Simultaneously, for all tasks during the experiment, a set of multimodal task-related time window features was constructed, as shown in the following formula: ,

[0067] Where N is the total number of tasks during the experiment.

[0068] S3. Using a pre-trained cross-modal noise joint modeling module, feature extraction is performed on multimodal preprocessed data to obtain cross-modal joint noise features.

[0069] In step S3, based on a multimodal dataset including various known noise labels, the noise joint modeling module is trained to extract cross-modal joint noise features caused by common noise sources in EEG and fNIRS signals. The noise sources involved include eye movement artifacts, electromyography artifacts, motion artifacts, and physiological noise, etc.

[0070] In this embodiment, the optimization objective of the joint noise modeling module is to jointly learn the noise features in the EEG and fNIRS signals to effectively reduce cross-modal interference. Joint feature extraction is performed on the multimodal signal of the i-th task, and its cross-modal joint noise features are represented as follows: ,

[0071] in, The noise joint modeling module extracts noise features caused by a common noise source by performing feature extraction on EEG and fNIRS signals.

[0072] Simultaneously, cross-modal noise consistency constraints are introduced. This enables the joint noise modeling module to learn the relevant features in the EEG and fNIRS signals caused by a common noise source, specifically as follows: ,

[0073] in, Representational paradigm, These are the single-modal noise feature representations extracted from the EEG signal and functional near-infrared signal under the same task i, respectively.

[0074] S4. Denoise the multimodal preprocessed data based on the cross-modal joint noise features, and map the denoised features to a unified feature representation space to obtain the multimodal denoised representation.

[0075] In the joint noise modeling process, increasing the noise type and the number of samples allows for the provision of more shared noise features, thereby improving the subsequent denoising effect.

[0076] Based on the extracted cross-modal joint noise features, the preprocessed signal is collaboratively denoised, and the denoised multimodal preprocessed data is mapped to a unified feature representation space. This eliminates the scale differences between different modal signals and ensures their comparability in terms of time and feature scale.

[0077] In this embodiment, for the i-th task, a multimodal denoising mapping function is defined for the EEG and fNIRS signals, specifically expressed as: ,

[0078] in, This represents a nonlinear normalization function, and the output is a noise suppression mask with the same dimension as the input signal. It is used to weighted suppress noise components in both the time and feature dimensions. For the features of the i-th task mode m, This means mapping the joint noise features to the suppression weight space corresponding to the m-mode signal. As a bias term, it is used to adjust the baseline level of the suppression intensity. This indicates element-wise multiplication.

[0079] After denoising, the denoised multimodal signals are further mapped to a unified feature representation space. Specifically, the unified representation mapping function is defined as follows: , ,

[0080] in, and Represents the modal feature mapping function. , ∈ It is used in the i-th task to map EEG signals or functional near-infrared signals to a unified feature space with consistent feature dimensions and alignable temporal structure.

[0081] S5. Based on the attention mechanism, perform cross-modal feature interaction modeling on the multimodal denoising representation to obtain fused features.

[0082] This embodiment uses an attention mechanism to perform cross-modal interactive modeling of the denoised representations of EEG and fNIRS signals. Through self-attention and cross-attention mechanisms, the correlation between EEG features and fNIRS features is dynamically modeled, ultimately generating a fused feature representation.

[0083] In this embodiment, intramodal self-attention modeling is performed on the unified feature representations of EEG and fNIRS, and the self-attention calculation process is expressed as follows: , , ,

[0084] in, , , For a trainable parameter matrix, , , These represent the query vector, key vector, and value vector, respectively. Let represent the feature matrix of mode m after denoising and mapping to the unified feature representation space under the i-th task.

[0085] The corresponding self-attention output is represented as: ,

[0086] After intramodal feature enhancement, a cross-attention mechanism is introduced to model the dependencies between features from different modalities. Taking querying fNIRS features using EEG features as an example, the cross-attention calculation process is as follows: , , ,

[0087] in, This represents the trainable parameter matrix that maps EEG modal features to cross-modal query vectors. This represents the trainable parameter matrix that maps fNIRS modal features to cross-modal key vectors. This represents the trainable parameter matrix that maps fNIRS modal features to cross-modal value vectors.

[0088] The corresponding cross-attention output is: ,

[0089] in, This represents the cross-modal fusion feature representation that integrates EEG and fNIRS signals in the i-th task. Represents the normalization function. This represents the cross-modal query vector for the i-th task. This represents the cross-modal key vector of the i-th task. This represents the cross-modal value vector of the i-th task. Here, is the transpose symbol, and d is the feature dimension of the key vector. This represents the normalization function.

[0090] By dynamically calculating the attention weights mentioned above, adaptive weighting of different modal features during the fusion process is achieved.

[0091] S6. Construct a corresponding decoding model based on the task type, decode the fused features, and output brain-computer interface control commands or task recognition results.

[0092] The system outputs brain-computer interface control commands or task recognition results based on different task types. The decoding process constructs a task recognition model based on the fused representation, which can output brain-computer interface control commands or task category labels.

[0093] In this embodiment, the cross-modal fusion feature representation obtained in step S5 is used. Based on different experimental paradigms or tasks, a decoding model is constructed to output corresponding brain-computer interface control commands or task recognition results. The decoding mapping function is defined as follows: ,

[0094] in, This represents the decoding function, which can be implemented using a fully connected neural network, a time-series classification model, or a regression model. This represents the decoding output result corresponding to the i-th task.

[0095] In discrete task recognition scenarios, the decoding output is a task category label: ,

[0096] in, This represents the index function for finding the maximum value.

[0097] In continuous state estimation or brain-computer interface control scenarios, the decoding output is a continuous control parameter or state variable used to drive external devices or human-computer interaction systems.

[0098] Example 2.

[0099] In some more specific embodiments, Example 2 provides an application of joint preprocessing and fusion decoding based on EEG and functional near-infrared signals in a nine-class SSVEP experiment.

[0100] This embodiment uses multimodal signal fusion decoding based on the nine-class SSVEP experiment as a specific task scenario to illustrate the practical application of the joint preprocessing and fusion decoding method of EEG and functional near-infrared signals proposed in this invention.

[0101] Application scenario description:

[0102] In this embodiment, the combined preprocessing and fusion decoding method proposed in this invention is applied to the collected EEG and fNIRS signals from the nine-category SSVEP experiment. The goal of the experiment is to use EEG and fNIRS signals to jointly identify nine different frequencies of visual stimuli and use the identification results to control external devices, such as robotic arms, drones, wheelchairs, or rehabilitation training robots.

[0103] This experiment aims to verify the performance of multimodal brain signal fusion in complex task environments, especially how to improve the decoding accuracy and robustness of brain-computer interface systems under conditions of multiple noise sources and signal instability. The experiment adopted the SSVEP experimental paradigm, where each stimulus corresponds to a specific frequency, and subjects need to complete the task under different frequency light spot flashes. This embodiment improves noise reduction and decoding capabilities by applying the joint preprocessing and fusion decoding method proposed in this invention.

[0104] The specific steps are as follows:

[0105] S1. Data Acquisition and Task Paradigm Design.

[0106] In this embodiment, the international 10-20 system is used to arrange the EEG acquisition electrodes, focusing on acquiring the occipital lobe and occipitoparietal junction channels, such as O1, Oz, O2, PO3, and PO4, at a acquisition frequency of 1000Hz. The functional near-infrared signal is acquired using a 16-channel fNIRS system, focusing on the occipital lobe, occipitoparietal junction, and frontal lobe channels, at a sampling frequency of 11Hz.

[0107] The SSVEP task paradigm is designed as follows:

[0108] The experiment required participants to focus on nine flashing squares at different frequencies (8.0Hz, 8.5Hz, 9.0Hz, 11.0Hz, 13.0Hz, 15.0Hz, 17.0Hz, 18.0Hz, and 19.0Hz) on a screen according to the experimental procedure. Each stimulus lasted for 5 seconds, with a 3-second interval between tasks. The system simultaneously recorded data from two modalities: electroencephalogram (EEG) signals and functional near-infrared spectroscopy (FIR) signals. The system ensured signal alignment based on the task trigger time, and the collected data was used for subsequent preprocessing and feature extraction.

[0109] S2. Construction of the joint noise modeling model.

[0110] In this embodiment, a multimodal training dataset containing known labels is used. Preferably, the noise labels include eye-tracking artifacts, electromyography artifacts, motion artifacts, etc., to train a joint noise modeling module. Specifically, this joint noise modeling module employs a shared encoder.

[0111] The goal of the joint noise modeling is to jointly learn the noise features in EEG and fNIRS signals to reduce cross-modal interference. By learning the joint noise features across modes, artifacts in EEG and fNIRS signals can be eliminated in a coordinated manner, thereby improving signal quality efficiently and effectively.

[0112] S3, Multimodal signal collaborative denoising and fusion decoding.

[0113] In this embodiment, the acquired data is first subjected to modal adaptive preprocessing. The preprocessing of EEG data includes bandpass filtering, notch filtering, and other operations; the preprocessing of fNIRS data includes bandpass filtering, Beer-Lambert law transformation, and other operations.

[0114] Noise features obtained through the joint noise modeling module are used to denoise the EEG and fNIRS signals. The denoised signals are then mapped to a unified feature representation space to eliminate scale differences between different modal signals and ensure their consistency across time and feature scales.

[0115] The cross-modal feature interaction module uses an attention mechanism to interactively model the denoised EEG and fNIRS features, learns the correlation between the two modal features, and the system can dynamically adjust the attention weights to adaptively weight different modal signals during the fusion process.

[0116] Decoding is performed based on fused features to output the corresponding task category label or control command.

[0117] In summary, this embodiment verifies the significant advantages of the joint preprocessing and fusion decoding method based on EEG and functional near-infrared spectroscopy (fNIRS) signals in the nine-class SSVEP experiment. The joint noise modeling module effectively removes common noise sources from both EEG and fNIRS signals, significantly improving signal quality. Especially under conditions of high noise, the system's decoding performance remains high, enhancing the adaptability and usability of the brain-computer interface system under different task states.

[0118] It should be noted that, for the sake of simplicity, the aforementioned method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously.

[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals, characterized in that, Includes the following steps: The system synchronously acquires raw sequence data during task execution, including EEG signals and functional near-infrared signals, records task triggering events, and performs time alignment on the raw sequence data based on the task triggering events to obtain aligned data. The aligned data is subjected to modal adaptive preprocessing to obtain multimodal preprocessed data; By utilizing a pre-trained noise joint modeling module, feature extraction is performed on multimodal preprocessed data to obtain cross-modal joint noise features; Based on cross-modal joint noise features, the multimodal preprocessed data is denoised, and the denoised multimodal preprocessed data is mapped to a unified feature representation space to obtain the multimodal denoised representation. Cross-modal feature interaction modeling is performed on the multimodal denoising representation based on an attention mechanism to obtain fused features; Construct a corresponding decoding model based on the task type, decode the fused features, and output brain-computer interface control commands or task recognition results.

2. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, The step of performing modal adaptive preprocessing on the aligned data to obtain multimodal preprocessed data specifically involves: Perform at least one of the following processing on the EEG signal: DC drift removal, bandpass filtering, notch filtering, rereference processing, eye movement artifact suppression, and electromyography artifact suppression; Perform at least one of the following processing steps on the functional near-infrared signal: bandpass filtering, optical density conversion, hemoglobin concentration calculation, baseline correction, and motion artifact correction; Based on the task triggering event within a preset time range, the short-term neural response time window length for the adapted EEG signal and the delayed blood flow response time window length for the adapted near-infrared signal are respectively extracted.

3. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 2, characterized in that, The short-term neural response time window length and the delayed blood flow response time window length are set with reference to the same task triggering event, and the short-term neural response time window length is shorter than the delayed blood flow response time window length.

4. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, The pre-training process of the noise joint modeling module includes: Using a multimodal dataset including various known noise labels, a joint noise modeling module is constructed. This module includes a shared feature extraction structure and performs joint feature extraction on the multimodal signals of the i-th task. The cross-modal joint noise features are represented as follows: , in, This indicates the joint noise modeling module. This represents the temporal window characteristics of the EEG signal under the i-th task. This represents the time window characteristics of the functional near-infrared signal under the i-th task; A cross-modal noise consistency constraint is introduced. By minimizing the differences between noise representations of different modes, cross-modal joint noise features are obtained. As shown in the following formula: , in, Representational paradigm, and These are the single-modal noise feature representations extracted from the EEG signal and functional near-infrared signal under the same task i, respectively.

5. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, The denoising process for multimodal preprocessed data based on cross-modal joint noise features is as follows: For the i-th task, a multimodal denoising mapping function is defined for EEG signals and functional near-infrared signals, as follows: , in, This represents the preprocessing of the input signal data in mode m for the i-th task. This represents a nonlinear normalization function, and the output is a noise suppression mask with the same dimension as the input signal. This means mapping the joint noise features to the suppression weight space corresponding to the modal signals. Indicates the bias term. This represents the denoised signal data of the i-th task in mode m. This indicates element-wise multiplication.

6. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, Define a unified feature representation space as follows: , , in, Represents the feature mapping function of EEG signals. This represents the feature mapping function of the near-infrared signal. ∈ It is used in the i-th task to map EEG signals or functional near-infrared signals to a unified feature space with consistent feature dimensions and aligned temporal structure.

7. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, The step of performing cross-modal feature interaction modeling on the multimodal denoising representation based on an attention mechanism to obtain fused features includes: The multimodal denoising representation is modeled with intramodal self-attention to obtain the self-attention output corresponding to each modality; We utilize a cross-attention mechanism to construct dependencies between cross-modal features and obtain cross-modal fusion feature representations.

8. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 7, characterized in that, The cross-attention mechanism is used to construct the dependency relationship of cross-modal features and obtain the cross-modal fusion feature representation, as shown in the following formula: , in, This represents the cross-modal fusion feature representation of the i-th task. This represents the cross-modal query vector for the i-th task. This represents the cross-modal key vector of the i-th task. This represents the cross-modal value vector of the i-th task. Here, is the transpose symbol, and d is the feature dimension of the key vector. This represents the normalization function.

9. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 1, characterized in that, The process of constructing a corresponding decoding model based on the task type, decoding the fused features, and outputting brain-computer interface control commands or task recognition results specifically involves: Decoding models are constructed based on different experimental paradigms or tasks to decode the fused features and output the corresponding brain-computer interface control commands or task recognition results. The decoding model is as follows: , in, Indicates the decoding function. This represents the cross-modal fusion feature representation of the i-th task. This represents the decoding output result corresponding to the i-th task.

10. The method for joint preprocessing and fusion decoding based on EEG and functional near-infrared signals according to claim 9, characterized in that, In discrete task recognition scenarios, the decoding output is a task category label: , in, This represents the index function for finding the maximum value. Represents the normalization function; In continuous state estimation or brain-computer interface control scenarios, the decoding output is a continuous control parameter or state variable.