Multi-modal electroencephalogram analysis model construction method, online processing method and system
By constructing a multimodal EEG analysis model, the feature extraction and fusion is performed using time convolution network, multi-view embedding layer and adaptive Transformer mechanism, and the features are encoded into a shared potential space to complete the missing data, solving the problems of multimodal data loss and incompleteness, and improving the robustness and accuracy of the system.
Patent Information
- Application Number
- CN202411908907.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
Multimodal BCI-VR systems have difficulty maintaining reliability and accuracy in the absence and incomplete multimodal data, and existing completion methods have failed to fully utilize the correlation between different modal data.
The multimodal EEG analysis model construction method is adopted to collect and preprocess multimodal signals, and feature extraction and fusion is performed using time convolution network, multi-view embedding layer and adaptive Transformer mechanism, and the features are encoded into the shared potential space to complete the missing data. The combined loss function is used to optimize the quality of the generated features, and the multi-scale masking strategy is used to enhance the robustness of the model.
The robustness, real-time and universality of the multimodal BCI-VR system in multimodal data processing is improved, and the performance of the system in applications such as emotion recognition, cognitive load monitoring and neurorehabilitation is enhanced, ensuring stability and accuracy in the absence of modalities or incomplete data.
Smart Images

Figure CN120046090A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a multimodal electroencephalogram analysis model, and also relates to a corresponding online processing method and an online processing system for multimodal electroencephalogram analysis, belonging to the technical field of brain-computer interfaces. Background Art
[0002] Brain-computer interface (BCI) technology originated in the 1970s, and its core goal is to establish a direct communication bridge between the human brain and a computer or other electronic devices. This technology converts brain activities into instructions that a computer can understand by collecting and analyzing biological signals such as electroencephalogram (EEG), realizing a non-muscular "thought-to-action" communication and control method. With the rapid development of virtual reality (VR) technology, the combination of BCI and VR has opened up new application fields. Especially in the integration of multimodal BCI and VR technologies, it provides new ideas for solving multimodal data fusion and real-time state recognition problems. A multimodal BCI system can collect multiple biological signals including EEG, electrodermal activity, and eye movement, and these signals reflect the physiological state of the user. In an immersive VR environment, the system can process and feedback the user's physiological signals in real time, thereby enhancing the user's virtual experience and interaction effect.
[0003] The multimodal BCI-VR system not only plays a role in enhancing the user experience, but also shows broad application potential in many fields such as assisted therapy, psychological intervention, education, and entertainment. By integrating multiple biological signals, the system can overcome the limitations of a single signal and improve the accuracy of recognition and the robustness of the system. However, in practical applications, multimodal data collection may be affected by factors such as environmental interference, poor sensor contact, or user movement, resulting in the loss or quality degradation of modal signals, which poses challenges to the reliability and accuracy of the system.
[0004] To solve the problems of multimodal data loss and incomplete data, researchers have proposed various methods to complement incomplete modal data. The simplest method is to directly discard the unavailable data part, but this method will cause information loss and reduce the reliability of the data. In contrast, the data filling method can more effectively improve the effectiveness of the data and reduce information loss. For example, Ngai et al. used linear interpolation to complement the noise data generated by blinking and rapid eye movement, but this method did not fully utilize the correlation between different modal data. Therefore, in order to maintain the reliability and accuracy of the multimodal BCI-VR system in the case of multimodal data loss and incompleteness, further research and technological innovation are needed to fully utilize the correlation between different modal data and improve the comprehensive performance of the system. Summary of the Invention
[0005] The primary technical problem to be solved by the present invention is to provide a method for constructing a multi-modal electroencephalogram analysis model.
[0006] Another technical problem to be solved by the present invention is to provide a method for online processing of multi-modal electroencephalogram analysis.
[0007] Another technical problem to be solved by the present invention is to provide an online processing system for multi-modal electroencephalogram analysis.
[0008] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0009] According to the first aspect of the embodiments of the present invention, a method for constructing a multi-modal electroencephalogram analysis model is provided, including the following steps:
[0010] S1: Collect multi-modal signals and preprocess the signals of each modality separately;
[0011] S2: Slice the preprocessed data according to a fixed time window to ensure the temporal consistency of the signals of all modalities;
[0012] S3: Use a temporal convolutional network to extract features from the sliced data to obtain features of each modality;
[0013] S4: Use a multi-view embedding layer to project the features of each modality into multiple subspaces to enrich the feature representation;
[0014] S5: Perform positional embedding and modality embedding, introduce a multi-head self-attention mechanism to enrich the feature representation, and introduce multiple modality experts based on the adaptive Transformer mechanism to adaptively select experts according to the input modality to capture specific information of each modality;
[0015] S6: Use a cross-modal attention mechanism for multi-modal fusion. After obtaining the fusion features of all modalities, aggregate the fusion features again;
[0016] S7: Encode the multi-modal features into a shared latent space to complete the features and map them to a unified discrete space in the case of modality or data loss;
[0017] S8: In the shared latent space, the decoder fuses the multi-modal features through the quantized latent representation and outputs the reconstructed original multi-modal feature representation; among them, a joint loss function is used to improve the quality of the generated features.
[0018] Preferably, the joint loss function is a weighted sum of a reconstruction loss, a classification loss, a quantization loss, a perceptual loss, and a multi-scale loss.
[0019] Preferably, the joint loss function is:
[0020]
[0021] Among them, the dynamic weight formula is adopted to calculate the weights λ corresponding to each loss 1 to λ 5 , where Indicator i represents the importance index corresponding to each loss.
[0022] Preferably, after the step S8, there is further a step S9: in the incomplete modality, the tuning adopts a multi-scale masking strategy to enhance the robustness of the model in the case of modality loss and local feature loss.
[0023] Preferably, the tuning in the incomplete modality includes the following sub-steps:
[0024] First, in the full-modal level mask, some modalities are randomly and completely masked to simulate the situation where some sensors are completely unavailable in actual applications;
[0025] Secondly, the time-scale mask randomly masks different lengths of time steps in the single-modal data to simulate the loss of some time segments;
[0026] Finally, the multi-modal local mask randomly masks the frequency bands or feature dimensions of a specific modality, so that the model depends on other features when a specific frequency band or feature is missing.
[0027] Preferably, during the tuning process, the mask ratio and the dropout ratio are gradually increased as the training process progresses.
[0028] Preferably, when the performance of the model tends to be stable at a certain mask ratio, the ratio increase is stopped.
[0029] Preferably, according to the input modality, the modality experts are adaptively selected to capture the specific information of the corresponding modality.
[0030] According to the second aspect of the embodiments of the present invention, there is provided a multi-modal electroencephalogram analysis online processing method, which is implemented based on the model obtained by the foregoing multi-modal electroencephalogram analysis model construction method, and includes the following steps:
[0031] S101: Using a multi-modal signal acquisition device to collect at least two modal signals of the electroencephalogram, electrocardiogram, skin electrical activity, and eye movement of the subject;
[0032] S102: Preprocessing the multi-modal signals;
[0033] S103: Extracting features from the multi-modal signals;
[0034] S104: Determine whether there is a situation of incomplete modal data. If so, proceed to step S105; if not, proceed to step S106;
[0035] S105: Perform feature fusion on the multi-modal signals, complete feature complementation in the incomplete modality, and then output the classification result;
[0036] S106: Perform feature fusion on the multi-modal signals and directly output the classification result.
[0037] According to the third aspect of the embodiments of the present invention, there is provided an online processing system for multi-modal electroencephalogram analysis, including a multi-modal signal acquisition device, a VR device, and an edge computing module. The edge computing module includes a processor and a memory, and the processor is coupled to the memory. Among them, the memory is used to store computer programs; the processor is used to run the computer programs stored in the memory and execute the above-mentioned multi-modal electroencephalogram analysis online processing method.
[0038] Compared with the prior art, the present invention improves the robustness, real-time performance, and universality of the brain-computer interface technology in multi-modal data processing by constructing a multi-modal electroencephalogram analysis model, its online processing method, and system. The present invention collects and preprocesses multi-modal biological signals, uses a temporal convolutional network and a multi-view embedding layer for feature extraction and rich expression, combines an adaptive Transformer mechanism and a cross-modal attention mechanism to achieve feature fusion, and encodes the features into a shared latent space to complement missing data. In addition, the quality of the generated features is optimized through a joint loss function, and a multi-scale masking strategy is adopted to enhance the robustness of the model in the case of modal loss, so as to provide more intelligent and convenient multi-modal data acquisition and processing in applications such as emotion recognition, cognitive load monitoring, and neurorehabilitation, and enhance the practicability and stability of the BC I-VR system. Description of the Drawings
[0039] Figure 1 It is a schematic diagram of the correspondence between the online processing method for multi-modal electroencephalogram analysis and the online processing system for multi-modal electroencephalogram analysis;
[0040] Figure 2 It is a schematic diagram of the general process of the method for constructing a multi-modal electroencephalogram analysis model in the first embodiment of the present invention;
[0041] Figure 3 It is a schematic diagram of the detailed process of the method for constructing a multi-modal electroencephalogram analysis model provided by the embodiments of the present invention;
[0042] Figure 4 It is a schematic diagram of the process of the online processing method for multi-modal electroencephalogram analysis in the second embodiment of the present invention. Detailed Embodiments
[0043] The technical content of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0044] The technical concept of the embodiments of the present invention lies in: by comprehensively collecting various physiological signals such as electroencephalogram (EEG), electrocardiogram (ECG), galvanic skin response (GSR), and eye movement, and using a temporal convolutional network, a multi-view embedding layer, position and modality embedding, and an adaptive Transformer mechanism, feature extraction and deep fusion of these multi-modal signals are achieved. The corresponding multi-modal EEG analysis model (abbreviated as the model) further aggregates features through a cross-modal attention mechanism and uses the VQ-VAE technology to encode multi-modal features into a shared latent space for feature completion when modalities or data are missing. In addition, the multi-modal EEG analysis model adopts a joint loss function, including reconstruction loss, classification loss, quantization loss, perceptual loss, and multi-scale loss, to improve the quality of the generated features, and combines a multi-scale masking strategy to enhance the robustness of the model in the face of modality and local feature losses. The present invention can be implemented on an edge computing module, reducing latency and improving the real-time performance and privacy security of data processing.
[0045] The First Embodiment
[0046] As Figure 1 shown, the first embodiment of the present invention provides a method for constructing a multi-modal EEG analysis model. This method uses an existing multi-modal signal acquisition device to collect signals. After alignment and transmission processing, these signals will be used for a multi-modal EEG analysis model, and then control a virtual display device to achieve virtual reality interaction. Specifically, the multi-modal signal acquisition device integrates a variety of sensors in the head area of the user, including electroencephalogram (EEG), electrocardiogram (ECG), galvanic skin response (GSR), and eye movement (EYE), and can collect rich multi-modal physiological signals. These signals provide multi-dimensional data support for user emotion recognition, attention monitoring, and stress management. The multi-modal EEG analysis model is a computational model obtained based on the method for constructing a multi-modal EEG analysis model provided by the embodiments of the present invention.
[0047] The method for constructing a multi-modal EEG analysis model provided by the embodiments of the present invention is divided into two main stages: pre-training in the complete modality and tuning in the incomplete modality. As Figure 2 and Figure 3As shown, the pre-training stage is carried out in the complete modality, including steps such as feature extraction, feature fusion, training of VQ-VAE (Vector Quantization Variational Autoencoder) in the complete modality, and reconstruction output. The purpose of this stage is to build a model that can operate efficiently when multi-modal data is complete, laying a foundation for subsequent tuning in the incomplete modality. Through these steps, the model can learn how to extract and fuse features from complete multi-modal signals, and perform effective feature completion and reconstruction when modal data is missing, thereby improving the robustness and adaptability of the model in practical applications. Next, the specific steps are described in detail:
[0048] S1: Collect multi-modal signals and preprocess the signals of each modality separately.
[0049] Preprocess the original EEG, GSR, ECG, and eye movement (EYE) data, so that the data of each modality undergoes different standardization and filtering processes to remove noise and enhance signal quality. Specifically, for EEG data, band-pass filtering is used to remove frequency components of non-EEG signals, such as power frequency interference, etc., and at the same time, independent component analysis (ICA) is used to remove eye movement and body movement artifacts, so as to more accurately capture EEG activities and reduce interference from non-EEG signals; for GSR and ECG data, low-pass filtering is used for smoothing processing, which can reduce high-frequency noise, such as electromyogram interference, etc., making the signals smoother and easier to analyze; for eye movement data, through standardization and denoising processing, the noise caused by head movement, etc., can be reduced, improving the accuracy and reliability of eye movement data.
[0050] S2: Slice the preprocessed data according to a fixed time window to ensure the temporal consistency of the signals of all modalities.
[0051] In an embodiment of the present invention, the slicing process includes the following sub-steps:
[0052] S21: Determine the time window size. Determine the same length of the time window for the data of all modalities. For example, each time window can be set to a 2-second window with a step size of 1 second.
[0053] S22: Align the timestamps of the data of all modalities. If the timestamps of the data are not precisely aligned, interpolation or rounding to the closest time window boundary may be required.
[0054] S23: Create an index array to identify which time window each data point belongs to. This can be achieved by calculating the difference between the timestamp of each data point and the start time of the time window.
[0055] S24: Slice and output the sliced data to ensure the temporal consistency of all modalities and provide support for subsequent feature extraction.
[0056] During the sharding process, according to the index array, the data is split into data segments for multiple time windows. For each time window, all data points within that window are extracted. Then, the data within each time window is output as a separate segment for subsequent processing or analysis. For data that does not completely fill a time window, options such as padding (e.g., padding with zeros, copying the last valid data point, etc.) or truncation can be selected.
[0057] S3: Use a temporal convolutional network to extract features from the sharded data to obtain electroencephalogram features, electrocardiogram features, galvanic skin response features, and eye movement features (collectively referred to as multi-modal features).
[0058] In the feature extraction stage, a temporal convolutional network (TCN) is used to process and analyze time series data of different modalities (sharded data of each modality, such as EEG, EYE, ECG, GSR shards). The TCN extracts spatio-temporal features of each modality from the multi-modal data and aligns these features in time to generate a feature representation with spatio-temporal information, thus providing support for subsequent fusion and analysis steps. Since the TCN can capture long-term dependencies in time series data, it enhances the model's ability to capture time dynamics.
[0059] S4: Use a multi-view embedding layer to project the multi-modal features into multiple subspaces to enrich the feature representation.
[0060] The multi-view embedding layer performs a linear transformation on the data of each modality to extract features and adjust the data dimensions to unify the representation of different modalities. Nonlinearity is introduced using activation functions (such as sigmoid, ReLU, etc.) so that the model can learn more complex features, and batch normalization is used to accelerate the training process and improve the generalization ability of the model. The multi-view embedding layer maps the features of different modalities into a shared low-dimensional space or multiple subspaces to capture and utilize complementary information between different modalities, which can enrich the feature representation and enable the model to better understand and process features from different modalities.
[0061] S5: Perform positional embedding and modality embedding, introduce the multi-head self-attention mechanism to enrich the feature representation, and introduce multiple modality experts based on the adaptive Transformer mechanism to adaptively select experts according to the input modality to capture specific information of each modality.
[0062] such as Figure 3As shown, after batch normalization, positional embedding (Pos_emb) and modality embedding (Mod_emb) are performed. Positional embedding marks the time steps in the time series data, ensuring that the model retains time information in the temporal encoding. Modality type embedding is used to mark the modality source of the features. Therefore, positional embedding provides the position information of each element in the sequence, which helps the model understand the order of elements in the sequence; modality embedding is used to distinguish data of different modalities, which can help the model distinguish the feature sources of different modalities when modalities are missing and utilize the complementary relationships between modalities for corresponding supplementation. For example, when only EEG and eye movement signals enter the model, according to the modality embedding, the corresponding attention expert models can be selected to extract features from the EEG and eye movement signals respectively, and their corresponding EEG CLS and EYE CLS are obtained. Then, through the cross-modal attention mechanism, the cls of EEG and EYE are fused across modalities to obtain the corresponding complementary information.
[0063] Through positional embedding and modality embedding, the model has obtained information about the position and source modality of each element in the sequence. Introducing the multi-head self-attention mechanism can further enhance these features by learning the correlations between different time points or different modalities to enrich the feature representation. The multi-head self-attention mechanism allows the model to process features of different modalities in parallel, improving the computational efficiency of the model and facilitating real-time detection.
[0064] In an embodiment of the present invention, for the four modalities of electroencephalogram (EEG), electrocardiogram (ECG), galvanic skin response (GSR), and eye movement (EYE), a method based on the adaptive Transformer architecture is adopted, and four modality experts are particularly introduced to replace the traditional standard feed-forward network (FFN). These four modality experts are EEG-FFN, EYE-FFN, ECG-FFN, and GSR-FFN respectively, which can adaptively select the most suitable expert for feature processing according to the input modality to capture the specific information of each modality and enhance the model's analysis ability for multi-modal data. For example, if the input is only EEG, the expert of EEG-FFN is used to encode the features. If the input contains multiple modalities, multiple experts are used to process the features of the corresponding modalities. This design refers to the Figure 3 model architecture shown in, aiming to improve the processing efficiency and accuracy of the model for different modality features.
[0065] Because the modal embeddings provide information about which modality the input data belongs to. This information can be utilized when adaptively selecting experts to determine which modality expert is most suitable for processing the current modal features. The positional embeddings help the model understand the temporal or spatial relationships of different modality data, thereby guiding the modality experts to better process these features. Moreover, since the interactions between different modalities are crucial for understanding complex phenomena, the positional embeddings help the model capture these interactions, especially for modalities that are closely related temporally or spatially. This combined embedding of position and modality enables the model to more accurately select the most appropriate multiple modality experts for processing, enabling the modality experts to work more effectively together, improving adaptability and cross-modal feature supplementation, and facilitating cross-modal feature fusion.
[0066] S6: Perform multimodal fusion using a cross-modal attention mechanism. After obtaining the fusion features of all modalities, these fusion features are aggregated again to obtain.
[0067] Use a dynamic cross-modal attention mechanism to identify and strengthen the interaction features between different modalities. By dynamically adjusting the importance between different modalities, the model can more effectively fuse data from different modalities. In this embodiment, the interaction features between each modality and other modalities are calculated respectively through the cross-modal attention mechanism, and the features of the guiding modality and the follower modality are merged together by weighted summation to obtain F i , and then the fusion features F of all modalities are obtained i After that, these fusion features are aggregated again to obtain F fusion .
[0068] Through the cross-modal attention mechanism, the model can capture the correlations between different modalities, thereby obtaining rich fusion features. Aggregating these features again can further extract and strengthen the useful information in these features, improving the model's ability to express key information. During the aggregation process, the aggregated features can further promote the complementarity between different modalities, enabling the model to learn the complementary relationships between different modality features, and thus better utilize these complementarities to improve performance.
[0069] S7: Encode the multimodal features into a shared latent space to complete the features and map them to a unified discrete space in case of modality or data loss.
[0070] In an embodiment of the present invention, VQ-VAE (Vector Quantized Variational Autoencoder) is used to train in the complete modality, and the output of the encoder is mapped to the nearest neighbor in a set of discrete "codebooks", thereby converting the high-dimensional continuous encoding into a low-dimensional discrete encoding, and encoding the multimodal features into a shared latent space. This can achieve feature completion in case of modality or data loss.
[0071] The shared latent space is a continuous latent space where high-dimensional features of all modalities are mapped to a low-dimensional space through convolutional encoders. This space allows data representations of different modalities to be compared and fused in the same dimension, thus enabling information sharing and feature complementation. In VQ-VAE, the continuous latent representation is mapped to a discrete codebook through quantization operations. This discrete space is composed of elements in the codebook, and each element represents a point in the latent space.
[0072] Specifically, first, the high-dimensional features of each modality are projected to a low-dimensional space using convolutional encoders to capture local features of the data and reduce computational complexity. The output of the convolutional encoder is processed through non-linear activation and batch normalization to ensure stable training. Then, the latent representations of each modality are mapped to a unified shared latent space through a shared projection matrix, facilitating information sharing and feature complementation. Subsequently, the quantization operation compares the continuous latent vectors output by the encoder with the vectors in the codebook (each entry e 1 、e 2 ......e n in the discrete codebook is a vector representing a specific point in the latent space) and selects the closest codebook vector as the quantized representation. In this way, by mapping the continuous latent representation to the discrete codebook and using the shared discrete codebook to ensure that the fused features of all modalities are mapped to a unified discrete space, VQ-VAE can learn the shared representation between modalities.
[0073] As Figure 3 shown, in the encoder-decoder structure of the Transformer model, Z represents the output of the encoder, that is, the feature representation after being processed by the encoder. and represent the shared feature representations, which are used for the encoder (e) and decoder (q) respectively. V2L represents a learnable transformation or mapping that converts the output Z of the encoder into a form suitable for the input of the decoder. In multimodal learning, the shared feature representation means that data of different modalities are mapped to a common feature space, which helps the model learn cross-modal associations.
[0074] S8: In the shared latent space, the decoder fuses multimodal features through the quantized latent representation and outputs the reconstructed original multimodal feature representation; among them, a joint loss function is used to improve the quality of the generated features.
[0075] In the shared latent space, the decoder outputs the reconstructed original multi-modal feature representation based on the multi-modal fusion features of the quantized latent representation. These features can be used for subsequent classification tasks. During the pre-training process, the reconstruction loss is calculated to ensure the high quality of the generated features, and the classification loss is calculated to ensure the quality of feature extraction and fusion. In addition, to improve the quality of the generated features, a joint loss function is adopted, through quantization loss, perceptual loss, and multi-scale discriminator, to evaluate the quality of the generated features from multiple perspectives at the high-level semantics and multi-scale resolutions.
[0076] In one embodiment of the present invention, the joint loss function is the weighted sum of the following five losses: 1) Reconstruction loss 2) Classification loss 3) Quantization loss 4) Perceptual loss 5) Multi-scale loss
[0077] The following introduces each loss function:
[0078] 1) Reconstruction loss
[0079] Among them, the fusion feature X fusion , X reconstructed represents the data reconstructed from the fusion feature X fusion . The loss function measures the difference between the reconstructed data and the original fusion feature, and usually uses the Euclidean distance (L2 norm) to calculate this difference.
[0080] 2) Classification loss
[0081] For the classification problem, the cross-entropy loss is used where N represents the number of samples, y i represents the true label of the i-th sample, the predicted probability of the i-th sample.
[0082] Adopt a classification enhancement mechanism to combine the latent representation z fusion of the fusion feature X fusion with the output of the classification head to strengthen the classification-related features: Among them, γ is a hyperparameter used to balance the weights between the two loss terms; represents the classification loss of the fusion feature X fusion . In the formula, the Euclidean distance term is used as a regularization term to encourage the classification head output f cls (z fusion ) to be close to the input feature z fusionClose, and during the backpropagation process, the parameters will be optimized simultaneously, enabling the classifier to not only complete the classification task but also ensure that its output f cls (z fusion ) has the smallest difference from the input z fusion .
[0083] Therefore, through this mechanism, the model will pay more attention to classification-related features during the classification task and will not lose useful information in the input features, thus achieving 1) preventing over-shift: by constraining the Euclidean distance, the classifier will not overly change the representation of the input features, maintaining the internal consistency between the input and the output; 2) enhancing feature robustness: strengthening the reconstruction ability of the classification head for the input features to ensure that the features learned by the model are discriminative; 3) feature interpretability: the proximity of the classification head output to the original features makes the model easier to understand and interpret.
[0084] 3) Quantization loss
[0085] The quantization loss realizes quantization alignment between modalities. When calculating the quantization loss, an inter-modal alignment constraint is added.
[0086]
[0087] Among them, sg(·) represents the stop-gradient operation, which will prevent backpropagation and ensure that certain branches do not update the gradient; it is used here to fix the components to achieve co-optimization between the encoder and the quantization operation. β is a balance coefficient used to adjust the weight of the quantization error term.
[0088] The quantization loss can optimize both the encoder (making it easier for the continuous representation to find the closest discrete codebook points) and the codebook (updating the center points of the codebook to better fit the distribution of the continuous representation).
[0089] 4) Perceptual loss
[0090] Through the perceptual loss, the consistency of the reconstructed features in the high-level semantic space can be constrained, focusing on the preservation of global semantic information. Use a pre-trained MLP to extract high-level feature representations φL(·), and compare the representation differences between the original input features and the reconstructed features at different levels to obtain
[0091] 5) Multi-scale loss
[0092] The multi-scale loss is used to evaluate the quality of the reconstructed features at different scales to ensure that information is consistent at multiple levels.
[0093]
[0094] In the above formula, X(k) represents the representation of the feature at the k-th scale (obtained by multi-scale pooling).
[0095] Finally, the combined loss function is obtained as follows:
[0096]
[0097] As can be seen from the above formula, the combined loss function introduces a dynamic weight adjustment mechanism, and the advantages are as follows: 1) Adaptive weight update: The weight of each loss term is dynamically allocated according to its index, ensuring that the model automatically focuses on the loss term that most needs to be optimized during training; 2) Multi-objective balance: The dynamic weights of different loss terms help to balance the reconstruction quality, classification performance, semantic consistency, and multi-scale reconstruction, avoiding over-optimization of a single objective; 3) Enhanced robustness: The dynamic weight mechanism can be adjusted in real time according to the task requirements and model performance, which helps to improve the robustness and generalization performance of the model in the case of missing modalities.
[0098] The dynamic weight adjustment mechanism adopts the dynamic weight formula to calculate the weights λ 1 to λ 5 . In the formula, Indicator i represents the importance index corresponding to each loss term, which is specifically as follows:
[0099] 1) The reconstruction loss index Indicator recon represents the reconstruction error of the fused feature, which is used to ensure the quality of the basic feature.
[0100] 2) The classification loss index Indicator cls is based on the classification accuracy Acc val on the validation set, reflecting the optimization requirements of the classification performance. That is, Indicator cls = 1 - Acc val . It can be seen that the lower the classification accuracy, the greater the weight of the classification loss.
[0101] 3) The quantization loss index Indicator quant represents the discreteness error of the latent representation, which is used to optimize the effectiveness of the latent space.
[0102]
[0103] 4) The perceptual loss index Indicator perceptual represents the consistency of the fused feature at the high-level semantics, which is measured by the value of the perceptual loss
[0104] 5) Multi-scale Discriminative Loss Indicator multi Measure the reconstruction error of features at different scales, that is
[0105] This joint loss function aims to comprehensively address the core challenges in multi-modal data processing by introducing a dynamic weight adjustment mechanism and multiple metric constraints. It mainly focuses on the following key requirements: feature reconstruction quality, improvement of classification performance, effectiveness of latent representations in the discrete space, preservation of high-level semantic information, and multi-scale feature reconstruction quality. By adaptively balancing the importance of each loss term, this design effectively enhances the fusion effect of multi-modal features and improves the robustness of the model in the case of missing modalities or incomplete data. This joint loss takes into account both global feature semantics and local detail expressions, ensuring that the generated features can maintain high-quality performance at different levels, providing an innovative and practical solution for complex multi-modal data processing tasks.
[0106] S9: In the tuning under incomplete modalities, a multi-scale masking strategy is adopted to enhance the robustness of the model in the case of missing modalities and local feature loss.
[0107] For the model constructed using complete modalities in the previous steps, tuning under incomplete modalities is required. For this purpose, a multi-scale masking strategy is adopted in this embodiment to enhance the robustness of the model in the case of missing modalities and local feature loss.
[0108] First, in the full-modal level mask, randomly completely mask some modalities (such as EEG or GSR) to simulate the situation where some sensors may be completely unavailable in actual applications. For key modalities (such as EEG and eye movement), the masking probability is set low to ensure the retention of core information; for auxiliary modalities, the masking probability is increased to prompt the model to gradually rely on the core modalities to complete feature complementation.
[0109] Second, the time-scale mask randomly masks different lengths of time steps in the single-modal data to simulate the loss of some time segments, enabling the model to maintain information integrity when processing the missing local temporal information.
[0110] Finally, the multi-modal local mask randomly masks specific frequency bands or feature dimensions of a specific modality, enabling the model to rely on other features when a specific frequency band or feature is missing, thereby enhancing the generalization ability at the feature level.
[0111] This multi-scale masking strategy comprehensively uses multiple masking methods, enabling the model to generate stable and complete multi-modal feature representations in the face of various modality and feature loss situations.
[0112] Moreover, in the initial stage of the optimization training, the masking ratio and the dropout ratio are relatively low so that the model can gradually adapt to the multi-modal information. As the training process progresses, the ratio is gradually increased to enable the model to have a stronger ability to handle multi-modal missing combinations. By gradually monitoring the feature completion and task performance of the model, the masking ratio is adjusted in a timely manner to avoid information loss caused by excessive masking. When the performance of the model stabilizes at a certain masking ratio, the ratio increase is stopped to ensure that the model reaches the optimal state between balancing information utilization and information loss.
[0113] Second Embodiment
[0114] The second embodiment of the present invention provides an online processing method for multi-modal electroencephalogram analysis, which is implemented based on the model obtained by the above multi-modal electroencephalogram analysis model construction method and is used for real-time testing when the subject wears a portable BCI device.
[0115] As Figure 4 shown, the online processing method for multi-modal electroencephalogram analysis provided by the second embodiment of the present invention at least includes the following steps.
[0116] S101: Use a multi-modal signal acquisition device to acquire at least two modal signals of the electroencephalogram, electrocardiogram, skin electrical activity, and eye movement of the subject;
[0117] S102: Preprocess the multi-modal signals;
[0118] S103: Extract features from the multi-modal signals;
[0119] S104: Determine whether there is a situation where the modal data is incomplete. If so, go to step S105; if not, go to step S106;
[0120] S105: Perform feature fusion on the multi-modal signals, and perform feature completion in the incomplete modality, and then output the classification result;
[0121] S106: Perform feature fusion on the multi-modal signals and directly output the classification result.
[0122] Among them, the aforementioned steps S102 to S106 are all implemented according to the steps introduced in the first embodiment. Among them, to determine whether there is a situation where the modal data is incomplete, one or all of the following methods can be used:
[0123] Method 1: In the signal acquisition stage, check the acquisition status of each modal signal to confirm whether all expected modal signals are successfully acquired. If it is found that the signal of a certain modality is missing or the acquisition is incomplete, this will be regarded as incomplete modal data.
[0124] Method 2: Evaluate the quality of the collected signals, including indicators such as the signal-to-noise ratio and stability of the signals. If the signal quality of a certain modality is lower than the preset threshold, further processing may be required or the data may be regarded as incomplete.
[0125] Method 3 (preferred solution): In the feature extraction stage, analyze whether the extracted features are complete. If it is found that the features of some modalities are missing or abnormal, this may indicate that there are problems with the original signals in these modalities.
[0126] The third embodiment
[0127] Based on the above multi-modal EEG analysis online processing method, the third embodiment of the present invention further provides a multi-modal EEG analysis online processing system, including a multi-modal signal acquisition device, a VR device, and an edge computing module. Since data processing and feature extraction can be performed on the local device (edge computing module) where the data is generated, rather than on a remote server, the latency can be reduced, thereby improving the real-time performance; it can also enhance privacy and security by directly processing the data on the edge computing module to reduce the risk of sensitive information being transmitted over the network.
[0128] The edge computing module includes one or more processors and a memory. Among them, the memory is coupled to the processor and is used to store one or more programs. When the program is executed by the processor, the processor implements the multi-modal EEG analysis online processing method in the above embodiment.
[0129] Among them, the processor is used to control the overall operation of the system to complete all or part of the steps of the above multi-modal EEG analysis online processing method. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. The memory is used to store various types of data to support the operation of the system. These data can include, for example, instructions for any application program or method operating on the system, as well as application program related data. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, etc.
[0130] In another exemplary embodiment, the present invention further provides a computer-readable storage medium including program instructions, which, when executed by a processor, implement the steps of the multi-modal EEG analysis online processing method in any one of the above embodiments. For example, the computer-readable storage medium may be the memory including the program instructions above, and the program instructions may be executed by the processor of the system to complete the above multi-modal EEG analysis online processing method and achieve the same technical effects as the above method.
[0131] In summary, the present invention improves the robustness, real-time performance, and universality of brain-computer interface technology in multi-modal data processing by constructing a multi-modal EEG analysis model, its online processing method, and system. The present invention collects and preprocesses multi-modal biological signals, uses a temporal convolutional network and a multi-view embedding layer for feature extraction and rich expression, combines an adaptive Transformer mechanism and a cross-modal attention mechanism to achieve feature fusion, and encodes the features into a shared latent space to complete missing data. In addition, the quality of the generated features is optimized through a joint loss function, and a multi-scale masking strategy is adopted to enhance the robustness of the model in the case of modal loss, thereby providing more intelligent and convenient multi-modal data collection and processing in applications such as emotion recognition, cognitive load monitoring, and neurorehabilitation, and enhancing the practicality and stability of the BCI-VR system.
[0132] It should be noted that the above-mentioned multiple embodiments are only examples. The technical solutions of each embodiment can be combined, and the order of each step can be changed, all within the protection scope of this patent.
[0133] The above has described in detail the multi-modal EEG analysis model construction method, online processing method, and system provided by the present invention. For those of ordinary skill in the art, any obvious modification made without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.
Claims
1. A method for constructing a multimodal EEG analysis model, characterized in that The following steps are involved: S1: Collect multi-modal signals and pre-process the signals of each mode respectively; S2: The preprocessed data is sliced according to fixed time windows to ensure the temporal consistency of signals of all modes; S3: Use the temporal convolutional network to extract features from the sliced data and obtain the features of each modality; S4: Use the multi-view embedding layer to project each modality feature into multiple subspaces to enrich the feature expression; S5: Perform position embedding and modality embedding, introduce a multi-head self-attention mechanism to enrich feature representation, and introduce multiple modality experts based on the adaptive Transformer mechanism to adaptively select experts according to the input modality to capture the specific information of each modality; S6: Use the cross-modal attention mechanism to perform multimodal fusion, and after obtaining the fusion features of all modalities, aggregate the fusion features again; S7: Encode multimodal features into a shared latent space to complete features and map them to a unified discrete space in the case of missing modalities or data; S8: In the shared latent space, the decoder fuses multimodal features through quantized latent representations and outputs the reconstructed original multimodal feature representations; wherein a joint loss function is used to improve the quality of the generated features.
2. The method for constructing a multimodal EEG analysis model according to claim 1, wherein: The joint loss function is a weighted sum of reconstruction loss, classification loss, quantization loss, perceptual loss and multi-scale loss.
3. The method for constructing a multimodal EEG analysis model according to claim 2, wherein: The joint loss function is: Among them, the dynamic weight formula is used To calculate the weights λ1 to λ5 corresponding to each loss, where Indicator i Represents the importance index corresponding to each loss.
4. The method for constructing a multimodal EEG analysis model as claimed in claim 3, characterized in that After step S8, the method further includes step S9: The fine-tuning under incomplete modality adopts a multi-scale masking strategy to enhance the robustness of the model in the case of modality loss and local feature loss.
5. The method for constructing a multimodal EEG analysis model according to claim 4, characterized in that: First, in the full modality level mask, some modalities are completely masked randomly to simulate the situation in real applications where some sensors may be completely unavailable; Second, the time scale mask randomly masks time steps of different lengths in single modality data to simulate the loss of some time segments; Finally, multimodal local masks randomly mask the frequency bands or feature dimensions of a specific modality, allowing the model to rely on other features when a specific frequency band or feature is missing.
6. The method for constructing a multimodal EEG analysis model according to claim 5, characterized in that: During the tuning process, the mask ratio and the discard ratio are gradually increased as the training progresses.
7. The method for constructing a multimodal EEG analysis model according to claim 6, wherein: When the performance of the model tends to be stable at a certain mask ratio, stop increasing the ratio.
8. The method for constructing a multimodal EEG analysis model according to claim 5, wherein: According to the input modality, modality experts are adaptively selected to capture the specific information of the corresponding modality.
9. A multimodal EEG analysis online processing method, based on the model obtained by the multimodal EEG analysis model construction method according to any one of claims 1 to 8, characterized in that The following steps are involved: S101: using a multimodal signal acquisition device to collect at least two modal signals of the subject's electroencephalogram, electrocardiogram, skin electrical activity, and eye movement; S102: preprocessing the multimodal signal; S103: extracting features from multimodal signals; S104: Determine whether the modal data is incomplete, if so, proceed to step S105; if not, proceed to step S106; S105: performing feature fusion on the multimodal signal, and completing the features under the incomplete mode, and then outputting the classification result; S106: Perform feature fusion on the multimodal signal and directly output the classification result.
10. A multimodal EEG analysis online processing system, characterized in that It includes a multimodal signal acquisition device, a VR device, and an edge computing module, wherein the edge computing module includes a processor and a memory, and the processor and the memory are coupled; Wherein, the memory is used to store computer programs; the processor is used to run the computer programs stored in the memory to implement the multimodal EEG analysis online processing method described in claim 9.
Citation Information
Cited By
Portable multi-mode brain-computer interface intelligent identification system
CN121070189A
A portable multimodal brain-computer interface intelligent recognition system
CN121070189B
Electroencephalogram signal processing method and device and computer equipment
CN121465609A
Brain heuristic multi-expert multi-modal emotion recognition method and system, equipment and medium
CN121524766A
A brain-inspired multi-expert multi-modal emotion recognition method, system, device and medium
CN121524766B