A method and system for intelligent EEG recognition based on multimodal data fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-03-31
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the performance of a single EEG signal is limited under different physiological states, and multimodal signal fusion methods fail to fully consider modal differences, resulting in information redundancy or loss of key features, which affects the model's discrimination ability.
By constructing a multimodal signal recognition model, employing multimodal representation learning within a time period and temporal context enhancement across time periods, and utilizing preprocessed multimodal signals for feature extraction and iterative updates, the recognition accuracy and efficiency are improved.
It achieves multimodal representation learning within a time period and temporal context enhancement across time periods, improving the accuracy and efficiency of EEG intelligent recognition through multimodal data fusion.
Smart Images

Figure CN122075014A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of EEG intelligent recognition based on multimodal data fusion, and in particular to a method and system for EEG intelligent recognition based on multimodal data fusion. Background Technology
[0002] Electroencephalography (EEG), as a non-invasive clinical monitoring technique, is widely used in the diagnosis of epilepsy, coma, brain tumors, and sleep disorder monitoring. However, due to limitations in acquisition equipment, the performance of a single EEG signal is limited under different physiological states. To improve recognition performance, more and more studies are considering multimodal signals (such as EOG and EMG). However, many methods employ static fusion strategies, which fail to fully consider the differences between different modalities. Because different modalities differ significantly in statistical characteristics, frequency distribution, and signal-to-noise ratio, simple feature splicing or unified feature space modeling often leads to information redundancy or loss of key features, thus affecting the model's discriminative ability. Summary of the Invention
[0003] This application aims to at least address the technical problems existing in the prior art. To this end, this application proposes a multimodal data fusion-based EEG intelligent recognition method and system, which can achieve multimodal representation learning within a time period and temporal context enhancement across time periods, thereby improving the accuracy and efficiency of multimodal data fusion-based EEG intelligent recognition.
[0004] The first aspect of this application provides a multimodal data fusion-based EEG intelligent recognition method, comprising the following steps: Acquire preprocessed multimodal signals of the target user within a preset time period, wherein the preprocessed multimodal signals include preprocessed EEG modal signals, preprocessed EEG modal signals, and preprocessed EMG modal signals; Based on the preset time window length and the preset time period, the preprocessed multimodal signal is divided to obtain several segmented multimodal signals for the time period to be identified. The multimodal signals, after being divided into several time periods to be identified, are input into a trained multimodal signal recognition model to obtain the recognition result of the target user for each of the time periods to be identified, as output by the trained multimodal signal recognition model. The training process of the trained multimodal signal recognition model includes: Obtain the multimodal signals after dividing several historical time periods and the true label value of the multimodal signals after dividing each historical time period; An initial multimodal signal recognition model is constructed. Based on the multimodal signals after the division of the several historical time periods, feature extraction is performed through the initial multimodal signal recognition model to obtain several single-modal features for each historical time period. The several single-modal features include EEG modal features, EOS modal features, and EMG modal features. Based on all the single-modal features, the fusion features for each historical time period are determined by the initial multimodal signal recognition model; Based on all the fused features and the true label values, the initial multimodal signal recognition model is iteratively updated to obtain the trained multimodal signal recognition model.
[0005] The multimodal data fusion-based EEG intelligent recognition method according to the embodiments of this application has at least the following beneficial effects: This application inputs multimodal signals from several time periods to be identified into a pre-trained multimodal signal recognition model to obtain the target user's recognition result for each time period. The training process of the pre-trained multimodal signal recognition model includes: acquiring multimodal signals from several historical time periods and the true label values of the multimodal signals from each historical time period; constructing an initial multimodal signal recognition model; extracting features from the multimodal signals from several historical time periods to obtain several single-modal features for each historical time period; determining the fusion features for each historical time period based on all single-modal features; and iteratively updating the initial multimodal signal recognition model based on all fusion features and the true label values to obtain a pre-trained multimodal signal recognition model. This achieves multimodal representation learning within time periods and temporal context enhancement across time periods, improving the accuracy and efficiency of EEG intelligent recognition based on multimodal data fusion.
[0006] A second aspect of this application provides a multimodal data fusion-based EEG intelligent recognition system, comprising: The data acquisition module is used to acquire preprocessed multimodal signals of the target user within a preset time period, wherein the preprocessed multimodal signals include preprocessed EEG modal signals, preprocessed EEG modal signals, and preprocessed EMG modal signals. The signal segmentation module is used to segment the preprocessed multimodal signal based on a preset time window length and the preset time period to obtain a number of segmented multimodal signals for a time period to be identified. The recognition result output module is used to input the multimodal signals after dividing the several time periods to be recognized into a trained multimodal signal recognition model, so as to obtain the recognition result of the target user in each of the time periods to be recognized, as output by the trained multimodal signal recognition model. The training process of the trained multimodal signal recognition model includes: Obtain the multimodal signals after dividing several historical time periods and the true label value of the multimodal signals after dividing each historical time period; An initial multimodal signal recognition model is constructed. Based on the multimodal signals after the division of the several historical time periods, feature extraction is performed through the initial multimodal signal recognition model to obtain several single-modal features for each historical time period. The several single-modal features include EEG modal features, EOS modal features, and EMG modal features. Based on all the single-modal features, the fusion features for each historical time period are determined by the initial multimodal signal recognition model; Based on all the fused features and the true label values, the initial multimodal signal recognition model is iteratively updated to obtain the trained multimodal signal recognition model.
[0007] This system inputs multimodal signals from several time periods to be identified into a pre-trained multimodal signal recognition model to obtain the target user's recognition result for each time period. The training process of the pre-trained multimodal signal recognition model includes: acquiring multimodal signals from several historical time periods and the true label values of the multimodal signals from each historical time period; constructing an initial multimodal signal recognition model; extracting features from the multimodal signals from the several historical time periods to obtain several single-modal features for each historical time period; determining the fusion features for each historical time period based on all single-modal features; and iteratively updating the initial multimodal signal recognition model based on all fusion features and the true label values to obtain a pre-trained multimodal signal recognition model. This achieves multimodal representation learning within time periods and temporal context enhancement across time periods, improving the accuracy and efficiency of EEG intelligent recognition based on multimodal data fusion.
[0008] A third aspect of this application provides a multimodal data fusion-based EEG intelligent recognition electronic device, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which are executed by the at least one control processor to enable the at least one control processor to perform the aforementioned multimodal data fusion-based EEG intelligent recognition method.
[0009] In a fourth aspect, this application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the above-described multimodal data fusion-based EEG intelligent recognition method.
[0010] It should be noted that the beneficial effects of the second to fourth aspects of this application with respect to the prior art are the same as the beneficial effects of the aforementioned multimodal data fusion EEG intelligent recognition system with respect to the prior art, and will not be elaborated here.
[0011] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0012] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating an embodiment of the EEG intelligent recognition method based on multimodal data fusion provided in this application; Figure 2 This is a schematic diagram of the structure of an embodiment of the multimodal data fusion-based EEG intelligent recognition system provided in this application; Figure 3 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation
[0013] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0014] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.
[0015] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0016] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0017] Electroencephalography (EEG), as a non-invasive clinical monitoring technique, is widely used in the diagnosis of epilepsy, coma, brain tumors, and sleep disorder monitoring. However, due to limitations in acquisition equipment, the performance of a single EEG signal is limited under different physiological states. To improve recognition performance, more and more studies are considering multimodal signals (such as EOG and EMG). However, many methods employ static fusion strategies, which fail to fully consider the differences between different modalities. Because different modalities differ significantly in statistical characteristics, frequency distribution, and signal-to-noise ratio, simple feature splicing or unified feature space modeling often leads to information redundancy or loss of key features, thus affecting the model's discriminative ability.
[0018] To address the aforementioned technical deficiencies, this application provides a method and system for intelligent EEG recognition based on multimodal data fusion.
[0019] Please see Figure 1 This is a flowchart illustrating a multimodal data fusion-based EEG intelligent recognition method provided in an embodiment of this application. This method is applied to an electronic device, which may be a server, etc. Figure 1 As shown, this multimodal data fusion-based EEG intelligent recognition method includes: Step S101: Obtain the preprocessed multimodal signals of the target user within a preset time period, wherein the preprocessed multimodal signals include preprocessed EEG modal signals, preprocessed EEG modal signals, and preprocessed EMG modal signals. The aforementioned preset time period can be a value set in advance according to actual needs.
[0020] Step S102: Based on the preset time window length and preset time period, the preprocessed multimodal signal is divided to obtain several segmented multimodal signals for the time period to be identified. The above-mentioned preset time window length can be a value that can be preset according to actual needs.
[0021] The aforementioned segmented multimodal signals may include segmented EEG modal signals, segmented EOS modal signals, and segmented EMG modal signals.
[0022] Step S103: Input the multimodal signals divided into several time periods to be identified into the trained multimodal signal recognition model to obtain the target user's recognition result for each time period output by the trained multimodal signal recognition model. The training process of the trained multimodal signal recognition model includes: Step S1031: Obtain the multimodal signals after the division of several historical time periods and the true label value of the multimodal signals after the division of each historical time period; In step S1031, the calculation process of obtaining the multimodal signals after the division of several historical time periods is similar to the calculation process of obtaining the multimodal signals after the division of several time periods to be identified, and will not be repeated here.
[0023] In step S1031, the above-mentioned acquisition of the true label value of the multimodal signal after the division of each historical time period can be obtained by manually labeling the multimodal signal after the division of each historical time period.
[0024] Step S1032: Construct an initial multimodal signal recognition model. Based on the multimodal signals after the division of several historical time periods, feature extraction is performed through the initial multimodal signal recognition model to obtain several single-modal features for each historical time period. Among them, the several single-modal features include EEG modal features, EEG modal features, and EMG modal features. The aforementioned initial multimodal signal recognition model may include a multi-scale dilated convolution module and a spatial convolution operation module.
[0025] Step S1033: Based on all single-modal features, determine the fusion features for each historical time period through the initial multimodal signal recognition model; Step S1034: Based on all fused features and real label values, iteratively update the initial multimodal signal recognition model to obtain a trained multimodal signal recognition model.
[0026] The above identification results can be physiological state category labels for the time period to be identified. These labels will vary in different application scenarios. For example, in the sleep staging task, the category labels may include the awake period (0), the sleep onset period (1), the light sleep period (2), the deep sleep period (3), and the REM sleep period (4) (it should be noted that the numerical labels here can be used for subsequent loss calculations); while in other tasks, such as epilepsy monitoring or acute disease diagnosis, the category labels may represent different pathological states or signal patterns and be represented in the form of numerical codes.
[0027] This application inputs multimodal signals from several time periods to be identified into a pre-trained multimodal signal recognition model to obtain the target user's recognition result for each time period. The training process of the pre-trained multimodal signal recognition model includes: acquiring multimodal signals from several historical time periods and the true label values of the multimodal signals from each historical time period; constructing an initial multimodal signal recognition model; extracting features from the multimodal signals from several historical time periods to obtain several single-modal features for each historical time period; determining the fusion features for each historical time period based on all single-modal features; and iteratively updating the initial multimodal signal recognition model based on all fusion features and the true label values to obtain a pre-trained multimodal signal recognition model. This achieves multimodal representation learning within time periods and temporal context enhancement across time periods, improving the accuracy and efficiency of EEG intelligent recognition based on multimodal data fusion.
[0028] In some embodiments, step S101 may include steps S201 to S204: Step S201: Obtain the initial EEG modal signal, initial EEG modal signal, initial EMG modal signal, initial EMG modal signal, sampling frequency of the initial EEG modal signal, sampling frequency of the initial EEG modal signal, and sampling frequency of the initial EMG modal signal within a preset time period for the target user; In step S201, the acquisition of the initial EEG modal signal, initial EOL modal signal, initial EMG modal signal, and initial EMG modal signal of the target user within a preset time period, and the sampling frequency of the initial EEG modal signal, the initial EOL modal signal, and the initial EMG modal signal can be achieved by collecting the initial EEG modal signal, initial EOL modal signal, and initial EMG modal signal of the target user within a preset time period through sensors, and recording the sampling frequency of the initial EEG modal signal, the initial EOL modal signal, and the initial EMG modal signal.
[0029] Step S202: Based on the sampling frequency of the initial EEG modal signal, the initial EOL modal signal, the initial EMG modal signal, and the preset target sampling rate, the initial EEG modal signal, the initial EOL modal signal, and the initial EMG modal signal are resampled using the following formula to obtain the aligned multimodal signal. The aligned multimodal signal includes the aligned EEG modal signal, the aligned EOL modal signal, and the aligned EMG modal signal.
[0030] in, For pre-acquired Total number of channels in the modality After alignment The first modality Signal from each channel, Modalities include EEG modality, EEG modality, and EMG modality. for Modal sampling frequency, To preset the target sampling rate, For the initial The first modality Signals from each channel; Step S203: Filter the aligned multimodal signal to obtain the filtered multimodal signal; Step S204: Perform Z-score normalization on the aligned multimodal signal to obtain the preprocessed multimodal signal within a preset time period.
[0031] This application improves the accuracy of EEG intelligent recognition by acquiring preprocessed multimodal signals of target users within a preset time period, providing more accurate data for subsequent model prediction.
[0032] In some embodiments, step S102 divides the preprocessed multimodal signal based on a preset time window length and a preset time period using the following formula to obtain a number of segmented multimodal signals for different time periods to be identified:
[0033] in, This represents the total number of time periods to be identified. The total duration of the preset time period, The preset time window length, For preprocessing The first modality Signal from each channel, After dividing the first time period to be identified The first modality Signal from each channel, After dividing the second time period to be identified The first modality Signal from each channel, For the first After dividing the time period to be identified The first modality Signals from each channel.
[0034] This application divides the preprocessed multimodal signal based on a preset time window length and a preset time period to obtain several segmented multimodal signals for identification time periods. This realizes multimodal data within the time period, providing more accurate data basis for subsequent model prediction, thereby improving the accuracy of EEG intelligent recognition based on multimodal data fusion.
[0035] In some embodiments, step S1032 may include steps S301 to S305: Step S301: Using a multi-scale dilated convolution module, dilated convolution operations are performed on the multimodal signals after dividing several historical time periods using the following formula to obtain the convolutional multimodal features for several historical time periods. The convolutional multimodal features include convolutional EEG modal features, convolutional EOS modal features, and convolutional EMG modal features:
[0036] in, The void ratio of the multi-scale dilated convolutional module. For the first The porosity of the multi-scale dilated convolutional module over a historical time period is [missing information]. After convolution The first modality Characteristics of each channel Let be a grouped convolution function with a dilation rate of d. For activation function, For the first After dividing the historical period The first modality Signals from each channel; Step S302: Based on the preset pooling factor, pool each convolutional multimodal feature to obtain pooled multimodal features for several historical time periods. The preset pooling factor mentioned above can be a value set in advance according to actual needs, and can take values of [value missing]. .
[0037] Step S303: After obtaining the time length of each pooled multimodal feature, the aligned multimodal features are determined by bilinear interpolation using the minimum value among all time lengths, according to the following formula. The aligned multimodal features include aligned EEG modal features, aligned EOS modal features, and aligned EMG modal features:
[0038] in, For the void ratio d, the first Pooling for a historical period The first modality Characteristics of each channel For the void ratio d, the first After aligning the historical time periods The first modality Characteristics of each channel It is the minimum value among all time lengths; Step S304: The spliced multimodal features are obtained by splicing and aligning the multimodal features using the following formula, wherein the spliced multimodal features include spliced EEG modal features, spliced EOS modal features, and spliced EMG modal features:
[0039] in, For the first splicing together historical time periods The first modality Characteristics of each channel The first one with a void ratio of 1 After aligning the historical time periods The first modality Characteristics of each channel The first one with a void ratio of 2 After aligning the historical time periods The first modality Characteristics of each channel The first one with a void ratio of 4 After aligning the historical time periods The first modality Characteristics of each channel The first one with a void ratio of 8 After aligning the historical time periods The first modality Characteristics of each channel; Step S305: Perform spatial convolution operation on the spliced multimodal features using the spatial convolution operation module to obtain several single-modal features for each historical time period.
[0040] This application utilizes multimodal signals divided into several historical time periods to decompose continuous, highly redundant original signals into local signal segments with temporal discriminativeness, preserving temporal dependent features. This reduces the computational complexity and feature confusion issues associated with directly processing long sequences. Furthermore, by extracting features from an initial multimodal signal recognition model, several single-modal features for each historical time period are obtained. This enables independent representation learning of different modal information, fully exploring the inherent patterns and discriminative characteristics of each modality, avoiding mutual interference of multimodal information in the early fusion stage, and providing more accurate data for subsequent feature fusion, thereby improving the accuracy of model recognition.
[0041] In some embodiments, step S1033 may include steps S401 to S403: Step S401: Based on all single-modal features, using the initial multimodal signal recognition model, perform temporal attention pooling using the following formula to obtain the global feature vector:
[0042]
[0043]
[0044] in, For the first A historical period Modal features, For the first The feature vectors of all channels at time 1000. For the first The first historical period Moment Raw scores of modal attention For the first The first historical period Moment Modal time weights, For the first The first historical period Moment Raw scores of modal attention The total number of moments for a single-modal feature. For the first A historical period Global feature vectors of a modality; Step S402: Based on the global feature vector, and using the initial multimodal signal recognition model, calculate the modulated global feature vector using the following formula:
[0045]
[0046]
[0047]
[0048]
[0049] in, For the first The total modal global feature vector for each historical time period For the first Global feature vectors of EEG modalities over a historical time period For the first Global feature vectors of electrooculography modalities over a historical time period. For the first Global feature vectors of electromyography modes over a historical time period For the first Stage embedding vectors for each historical time period, This is the trainable first weight matrix of the initial multimodal signal recognition model. This is the trainable second weight matrix of the initial multimodal signal recognition model. This is the learnable first bias vector of the initial multimodal signal recognition model. This is the learnable second bias vector of the initial multimodal signal recognition model. For the first A historical period The dynamic scaling factor of the modality For the first A historical period Modal shift factor, For the initial trainable multimodal signal recognition model The first parameter of the mode, For the initial trainable multimodal signal recognition model The second parameter of the mode, For the initial trainable multimodal signal recognition model The third parameter of the mode, For the initial trainable multimodal signal recognition model The fourth parameter of the mode, For the first A historical period Global feature vector after modal modulation; Step S403: Based on the modulated global feature vector, and through the initial multimodal signal recognition model, calculate the fused features for each historical time period using the following formula:
[0050] in, For the first The characteristics of the fusion of historical time periods For the first Global feature vectors of EEG modalities over a historical time period For the first Modulated global feature vectors of electrooculography modes over a historical time period For the first Modulated global feature vectors of electromyographic modes for a historical time period.
[0051] This application improves the recognition accuracy and interpretability of the initial multimodal signal recognition model under complex physiological conditions by fusing single-modal multiscale features from different modalities and dynamically adjusting the contribution of each modality (trainable weight matrix and learnable bias vector) according to the current physiological state (single-modal features).
[0052] In some embodiments, step S1034 may include steps S501 to S504: Step S501: Perform average pooling on the fused features to obtain the average pooled fused features for each historical time period; Step S502: Based on the fused features after average pooling, the predicted label value for each historical time period is determined using the following formula through the initial multimodal signal recognition model:
[0053]
[0054] in, For the first The hidden state of a historical period For the first The hidden state of a historical period For the first Average pooling fusion features over a historical time period This is the trainable third weight matrix of the initial multimodal signal recognition model. This is the learnable third bias vector of the initial multimodal signal recognition model. For the first Predicted label values for each historical time period; Step S503: Determine the first loss value based on the predicted label value and the true label value; Step S504: If the first loss value is less than the preset loss threshold, the initial multimodal signal recognition model is used as the trained multimodal signal recognition model.
[0055] The aforementioned preset loss threshold can be a value set in advance according to actual needs.
[0056] In some embodiments, this application may further include: when the first loss value is greater than or equal to a preset loss threshold, iteratively updating the initial multimodal signal recognition model according to the first loss value using a backpropagation algorithm (which may include iteratively updating the trainable weight matrix and learnable bias vector of the initial multimodal signal recognition model) until the number of iterations reaches a preset number of training rounds set according to actual needs, thereby obtaining a trained multimodal signal recognition model.
[0057] This application iteratively updates the initial multimodal signal recognition model based on all fused features and real label values to obtain a trained multimodal signal recognition model that can capture the contextual information of time series. This enables multimodal representation learning within a time period and temporal context enhancement across time periods, thereby improving the accuracy and efficiency of EEG intelligent recognition based on multimodal data fusion.
[0058] In some embodiments, the first loss value is calculated using the following formula:
[0059] in, The first loss value, This represents the total number of historical time periods. For the first The true label value for a historical period This application calculates the first loss value for all historical time periods using the cross-entropy loss function, providing more accurate data for the iterative update of the initial multimodal signal recognition model, thereby improving the accuracy of EEG intelligent recognition based on multimodal data fusion.
[0060] In some embodiments, this application may further include: Based on a pre-set sliding step size according to actual needs, a multimodal signal is generated through a sliding window for each historical time period within a pre-set historical time period, wherein the sliding step size is less than the preset time window length.
[0061] This application enhances data utilization by using a sliding window, which can increase the number of training samples for the initial multimodal signal recognition model.
[0062] Additionally, refer to Figure 2 One embodiment of this application provides a multimodal data fusion-based EEG intelligent recognition system, including a data acquisition module 1100, a signal segmentation module 1200, and a recognition result output module 1300, wherein: The data acquisition module 1100 is used to acquire the preprocessed multimodal signals of the target user within a preset time period, wherein the preprocessed multimodal signals include preprocessed EEG modal signals, preprocessed EEG modal signals and preprocessed EMG modal signals; The signal segmentation module 1200 is used to segment the preprocessed multimodal signal based on a preset time window length and a preset time period to obtain a number of segmented multimodal signals for a time period to be identified. The recognition result output module 1300 is used to input the multimodal signals divided into several time periods to be recognized into the trained multimodal signal recognition model, so as to obtain the recognition result of the target user in each time period to be recognized by the trained multimodal signal recognition model. The training process of the trained multimodal signal recognition model includes: Obtain the multimodal signals after dividing several historical time periods and the true label value of the multimodal signals after dividing each historical time period; An initial multimodal signal recognition model is constructed. Based on the multimodal signals after the division of several historical time periods, feature extraction is performed through the initial multimodal signal recognition model to obtain several single-modal features for each historical time period. Among them, the several single-modal features include EEG modal features, EOS modal features, and EMG modal features. Based on all single-modal features, the fusion features for each historical time period are determined through an initial multimodal signal recognition model; Based on all fused features and true label values, the initial multimodal signal recognition model is iteratively updated to obtain a trained multimodal signal recognition model.
[0063] This system inputs multimodal signals from several time periods to be identified into a pre-trained multimodal signal recognition model to obtain the target user's recognition result for each time period. The training process of the pre-trained multimodal signal recognition model includes: acquiring multimodal signals from several historical time periods and the true label values of the multimodal signals from each historical time period; constructing an initial multimodal signal recognition model; extracting features from the multimodal signals from the several historical time periods to obtain several single-modal features for each historical time period; determining the fusion features for each historical time period based on all single-modal features; and iteratively updating the initial multimodal signal recognition model based on all fusion features and the true label values to obtain a pre-trained multimodal signal recognition model. This achieves multimodal representation learning within time periods and temporal context enhancement across time periods, improving the accuracy and efficiency of EEG intelligent recognition based on multimodal data fusion.
[0064] It should be noted that the system embodiments described above are based on the same inventive concept as the method embodiments described above. Therefore, the relevant content of the method embodiments described above is also applicable to the system embodiments described above, and will not be repeated here.
[0065] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations. The acquisition, storage, use and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.
[0066] Figure 3 This illustration shows a schematic diagram of the hardware structure for multimodal data fusion-based EEG intelligent recognition provided in an embodiment of this application.
[0067] The EEG intelligent recognition device that integrates multimodal data fusion may include a processor 301 and a memory 302 storing computer program instructions.
[0068] Specifically, the processor 301 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0069] Memory 302 may include mass storage for data or instructions. For example, and not limitingly, memory 302 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 302 may include removable or non-removable (or fixed) media. Where appropriate, memory 302 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 302 is non-volatile solid-state memory.
[0070] In some embodiments, memory 302 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Thus, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.
[0071] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any of the multimodal data fusion EEG intelligent recognition methods in the above embodiments.
[0072] In one example, the multimodal data fusion-based EEG intelligent recognition device may further include a communication interface 303 and a bus 310. For example, Figure 3 As shown, the processor 301, memory 302, and communication interface 303 are connected through bus 310 and complete communication with each other.
[0073] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0074] Bus 310 includes hardware, software, or both, that couples together components of an EEG intelligent recognition device that fuses multimodal data. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 310 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0075] This multimodal data fusion-based EEG intelligent recognition device can execute the multimodal data fusion-based EEG intelligent recognition method described in this application embodiment based on a three-dimensional design model, thereby achieving a combination of... Figure 1 and Figure 2 The paper describes a multimodal data fusion-based EEG intelligent recognition method and system.
[0076] Furthermore, in conjunction with the multimodal data fusion-based EEG intelligent recognition method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the multimodal data fusion-based EEG intelligent recognition methods in the above embodiments.
[0077] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0078] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0079] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0080] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0081] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A multi-modal data fusion electroencephalogram intelligent recognition method, characterized in that, The multi-modal data fusion electroencephalogram intelligent recognition method comprises: obtaining preprocessed multi-modal signals of a target user within a preset time period, wherein the preprocessed multi-modal signals comprise preprocessed electroencephalogram modal signals, preprocessed electrooculogram modal signals, and preprocessed electromyogram modal signals; dividing the preprocessed multi-modal signals based on a preset time window length and the preset time period to obtain divided multi-modal signals of a plurality of to-be-recognized time periods; inputting the divided multi-modal signals of the plurality of to-be-recognized time periods into a trained multi-modal signal recognition model to obtain recognition results of the target user in each of the to-be-recognized time periods output by the trained multi-modal signal recognition model, wherein a training process of the trained multi-modal signal recognition model comprises: obtaining divided multi-modal signals of a plurality of historical time periods and true label values of the divided multi-modal signals of each of the historical time periods; constructing an initial multi-modal signal recognition model, performing feature extraction on the divided multi-modal signals of the plurality of historical time periods through the initial multi-modal signal recognition model to obtain a plurality of single-modal features of each of the historical time periods, wherein the plurality of single-modal features comprise electroencephalogram modal features, electrooculogram modal features, and electromyogram modal features; determining fusion features of each of the historical time periods based on all the single-modal features through the initial multi-modal signal recognition model; iteratively updating the initial multi-modal signal recognition model based on all the fusion features and the true label values to obtain the trained multi-modal signal recognition model. 2.The multi-modal data fusion electroencephalogram intelligent recognition method according to claim 1, characterized in that, The method comprises the following steps: obtaining initial electroencephalogram modal signals, initial electrooculogram modal signals, initial electromyogram modal signals, a sampling frequency of the initial electroencephalogram modal signals, a sampling frequency of the initial electrooculogram modal signals, and a sampling frequency of the initial electromyogram modal signals of the target user within a preset time period; resampling the initial electroencephalogram modal signals, the initial electrooculogram modal signals, and the initial electromyogram modal signals based on the sampling frequencies of the initial electroencephalogram modal signals, the initial electrooculogram modal signals, the initial electromyogram modal signals, and a preset target sampling rate through the following formula to obtain aligned multi-modal signals, wherein the aligned multi-modal signals comprise aligned electroencephalogram modal signals, aligned electrooculogram modal signals, and aligned electromyogram modal signals: in, For pre-acquired Total number of channels in the modality After alignment The first modality Signal from each channel, Modalities include EEG modality, EEG modality, and EMG modality. for Modal sampling frequency, To preset the target sampling rate, For the initial The first modality Signals from each channel; filtering the aligned multi-modal signals to obtain filtered multi-modal signals; performing Z-score standardization on the aligned multi-modal signals to obtain the preprocessed multi-modal signals within the preset time period. 3.The multi-modal data fusion electroencephalogram intelligent recognition method of claim 2, wherein, The divided multi-modal signals comprise divided electroencephalogram modal signals, divided electrooculogram modal signals, and divided electromyogram modal signals, and the preprocessed multi-modal signals are divided based on a preset time window length and the preset time period through the following formula to obtain divided multi-modal signals of a plurality of to-be-recognized time periods: wherein, is a total number of time periods to be identified, is a total duration of preset time periods, is a preset time window length, is a signal of the first channel of the modality after pre-processing, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified, is a signal of the first channel of the modality after division of the first time period to be identified, is a signal of the first channel of the modality after division of the second time period to be identified.
4. The multi-modal data fusion electroencephalogram intelligent recognition method according to claim 3, characterized in that, The initial multi-modal signal recognition model includes a multi-scale hollow convolution module and a spatial convolution operation module. The initial multi-modal signal recognition model is used for feature extraction based on the divided multi-modal signals of the historical time periods, to obtain single-modal features of each historical time period, including: The multi-scale hollow convolution module is used to perform hollow convolution operation on the divided multi-modal signals of the historical time periods according to the following formula, to obtain convolution multi-modal features of the historical time periods, wherein the convolution multi-modal features include convolution electroencephalogram modal features, convolution electrooculogram modal features and convolution electromyography modal features: wherein, is a dilation rate of a multiscale atrous convolution module, is a dilation rate of a multiscale atrous convolution module for a th historical time period, is a signal of a th channel of a modality after a convolution with a th historical time period, is a grouped convolution function with a dilation rate of d, is an activation function, is a signal of a th channel of a modality after a division into a th historical time period, th channel of a modality. The convolution multi-modal features are pooled based on a preset pooling multiple, to obtain pooled multi-modal features of the historical time periods; In the case of obtaining the time length of each pooled multi-modal feature, the minimum value in all the time lengths is determined, and the aligned multi-modal features are determined based on all the pooled multi-modal features by bilinear interpolation according to the following formula, wherein the aligned multi-modal features include aligned electroencephalogram modal features, aligned electrooculogram modal features and aligned electromyography modal features: wherein, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, a feature of a i-th channel of a i-th historical time period with a hole rate of d, is the minimum value in all time lengths; The aligned multi-modal features are spliced according to the following formula, to obtain spliced multi-modal features, wherein the spliced multi-modal features include spliced electroencephalogram modal features, spliced electrooculogram modal features and spliced electromyography modal features: wherein, is a concatenation of features of a th channel of a modal for a th historical time period, is a concatenation of features of a th channel of a modal for a th historical time period with a void rate of 1, is a concatenation of features of a th channel of a modal for a th historical time period with a void rate of 2, is a concatenation of features of a th channel of a modal for a th historical time period with a void rate of 4, is a concatenation of features of a th channel of a modal for a th historical time period with a void rate of 8; The spatial convolution operation module is used to perform spatial convolution operation on the spliced multi-modal features, to obtain the single-modal features of each historical time period.
5. The multi-modal data fusion electroencephalogram intelligent recognition method according to claim 4, characterized in that, The initial multi-modal signal recognition model is used to determine fusion features of each historical time period based on all the single-modal features, including: The initial multi-modal signal recognition model is used to perform time attention pooling based on all the single-modal features according to the following formula, to obtain a global feature vector: in, For the first A historical period Modal features, For the first The feature vectors of all channels at time 1000. For the first The first historical period Moment Raw scores of modal attention For the first The first historical period Moment Modal time weights, For the first The first historical period Moment Raw scores of modal attention The total number of moments for a single-modal feature. For the first A historical period Global feature vectors of a modality; The initial multi-modal signal recognition model is used to calculate a modulated global feature vector based on the global feature vector according to the following formula: wherein, is a total modal global feature vector for the th historical time period, is a global feature vector for the electroencephalography modality for the th historical time period, is a global feature vector for the electrooculography modality for the th historical time period, is a global feature vector for the electromyography modality for the th historical time period, is a stage embedding vector for the th historical time period, is a first trainable weight matrix of the initial multi-modal signal recognition model, is a second trainable weight matrix of the initial multi-modal signal recognition model, is a first learnable bias vector of the initial multi-modal signal recognition model, is a second learnable bias vector of the initial multi-modal signal recognition model, is a dynamic scaling factor for the th modality for the th historical time period, is an offset factor for the th modality for the th historical time period, is a first trainable parameter of the initial multi-modal signal recognition model for the th modality, is a second trainable parameter of the initial multi-modal signal recognition model for the th modality, is a third trainable parameter of the initial multi-modal signal recognition model for the th modality, is a fourth trainable parameter of the initial multi-modal signal recognition model for the th modality, is a modulated global feature vector for the th modality for the th historical time period; The initial multi-modal signal recognition model is used to calculate the fusion features of each historical time period based on the modulated global feature vector according to the following formula: wherein, is a fusion feature for the th historical time period, is a modulated global feature vector for the electroencephalography modality for the th historical time period, is a modulated global feature vector for the electrooculography modality for the th historical time period, is a modulated global feature vector for the electromyography modality for the th historical time period.
6. The multi-modal data fusion electroencephalogram intelligent recognition method according to claim 5, characterized in that, The initial multi-modal signal recognition model is iteratively updated based on all the fusion features and the true label values, to obtain the trained multi-modal signal recognition model, including: The fusion features are averaged and pooled to obtain averaged and pooled fusion features of each historical time period; The initial multi-modal signal recognition model is used to determine a predicted label value of each historical time period based on the averaged and pooled fusion features according to the following formula: in, For the first The hidden state of a historical period For the first The hidden state of a historical period For the first Average pooling fusion features over a historical time period This is the trainable third weight matrix of the initial multimodal signal recognition model. This is the learnable third bias vector of the initial multimodal signal recognition model. For the first Predicted label values for each historical time period; A first loss value is determined based on the predicted label value and the true label value; In the case that the first loss value is less than a preset loss threshold, the initial multi-modal signal recognition model is used as the trained multi-modal signal recognition model.
7. The multi-modal data fusion electroencephalogram intelligent recognition method according to claim 6, characterized in that, The first loss value is calculated according to the following formula: wherein, is a first loss value, is a total number of historical time periods, is a true label value for the th historical time period.
8. A multi-modal data fusion electroencephalogram intelligent recognition system, characterized in that, The multi-modal data fusion electroencephalogram intelligent recognition system includes: The data acquisition module is configured to acquire preprocessed multi-modal signals of a target user within a preset time period, wherein the preprocessed multi-modal signals include preprocessed electroencephalogram signals, preprocessed electrooculogram signals, and preprocessed electromyogram signals; The signal division module is configured to divide the preprocessed multi-modal signals based on a preset time window length and the preset time period to obtain divided multi-modal signals of a plurality of to-be-identified time periods; The identification result output module is configured to input the divided multi-modal signals of the plurality of to-be-identified time periods into a trained multi-modal signal identification model to obtain identification results of the target user in each of the to-be-identified time periods output by the trained multi-modal signal identification model, wherein a training process of the trained multi-modal signal identification model includes: acquiring divided multi-modal signals of a plurality of historical time periods and true label values of the divided multi-modal signals of each of the historical time periods; constructing an initial multi-modal signal identification model, performing feature extraction on the divided multi-modal signals of the plurality of historical time periods through the initial multi-modal signal identification model to obtain a plurality of single-modal features of each of the historical time periods, wherein the plurality of single-modal features include electroencephalogram features, electrooculogram features, and electromyogram features; determining fusion features of each of the historical time periods based on all the single-modal features through the initial multi-modal signal identification model; performing iterative updating on the initial multi-modal signal identification model based on all the fusion features and the true label values to obtain the trained multi-modal signal identification model.
9. An electroencephalogram intelligent recognition electronic device of multi-modal data fusion, characterized in that, The at least one control processor and the memory connected in communication with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform a multi-modal data fusion electroencephalogram intelligent identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a multi-modal data fusion electroencephalogram intelligent identification method according to any one of claims 1 to 7.