Multi-modal information analysis method and device

Through the multimodal information analysis method, combined with EEG, demographic and clinical information, and using evidence theory and neural network model, the limitations of information fusion and analysis in the existing technology are solved, and more accurate and reliable feature analysis is achieved.

CN120337127APending Publication Date: 2025-07-18BEIJING ANDING HOSPITAL CAPITAL MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385216.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When existing machine learning algorithms process EEG data, it is difficult to effectively integrate and analyze multiple information, resulting in limitations in revealing deep-level features.

Method used

Multimodal information analysis method is adopted to obtain EEG information, demographic information and clinical information, multimodal fusion is performed using evidence theory methods, and feature analysis is used using neural network models to determine the optimal feature analysis results.

Benefits of technology

It improves the accuracy and reliability of information analysis, makes full use of complementary information from different data sources, and optimizes the operating efficiency and analysis process of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337127A_ABST
    Figure CN120337127A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal information analysis method and device. The method comprises the following steps: acquiring electroencephalogram information, demographic information and clinical information of a subject; extracting features in the electroencephalogram information, the demographic information and the clinical information to obtain electroencephalogram features, demographic features and clinical features of the subject; within a preset number of rounds, performing multi-modal fusion on the electroencephalogram features, the demographic features and the clinical features by using an evidence theory method, and performing feature analysis on the features subjected to the current multi-modal fusion by using a neural network model; and when a preset round number is reached, determining an optimal feature analysis result as a final analysis result. According to the technical scheme provided by the invention, complementary information in different data sources is fully analyzed by utilizing a multi-modal feature fusion technology, and the accuracy and reliability of a final analysis result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more specifically, to a multi-modal information parsing method and apparatus. Background Art

[0002] In recent years, with the development of neuroscience and data analysis technologies, electroencephalogram (EEG) data has become a means to assist in understanding mental illnesses. As a technology capable of real-time monitoring of brain activities, EEG provides researchers with a new perspective to observe and analyze the neurophysiological characteristics of subjects.

[0003] However, current research mainly focuses on using traditional machine learning algorithms to extract features and classify EEG data, and these methods have limitations in dealing with complex data and revealing deep features.

[0004] Therefore, how to design a multi-modal information parsing solution that can more effectively and accurately fuse and analyze various information about subjects has become a problem to be solved in this field. Summary of the Invention

[0005] In view of this, in the first aspect, this application proposes a multi-modal information parsing method, which includes: obtaining the electroencephalogram information, demographic information, and clinical information of a subject;

[0006] extracting the features in the electroencephalogram information, demographic information, and clinical information to obtain the electroencephalogram features, demographic features, and clinical features of the subject;

[0007] within a preset number of rounds, using the evidence theory method to perform multi-modal fusion on the electroencephalogram features, demographic features, and clinical features, and using a neural network model to perform feature parsing on the features after the current multi-modal fusion;

[0008] when the preset number of rounds is reached, determining the optimal feature parsing result as the final parsing result.

[0009] Preferably, before obtaining the electroencephalogram information of the subject, the method further includes:

[0010] collecting the electroencephalogram signals of the subject and preprocessing the electroencephalogram signals; the preprocessing includes: electrode screening, rereferencing, filtering, downsampling, segment marking, bad segment removal, and independent component analysis.

[0011] Preferably, extracting the features in the electroencephalogram signals includes:

[0012] extracting the power spectral density features, time domain features, and frequency domain features in the electroencephalogram signals.

[0013] Further preferably, extracting the power spectral density features from the electroencephalogram signals includes:

[0014] Calculating the power spectral density corresponding to each channel in each frequency band by using the Welch algorithm to obtain the power spectral density features in the electroencephalogram signals.

[0015] Further preferably, after extracting the features in the electroencephalogram information, demographic information, and clinical information to obtain the electroencephalogram features, demographic features, and clinical features of the subject, the method further includes:

[0016] Screening the extracted features by using the principal component analysis method and / or the variance threshold method.

[0017] Preferably, within a preset number of rounds, performing multimodal fusion on the electroencephalogram features, demographic features, and clinical features by using the evidence theory method, including:

[0018] Determining the fusion weights of the electroencephalogram features, demographic features, and clinical features in each round of multimodal fusion by using the evidence theory method, and performing multimodal fusion on the electroencephalogram features, demographic features, and clinical features accordingly.

[0019] Further preferably, if the feature analysis result corresponding to the fusion weight in the current round is better than the feature analysis result corresponding to the fusion weight in the previous round, then adjusting the fusion weight in the next round according to the value of the fusion weight in the current round.

[0020] Further preferably, the method further includes:

[0021] Adjusting the fusion weights in the multimodal fusion by using the regularization algorithm.

[0022] Preferably, the method further includes:

[0023] Constructing the neural network model by using the support vector machine algorithm and the random forest algorithm, and training the neural network model by using the 5-fold cross-validation method, the leave-one-out cross-validation method, and the callback function.

[0024] In a second aspect, an embodiment of the present invention further provides a multimodal information analysis device, where the device includes:

[0025] An information acquisition module, configured to acquire the electroencephalogram information, demographic information, and clinical information of a subject;

[0026] A feature extraction module, configured to extract the features in the electroencephalogram information, demographic information, and clinical information to obtain the electroencephalogram features, demographic features, and clinical features of the subject;

[0027] The fusion analysis module is configured to perform multimodal fusion on the electroencephalogram features, demographic features, and clinical features by using the evidence theory method within a preset number of rounds, and perform feature analysis on the features after the current multimodal fusion by using a neural network model;

[0028] The confirmation module is configured to determine the optimal feature analysis result as the final analysis result when the preset number of rounds is reached.

[0029] In the multimodal information analysis method provided in this application, within a preset number of rounds, multimodal fusion is performed on electroencephalogram features, demographic features, and clinical features by using the evidence theory method, and feature analysis is performed on the features after the current multimodal fusion by using a neural network model. When the preset number of rounds is reached, the optimal feature analysis result is determined as the final analysis result. This application makes full use of the multimodal feature fusion technology to fully integrate the complementary information in different data sources, improving the accuracy and reliability of the final analysis result.

[0030] Other features and advantages of this application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof are used to explain this application. In the drawings:

[0032] Figure 1 is a flowchart of the multimodal information analysis method of the preferred embodiment of the application;

[0033] Figure 2 is a comparison diagram of the electroencephalogram signal before and after preprocessing in the preferred embodiment of the application;

[0034] Figure 3 is a schematic diagram of the multimodal information analysis device of the preferred embodiment of the application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The technical solutions of this application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0036] First, this application proposes a multimodal information analysis method, as Figure 1 shown, including steps 110-140:

[0037] Step 110: Obtain the electroencephalogram information, demographic information, and clinical information of the subject;

[0038] Specifically, for the electroencephalogram information, it can be obtained through relevant electroencephalogram signal acquisition devices. For the demographic information, it can include basic personal information such as the age, gender, and education level of the subject. For the clinical information, it can include information such as the subject's physical disease history and marital status.

[0039] In a specific embodiment, after collecting the electroencephalogram (EEG) signals of a subject, it is necessary to preprocess the EEG signals to minimize or eliminate the influence of artifacts. The comparison of the states before and after signal preprocessing is shown as Figure 2 follows. The preprocessing includes: electrode screening, rereferencing, filtering, downsampling, segment marking, rejection of bad segments, and independent component analysis (ICA).

[0040] Among them, electrode screening includes locating electrode points and rejecting useless electrodes. Locating electrode points can ensure that the position of each electrode strictly follows the international 10 - 20 system to guarantee the spatial localization accuracy of the signals and the comparability of subsequent analyses. Rejecting useless electrodes can, after electrode positioning, identify and reject electrodes of non - EEG signals, such as electrooculogram (EOG), electrocardiogram (ECG), and electromyogram (EMG) signals, etc., to eliminate the influence of these interferences on the interpretation of EEG activities.

[0041] In rereferencing, the bilateral mastoids are used as reference points, or techniques such as common average reference and zero reference are adopted to reduce the influence of the reference electrode and improve the stability of the signals.

[0042] In filtering, by setting high - pass, low - pass, and notch filters, high - frequency noise and low - frequency drift in the signals, as well as interferences at specific frequencies, such as the frequency of the power supply line, can be effectively removed.

[0043] In downsampling, the sampling rate can be reduced according to research needs to reduce the data volume and improve processing efficiency, but it must be ensured that important frequency components are not lost during this process.

[0044] In segment marking, the EEG signals can be marked based on events to analyze the EEG activities under specific events or conditions.

[0045] In rejecting bad segments, bad segments with poor signal quality are rejected, and interpolation or rejection is performed on problematic electrodes to ensure the integrity of the signals.

[0046] In independent component analysis, the independent components in the EEG signals, including brain - source and non - brain - source components, can be separated, thereby effectively identifying and rejecting noise components such as EOG and ECG.

[0047] Step 120: Extract the features from the EEG information, demographic information, and clinical information to obtain the EEG features, demographic features, and clinical features of the subject;

[0048] In a specific embodiment, for the feature extraction of electroencephalogram (EEG) information, it includes extracting features such as power spectral density features, time-domain features (such as peak frequency, amplitude and duration of waveforms), frequency-domain features, and connection characteristics of EEG signals as EEG features. Among them, when extracting the power spectral density features of EEG signals, the Welch algorithm can be used to calculate the power spectral density corresponding to each channel in each frequency band.

[0049] For the feature extraction of demographic information, it includes extracting key demographic features such as age, gender, educational background, etc. as demographic features.

[0050] For the feature extraction of clinical information, it includes extracting scores such as the medical history of the subject, drug use records, Hamilton Depression Rating Scale (HDRS-17), Hamilton Anxiety Rating Scale (HAMA), and Young Mania Rating Scale (YMRS) as clinical features.

[0051] Moreover, after extracting the above features, the principal component analysis (PCA) method and / or variance threshold method can be used to screen the extracted features, aiming to make the data dimension sent to the neural network not so large, but without losing the main feature data at the same time. Step 130: Within a preset number of rounds, use the evidence theory method to perform multimodal fusion on EEG features, demographic features, and clinical features, and use the neural network model to perform feature analysis on the features after the current multimodal fusion;

[0052] Specifically, first, for multimodal fusion, when this application uses a neural network to analyze information, the data it analyzes is the data after multimodal feature fusion, which combines the feature information of three modalities: demographic information, clinical information, and EEG information. The multimodal feature fusion technology can make full use of the complementary information in different data sources to improve the accuracy and reliability of subsequent analysis.

[0053] This application uses the evidence theory method to quantify the uncertainty and information quality of different modal features, so that different modal features are fused under this method. The fused features then enter the neural network model for analysis to obtain the feature analysis result. Moreover, both fusion and analysis are performed in multiple rounds. Within the preset number of rounds, EEG features, demographic features, and clinical features will undergo multiple rounds of fusion and analysis, and the feature analysis results after each round of fusion and analysis will be recorded.

[0054] Use the evidence theory method to assign corresponding fusion weights to EEG features, demographic features, and clinical features, and fuse each modal feature according to the fusion weights of different modal features.

[0055] In a specific embodiment, the fusion weights are randomly assigned, and in each round of fusion, the system records the fusion weights with the optimal feature parsing result at the current time.

[0056] In a more optimal embodiment, when the fusion weights are assigned, they are adjusted according to the current optimal feature parsing result.

[0057] More specifically, in the first round of fusion, initial fusion weights are assigned to the electroencephalogram features, demographic features, and clinical features. Under this initial fusion weight, the neural network model will parse out the feature parsing result of the first round, and this result will be automatically recorded. In subsequent rounds of fusion, if the feature parsing result corresponding to the fusion weight of the current round is better than the feature parsing result corresponding to the fusion weight of the previous round, then the system will automatically record the fusion weight of the current round and adjust the fusion weight in the next round according to the value of the fusion weight of the current round, so as to optimize the operation efficiency of the system, accelerate model convergence, and thus accelerate the parsing process. For example, assume that in the first round, the fusion weights of the electroencephalogram features, demographic features, and clinical features are "0.3", "0.3", and "0.4" respectively. In the second round, the fusion weights of the electroencephalogram features, demographic features, and clinical features are "0.4", "0.3", and "0.3" respectively, and the feature parsing result of the second round is better than that of the first round. Then in the third round, the weight of the electroencephalogram features is increased accordingly, the weight of the demographic features remains unchanged, and the weight of the clinical features is decreased accordingly, and so on until the preset number of rounds is reached.

[0058] In a more optimal embodiment, in order to avoid overfitting during the feature fusion process, the present application also uses a regularization algorithm to optimize and regulate the fusion weights in multimodal fusion. By adding a penalty term for the fusion weights to the loss function, the complexity of the model is restricted, thereby improving the generalization ability. The regularization algorithm can include L1 and L2 regularization algorithms, and the present application does not make a limitation.

[0059] Second, for the neural network model, the present application uses the Support Vector Machine (SVM) algorithm and the Random Forests (RF) algorithm to construct the neural network model. The architecture of this neural network model can include a fully connected neural network, a recurrent neural network, a long short-term memory network, and a Transformer network. Preferably, it only includes a long short-term memory network and a Transformer network. As long as the network can meet the requirements of parsing features, the present application does not make a limitation.

[0060] Moreover, when training the neural network model, the present application adopts the 5-Fold cross-validation method, the Leave-One-Out Cross-Validation (LOO-CV) method, and the Call Back function to improve the model accuracy.

[0061] Among them, the 5-Fold cross-validation method randomly divides the dataset into five subsets of equal size. In each experiment, one subset is used as the test set, and the remaining four are used as the training set. This process is repeated five times, and each subset takes turns as the test set once.

[0062] The Leave-One-Out Cross-Validation method is that each sample in the dataset takes turns as the test set, and all the remaining samples are used as the training set. Specifically, in a dataset containing N samples, this method will perform N experiments. Each time, one sample is selected as the test set, and the remaining N - 1 samples are used as the training set to evaluate the model performance.

[0063] The Call Back function can specifically include the EarlyStopping function and the ReduceLROnPlateau function. The EarlyStopping function allows the model to monitor a specific metric on the validation set, such as loss or accuracy, and terminate the training early when the metric has not improved for a consecutive number of training epochs. This "patience" parameter defines the number of epochs to monitor, providing overfitting protection and computational resource savings for the model. The ReduceLROnPlateau function acts as an adaptive learning rate scheduler, dynamically adjusting the learning rate by monitoring the performance metric on the validation set. If the performance of the model does not show improvement within a certain number of training epochs, the learning rate will be reduced to carefully adjust the model weights, which helps the model jump out of local minima and continue to optimize. The reduction factor and patience parameter of this method allow researchers to customize the specific behavior of the learning rate reduction to optimize the training process and achieve more accurate model predictions.

[0064] Step 140: When the preset number of rounds is reached, determine the optimal feature parsing result as the final parsing result.

[0065] Specifically, when the number of times of fusion and parsing reaches the preset number of rounds, the fusion and parsing of features terminate, and an optimal feature parsing result is determined among all rounds, and this result is the final parsing result.

[0066] In the multi-modal information parsing method provided by this application, within a preset number of rounds, the Dempster-Shafer theory method is used to perform multi-modal fusion on electroencephalogram features, demographic features, and clinical features, and a neural network model is used to perform feature parsing on the features after the current multi-modal fusion. When the preset number of rounds is reached, the optimal feature parsing result is determined as the final parsing result. This application utilizes multi-modal feature fusion technology to fully integrate complementary information from different data sources, improving the accuracy and reliability of the final parsing result.

[0067] In addition, this application also provides a multi-modal information parsing device, as Figure 3 shown. The device includes:

[0068] An information acquisition module 310, configured to acquire the electroencephalogram information, demographic information, and clinical information of a subject;

[0069] A feature extraction module 320, configured to extract the features in the electroencephalogram information, demographic information, and clinical information to obtain the electroencephalogram features, demographic features, and clinical features of the subject;

[0070] A fusion and parsing module 330, configured to perform multi-modal fusion on the electroencephalogram features, demographic features, and clinical features using the Dempster-Shafer theory method within a preset number of rounds, and perform feature parsing on the features after the current multi-modal fusion using a neural network model;

[0071] A confirmation module 340, configured to determine the optimal feature parsing result as the final parsing result when the preset number of rounds is reached.

[0072] Other preferred embodiments of the multi-modal information parsing device disclosed in this application and the achievable technical effects are the same as those of the above multi-modal information parsing method, and will not be elaborated here.

[0073] The preferred embodiments of this application have been described in detail above. However, this application is not limited to the specific details in the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application.

[0074] In addition, it should be noted that, in the case of no contradiction, the various specific technical features described in the above specific embodiments can be combined in any suitable manner. To avoid unnecessary repetition, this application will not separately describe various possible combination methods.

[0075] Furthermore, any combination can be made between various different embodiments of this application, as long as it does not violate the idea of this application, and it should also be regarded as the content disclosed in this invention.

Claims

1. A multimodal information parsing method, characterized in that, The method includes: Obtaining electroencephalogram information, demographic information, and clinical information of a subject; Extracting features from the electroencephalogram information, demographic information, and clinical information to obtain electroencephalogram features, demographic features, and clinical features of the subject; Within a preset number of rounds, using the evidence theory method to perform multimodal fusion on the electroencephalogram features, demographic features, and clinical features, and using a neural network model to perform feature analysis on the features after the current multimodal fusion; When the preset number of rounds is reached, determining the optimal feature analysis result as the final analysis result.

2. The method according to claim 1, wherein Before obtaining the electroencephalogram information of the subject, the method further includes: Collecting the electroencephalogram signals of the subject and preprocessing the electroencephalogram signals; the preprocessing includes: electrode screening, rereferencing, filtering, downsampling, segment marking, bad segment removal, and independent component analysis.

3. The method according to claim 1, wherein Extracting features from the electroencephalogram signals, including: Extracting power spectral density features, time-domain features, and frequency-domain features from the electroencephalogram signals.

4. The method according to claim 3, wherein Extracting the power spectral density features from the electroencephalogram signals, including: Using the Welch algorithm to calculate the power spectral density corresponding to each channel in each frequency band to obtain the power spectral density features in the electroencephalogram signals.

5. The method according to claim 3, wherein After extracting features from the electroencephalogram information, demographic information, and clinical information to obtain electroencephalogram features, demographic features, and clinical features of the subject, the method further includes: Using the principal component analysis method and / or the variance threshold method to screen the extracted features.

6. The method according to claim 1, wherein Within a preset number of rounds, using the evidence theory method to perform multimodal fusion on the electroencephalogram features, demographic features, and clinical features, including: Using the evidence theory method to determine the fusion weights of the electroencephalogram features, demographic features, and clinical features in each round of multimodal fusion, and performing multimodal fusion on the electroencephalogram features, demographic features, and clinical features according to the fusion weights.

7. The method according to claim 6, characterized in that If the feature analysis result corresponding to the fusion weight in the current round is better than the feature analysis result corresponding to the fusion weight in the previous round, then adjust the fusion weight in the next round according to the value of the fusion weight in the current round.

8. The method according to claim 6 or 7, characterized in that The method further includes: Using a regularization algorithm to adjust the fusion weights in multimodal fusion.

9. The method according to claim 1, wherein The method further includes: Using the support vector machine algorithm and the random forest algorithm to construct the neural network model, and using the 5-fold cross-validation method, the leave-one-out cross-validation method, and the callback function to train the neural network model.

10. A multimodal information parsing device, characterized in that, The device includes: An information acquisition module configured to obtain electroencephalogram information, demographic information, and clinical information of a subject; A feature extraction module configured to extract features from the electroencephalogram information, demographic information, and clinical information to obtain electroencephalogram features, demographic features, and clinical features of the subject; A fusion analysis module configured to, within a preset number of rounds, use the evidence theory method to perform multimodal fusion on the electroencephalogram features, demographic features, and clinical features, and use a neural network model to perform feature analysis on the features after the current multimodal fusion; A confirmation module configured to, when the preset number of rounds is reached, determine the optimal feature analysis result as the final analysis result.