Alzheimer's disease multi-mode classification method based on electroencephalogram, brain image and genetic information

Through a multimodal classification method based on EEG, brain imaging and genetic information, various data characteristics are extracted and fused, and the problem of insufficient recognition accuracy in early detection of Alzheimer's disease is solved, achieving higher accuracy and speed of disease discrimination.

CN120011879APending Publication Date: 2025-05-16NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510065021.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

How to effectively utilize the synergistic effect between EEG, brain imaging and genetic information to improve the recognition accuracy of Alzheimer's disease, especially in early detection.

Method used

A multimodal classification method is proposed, which extracts their respective features by obtaining and preprocessing EEG, brain images and gene data, and performs standardization, feature selection and fusion, and finally disease identification is performed based on fusion characteristics.

Benefits of technology

It improves the recognition accuracy of Alzheimer's disease, provides a richer and comprehensive information analysis process, enhances the accuracy and speed of disease diagnosis, and achieves the wide applicability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011879A_ABST
    Figure CN120011879A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing and recognition, and provides an Alzheimer's disease multi-mode classification method based on electroencephalogram, brain image and genetic information, and the method comprises the steps: obtaining electroencephalogram data, brain image data and genetic data of a subject, and carrying out the corresponding preprocessing operation; performing feature extraction on the preprocessed electroencephalogram data, brain image data and gene data to obtain electroencephalogram signal features, brain image features and AD susceptibility features of the subject; respectively carrying out corresponding standardization processing on the features of the three different modes of the subject, then carrying out feature selection, eliminating irrelevant features and redundant features, and finally carrying out feature fusion to obtain fused features; according to the method, binary classification and multi-classification are carried out on the basis of the features of the three different modes after feature fusion and feature selection, and a disease discrimination result of the subject is obtained on the basis of a binary classification result and a multi-classification result, so that the Alzheimer's disease recognition precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing and recognition technology, and in particular to a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information. Background Art

[0002] Alzheimer's Disease (AD) is a complex neurodegenerative disease that poses a major challenge to global public health. As the disease worsens, patients will experience mental and cognitive disorders, memory loss, and behavioral changes, which seriously affect the patient's ability to live a normal life and have a profound impact on both individuals and society. AD is a progressive disease, and mild cognitive impairment (MCI) is a state between healthy aging and AD, which can be considered the early stage of AD. Nearly 10% to 15% of MCI patients are converted to AD patients each year. Therefore, accurate diagnosis, including MCI and even earlier stages, is crucial to slowing the progression of the disease and minimizing the prevalence and incidence.

[0003] Due to the complexity of AD's manifestations, a single biomarker may not provide enough information to accurately determine the severity of the disease, so a multimodal biomarker system is needed. In addition to integrating different data modalities, an important consideration is that the relationships and constraints between the data modalities are often clinically meaningful in clinical practice. In the many studies using AD biomarkers, biomarkers such as neuroimaging, neurophysiology, and genetics contain rich information about the active pathophysiological processes in AD and play an important role in the detection of precursor clues to the disease before symptoms appear.

[0004] The increasing development of neuroimaging has brought new vitality to the study of spatial information of human brain structure. Most of the studies on multimodal biomarkers for early detection rely on structural neuroimaging tools, and only a few have studied the integration of structural neuroimaging and electroneurophysiology. Electroencephalography (EEG) is a low-cost, non-invasive, portable technology with high temporal resolution and has become a well-known neurophysiological biomarker for AD classification. EEG reflects the superposition of electromagnetic fields generated by the interaction of cortical neurons at the macroscopic level, and the behavior of the underlying neuronal population can be studied indirectly through EEG. Compared with healthy controls, MCI / AD EEG patterns are significantly altered, manifested by slowing of EEG rhythm, loss of synchronization between electrode pairs, and loss of complexity. Such EEG abnormalities are highly correlated with cognitive impairment and AD, making EEG a promising marker for early detection of AD, and more and more studies are devoted to extracting relevant features from MCI / AD EEG signals. In addition, EEG can provide real-time, non-invasive dynamic assessment of the brain, which can complement the static information provided by magnetic resonance imaging (MRI). By capturing ongoing neural activity, EEG can detect early functional changes that precede structural changes detectable by MRI, thereby improving the accuracy of AD identification.

[0005] Genetic factors play a crucial role in the pathogenesis of AD, as evidenced by genome-wide association studies that have identified specific genetic variants associated with AD susceptibility. Single nucleotide polymorphisms (SNPs) in genes such as APOE, APOC1, and TOMM40 have been associated with an increased risk of AD, while the polygenic risk score (PRS) represents the overall effect of multiple genetic variants on predicting AD risk. Integrating genetic data into AD classification models can provide valuable insights into disease mechanisms and help identify individuals at high risk for AD. In addition, the combination of genetic data and neuroimaging features has been shown to improve the accuracy of AD classification, suggesting that there may be synergies between these modalities.

[0006] However, how to utilize the synergy between the temporal information of EEG signals, the spatial information of imaging phenotypes, and genetic information, while fully considering the heterogeneity of the data and the robustness of the model, is a major challenge in AD and its early detection. No research team has yet proposed an effective solution. Summary of the invention

[0007] In view of this, an embodiment of the present application proposes a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information, which realizes the characteristic analysis of brain time, space and genetic information through the characteristics of three different modalities, thereby effectively improving the recognition accuracy of Alzheimer's disease.

[0008] In a first aspect, an embodiment of the present application proposes a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information, the method comprising the following steps: obtaining electroencephalogram data, brain image data and genetic data of a subject, and performing corresponding preprocessing operations on the acquired electroencephalogram data, brain image data and genetic data of the subject, respectively; performing feature extraction on the preprocessed electroencephalogram data to obtain the subject's electroencephalogram signal features, performing feature extraction on the preprocessed brain image data to obtain the subject's brain image features, and performing feature extraction on the preprocessed genetic data to obtain the subject's AD susceptibility features; performing corresponding standardization processing on the features of the three different modalities of the subject, respectively, and then performing feature selection on the features of the three different modalities after the standardization processing to exclude irrelevant features and redundant features, and finally performing feature fusion on the features of the three different modalities after feature selection to obtain fused features; performing multiple binary classifications and multi-classifications based on the fused features and the features of the three different modalities after feature selection, and obtaining a disease discrimination result for the subject based on the results of the multiple binary classifications and the results of the multi-classifications.

[0009] Optionally, corresponding preprocessing operations are performed on the acquired EEG data, brain imaging data and gene data of the subject, respectively, including:

[0010] The EEG data of the subjects were band-pass filtered between 0.5 Hz and 55 Hz, and then filtered using a 50 Hz notch filter to eliminate power line interference. Next, the filtered EEG data were subjected to wavelet independent component analysis to eliminate unnecessary eye movement and muscle activity-related artifacts in multiple channels. Finally, the analyzed EEG data were segmented into a series of non-overlapping time period data using a 5-second time window, and the main frequency components were extracted from each time period data to obtain the preprocessed EEG data;

[0011] The brain imaging data of the subjects were subjected to voxel-based morphological analysis using SPM software. All scanned images were aligned with the T1-weighted template image, segmented into gray matter, white matter, and cerebrospinal fluid maps, and standardized to 2 × 2 × 2 mm in the standard space of MNI. 3 The voxels were smoothed using a FWHM kernel of 8 mm, and finally the VBM-sMRI scan images were registered to the same MNI space to obtain the preprocessed brain imaging data;

[0012] The genetic data of the subjects were genotyped using the human 610-Quad or OmniExpress Array, and then standard quality control including deletion rate, minor allele frequency and Harwin equilibrium filtering was performed to filter low-quality variant sites and samples. Finally, IMPUTE2 software was used for gene filling to obtain preprocessed genetic data.

[0013] Optionally, the EEG signal features are specifically composed of four feature parameters: power spectrum density, Hjorth parameter, sample entropy and time-frequency feature. The power spectrum density specifically includes absolute PSD feature and relative PSD feature. The Hjorth parameter specifically includes activity parameter H Activity , mobility parameter H Mobility and complexity parameter H Complexity , extract features from the preprocessed EEG data to obtain the EEG signal features of the subject, including:

[0014] The absolute PSD of the preprocessed EEG data is calculated using the Welch periodogram method to obtain the absolute PSD feature PSD ab , PSD ab The calculation is achieved through the following formula:

[0015]

[0016] G n =a n ×exp[-(Fc) 2 / 2w 2 ];

[0017] L = b-log(K+F x );

[0018] Where N represents the total number of peaks extracted from the power spectrum of the preprocessed EEG data, G n represents the Gaussian function fitting the nth peak, L represents the non-periodic or background component, a n represents the peak power of the nth peak, F represents the input frequency vector, c represents the center frequency, w represents the Gaussian standard deviation, b represents the broadband offset, χ represents the exponential parameter, and K represents the bending parameter for controlling the bending of the non-periodic component;

[0019] Calculate the ratio between the PSD of the specific frequency band and the total frequency band of the power spectrum of the preprocessed EEG data, and use the calculated ratio as the relative PSD feature PSD re ;

[0020] Let z(t) denote the time series of the preprocessed EEG data of each channel, and the activity parameter H Activity , mobility parameter H Mobility and complexity parameter HComplexity , calculated by the following formula:

[0021] H Activity =var[z(t)];

[0022]

[0023] H Complexity =H Mobility [dz(t) / dt] / H Mobility [y(t)];

[0024] Where var(·) represents the variance operation, dz(t) / dt represents the first-order derivative of z(t), and y(t) represents a pure sine wave;

[0025] Sample entropy is defined as the negative logarithm of whether m consecutive data points of two similar sequences remain similar at the m+1th point. Sample entropy is calculated by the following formula:

[0026] SE(m, r)=-ln[C m+1 (r) / C m (r)];

[0027] Among them, r is the preset amplitude threshold used to separate the similarity between two vectors, C m+1 (r) represents the number of vector pairs whose vector distance is less than r, C m (r) represents the total number of vector pairs, SE(m, r) represents the sample entropy;

[0028] The time-frequency characteristics are calculated by the following formula:

[0029]

[0030] Wherein, ω(t) is a preset window function, specifically a Hann window or Gaussian window centered on zero, and STFT{z(t)} represents the time-frequency characteristics.

[0031] Optionally, feature extraction is performed on the preprocessed brain image data to obtain brain image features of the subject, including:

[0032] The whole brain was subsampled based on the subject's brain imaging features, and 116 imaging measurements of regions of interest were generated based on the MarsBaR automatic anatomical labeling atlas, including mean gray matter density on structural MRI, amyloid values ​​on AV45 scans, and glucose utilization on FDG scans;

[0033] These imaging measures were adjusted using regression weights from healthy individuals to eliminate the effects of baseline age, sex, education, and handedness, ultimately yielding the subjects' brain imaging characteristics.

[0034] Optionally, feature extraction is performed on the preprocessed gene data to obtain AD susceptibility features of the subject, including:

[0035] All SNPs in the preprocessed genetic data were subjected to binary Logistic GWAS analysis using Plink software. The phenotype of the GWAS analysis was AD diagnosis status, and the genetic effect value of each gene locus of the subjects was obtained. The top 100 SNPs were then screened as feature candidates based on the genetic effect value.

[0036] PRSice-2 software was used to calculate the PRS of all feature candidates to obtain the AD susceptibility characteristics of the subjects.

[0037] Optionally, the features of the three different modalities of the subject are respectively standardized, and then feature selection is performed on the features of the three different modalities after the standardization, and irrelevant features and redundant features are excluded. Finally, feature fusion is performed on the features of the three different modalities after feature selection to obtain fused features, including:

[0038] The features of the three different modes of the subjects are scaled to the range of [-1, 1] respectively to obtain the features of the three different modes after standardization; wherein the features of each mode are scaled according to their maximum absolute value;

[0039] The Spearman correlation index is used to perform feature selection on the features of the three different modes after the standardization process. First, the features whose Spearman correlation index is lower than the preset first correlation threshold are retained to exclude irrelevant features and redundant features. Then, the features whose Spearman correlation index is higher than the preset second correlation threshold are retained to obtain the features of the three different modes after feature selection.

[0040] The features of the three different modalities after feature selection are connected by feature matrices to obtain fusion features.

[0041] Optionally, the binary classification includes binary classification of AD and HC, binary classification of AD and MCI, and binary classification of MCI and HC, and the multi-classification specifically includes multi-classification of AD, HC, and MCI;

[0042] Multiple binary and multi-classifications based on the features of three different modalities after fusion features and feature selection are achieved through a pre-trained classifier. The classifier consists of an input layer, a hidden layer and a softmax classification output layer. The size of the hidden layer is selected by evaluating the minimum reconstruction error.

[0043] The embodiment of the present application proposes a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information. The diagnosis and classification of Alzheimer's patients are carried out by combining electroencephalogram data, brain imaging data and genetic data. The characteristics of the brain's time, space and genetic information are analyzed by the characteristics of three different modalities. Compared with the method of using a single modality for analysis, it has a richer and more comprehensive information analysis process and a more complete dimension. The method is non-invasive and repeatable. Through different training processes for different populations, it can achieve wide applicability to different populations. The method uses different preprocessing and feature extraction methods for data of different modalities to ensure that the data of each modality can retain their own feature data and can be better integrated. The method performs AD susceptibility analysis of the subject based on the genetic data and genetic information of the subject, and well retains the important parts of the genetic data and genetic information, thereby effectively improving the accuracy of the subject's disease discrimination. The method well excludes irrelevant and redundant features through feature selection, reduces the feature dimension, and effectively improves the speed of disease discrimination of the subject.

[0044] In a second aspect, an embodiment of the present application proposes a multimodal classification system for Alzheimer's disease based on electroencephalogram, brain image and genetic information, the system comprising: a data acquisition module, a preprocessing module, a feature extraction module, a feature selection fusion module, and a discrimination module; the data acquisition module is used to acquire the electroencephalogram data, brain image data and genetic data of the subject; the preprocessing module is used to perform corresponding preprocessing operations on the acquired electroencephalogram data, brain image data and genetic data of the subject respectively; the feature extraction module is used to perform feature extraction on the preprocessed electroencephalogram data to obtain the subject's electroencephalogram signal features, and perform feature extraction on the preprocessed brain image data. The brain image features of the subject are obtained, and the features of the preprocessed genetic data are extracted to obtain the AD susceptibility features of the subject; the feature selection and fusion module is used to perform corresponding standardization processing on the features of the three different modalities of the subject respectively, and then perform feature selection on the features of the three different modalities after the standardization processing to exclude irrelevant features and redundant features, and finally perform feature fusion on the features of the three different modalities after feature selection to obtain fused features; the discrimination module is used to perform multiple binary classifications and multi-classifications based on the fused features and the features of the three different modalities after feature selection, and obtain the disease discrimination results of the subject based on the results of multiple binary classifications and multi-classifications.

[0045] In a third aspect, an embodiment of the present application proposes a terminal / server / electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information as described in the first aspect above.

[0046] In a fourth aspect, an embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information as described in the first aspect above.

[0047] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the drawings required for use in the embodiments of the present application or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 is a flowchart of a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information provided in one embodiment of the present application;

[0050] Figure 2 is a schematic diagram of a specific implementation of a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information provided in an embodiment of the present application;

[0051] Figure 3 is a schematic diagram of the structure of a multimodal classification system for Alzheimer's disease based on electroencephalogram, brain image and genetic information provided in another embodiment of the present application;

[0052] Figure 4 It is a structural schematic diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. In the various embodiments of the present application, many technical details are proposed in order to make the reader better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is only for the convenience of description, and the specific implementation mode of the present application should not constitute any limitation. The various embodiments can be combined with each other and referenced to each other under the premise of no contradiction.

[0054] An embodiment of the present application proposes a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information, which is applied to an electronic device, wherein the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described using the server as an example. The implementation details of the multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information proposed in this embodiment are specifically described below. The following content is only the implementation details provided for ease of understanding and is not necessary for the implementation of this solution.

[0055] The specific process of the multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information proposed in this embodiment can be as follows: Figure 1 As shown, including:

[0056] Step 101, obtaining the EEG data, brain imaging data and gene data of the subject, and performing corresponding preprocessing operations on the obtained EEG data, brain imaging data and gene data of the subject respectively.

[0057] In the specific implementation, the server needs to obtain the subject's EEG data, brain imaging data and genetic data, and then perform corresponding preprocessing operations on the EEG data, brain imaging data and genetic data to obtain preprocessed EEG data, preprocessed brain imaging data, and preprocessed genetic data. EEG data needs to be collected using professional EEG acquisition equipment, brain imaging data can be collected by a head CT machine, and genetic data can be collected by blood tests.

[0058] In one example, when performing corresponding preprocessing operations on the EEG data of the subject, the EEG data of the subject first needs to be bandpass filtered between 0.5Hz and 55Hz, and then filtered using a 50Hz notch filter to eliminate power line interference. Next, the filtered EEG data is subjected to wavelet independent component analysis to eliminate unnecessary eye movement and muscle activity-related artifacts in multiple channels. Finally, the analyzed EEG data is divided into a series of non-overlapping time period data using a 5-second time window, and the main frequency components are extracted from each time period data to obtain the preprocessed EEG data. Among them, the main frequency components extracted by the server include delta component (1Hz to 4Hz), theta component (4Hz to 8Hz), alpha component (8Hz to 13Hz), beta component (13Hz to 30Hz), and gamma component (30Hz to 45Hz). The preprocessing operation of segmented processing deeply considers the possible stability of the representative information of a specific time period of the EEG data. In order to reduce the complexity of the data and reduce the computing and storage requirements while ensuring that the main features of the data are retained, the preprocessing of the EEG data needs to further perform segmented approximate aggregation operations on each channel and each time window of the EEG signal, and finally obtain a data set of M segments and N time points for the subject.

[0059] In one example, when performing corresponding preprocessing operations on the brain imaging data of the subject, the SPM software is first used to perform voxel-based morphological analysis on the brain imaging data of the subject. All scanned images are aligned with the T1-weighted template image and segmented into gray matter, white matter, and cerebrospinal fluid images. Then, the images are standardized to 2×2×2mm in the standard space of MNI. 3 The voxels were smoothed using a FWHM kernel of 8 mm. Finally, the VBM-sMRI scan images were registered to the same MNI space to obtain the preprocessed brain image data.

[0060] In one example, when the server performs corresponding preprocessing operations on the subject's genetic data, it first needs to use the human 610-Quad or OmniExpress Array to genotype the subject's genetic data. Then, standard quality control including deletion rate, minor allele frequency and Harwin equilibrium filtering is performed to filter low-quality variant sites and samples. Finally, IMPUTE2 software is used to perform gene filling to obtain preprocessed genetic data. The reference sample can be selected from the widely used 1000G database (1000Genome Project).

[0061] In one example, when performing model training, the server needs to select three groups of people: AD subjects, MCI subjects, and HC (Health Control) subjects, and then obtain the EEG data, brain imaging data, and genetic data of these three groups of subjects respectively.

[0062] Step 102, feature extraction is performed on the preprocessed EEG data to obtain the EEG signal features of the subject, feature extraction is performed on the preprocessed brain image data to obtain the brain image features of the subject, and feature extraction is performed on the preprocessed gene data to obtain the AD susceptibility features of the subject.

[0063] In the specific implementation, after the server completes the preprocessing of the data of three different modalities, it can perform feature extraction respectively: perform feature extraction on the preprocessed EEG data to obtain the subject's EEG signal characteristics, perform feature extraction on the preprocessed brain image data to obtain the subject's brain image characteristics, and perform feature extraction on the preprocessed genetic data to obtain the subject's AD susceptibility characteristics.

[0064] In one example, the EEG signal features are specifically composed of four feature parameters: power spectrum density, Hjorth parameter, sample entropy, and time-frequency feature. The power spectrum density specifically includes absolute PSD feature and relative PSD feature. The Hjorth parameter specifically includes activity parameter H Activity , mobility parameter H Mobility and complexity parameter H Complexity .

[0065] The server first calculates the absolute PSD of the preprocessed EEG data using the Welch periodogram method to obtain the absolute PSD feature PSD ab , PSD ab The calculation is achieved through the following formula:

[0066]

[0067] G n =a n ×exp[-(Fc) 2 / 2w 2 ];

[0068] L = b-log(K+F x );

[0069] Where N represents the total number of peaks extracted from the power spectrum of the preprocessed EEG data, G n represents the Gaussian function fitting the nth peak, L represents the non-periodic or background component, L is modeled based on the Lorentzian function, and a nrepresents the peak power of the nth peak, expressed as log10P n Remember, P n represents the power value, F represents the input frequency vector, c represents the center frequency in Hz, w represents the Gaussian standard deviation, b represents the broadband offset, χ represents the exponential parameter, and K represents the bending parameter for controlling the bending of the non-periodic component.

[0070] Subsequently, the ratio between the PSD of the specific frequency band and the total frequency band of the power spectrum of the preprocessed EEG data is calculated, and the calculated ratio is used as the relative PSD feature PSD re .

[0071] Let z(t) denote the time series of the preprocessed EEG data of each channel, and the activity parameter H Activity , mobility parameter H Mobility and complexity parameter H Complexity , calculated by the following formula:

[0072] H Activity =var[z(t)];

[0073]

[0074] H Complexity =H Mobility [dz(t) / dt] / H Mobility [y(t)];

[0075] Where var(·) represents the variance operation, dz(t) / dt represents the first-order derivative of z(t), and y(t) represents a pure sine wave.

[0076] That is, H Activity It represents the signal power and variance of the time series, and further represents the surface of the power spectrum in the frequency domain. Mobility Represents the ratio of the mean frequency or standard deviation of the power spectrum, defined as the variance of the first derivative of z(t) divided by the square root of the variance of z(t). Complexity Indicates the change in the signal's frequency. This parameter compares the signal's similarity to a pure sine wave and converges to 1 if the signal is more similar.

[0077] Sample entropy is defined as the negative logarithm of whether m consecutive data points of two similar sequences remain similar at the m+1th point. Sample entropy is calculated by the following formula:

[0078] SE(m, r)=-ln[C m+1 (r) / C m (r)];

[0079] Among them, r is the preset amplitude threshold used to separate the similarity between two vectors, Cm+1 (r) represents the number of vector pairs whose vector distance is less than r, C m (r) represents the total number of vector pairs, and SE(m, r) represents the sample entropy.

[0080] Time-frequency features refer to how the frequency components of a signal change over time by analyzing it through STFT. Specifically, STFT divides a longer signal into shorter segments of the same length and calculates the Fourier transform of each shorter segment. The time-frequency features are calculated using the following formula:

[0081]

[0082] Wherein, ω(t) is a preset window function, specifically a Hann window or Gaussian window centered on zero, and STFT{z(t)} represents the time-frequency characteristics.

[0083] In one example, when the server extracts features from the preprocessed brain imaging data, it first subsamples the entire brain based on the subject's brain imaging features, and generates imaging measurements of 116 regions of interest based on the MarsBaR automatic anatomical labeling atlas, including the average gray matter density of structural MRI, the amyloid value of AV45 scans, and the glucose utilization of FDG scans. Then, using regression weights from healthy individuals (i.e., HC subjects), these imaging measurements are adjusted to eliminate the effects of baseline age, gender, education level, and handedness, and finally the subject's brain imaging features are obtained.

[0084] In one example, when the server extracts features from preprocessed genetic data, it first uses Plink software to perform a binary Logistics GWAS analysis on all SNPs in the preprocessed genetic data. The phenotype of the GWAS analysis is the AD diagnostic status (0 for HC, 1 for AD), thereby obtaining the genetic effect value of each gene locus of the subject, and then screening the top 100 SNPs as feature candidates based on the genetic effect value. Finally, the PRSice-2 software is used to calculate the PRS of all feature candidates to obtain the AD susceptibility characteristics of the subject.

[0085] In step 103, the features of the three different modalities of the subject are respectively standardized, and then feature selection is performed on the features of the three different modalities after the standardization processing to exclude irrelevant features and redundant features. Finally, feature fusion is performed on the features of the three different modalities after the feature selection to obtain fused features.

[0086] In the specific implementation, after the server completes the EEG signal features, brain image features and AD susceptibility features, it is necessary to perform corresponding standardization processing on the features of the three different modalities of the subjects respectively, and then perform feature selection on the features of the three different modalities after standardization to exclude irrelevant features and redundant features, and finally perform feature fusion on the features of the three different modalities after feature selection to obtain fused features.

[0087] In one example, the server scales the features of the three different modalities of the subject to the range of [-1, 1] respectively, and obtains the features of the three different modalities after standardization; wherein the features of each modality are scaled according to their own maximum absolute value. The Spearman correlation index (denoted by ρ) is then used to perform feature selection on the features of the three different modalities after standardization. First, the features whose Spearman correlation index is lower than the preset first correlation threshold (generally set to 0.8) are retained to exclude irrelevant features and redundant features, and then the features whose Spearman correlation index is higher than the preset second correlation threshold (generally set to 0.15) are retained to obtain the features of the three different modalities after feature selection. Finally, the features of the three different modalities after feature selection are connected by feature matrix to obtain fusion features.

[0088] Step 104 , performing multiple binary classifications and multi-classifications based on the fused features and the features of the three different modalities after feature selection, and obtaining disease discrimination results for the subject based on the results of the multiple binary classifications and multi-classifications.

[0089] In the specific implementation, after obtaining the fused features, the server needs to perform multiple binary classifications and multi-classifications based on the fused features and the features of the three different modalities after feature selection, and then obtain the disease identification results for the subject based on the results of multiple binary classifications and multi-classifications.

[0090] In an example, the binary classification includes the binary classification of AD and HC, the binary classification of AD and MCI, and the binary classification of MCI and HC, and the multi-classification includes the multi-classification of AD, HC, and MCI.

[0091] In one example, multiple binary and multi-classifications based on the features of three different modalities after fusion and feature selection are implemented by a pre-trained classifier, which consists of an input layer, a hidden layer, and a softmax classification output layer. The size of the hidden layer is selected by evaluating the minimum reconstruction error. The classifier is trained for 1000 epochs until the cross entropy loss function converges.

[0092] In one example, a server can quantify the effectiveness of a classifier using the following metrics:

[0093] PRECISION(prc)=TP / (TP+FP);

[0094] RECALL(rec)=TP / (TP+FN);

[0095] F-SCORE(fsc)=2(PC×RC) / (PC+RC);

[0096] ACCURACY(acc)=(TP+TN) / (fP+TN+FP+FN);

[0097] Among them, TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative, respectively. In addition, the server used a K-fold cross-validation procedure (K = 10), in which 70% of the data was used as a training set in each fold and the remaining 30% of the data was used as a test set. Therefore, since the overall classification performance was evaluated by estimating the average (avg) and standard deviation (std) of 10 folds, all results were expressed as average ± standard deviation, i.e., avg(acc) ± std(acc).

[0098] In this embodiment, the diagnosis and classification of Alzheimer's patients are carried out by combining EEG data, brain imaging data and genetic data, and the characteristic analysis of the brain's time, space and genetic information is realized through the characteristics of three different modalities. Compared with the method of using a single modality for analysis, it has a richer and more comprehensive information analysis process and a more complete dimension. This method is non-invasive and repeatable. Through different training for different populations, it can achieve wide applicability to different populations. This method uses different preprocessing and feature extraction methods for data of different modalities to ensure that the data of each modality can retain their own feature data and can be better integrated. This method analyzes the AD susceptibility of the subject based on the genetic data and genetic information of the subject, and well retains the important parts of the genetic data and genetic information, thereby effectively improving the accuracy of the subject's disease identification. This method eliminates irrelevant and redundant features through feature selection, reduces the feature dimension, and effectively improves the speed of disease identification of the subject.

[0099] The step division of the above methods is only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0100] Another embodiment of the present application proposes a multimodal classification system for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information. The implementation details of the multimodal classification system for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information proposed in this embodiment are described in detail below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this example.

[0101] Figure 3 20 is a schematic diagram of the structure of a multimodal classification system for Alzheimer's disease based on electroencephalogram, brain image and genetic information proposed in this embodiment. The system includes: a data acquisition module 201, a preprocessing module 202, a feature extraction module 203, a feature selection and fusion module 204, and a discrimination module 205.

[0102] The data acquisition module 201 is used to acquire the EEG data, brain imaging data and gene data of the subject.

[0103] The preprocessing module 202 is used to perform corresponding preprocessing operations on the acquired EEG data, brain image data and gene data of the subject.

[0104] The feature extraction module 203 is used to perform feature extraction on the preprocessed EEG data to obtain the EEG signal features of the subject, perform feature extraction on the preprocessed brain image data to obtain the brain image features of the subject, and perform feature extraction on the preprocessed gene data to obtain the AD susceptibility features of the subject.

[0105] The feature selection and fusion module 204 is used to perform corresponding standardization processing on the features of the three different modalities of the subject respectively, and then perform feature selection on the features of the three different modalities after the standardization processing to exclude irrelevant features and redundant features, and finally perform feature fusion on the features of the three different modalities after the feature selection to obtain fused features.

[0106] The discrimination module 205 is used to perform multiple binary classifications and multi-classifications based on the fused features and the features of the three different modalities after feature selection, and obtain disease discrimination results for the subject based on the results of the multiple binary classifications and multi-classifications.

[0107] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0108] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in conjunction with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiments.

[0109] Another embodiment of the present application provides an electronic device, whose specific structure is as follows: Figure 4 As shown, it includes: at least one processor 301; and a memory 302 that is communicatively connected to the at least one processor 301; wherein the memory 302 stores instructions that can be executed by the at least one processor 301, and the instructions are executed by the at least one processor 301 so that the at least one processor 301 can execute a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information as described in the above-mentioned method embodiments.

[0110] Among them, the memory and the processor can be connected in a bus manner, and the bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and will not be further described in this article. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be one component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium through an antenna, and further, the antenna also receives data and transmits the data to the processor.

[0111] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0112] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information as described in the above embodiments.

[0113] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (such as a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM (Read-Only Memory), RAM (Random Access Memory), disk or optical disk and other media that can store program codes.

[0114] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. A multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information, characterized in that: The method comprises: Acquiring electroencephalogram data, brain imaging data and gene data of the subject, and performing corresponding preprocessing operations on the acquired electroencephalogram data, brain imaging data and gene data of the subject respectively; Performing feature extraction on the preprocessed EEG data to obtain the EEG signal features of the subject, performing feature extraction on the preprocessed brain image data to obtain the brain image features of the subject, and performing feature extraction on the preprocessed gene data to obtain the AD susceptibility features of the subject; The features of the three different modalities of the subjects are respectively standardized, and then feature selection is performed on the features of the three different modalities after standardization, and irrelevant and redundant features are excluded. Finally, the features of the three different modalities after feature selection are fused to obtain fused features; Based on the fusion features and the features of the three different modalities after feature selection, multiple binary classifications and multi-classifications are performed, and based on the results of the multiple binary classifications and multi-classifications, the disease discrimination results of the subjects are obtained.

2. The multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to claim 1, characterized in that: The obtained EEG data, brain imaging data and gene data of the subject are respectively preprocessed by: The EEG data of the subjects were band-pass filtered between 0.5 Hz and 55 Hz, and then filtered using a 50 Hz notch filter to eliminate power line interference. Next, the filtered EEG data were subjected to wavelet independent component analysis to eliminate unnecessary eye movement and muscle activity-related artifacts in multiple channels. Finally, the analyzed EEG data were segmented into a series of non-overlapping time period data using a 5-second time window, and the main frequency components were extracted from each time period data to obtain the preprocessed EEG data; The brain imaging data of the subjects were subjected to voxel-based morphological analysis using SPM software. All scanned images were aligned with the T1-weighted template image, segmented into gray matter, white matter, and cerebrospinal fluid maps, and standardized to 2 × 2 × 2 mm in the standard space of MNI. 3 The voxels were smoothed using a FWHM kernel of 8 mm, and finally the VBM-sMRI scan images were registered to the same MNI space to obtain the preprocessed brain imaging data; The genetic data of the subjects were genotyped using the human 610-Quad or OmniExpress Array, and then standard quality control including deletion rate, minor allele frequency and Harwin equilibrium filtering was performed to filter low-quality variant sites and samples. Finally, IMPUTE2 software was used for gene filling to obtain preprocessed genetic data.

3. The multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to claim 2, characterized in that: The EEG signal features are specifically composed of four characteristic parameters: power spectrum density, Hjorth parameter, sample entropy and time-frequency characteristics. The power spectrum density specifically includes absolute PSD features and relative PSD features. The Hjorth parameter specifically includes activity parameter H Activity , mobility parameter H Mobility and complexity parameter H Complexity , extract features from the preprocessed EEG data to obtain the EEG signal features of the subject, including: The absolute PSD of the preprocessed EEG data is calculated using the Welch periodogram method to obtain the absolute PSD feature PSD ab , PSD ab The calculation is achieved through the following formula: G n =a n ×exp[-(F-c) 2 / 2w 2 ]; L=b-log(K+F χ ); Where N represents the total number of peaks extracted from the power spectrum of the preprocessed EEG data, G n represents the Gaussian function fitting the nth peak, L represents the non-periodic or background component, a n represents the peak power of the nth peak, F represents the input frequency vector, c represents the center frequency, w represents the Gaussian standard deviation, b represents the broadband offset, χ represents the exponential parameter, and K represents the bending parameter for controlling the bending of the non-periodic component; Calculate the ratio between the PSD of the specific frequency band and the total frequency band of the power spectrum of the preprocessed EEG data, and use the calculated ratio as the relative PSD feature PSD re ; Let z(t) denote the time series of the preprocessed EEG data of each channel, and the activity parameter H Activity , mobility parameter H Mobility and complexity parameter H Complexity , calculated by the following formula: H Activity =var[z(t)]; H Complexity =H Mobility [dz(t) / dt] / H Mobility [y(t)]; Where var(·) represents the variance operation, dz(t) / dt represents the first-order derivative of z(t), and y(t) represents a pure sine wave; Sample entropy is defined as the negative logarithm of whether m consecutive data points of two similar sequences remain similar at the m+1th point. Sample entropy is calculated by the following formula: SE(m,r)=-ln[C m+1 (r) / C m (r)]; Among them, r is the preset amplitude threshold used to separate the similarity between two vectors, C m+1 (r) represents the number of vector pairs whose vector distance is less than r, C m (r) represents the total number of vector pairs, SE(m,r) represents the sample entropy; The time-frequency characteristics are calculated by the following formula: Wherein, ω(t) is a preset window function, specifically a Hann window or Gaussian window centered on zero, and STFT{z(t)} represents the time-frequency characteristics.

4. The multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to claim 2, characterized in that: Feature extraction is performed on the preprocessed brain image data to obtain the subject's brain image features, including: The whole brain was subsampled based on the subject's brain imaging features, and 116 imaging measurements of regions of interest were generated based on the MarsBaR automatic anatomical labeling atlas, including mean gray matter density on structural MRI, amyloid values ​​on AV45 scans, and glucose utilization on FDG scans; These imaging measures were adjusted using regression weights from healthy individuals to eliminate the effects of baseline age, sex, education, and handedness, ultimately yielding the subjects' brain imaging characteristics.

5. The multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to claim 2, characterized in that: Feature extraction is performed on the preprocessed gene data to obtain the AD susceptibility characteristics of the subjects, including: All SNPs in the preprocessed genetic data were subjected to binary Logistic GWAS analysis using Plink software. The phenotype of the GWAS analysis was AD diagnosis status, and the genetic effect value of each gene locus of the subjects was obtained. The top 100 SNPs were then screened as feature candidates based on the genetic effect value. PRSice-2 software was used to calculate the PRS of all feature candidates to obtain the AD susceptibility characteristics of the subjects.

6. A multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to any one of claims 1 to 5, characterized in that: The three different modal features of the subject are respectively standardized, and then feature selection is performed on the three different modal features after standardization, and irrelevant and redundant features are excluded. Finally, the three different modal features after feature selection are fused to obtain fused features, including: The features of the three different modes of the subjects are scaled to the range of [-1, 1] respectively to obtain the features of the three different modes after standardization; wherein the features of each mode are scaled according to their maximum absolute value; The Spearman correlation index is used to perform feature selection on the features of the three different modes after the standardization process. First, the features whose Spearman correlation index is lower than the preset first correlation threshold are retained to exclude irrelevant features and redundant features. Then, the features whose Spearman correlation index is higher than the preset second correlation threshold are retained to obtain the features of the three different modes after feature selection. The features of the three different modalities after feature selection are connected by feature matrices to obtain fusion features.

7. A multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information according to any one of claims 1 to 5, characterized in that: Binary classification includes binary classification of AD and HC, binary classification of AD and MCI, binary classification of MCI and HC, and multi-classification includes multi-classification of AD, HC and MCI; Multiple binary and multi-classifications based on the features of three different modalities after fusion features and feature selection are achieved through a pre-trained classifier. The classifier consists of an input layer, a hidden layer and a softmax classification output layer. The size of the hidden layer is selected by evaluating the minimum reconstruction error.

8. A multimodal classification system for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information, characterized in that: The system comprises: A data acquisition module, used to acquire the subject's electroencephalogram data, brain imaging data and genetic data; A preprocessing module, used to perform corresponding preprocessing operations on the acquired EEG data, brain imaging data and gene data of the subject respectively; A feature extraction module is used to extract features from the preprocessed EEG data to obtain the EEG signal features of the subject, to extract features from the preprocessed brain image data to obtain the brain image features of the subject, and to extract features from the preprocessed gene data to obtain the AD susceptibility features of the subject; The feature selection and fusion module is used to perform corresponding standardization processing on the features of the three different modalities of the subject respectively, and then perform feature selection on the features of the three different modalities after the standardization processing, exclude irrelevant features and redundant features, and finally perform feature fusion on the features of the three different modalities after feature selection to obtain fusion features; The discrimination module is used to perform multiple binary classifications and multi-classifications based on the features of three different modalities after fusion features and feature selection, and obtain disease discrimination results for the subject based on the results of multiple binary classifications and multi-classifications.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain imaging and genetic information as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a multimodal classification method for Alzheimer's disease based on electroencephalogram, brain image and genetic information as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Health risk assessment method based on AI multiple modes

    CN120613137A