A disease classification method based on electroencephalogram signals and related equipment

By employing a two-stage feature selection process involving feature aggregation and prior knowledge guidance in resting-state EEG data, the instability of feature selection under high-dimensional, small-sample conditions was addressed, thereby improving the stability of cross-mechanism biomarker combinations and the accuracy of disease classification.

CN122490232APending Publication Date: 2026-07-31CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing resting-state EEG recognition schemes are unstable in feature screening under high-dimensional, small-sample conditions, lack medical interpretability, are difficult to form cross-mechanism biomarker combinations, and lack systematic introduction of disease-related knowledge.

Method used

By acquiring resting-state EEG data for training, candidate features are extracted and aggregated at the window, channel, or brain region level. Feature matching and hierarchical subsampling are performed in conjunction with prior knowledge sources. A sparse-constrained linear classifier is used to filter features, and a two-stage feature selection is carried out with the assistance of a large model to construct a disease classification model.

Benefits of technology

It improves the stability and interpretability of EEG biomarker recognition, and enhances the accuracy and reproducibility of disease classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490232A_ABST
    Figure CN122490232A_ABST
Patent Text Reader

Abstract

This invention provides a disease classification method and related equipment based on electroencephalogram (EEG) signals. It extracts candidate feature sets from training resting-state EEG data and aggregates these sets at the window, channel, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. A prior knowledge source related to the target disease is constructed. Based on the prior knowledge source and the feature description information table, knowledge matching is performed on each candidate feature in the candidate feature set to obtain auxiliary guidance information. Based on weight values, a subset of candidate features are selected from the candidate feature set to obtain a first-stage candidate biomarker set and a remaining candidate feature set. Based on the auxiliary guidance information, further selection is performed on the remaining candidate feature set to obtain a second-stage supplementary selection set. This improves the stability and interpretability of EEG biomarker recognition, contributing to enhanced disease classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart medical technology, and in particular to a disease classification method and related equipment based on electroencephalogram (EEG) signals. Background Technology

[0002] Electroencephalogram (EEG) signals can reflect the state of brain neural activity and have advantages such as being non-invasive, having high temporal resolution, relatively low cost, and being repeatable. In recent years, they have become an important research subject for the objective assessment and auxiliary identification of neuropsychiatric diseases. International research generally believes that resting-state EEG has high potential for biomarker discovery in areas such as depressive disorders, bipolar disorder, anxiety disorders, and schizophrenia spectrum disorders. Related research is gradually shifting from traditional empirical interpretation to quantitative analysis, machine learning identification, and interpretable modeling.

[0003] From the perspective of existing technical processes, resting-state EEG recognition schemes typically follow the basic steps of "signal preprocessing - candidate feature extraction - feature screening - classification modeling - result interpretation". Among them, preprocessing generally includes rereference, bandpass filtering, power frequency removal, artifact removal, segmentation, and quality control; candidate feature extraction has gradually expanded from early quantitative EEG indicators such as band power, peak frequency, and frontal asymmetry to multi-dimensional characterization such as time-domain statistics, frequency-domain power spectrum, time-frequency features, micro-state parameters, nonlinear complexity indicators, functional connectivity indicators, background spectrum parameters, and source space or brain network features.

[0004] In international research, techniques for patient identification based on resting-state EEG can be broadly categorized into three types: The first type involves statistical analysis based on single or limited EEG indicators, such as comparing differences in specific frequency band power, frontal lobe asymmetry, spectral peak parameters, or microstate parameters between patient and healthy control groups. The second type utilizes multi-domain candidate features and traditional machine learning classifiers, extracting multi-domain features and then classifying them using statistical tests, feature ranking, L1 regularization, recursive feature elimination, or packaged screening. The third type is based on deep learning or end-to-end modeling, directly learning discriminative patterns from raw EEG or its transformed representations. Among these, the second type remains the mainstream approach in current EEG biomarker research.

[0005] Despite the progress made by the aforementioned technical approaches, several common shortcomings remain in the task of identifying high-dimensional EEG biomarkers in resting-state EEG. First, under high-dimensional, small-sample conditions, the number of candidate EEG features is usually significantly higher than the number of samples. Furthermore, different feature families exhibit differences in dimensionality, distribution, and statistical dominance. Therefore, the screening results obtained from a single ranking, single-round modeling, or single-model weighting are easily affected by sample perturbations and fluctuate, making it difficult to form a repeatable and verifiable stable feature set. Second, when multiple EEG features participate in the screening, a statistically dominant feature family often maintains its dominance for a long period, causing the final results to focus on a single mechanism dimension and making it difficult to form a cross-mechanism EEG biomarker combination that combines information complementarity and mechanism coverage.

[0006] Furthermore, existing technologies still fall short in utilizing disease-related clinical rules, pathological mechanisms, brain region function, and the physiological significance of frequency bands. Most methods remain purely data-driven screening, lacking mechanisms to systematically incorporate clinical knowledge, literature evidence, and historical analytical experience into the candidate feature screening process. Although large models can be used for semantic understanding and knowledge integration, existing research indicates that they may still suffer from inconsistencies and illusions in medical text analysis and high-risk tasks. Therefore, they are more suitable as tools for knowledge organization, semantic parsing, and screening guidance, rather than directly replacing the final feature determination or patient classification decision-making modules.

[0007] Based on the aforementioned research status both domestically and internationally, it can be seen that while existing resting-state EEG recognition schemes possess multi-domain feature extraction and classification modeling capabilities, they still lack a unified technical approach that can simultaneously ensure the stability of high-dimensional feature screening, the complementarity of cross-mechanism features, and medical interpretability. Particularly in the task of classifying patients with mental disorders, there is a lack of an EEG biomarker recognition method that uses a two-stage feature screening and patient classification as the main thread, and appropriately incorporates disease-related prior knowledge and large-scale model assistance to improve the stability, discriminative power, and interpretability of biomarker combinations while ensuring the verifiability and repeatability of the final feature determination process. Summary of the Invention

[0008] This invention provides a disease classification method and related equipment based on electroencephalogram (EEG) signals, with the aim of improving the accuracy of disease classification.

[0009] To achieve the above objectives, the present invention provides a disease classification method based on electroencephalogram (EEG) signals, comprising: Step 1: Acquire resting-state EEG data for training. The resting-state EEG data for training includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy individuals with their eyes closed and open. Step 2: Extract candidate feature sets from the training resting-state EEG data, and aggregate the candidate feature sets at the window level, channel level, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. Step 3: Construct a prior knowledge source related to the target disease, and perform knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information; Step 4: Perform hierarchical subsampling on the training sample set consisting of the feature vectors and class labels corresponding to the candidate feature set to obtain multiple subsamples. After preprocessing each subsample, train a linear classifier with sparse constraints. Step 5: During the training process of the linear classifier, determine the weight value of each candidate feature in each round of the linear classifier in the candidate feature set, and select some candidate features based on the weight values ​​to obtain the first-stage candidate biomarker set and the remaining candidate feature set. Step 6: Based on the auxiliary guidance information, select again from the remaining candidate feature set to obtain the second-stage supplementary selection set, and combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain the biomarker combination; Step 7: Train the selected classification model based on the feature vectors corresponding to the biomarker combinations to obtain the disease classification model; Step 8: Extract features from the resting-state EEG data of the object to be classified to obtain the target biomarker combination, and input the target feature vector corresponding to the target biomarker combination into the disease classification model for classification to obtain the disease classification result.

[0010] Furthermore, step 1 includes: Raw resting-state EEG signals were collected from disease patients with their eyes closed and open, and raw resting-state EEG signals were collected from healthy individuals with their eyes closed and open as raw resting-state EEG data. The raw resting-state EEG data were preprocessed to obtain resting-state EEG data for training.

[0011] Furthermore, the candidate feature set includes one or more of the following: time-domain features, frequency-domain features, spatial-domain features, time-frequency-domain features, and decomposition-domain features.

[0012] Furthermore, the feature description information should include at least the feature name, state conditions, mechanism category, statistical expression method, channel or brain region, relevant frequency band, and potential physiological significance.

[0013] Furthermore, sources of prior knowledge include one or more of the following: clinical guidelines, literature knowledge, descriptions of pathological mechanisms, expert experience, and conclusions from historical analyses.

[0014] Furthermore, based on prior knowledge sources and feature description information tables, knowledge matching is performed on each candidate feature in the candidate feature set to obtain auxiliary guidance information, including: Based on a pre-constructed terminology dictionary, brain region mapping table, frequency band mapping table, and feature category rules, the feature description information in the feature description information table is matched according to rules to obtain auxiliary guidance information; When auxiliary guidance information cannot be obtained through rule matching, a large model is introduced to perform semantic parsing on various prior knowledge in the prior knowledge source and feature description information in the feature description information table to obtain auxiliary guidance information for each candidate feature in the candidate feature set. The auxiliary guidance information includes category priority, brain region attention direction, frequency band attention direction, candidate feature grouping relationship, redundancy suppression rules, and complementary selection prompts.

[0015] The present invention also provides a disease classification device based on electroencephalogram (EEG) signals, comprising: The acquisition module is used to acquire resting-state EEG data for training. The resting-state EEG data for training includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy individuals with their eyes closed and open. The extraction module is used to extract candidate feature sets from training resting-state EEG data, and to aggregate the candidate feature sets at the window level, channel level or brain region level to generate feature description information for each candidate feature, forming a feature description information table. The knowledge matching module is used to construct a prior knowledge source related to the target disease. Based on the prior knowledge source and the feature description information table, it performs knowledge matching on each candidate feature in the candidate feature set to obtain auxiliary guidance information. The subsampling module is used to perform hierarchical subsampling on the training sample set consisting of the feature vectors and class labels corresponding to the candidate feature set, to obtain multiple subsamples. After preprocessing each subsample, a linear classifier with sparse constraints is trained. The first screening module is used to determine the weight value of each candidate feature in the candidate feature set in each round of the linear classifier during the training process of the linear classifier, and to screen out some candidate features based on the weight values ​​in the candidate feature set to obtain the first-stage candidate biomarker set and the remaining candidate feature set. The second screening module is used to select again from the remaining candidate feature set based on the auxiliary guidance information to obtain the second-stage supplementary selection set, and to combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain the biomarker combination. The training module is used to train the selected classification model based on the feature vectors corresponding to the combination of biomarkers to obtain the disease classification model; The classification module is used to extract features from the resting-state EEG data of the object to be classified, obtain the target biomarker combination, and input the target feature vector corresponding to the target biomarker combination into the disease classification model for classification to obtain the disease classification result.

[0016] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a disease classification method based on electroencephalogram (EEG) signals.

[0017] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a disease classification method based on electroencephalogram (EEG) signals.

[0018] The above-described solution of the present invention has the following beneficial effects: Compared with existing technologies, this invention extracts candidate feature sets from training resting-state EEG data and aggregates these candidate feature sets at the window, channel, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. It then constructs a prior knowledge source related to the target disease, performs knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table, and obtains auxiliary guidance information. Based on weight values, it filters out some candidate features from the candidate feature set to obtain a first-stage candidate biomarker set and a remaining candidate feature set. Based on the auxiliary guidance information, it further selects from the remaining candidate feature set to obtain a second-stage supplementary selection set. This improves the stability and interpretability of EEG biomarker recognition and helps enhance the accuracy of disease classification.

[0019] Other beneficial effects of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating an embodiment of the present invention; Figure 2 This is a schematic diagram showing the distribution of biomarker combinations in the BP and HC groups under open-eye resting state conditions in an embodiment of the present invention; Figure 3 This is a schematic diagram of the disease classification device in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the terminal device in an embodiment of the present invention. Detailed Implementation

[0021] To make the technical problems, solutions, and advantages of this invention clearer, a detailed description will be provided below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] In the description of this invention, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0023] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0024] This invention addresses existing problems by providing a disease classification method and related equipment based on electroencephalogram (EEG) signals.

[0025] like Figure 1 As shown, embodiments of the present invention provide a disease classification method based on electroencephalogram (EEG) signals, including: Step 1: Acquire resting-state EEG data for training. The resting-state EEG data for training includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy individuals with their eyes closed and open. Step 2: Extract candidate feature sets from the training resting-state EEG data, and aggregate the candidate feature sets at the window level, channel level, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. Step 3: Construct a prior knowledge source related to the target disease, and perform knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information; Step 4: Perform hierarchical subsampling on the training sample set consisting of the feature vectors and class labels corresponding to the candidate feature set to obtain multiple subsamples. After preprocessing each subsample, train a linear classifier with sparse constraints. Step 5: During the training process of the linear classifier, determine the weight value of each candidate feature in each round of the linear classifier in the candidate feature set, and select some candidate features based on the weight values ​​to obtain the first-stage candidate biomarker set and the remaining candidate feature set. Step 6: Based on the auxiliary guidance information, select again from the remaining candidate feature set to obtain the second-stage supplementary selection set, and combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain the biomarker combination; Step 7: Train the selected classification model based on the feature vectors corresponding to the biomarker combinations to obtain the disease classification model; Step 8: Extract features from the resting-state EEG data of the object to be classified to obtain the target biomarker combination, and input the target feature vector corresponding to the target biomarker combination into the disease classification model for classification to obtain the disease classification result.

[0026] Specifically, step 1 includes: Raw resting-state EEG signals were collected from disease patients with their eyes closed and open, and raw resting-state EEG signals were collected from healthy individuals with their eyes closed and open as raw resting-state EEG data. The raw resting-state EEG data were preprocessed to obtain resting-state EEG data for training.

[0027] In this embodiment of the invention, the preprocessing process includes bandpass filtering, artifact removal, bad derivative processing, rereference, segmentation, and normalization, as detailed below: First, set the sampling rate The frequency was uniformly set to 250Hz, and raw resting-state EEG data was read and electrode position information was loaded. The raw resting-state EEG data was then subjected to bandpass filtering, expressed as follows: ; in, This represents the filtered resting-state EEG data. This represents a pre-filtering operator, which may include bandpass filtering and power frequency notch filtering. Represents the resampling operator. Indicates the first The first subject Raw resting-state EEG data from each channel. , Indicates the number of subjects. , Indicates the number of EEG channels; In the bad lead processing stage, the results of flat leads, deviated leads, low correlation leads, high-frequency noise leads, low signal-to-noise ratio leads, and RANSAC abnormal leads are combined to obtain the bad lead set. Non-destructive channels are set as For non-bad-conducting channels, their pre-filtered signals remain unchanged. For bad-conducting channels, each bad-conducting channel in the bad-conducting channel set is reconstructed using a weighted average of its neighboring non-bad-conducting channels. The expression is as follows: ; ; in, This represents the resting-state EEG data after negative guidance processing. Indicates the first Among the subjects, the bad conduction channel The set of non-destructive channels in spatial proximity This represents the interpolation weights determined by the spatial location of the electrodes. ; The resting-state EEG data after bad-conduction processing undergoes rereference processing, expressed as follows: ; in, Indicates a reference result. This indicates the number of valid channels participating in the reference. Then, conventional independent component analysis was used to remove artifacts such as electrooculography (EOG), electromyography (EMG), electrocardiography (ECG), and power line noise, resulting in clean EEG data. For example, let's assume... Given the ICA unmixing matrix, the independent components are represented as: ; in, Indicates the first The subjects were at time Multi-channel signals; Record No. The set of independent components that were identified as artifacts by the subjects was: Construct a diagonal matrix for: ; set up Given an ICA mixing matrix, the clean EEG signal after artifact removal is: ; For clean EEG data, the closed-eye and open-eye segments are segmented based on experimental event labels to obtain closed-eye and open-eye segments. Transition buffers can be set at the beginning and end of these segments to reduce the impact of state transitions. For example, let the state set be... , define the first The subjects were in a state of... The effective time interval is denoted as:

[0028] in, and These represent the start and end times of the segment determined by the event markers. Indicates the transition buffer duration for cropping the beginning and end of a segment; Therefore, the state segment signal is obtained as follows: ; in, Indicates the first The subjects were in a state of... The number of valid time sampling points retained after segmentation, cropping, and resampling. - ; Finally, the segmented channel signal is standardized, and the expression is: ; in, This represents standardized resting-state EEG data. The mean of the corresponding signal is expressed as: ; The variance of the corresponding signal is expressed as: ; To prevent extremely small constants with a denominator of zero; After the above processing, the first The subjects were in a state of... Effective resting-state EEG data can be represented as follows: .

[0029] Specifically, the candidate feature set includes one or more of the following: time-domain features, frequency-domain features, spatial-domain features, time-frequency-domain features, and decomposition-domain features, where the first... The subjects were in a state of... The next Each candidate feature can be represented as , , This indicates the number of candidate features.

[0030] In this embodiment of the invention, the time-domain features include mean, median, standard deviation, variance, root mean square, peak-to-peak value, interquartile range, median absolute deviation, line length, zero crossover rate, skewness, kurtosis, Hjorth parameter, Higuchi fractal dimension, DFA exponent, and autoregressive features, etc. Frequency domain features are used to characterize the oscillatory activity of EEG signals in different frequency bands and can be obtained based on Welch power spectrum estimation. For example, let the first... The first subject Each channel in the frequency band The absolute power and relative power on are respectively and .in Indicates the first The first subject Each channel in the frequency band Absolute power on Corresponding frequency band The relative power on, Indicates the first The first subject The power spectral density of each channel , and Indicates frequency band The lower and upper limits, and Indicates the frequency range for power spectrum analysis; Further extraction is also possible The frequency domain features are summarized as the ratio, individual alpha peak frequency, alpha peak sharpness, spectral entropy, left and right hemisphere asymmetry, and brain regions such as the frontal and posterior regions. Spatial domain features can be obtained through methods such as CSP, RCSP, CSSP, CSSSP, or FBCSP; Time-frequency domain features are extracted using short-time Fourier transform or wavelet transform to extract energy or entropy features at different time slices and frequency bands; Features of the decomposition domain can be obtained through wavelet packet decomposition, EMD / EEMD / CEEMDAN, LMD style decomposition, and adaptive Hermite approximation.

[0031] Specifically, the feature description information should include at least the feature name, state conditions, mechanism category, statistical expression method, channel or brain region, relevant frequency band, and potential physiological significance.

[0032] In this embodiment of the invention, for any candidate feature Its feature description information It can be represented as: ; in, Indicates the feature name, Indicates state conditions. Indicates the mechanism category, Indicates a channel or brain region. Indicates the relevant frequency band, Indicates statistical expression methods, It indicates potential physiological significance.

[0033] For ease of explanation, the sample set in this embodiment of the invention is referred to as... ,in, Let represent the candidate feature vector of the i-th subject. Represent its category label; denoted as candidate feature set as , Represents the total number of candidate features, and the sample set. The data is derived from feature description tables formed after preprocessing, segmentation, multi-domain feature extraction, and aggregation of each subject, rather than the original time-point data.

[0034] Specifically, for each subject, preprocessing and multi-domain feature extraction were performed under conditions of closed eyes and / or open eyes. Then, all candidate features corresponding to the same subject were concatenated into a one-dimensional feature vector according to uniform column names. ; and construct a label vector from its category labels. This is for use in subsequent stable selection and classification modeling.

[0035] Specifically, prior knowledge sources include one or more of the following: clinical guidelines, literature knowledge, descriptions of pathological mechanisms, expert experience, and conclusions from historical analyses.

[0036] Specifically, based on prior knowledge sources and feature description information tables, knowledge matching is performed on each candidate feature in the candidate feature set to obtain auxiliary guidance information, including: Based on the pre-constructed terminology dictionary, brain region mapping table, frequency band mapping table, and feature category rules, the feature description information in the feature description information table is matched according to rules to obtain auxiliary guidance information. For example, when the mechanism category, brain region, or frequency band of a candidate feature is consistent with the corresponding field in a knowledge entry in the prior knowledge source or the terminology is pre-defined synonym set, a matching relationship is established. When auxiliary guidance information cannot be obtained through rule matching, a large model is introduced to perform semantic parsing on various prior knowledge in the prior knowledge source and feature description information in the feature description information table to obtain auxiliary guidance information for each candidate feature in the candidate feature set. The auxiliary guidance information includes category priority, brain region attention direction, frequency band attention direction, candidate feature grouping relationship, redundancy suppression rules, and complementary selection prompts.

[0037] Specifically, in this embodiment of the invention, the input to the large model is limited to: the target disease name, structured candidate feature description, knowledge text to be parsed, and a preset output field template. The output is limited to structured auxiliary guidance information. At the same time, in order to ensure the usability of the results, the output of the large model also needs to undergo field integrity verification, terminology standardization mapping, and rule consistency verification. Results that fail the verification will not directly enter the subsequent screening process.

[0038] In this embodiment of the invention, for any candidate feature Its auxiliary guidance information can represent for: ; in, Indicates category priority. Indicates the direction of attention in the brain region. Indicates the frequency band of interest. Indicates the grouping relationship of candidate features. This represents the redundancy suppression rule. This indicates a complementary supplementary selection suggestion.

[0039] It should be noted that, in the embodiments of the present invention, when performing hierarchical subsampling, it is necessary to maintain the basic consistency of the proportion of each category. For each subsample, missing values ​​are first filled in, then standardized, and then a linear classifier with sparse constraints is trained.

[0040] Specifically, during the training process of a linear classifier, it is assumed that... Round-by-round sub-sampling, in the first In the round, candidate features In the The selected features in the training round are denoted as Determine the weight value of each candidate feature in each round of the linear classifier from the candidate feature set. When the corresponding weights in the sparse classifier in this round When it is not zero, it is denoted as Otherwise, remember Then the stability score of the candidate feature It can be represented as: Based on stability score A subset of candidate features were selected from the candidate feature set to form the first-stage candidate biomarker set. and remaining candidate feature set .

[0041] It should be noted that the second phase of the supplementary selection set The features in the data are derived from the remaining candidate feature sets that combine auxiliary guidance information from different mechanism categories, different brain regions, different frequency bands, or different statistical expression methods. The supplementary selection, the characteristics of the supplementary selection and the set of candidate biomarkers in the first stage. The features in the samples are complementary, and the first-stage candidate biomarker set is included. Second-stage supplementary selection set By combining them, a combination of biomarkers can be obtained. .

[0042] In this embodiment of the invention, the classification model is preferably a linear classification model with good interpretability, but logistic regression, tree model or other supervised learning models may also be used.

[0043] In this embodiment of the invention, when the feature vector corresponding to the biomarker combination is used as the input to the classification model, the decision function of the classification model can be expressed as: ; in, This indicates the patient category determination result. This represents the feature vector corresponding to a combination of biomarkers. This represents the parameter vector of the classification model. Indicates the bias term; In summary, the first The category prediction results for the subjects are as follows: .

[0044] In order to evaluate whether the model performance is statistically significant, this invention performs label permutation test or other statistical significance test on the classification results. At the same time, it combines model weight analysis, feature contribution analysis and feature semantic information to interpret the direction of action, importance and corresponding brain region-frequency band-mechanism meaning of the finally selected features.

[0045] This invention uses resting-state electroencephalogram (EEG) signals with eyes open as specific experimental data to obtain, as follows: Figure 2 The biomarker combinations shown exhibit discriminative distribution differences between patients (BP) and healthy individuals (HC). The temporal, spatial, and frequency domain features all demonstrate a certain degree of inter-group discriminative ability, indicating that the biomarker combinations screened in the embodiments of the present invention not only have stability but also good interpretability and practical discriminative value.

[0046] To further verify the interpretability of the selected features, this embodiment of the invention performs SHAP interpretation analysis on the model. The analysis results are as follows: For the average sample of the BP group, all selected features make positive contributions to the model output, jointly driving the model to output in the BP category; for the average sample of the HC group, the above features generally show negative contributions, causing the model output to change in the direction of the HC category. The above results show that the features selected by this embodiment of the invention can not only effectively distinguish between patients and healthy controls, but also have good directional consistency and mechanism interpretability, verifying the feasibility of the method provided by this embodiment of the invention.

[0047] To verify the effectiveness of the embodiments of the present invention, a five-fold cross-validation experiment was conducted based on the electroencephalogram (EEG) data of 15 healthy control subjects and 15 subjects with bipolar disorder. The results are shown in Table 1 below: Table 1. Classification results of five-fold cross-validation

[0048] As can be seen from Table 1 above, the method provided by the embodiments of the present invention achieves AUC=0.822 and ACC=0.733 under the final feature combination condition. Compared with the classification results of each single feature, it shows better comprehensive discrimination ability and has better classification effect compared with KNN, random forest and other comparative methods. This shows that the two-stage feature screening and feature combination strategy proposed by the embodiments of the present invention can extract stable and effective discrimination information from EEG data and has practical feasibility.

[0049] Compared with existing technologies, the embodiments of the present invention extract candidate feature sets from training resting-state EEG data, and aggregate the candidate feature sets at the window level, channel level, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. A prior knowledge source related to the target disease is constructed, and knowledge matching is performed on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information. Based on weight values, a portion of candidate features are selected from the candidate feature set to obtain a first-stage candidate biomarker set and a remaining candidate feature set. Based on the auxiliary guidance information, further selection is performed on the remaining candidate feature set to obtain a second-stage supplementary selection set. This improves the stability and interpretability of EEG biomarker recognition, and helps to improve the accuracy of disease classification.

[0050] Corresponding to the disease classification method based on electroencephalogram (EEG) signals described in the above embodiments, such as Figure 3 As shown, this embodiment of the invention also provides a disease classification device 100 based on electroencephalogram (EEG) signals, the disease classification device 100 comprising: The acquisition module 101 is used to acquire resting-state EEG data for training. The resting-state EEG data for training includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy people with their eyes closed and open. The extraction module 102 is used to extract candidate feature sets from training resting-state EEG data, and to aggregate the candidate feature sets at the window level, channel level or brain region level to generate feature description information for each candidate feature and form a feature description information table. The knowledge matching module 103 is used to construct a prior knowledge source related to the target disease, and to perform knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information. The subsampling module 104 is used to perform hierarchical subsampling on the resting-state EEG data for training to obtain multiple subsamples. After preprocessing each subsample, a linear classifier with sparse constraints is trained. The first screening module 105 is used to determine the weight value of each candidate feature in the candidate feature set in each round of the linear classifier during the training process of the linear classifier, and to screen out some candidate features in the candidate feature set based on the weight value to obtain the first stage candidate biomarker set and the remaining candidate feature set. The second screening module 106 is used to select again from the remaining candidate feature set based on the auxiliary guidance information to obtain the second-stage supplementary selection set, and to combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain a biomarker combination. Training module 107 is used to train the selected classification model based on the feature vector corresponding to the combination of biomarkers to obtain the disease classification model; The classification module 108 is used to extract features from the resting-state EEG data of the object to be classified, obtain a combination of target biomarkers, and input the target feature vector corresponding to the combination of target biomarkers into the disease classification model for classification to obtain the classification result.

[0051] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0052] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0053] This invention also provides a terminal device, such as... Figure 4 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 4The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the above-described disease classification method based on electroencephalogram (EEG) signals.

[0054] The terminal device D10 can be a desktop computer, laptop, handheld computer, server, server cluster, or cloud server, etc. This terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art will understand that... Figure 4 This is merely an example of terminal device D10 and does not constitute a limitation on terminal device D10. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0055] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0056] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0057] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0058] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0059] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a disease classification method based on electroencephalogram (EEG) signals.

[0060] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a building device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0061] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A disease classification method based on electroencephalogram (EEG) signals, characterized in that, include: Step 1: Acquire resting-state EEG data for training. The resting-state EEG data for training includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy individuals with their eyes closed and open. Step 2: Extract candidate feature sets from the training resting-state EEG data, and aggregate the candidate feature sets at the window level, channel level, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. Step 3: Construct a prior knowledge source related to the target disease; perform knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information. Step 4: Perform hierarchical subsampling on the training sample set consisting of the feature vectors and class labels corresponding to the candidate feature set to obtain multiple subsamples. After preprocessing each subsample, train a linear classifier with sparse constraints. Step 5: During the training process of the linear classifier, determine the weight value of each candidate feature in each round of the linear classifier in the candidate feature set, and select some candidate features based on the weight values ​​in the candidate feature set to obtain the first stage candidate biomarker set and the remaining candidate feature set; Step 6: Based on the auxiliary guidance information, select again from the remaining candidate feature set to obtain the second-stage supplementary selection set, and combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain a biomarker combination; Step 7: Train the selected classification model based on the feature vector corresponding to the biomarker combination to obtain the disease classification model; Step 8: Extract features from the resting-state EEG data of the object to be classified to obtain the target biomarker combination, and input the target feature vector corresponding to the target biomarker combination into the disease classification model for classification to obtain the disease classification result.

2. The disease classification method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, Step 1 includes: Raw resting-state EEG signals were collected from disease patients with their eyes closed and open, and from healthy individuals with their eyes closed and open, as raw resting-state EEG data. The raw resting-state EEG data is preprocessed to obtain resting-state EEG data for training.

3. The disease classification method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, The candidate feature set includes one or more of the following: time-domain features, frequency-domain features, spatial-domain features, time-frequency-domain features, and decomposition-domain features.

4. The disease classification method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, The feature description information includes at least the feature name, state condition, mechanism category, statistical expression method, channel or brain region, relevant frequency band, and potential physiological significance.

5. The disease classification method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, The prior knowledge sources include one or more of the following: clinical guidelines, literature knowledge, descriptions of pathological mechanisms, expert experience, and historical analysis conclusions.

6. The disease classification method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, Based on the prior knowledge source and the feature description information table, knowledge matching is performed on each candidate feature in the candidate feature set to obtain auxiliary guidance information, including: Based on a pre-constructed terminology dictionary, brain region mapping table, frequency band mapping table, and feature category rules, the feature description information in the feature description information table is matched according to rules to obtain auxiliary guidance information; When the auxiliary guidance information cannot be obtained through rule matching, a large model is introduced to perform semantic parsing on various prior knowledge in the prior knowledge source and the feature description information in the feature description information table to obtain the auxiliary guidance information for each candidate feature in the candidate feature set. The auxiliary guidance information includes category priority, brain region attention direction, frequency band attention direction, candidate feature grouping relationship, redundancy suppression rules, and complementary selection prompts.

7. A disease classification device based on electroencephalogram (EEG) signals, characterized in that, include: The acquisition module is used to acquire resting-state EEG data for training, which includes resting-state EEG signals of patients with diseases with their eyes closed and open, and resting-state EEG signals of healthy individuals with their eyes closed and open. The extraction module is used to extract candidate feature sets from the training resting-state EEG data, and to aggregate the candidate feature sets at the window level, channel level, or brain region level to generate feature description information for each candidate feature, forming a feature description information table. The knowledge matching module is used to construct a prior knowledge source related to the target disease, and to perform knowledge matching on each candidate feature in the candidate feature set based on the prior knowledge source and the feature description information table to obtain auxiliary guidance information. The subsampling module is used to perform hierarchical subsampling on the training sample set consisting of the feature vectors and class labels corresponding to the candidate feature set to obtain multiple subsamples. After preprocessing each subsample, a linear classifier with sparse constraints is trained. The first screening module is used to determine the weight value of each candidate feature in the candidate feature set in each round of the linear classifier during the training process of the linear classifier, and to screen out some candidate features in the candidate feature set based on the weight value to obtain the first stage candidate biomarker set and the remaining candidate feature set. The second screening module is used to select again from the remaining candidate feature set based on the auxiliary guidance information to obtain a second-stage supplementary selection set, and to combine the first-stage candidate biomarker set and the second-stage supplementary selection set to obtain a biomarker combination. The training module is used to train the selected classification model based on the feature vector corresponding to the biomarker combination to obtain the disease classification model; The classification module is used to extract features from the resting-state EEG data of the object to be classified, obtain a combination of target biomarkers, and input the target feature vector corresponding to the combination of target biomarkers into the disease classification model for classification to obtain the disease classification result.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the disease classification method based on electroencephalogram (EEG) signals as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the disease classification method based on electroencephalogram (EEG) signals as described in any one of claims 1 to 6.