Sleep staging system based on multi-threshold neighborhood extremum mode statistics

By combining the time domain, frequency domain and marked signal position characteristics, SMNE features are extracted, and using the Gray Wolf Optimizer and Random Forest Classifier, the problem of low accuracy of the single-channel EEG signal automatic sleep staging system in the existing technology is solved, and higher sleep staging accuracy and classification performance are achieved.

CN119924774AActive Publication Date: 2025-05-06HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202411922704.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing automatic sleep staging system with single-channel EEG signals is not very accurate and it is difficult to effectively classify sleep stages.

Method used

A sleep staging system based on multi-threshold neighborhood extreme mode statistics is adopted to extract SMNE features by combining the time domain, frequency domain and marked signal position characteristics, and the classification performance is improved by using the Gray Wolf Optimizer and the Random Forest Classifier.

Benefits of technology

It improves the accuracy of sleep staging, enhances the accuracy of classification, and can effectively track and reproduce pattern changes of EEG signals in different sleep stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119924774A_ABST
    Figure CN119924774A_ABST
Patent Text Reader

Abstract

The invention discloses a sleep staging system based on multi-threshold neighborhood extremum mode statistics. The sleep staging system comprises an electroencephalogram signal collector, a signal preprocessor, a multi-threshold neighborhood extremum time domain feature extractor, a grey wolf optimizer and a random forest classifier. Wherein the pre-processing is used for carrying out preliminary filtering and 30s segmentation on selected electroencephalogram signals. Dividing the single-channel electroencephalogram signals into two types according to the duration of the annotation of the single-channel electroencephalogram signals: an original signal segment and a signal segment needing data enhancement; secondly, performing data enhancement on the specified fragment, and selecting an extreme value of the data fragment after data enhancement; and then calculating neighborhood difference values of extreme values, and carrying out multi-threshold and multi-state processing on a five-classification-stage matrix obtained according to the difference values so as to obtain a multi-threshold neighborhood extreme value mode statistical code. A threshold in the multi-threshold processing is determined by a grey wolf optimizer. And finally, inputting the obtained feature matrix into a random forest classifier. Based on EEG signals, the sleep automatic staging capability is enhanced, and a set of complete system containing a visual program from electroencephalogram collection to sleep analysis is designed and manufactured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of feature extraction and detection, and relates to a sleep staging system based on multi-threshold neighborhood extreme value pattern statistics. Background Art

[0002] Accurate sleep stage classification is a prerequisite for evaluating sleep quality. Currently, polysomnography (PSG) is considered the gold standard for evaluating sleep quality, and electroencephalogram (EEG), electromyography (EMG), electrocardiogram (ECG), and electrooculogram (EOG) are recorded simultaneously when evaluating sleep quality. Then, sleep experts reconstruct the hypnogram by visually observing physiological signals for 30 seconds. The two popular gold standards for sleep staging are the R&K (Rechtschaffen & Kales, R&K) and the 2007 American Academy of Sleep Medicine (AASM) rules, which are used to analyze sleep records. In particular, PSG or single-channel EEG recordings are usually divided into 30-second segments, each of which is manually reviewed by a sleep expert and then classified into one of five stages: wake (W), rapid eye movement (REM), and three non-REM stages (N1, N2, and N3).

[0003] The process of sleep staging is usually signal acquisition, signal preprocessing, feature extraction, feature selection, and classification model learning and evaluation. The quality of features directly determines the effect of classification detection, so feature extraction and feature selection are very important parts. The use of symbolic analysis for time series data analysis in sleep staging research has attracted more and more attention. Symbolic analysis provides advantages such as improving discovery efficiency, reducing sensitivity to measurement noise, and distinguishing between specific and general types of models. Each point in a time series is closely related to time, resulting in high-dimensional features. Before processing a time series, it is usually necessary to modify it with a representation method. One of the basic methods of symbolic analysis of time series is symbolic aggregate approximation (SAX). SAX has the advantages of simple calculation, high efficiency, and allows dimensionality reduction. Different data are mapped to the same dimension through SAX, and each value under this dimension has the same practical meaning. This allows the data to be fused simply and effectively. It provides a comprehensive description of the dynamic system by converting the dynamic system into a new representation space that retains the most important time information, assigning symbols corresponding to the system state, and extracting valuable information from this new space. As a mathematical tool, it is an effective means to evaluate the complexity or irregularity of biomedical records. However, the existing automatic sleep staging system for single-channel EEG signals has the technical problem of low accuracy.

[0004] Therefore, in view of the technical defects of the prior art, it is necessary to propose a solution to solve the technical problems of the prior art. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention proposes a sleep staging system based on multi-threshold neighborhood polar pattern statistics, which improves classification performance and enhances accuracy by combining the position characteristics of time domain, frequency domain and marked signals.

[0006] In order to solve the technical problems existing in the prior art, the technical solution of the present invention is as follows:

[0007] A sleep staging system based on multi-threshold neighborhood extreme value mode statistics (SMNE) includes an electroencephalogram signal collector, a signal preprocessor, a multi-threshold neighborhood extreme value feature extractor, a gray wolf optimizer and a random forest classifier, wherein:

[0008] The EEG signal collector is used to obtain single-channel EEG signals (Fpz-Cz channel), that is, raw EEG data;

[0009] The signal preprocessor is used to receive the original EEG data of the EEG signal collector and perform data preprocessing, including filtering and data enhancement of single-channel EEG signals; wherein, the collected original EEG data is subjected to 0.5Hz to 40Hz bandpass filtering, and the processed EEG signal is segmented into segments; in data enhancement, the signal-to-noise ratio data of the specified segment is enhanced according to the duration of the segment annotation after segmentation.

[0010] The multi-threshold neighborhood extreme value feature extractor is used to extract multi-threshold neighborhood extreme value features from the preprocessed signal; first, the extreme value of the data segment after data enhancement is calculated, and then the neighborhood difference of the extreme value is calculated, and the difference is marked with five categories of states, and the SMNE code is obtained by analyzing the multi-mode encoding of the state.

[0011] The Gray Wolf optimizer determines the optimal thresholds for signal-to-noise ratio enhancement and SMNE feature extraction through multiple iterations.

[0012] The random forest classifier is used to input the best feature subset for classification and obtain the final classification result.

[0013] As a further improvement, the EEG signal collector uses dry electrodes to collect EEG signals and transmits the data to a receiving device via a Bluetooth module.

[0014] As a further improvement, the time span of each EEG signal segment is 30 seconds.

[0015] As a further improvement scheme, the signal preprocessor is used to filter out interference components in the EEG signal and eliminate the problem of data unevenness of the EEG signal.

[0016] As a further improvement, data augmentation optimizes the physiological characteristics of the uneven number of segments of the EEG signal. The steps are:

[0017] S1: Filter the collected EEG signal;

[0018] S2: Based on S1, the filtered signal is shifted back to obtain a new 30s segment;

[0019] S3: Based on S2, adjacent segments with the same sleep stage are marked as collective segments;

[0020] S4: Based on S3, the signal-to-noise ratio of the set segments is calculated, and the absolute value is taken to obtain the data-enhanced segments;

[0021] S5: The extracted new segment and the original data set are used as the preprocessed signal segment.

[0022] As a further improvement, the SMNE feature extracts the pattern change statistics of extreme values ​​and accurately classifies the sleep signal. The steps are:

[0023] S10: Extract local extreme values: including position information and amplitude information.

[0024] S20: Based on S10, the obtained extreme values ​​are sorted to determine the number and distribution of the five states, ensuring that each extreme value is marked with a state.

[0025] S30: Based on S20, multi-threshold encoding is performed on the pattern change between the marked state neighborhoods;

[0026] S40: Based on S30, the threshold coding is input into five types of weight values, and weighted calculation is performed to obtain polymorphic coding.

[0027] S50: The extracted polymorphic coding, position features and frequency domain features are used as the SMNE total feature set.

[0028] As a further improvement, the obtained feature subset was put into the random forest classifier for classification. The EEG signal was divided into training set and test set by the ten-fold cross validation method. The optimal sleep stage prediction model was established based on the training set data to distinguish W, REM and N1, N2 and N3 data, and the accuracy, F1 score, sensitivity and kappa value were calculated.

[0029] Compared with the prior art, the present invention has the following technical effects:

[0030] (1) The present invention can collect EEG signals in real time and has a small size, making professional sleep EEG staging at home possible;

[0031] (2) The EEG signals are processed by using signal overlap and joint signal-to-noise ratio data enhancement algorithm to solve the problem of sample data imbalance;

[0032] (3) SMNE feature is a new time-domain feature for automatic sleep staging based on the pattern changes of sleep EEG signals. This feature can effectively track and reproduce the pattern changes of EEG signals in different sleep stages by encoding and quantifying local extreme values.

[0033] (4) Combine the time domain, frequency domain and position characteristics of the marked signal to improve classification performance and enhance accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a principle block diagram of a sleep staging system based on multi-threshold neighborhood extreme value pattern statistics.

[0035] Figure 2 This is a flowchart of the onset prediction of the sleep staging system based on multi-threshold neighborhood extreme value pattern statistics in an embodiment of the present invention.

[0036] Figure 3 Schematic diagram of the four steps for creating eigenvectors for multi-threshold neighborhood extreme mode statistics.

[0037] Figure 4 are samples of all possible modes in five states when the multi-threshold neighborhood extreme mode statistical weight layer (w3,:). DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0039] On the contrary, the present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims. Further, in order to make the public have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. Those skilled in the art can fully understand the present invention without the description of these details.

[0040] Accurate sleep stage classification is a prerequisite for evaluating sleep quality. According to the 2007 American Academy of Sleep Medicine rules, single-channel EEG recordings are usually divided into 30-second segments, each of which is manually reviewed by a sleep specialist and then classified into one of the five stages: wakefulness, REM, and three non-REM stages.

[0041] In order to improve the accuracy of sleep staging, the present invention proposes a sleep staging system based on multi-threshold neighborhood extreme value pattern statistics, see Figure 1 , shown as its principle block diagram, including EEG signal collector, signal preprocessor, multi-threshold neighborhood extreme value feature extractor, gray wolf optimizer and random forest classifier, among which,

[0042] The EEG signal collector is used to obtain single-channel EEG signals, i.e., raw EEG data;

[0043] The signal preprocessor is used to receive the original EEG data of the EEG signal collector and perform data preprocessing, including filtering and data enhancement of single-channel EEG signals; wherein, the collected original EEG data is subjected to 0.5Hz-40Hz bandpass filtering, and the processed EEG signal is segmented into segments; in data enhancement, the signal-to-noise ratio data of the specified segment is enhanced according to the duration of the segment annotation after segmentation;

[0044] The multi-threshold neighborhood extreme value feature extractor is used to extract multi-threshold neighborhood extreme value features from the signal preprocessed by the signal preprocessor; wherein, the extreme value of the data segment after data enhancement is first calculated, and then the neighborhood difference of the extreme value is calculated, and the difference is marked with five categories of states, and the multi-threshold neighborhood extreme value mode statistical total feature set is obtained by analyzing the multi-mode encoding of the state;

[0045] The gray wolf optimizer is connected with the multi-threshold neighborhood extreme value feature extractor to determine the signal-to-noise ratio enhancement and multi-threshold neighborhood extreme value pattern statistical features through multiple iterations to extract the optimal threshold;

[0046] The random forest classifier is used to classify the best input feature subset to obtain the final classification result.

[0047] In the above technical solution, symbolic techniques such as multi-threshold neighborhood extreme mode statistics are used to compare and track the dynamics of sleep EEG signals in different sleep stages. One of the main advantages of multi-threshold neighborhood extreme mode statistics is that it can detect changes in EEG signals that transition between sleep stages, using a method of symbolic assignment of dynamic patterns found in single-channel EEG signals. This method makes it possible to detect changes in EEG signals that transition between sleep stages, and trend detection and quantification of changes become possible to improve the efficiency of analyzing sleep disorders.

[0048] In the above technical solution, classification is required after feature extraction, but the setting of different threshold values ​​in signal-to-noise ratio enhancement and polymorphic coding may reduce the classification performance of the model, so it is considered to use an optimizer to confirm the optimal threshold. The present invention uses the Gray Wolf Optimizer to obtain the optimal threshold, and experiments show that it has a better classification effect.

[0049] The present invention adopts the above technical solution and uses the encoded time domain features for the first time to automatically classify sleep stages, effectively improving the accuracy of the automatic sleep staging system based on single-channel EEG signals and solving the problem of sleep stage classification.

[0050] See also Figure 2 , as shown in the present invention, a sleep staging algorithm based on multi-threshold neighborhood extreme value pattern statistics is proposed, wherein the signal preprocessor is used to receive the original EEG data of the electroencephalogram signal collector and perform data preprocessing, and the data processing process is as follows:

[0051] S1: Perform 0.5 Hz to 40 Hz bandpass filtering on the collected raw EEG data, and divide the processed EEG signal into segments.

[0052] S2: Signal shift, shift the filtered signal back five sampling points to get a new 30s segment x i+1 , which is marked as the same sleep period as the old signal;

[0053] S3: After the newly obtained shifted signal, the adjacent segments with the same sleep stage are marked as signal enhancement set segments, and the signal enhancement set segments are overlapped. The design randomly selects a starting point p and uses the following 30s as a data enhancement segment for signal-to-noise ratio enhancement, marked as x i+2 ;

[0054] S4: Calculate the signal-to-noise ratio of the set of segments and take the absolute value to obtain the segment after data enhancement;

[0055] The calculation method of the signal-to-noise ratio is shown in the following formula 1:

[0056]

[0057] Where P signal is the power of the signal, P noise is the power of the noise, log10 represents the logarithm with base 10, A signal Indicates the amplitude of the signal, A noise represents the amplitude of the noise. Therefore, this formula can be used to mark EEG signals that need to improve the signal-to-noise ratio performance. If the calculated signal-to-noise ratio of this segment is lower than the preset threshold s1 obtained by GWO, its absolute value is assigned as the enhancement value, and a new 30s segment x is obtained. i+3 .

[0058] S5: The extracted new segment and the original data set are used as the preprocessed signal segment.

[0059] The multi-threshold neighborhood extreme value time domain feature extractor extracts features from single-channel EEG signals in different sleep stages. The data processing process is as follows:

[0060] S10: Extract local extreme values: including position information labelE(t) and amplitude information E(t). The extraction method is shown in Algorithm 1 below:

[0061]

[0062] And the information obtained is calculated to obtain the extreme value difference and position difference, and D represents the difference between the two extreme values. The formula for adjacent extreme differences is:

[0063]

[0064] Since the poles are marked as position information in the previous text, a domain difference P can be constructed using the position labels to capture the change in the pole positions between domains. The formula for calculating the position difference P can be expressed as:

[0065]

[0066] P reflects the temporal changes of the poles and focuses on the intensity and overall trend of these changes.

[0067] S20: Based on S10, using the extreme value E obtained in the initial step, each extreme value E is arranged by order of magnitude to form a histogram, thereby obtaining an estimate of the global amplitude distribution. The specific process of division is as follows Figure 3 shown.

[0068] S30: Based on S20, when quantizing the difference threshold, the present invention defines five modes, representing different modes of two adjacent symbols in the S(t) state. The present invention uses GWO to obtain the thresholds b1 and b2 for judging the state, and the domain difference corresponding to each element in the five sub-matrices of the state S is shown as follows:

[0069]

[0070] However, some dynamic patterns of EEG signals do not occur in real life. Figure 4 All possible modes are listed in .

[0071] S40: Based on S30, the five states are weighted and multi-state encoded to obtain SMNE features. Due to the close relationship between EEG signals, the weight calculation is divided into 5 weight layers ω n (n=1:5), such as Figure 3 As shown. The ω1 weighted layer consists of a single multi-threshold code with a weight of 0.001. ω2 consists of two multi-threshold codes, the weight of the latter does not change to 0.001, and the weight of the former is 0.01. And so on, 5 weighted layers are formed. These 5 layers are obtained by combining efficiency and accuracy in experiments. This weighting method can restore and predict the signal pattern of each extreme point as clearly as possible, which plays a key role in determining the sleep mode. The specific possible signal patterns are as follows Figure 4 shown.

[0072] S50: The extracted polymorphic coding, position features and frequency domain features are used as the SMNE total feature set.

[0073] Traditional spectrum analysis involves Fourier transforming the signal to achieve the purpose of analysis. In this method, the root mean square frequency (RMSF) is used as the feature extracted in the frequency domain. Frequency domain analysis examines the characteristics of the signal from the perspective of frequency. In signal analysis, time domain analysis and frequency domain analysis complement each other. The method of RMSF extraction is as follows:

[0074] First, the input signal sequence x iPerform an n-point discrete Fourier transform (DFT) and calculate according to equation 6:

[0075]

[0076] Where F(f) is the output with frequency f, is the rotation factor, and from formula 7 we can get:

[0077]

[0078] Where N is x i The length of the extracted frequency signal is then used to extract the root mean square frequency, which is obtained from formula 8:

[0079]

[0080] The gray wolf optimizer is used to determine the optimal threshold through multiple iterations. GWO imitates the hunting method of gray wolves in nature. The wolf pack consists of four levels: α wolf, β wolf, δ wolf and ω wolf. β wolf supports α wolf in hunting prey. δ wolf is at the third level of the dominance hierarchy. The remaining wolves are called ω wolves, which are completely dominated by α, β and δ wolves. Once the prey is found, β and δ wolves will surround and attack the prey under the command of α wolf. The position of the wolf in the wolf pack can be expressed by the thresholds of s1, b1 and b2, as shown in Formula 9:

[0081] W(i w ,j w )=(s1(i w ,j w ),b1(i w ,j w ),b2(i w ,j w )) (9)

[0082] Among them, s1, b1 and b2 represent the thresholds of SNR and SMNE features respectively, i w ∈{1,2,...,I max}, I max is the maximum number of iterations, and j w ∈{1,2,...,N max},N max It is the number of wolves in the pack.

[0083]

[0084] At the beginning of GWO, the eigenvalues ​​of individuals are usually generated by random initialization, and the fitness value of each individual population is calculated according to the fitness function. Since the combination of s1, b1 and b2 will affect the detection accuracy, the accuracy of SMNE algorithm detection is used as the fitness function to evaluate the performance of the convergence factor, which is given by

[0085] Formula 10 yields:

[0086]

[0087] Where N t (i w ,j w ) and N m (i w ,j w ) represent the number of correct sleep stages and the number of incorrect sleep stages detected correctly at this location, respectively. Then, the highest ACC value is selected as the adaptation value.

[0088] The random forest classifier is used to input the best feature subset for classification and obtain the final classification result.

[0089] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A sleep staging system based on multi-threshold neighborhood extreme value pattern statistics, characterized in that: It includes EEG signal collector, signal preprocessor, multi-threshold neighborhood extreme value feature extractor, gray wolf optimizer and random forest classifier, among which, The EEG signal collector is used to obtain single-channel EEG signals, i.e., raw EEG data; The signal preprocessor is used to receive the original EEG data of the EEG signal collector and perform data preprocessing, including filtering and data enhancement of single-channel EEG signals; wherein, the collected original EEG data is subjected to 0.5Hz-40Hz bandpass filtering, and the processed EEG signal is segmented into segments; in data enhancement, the signal-to-noise ratio data of the specified segment is enhanced according to the duration of the segment annotation after segmentation; The multi-threshold neighborhood extreme value feature extractor is used to extract multi-threshold neighborhood extreme value features from the signal preprocessed by the signal preprocessor; wherein, the extreme value of the data segment after data enhancement is first calculated, and then the neighborhood difference of the extreme value is calculated, and the difference is marked with five categories of states, and the multi-threshold neighborhood extreme value mode statistical total feature set is obtained by analyzing the multi-mode encoding of the state; The gray wolf optimizer is connected with the multi-threshold neighborhood extreme value feature extractor to determine the signal-to-noise ratio enhancement and multi-threshold neighborhood extreme value pattern statistical features through multiple iterations to extract the optimal threshold; The random forest classifier is used to classify the best input feature subset to obtain the final classification result.

2. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 1, characterized in that: The EEG signal collector uses dry electrodes to collect EEG signals, transmits the data to a receiving device via a Bluetooth module, and performs filtering on the receiving device.

3. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 2, characterized in that: The time span of each EEG signal segment is 30 s.

4. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 3, characterized in that: The signal preprocessor obtains a set of preprocessed segments, which are used to filter out low-frequency and high-frequency interference components in the EEG signal, and optimize the physiological characteristics of the imbalanced number of segments of the EEG signal by using data enhancement.

5. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 3, characterized in that: The process of data preprocessing by the signal preprocessor is as follows: S1: Filter the collected EEG signal; S2: Based on S1, the filtered signal is shifted back to obtain a new 30s segment; S3: Based on S2, adjacent segments with the same sleep stage are marked as collective segments; S4: Based on S3, the signal-to-noise ratio of the set segments is calculated, and the absolute value is taken to obtain the data-enhanced segments; S5: The extracted new segment and the original data set are used as the preprocessed signal segment.

6. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 1, characterized in that: The multi-threshold neighborhood extreme value feature extractor extracts multi-threshold neighborhood extreme value features and extracts the pattern change statistics of the extreme values ​​to accurately classify the sleep signal. The processing process is as follows: S10: extracting local extreme values: including position information and amplitude information; S20: on the basis of S10, the obtained extreme values ​​are sorted to determine the number and distribution of the five states, ensuring that each extreme value is marked with a state; S30: Based on S20, multi-threshold encoding is performed on the pattern change between the marked state neighborhoods; S40: Based on S30, the threshold coding is input into five types of weight values, and weighted calculation is performed to obtain polymorphic coding; S50: The extracted polymorphic coding, position features and frequency domain features are used as a total feature set of multi-threshold neighborhood extreme value pattern statistics.

7. The sleep staging system based on multi-threshold neighborhood extreme value pattern statistics according to claim 6, characterized in that: The obtained feature subset is input into the random forest classifier for classification, in which the EEG signal is divided into training set and test set using the ten-fold cross validation method.

Citation Information

Patent Citations

  • COG (Chip-On-Glass) offset detection method based on extreme value difference statistical characteristic

    CN104952081A

  • Industrial Internet of Things high-frequency data compression method based on time sequence segmentation and clustering

    CN115459782A

Cited By

  • Rodent automatic sleep staging method and system and storage medium

    CN120257176A