Sleep stage classification method based on multi-scale feature contrast learning

By adopting a multi-scale feature comparison learning method in sleep stage classification, combining the comparative representation learning and pyramid time context learning modules, the problems of insufficient recognition ability and category imbalance in the boundary transition stage in the existing technology are solved, and higher classification accuracy and robustness are achieved.

CN120154301AActive Publication Date: 2025-06-17NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510283841.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing sleep staging method has insufficient recognition capabilities during the boundary transition stage, serious category imbalance problem, high dependence on manual feature extraction, and insufficient utilization of time context information in deep learning models.

Method used

The sleep stage classification method based on multi-scale feature comparison learning is adopted. By comparing and characterizing the combination of learning module, pyramid time context learning module and classification module, the characteristic representation of the EEG signal is optimized, the dynamic change relationship of the EEG signal on different time scales is modeled, and the sleep stage classification is performed.

Benefits of technology

It significantly improves the classification performance of a few classes and boundary transition stages, provides higher classification accuracy and robustness, and especially shows significant advantages in dealing with the problems of fuzzy boundary and category imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120154301A_ABST
    Figure CN120154301A_ABST
Patent Text Reader

Abstract

The invention discloses a sleep stage classification method based on multi-scale feature contrast learning. The method comprises the steps that electroencephalogram signals to be classified are acquired; preprocessing the electroencephalogram signal to obtain an enhanced EEG signal; inputting the EEG signal into the trained multi-scale feature comparison representation learning sleep network model, and outputting a final classification result by using the model; the multi-scale feature comparative representation learning sleep network model comprises a comparative representation learning module, a pyramid time context learning module and a classification module which are connected in sequence; a difficult sample mining unit is arranged in the comparison characterization learning module and is used for mining boundary samples and samples with low classification accuracy from a training data set, introducing a sample weighting mechanism into a supervised comparison loss function, setting weights of the samples and limiting the range of the weights through a maximum weight limiting mechanism; and finally, optimizing network parameters in the contrast representation learning module through a supervised contrast loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of biomedical signal processing and deep learning, and particularly to a sleep stage classification method based on multi-scale feature contrast learning. Background Art

[0002] Sleep is one of the most important physiological activities of humans. Sufficient sleep not only supports the normal functions of various physiological processes of the human body, but also optimizes learning, memory, attention, mood, and decision-making abilities. However, sleep disorders and insufficient sleep are widespread globally, constituting a serious public health problem. Polysomnography (PSG) is considered the gold standard for monitoring sleep states, providing key sleep staging information by measuring multiple physiological signals (including electroencephalogram (EEG), electrooculogram (EOG), electromyogram (EMG), and electrocardiogram (ECG)). Although PSG remains the gold standard in sleep staging, its manual interpretation process is time-consuming and labor-intensive, and there may be inconsistencies among different experts. In addition, the accuracy of manual classification depends to a large extent on the training of technicians, which is particularly impractical in resource-limited environments.

[0003] Therefore, automatic sleep stage classification technologies based on machine learning or deep learning have become an increasing research focus. These methods not only significantly reduce the need for manual intervention, but also improve the efficiency and accuracy of sleep staging. Traditional machine learning-based methods usually rely on manually extracting features, which is time-consuming and labor-intensive, and manual features often perform poorly or unreliably when dealing with new data. To overcome this problem, deep learning-based methods have gradually replaced traditional methods, capable of automatically learning effective features from data, demonstrating stronger generalization ability and higher classification performance.

[0004] In recent years, representation learning methods for EEG have emerged continuously, providing more effective EEG data representations and achieving remarkable results in various tasks. However, existing self-supervised or unsupervised contrast learning methods have not fully utilized the rich annotated PSG data, limiting the improvement of classification accuracy. In addition, it is difficult to effectively capture the temporal context of different sleep stages only relying on contrast learning. Therefore, improving the accuracy of sleep staging, especially in solving the problems of boundary ambiguity and class imbalance, remains a key task. Summary of the Invention

[0005] The object of the present invention is to provide a sleep stage classification method based on multi-scale feature contrast learning to solve the problems of insufficient recognition ability for boundary transition stages, serious class imbalance problems, high dependence on manual feature extraction, and insufficient utilization of temporal context information by deep learning models in existing sleep staging methods.

[0006] To achieve the above tasks, the present invention adopts the following technical solutions:

[0007] A sleep stage classification method based on multi-scale feature contrast learning, comprising:

[0008] Obtain the electroencephalogram (EEG) signal to be classified; preprocess the EEG signal to obtain the enhanced EEG signal; input the EEG signal into the trained multi-scale feature contrast representation learning sleep network model, and use the model to output the final classification result;

[0009] The multi-scale feature contrast representation learning sleep network model includes a contrastive representation learning module, a pyramid time context learning module, and a classification module connected in sequence, wherein:

[0010] The contrastive representation learning module is used to optimize the feature representation of the EEG signal, and obtain the contrastive learning embedded features through the contrastive learning strategy; the pyramid time context learning module is used to model the dynamic change relationship of the EEG signal at different time scales according to the embedded features, and obtain the fused features with temporal consistency; the classification module classifies the sleep stages based on the fused features and finally outputs the corresponding sleep stage classification results;

[0011] A hard sample mining unit is set in the contrastive representation learning module, which is used to mine boundary samples and samples with low classification accuracy from the training dataset, introduce a sample weighting mechanism into the supervised contrastive loss function, set the weights of the samples, and limit the range of the weights through the maximum weight limit mechanism; finally, the network parameters in the contrastive representation learning module are optimized through the supervised contrastive loss function.

[0012] Further, when constructing the training dataset, only the single-channel EEG signal is retained for the obtained EEG signal, all EEG signals are resampled to a preset frequency, and band-pass filtering is used to remove the low-frequency drift and high-frequency noise in the EEG signal; an artifact removal method is adopted to reduce the interference of eye movement and muscle artifacts on the EEG signal; according to the AASM standard, the N3 and N4 stages are merged, and the EEG signals with incorrect labels are removed; a preset length is intercepted for each EEG signal as a sample, and the label of the EEG signal is the corresponding sleep stage.

[0013] Further, the contrastive representation learning module includes a data augmentation unit, a momentum encoder, and a projection network;

[0014] In the data augmentation unit, the input samples are enhanced by using the data augmentation strategy to generate the corresponding positive and negative samples; the samples and the corresponding positive and negative samples are jointly used as the enhanced samples;

[0015] The enhanced samples are input into the momentum encoder and successively pass through a convolutional layer, a batch normalization layer, an average pooling layer, and an activation function to learn complex and abstract feature representations;

[0016] The projection network is used to map the feature representations extracted by the momentum encoder into the embedding space for contrastive learning to obtain the corresponding embedded features; where a multi-layer perceptron MLP with a single hidden layer is used as the projection network.

[0017] Furthermore, the data augmentation strategy is as follows:

[0018] Positive samples: Positive samples are generated by applying various data augmentation operations that are semantically invariant to the samples; the positive samples have the same label as the current sample;

[0019] Negative sample pairs: Samples with different labels from the current sample are randomly selected from the training dataset as negative samples.

[0020] Furthermore, the embedded feature z is expressed as:

[0021]

[0022] where z' = P(h t ; w P ; b P ), h t is the feature representation extracted by the momentum encoder, P(·) is the projection network, w P , b P are the weights and biases of the projection network respectively, is the real number space.

[0023] Furthermore, a hard sample mining unit is set in the contrastive representation learning module for mining boundary samples and low classification accuracy samples from the training dataset, specifically:

[0024] Boundary samples at the sleep stage transition points are identified through the labels of the samples, thus forming a boundary sample set, which is defined as:

[0025] Β = {i|y true [i] ≠ y true [i + 1]}

[0026] where Β is the boundary sample set, y true [i] is the label of the sleep stage corresponding to the i-th sample; the i-th sample that satisfies the condition y true [i] ≠ y true [i + 1] is used as the boundary sample;

[0027] Low classification accuracy samples of classes below the threshold θ accuracy are defined as:

[0028] A = {i | y true [i] = c, p o,c < θ accuracy}

[0029] Among them, A is a set of samples with low classification accuracy, c represents the category of a sleep stage, and p o,c represents the probability that the i-th sample belongs to category c predicted by the multi-scale feature contrast representation learning sleep network model; the samples that satisfy y true [i] = c, p o,c < θ accuracy are used as samples with low accuracy.

[0030] Furthermore, a sample weighting mechanism is introduced into the supervised contrast loss function, the weights of the samples are set, and the range of the weights is limited by the maximum weight limit mechanism, including:

[0031] A sample weighting mechanism is introduced into the supervised contrast loss function, and the weight w of each sample i is defined as:

[0032]

[0033] where the adjustable parameters λ a and λ b are set to an initial value of 1;

[0034] A maximum weight limit mechanism is introduced to limit the weight value within a reasonable range:

[0035] w' i = min(w i , w max )

[0036] where w' i represents the weight of the i-th sample after the maximum weight limit, and w max is a hyperparameter used to limit the maximum value of the weight;

[0037] The supervised contrast loss function L sc is specifically:

[0038]

[0039] Among them, k and a represent the k-th and a-th samples, i ≠ k and i ≠ a, and the i-th sample and the k-th sample have the same label, z i , z k , z a are the embedded features of the i-th, k-th, and a-th samples respectively; N bDenote the number of positive / negative samples. Let \(N(i)\) be the index set of all positive samples of sample \(i\), and \(|N(i)|\) be its quantity; \(\tau\) is a preset temperature parameter.

[0040] Furthermore, the pyramid temporal context learning module includes a multi-scale pyramid pooling network and a convolutional network. The embedded features of the sample are first input into the convolutional network to obtain output features. Then, the output features are input into the multi-scale pyramid pooling network, and the temporal features output by all layers of the multi-scale pyramid pooling network are concatenated to capture features containing temporal scale information of short-term, medium-term, and long-term respectively, so as to form a fused feature with temporal consistency.

[0041] Furthermore, the fused feature is input into the self-attention mechanism of the classification module, and the output of the self-attention mechanism is input into the fully connected layer to obtain the sleep stage classification result corresponding to the sample, that is, the probability that the sample belongs to a certain category predicted by the multi-scale feature contrast representation learning sleep network model.

[0042] A terminal device includes a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, the sleep stage classification method based on multi-scale feature contrast learning is implemented.

[0043] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the sleep stage classification method based on multi-scale feature contrast learning is implemented.

[0044] Compared with the prior art, the present invention has the following technical features:

[0045] By combining contrastive representation learning, hard sample mining, and multi-scale pyramid pooling strategies, the present invention comprehensively captures the dynamic features of sleep EEG signals at different time scales, effectively improving the classification performance of minority classes and boundary transition stages. The newly added boundary error rate index provides a multi-dimensional evaluation framework for evaluating the prediction accuracy of the model at the stage transition point. Experimental results show that the proposed MFCR-SleepNet performs excellently on multiple public datasets, and its accuracy and robustness are better than existing mainstream models, especially showing significant advantages in dealing with boundary ambiguity and class imbalance problems. Description of the Drawings

[0046] Figure 1 It is a schematic diagram of the overall architecture of the multi-scale feature contrast representation learning sleep network model (MFCR-SleepNet);

[0047] Figure 2 It is a detailed structure diagram of the pyramid-based temporal context learning (PTCL) module;

[0048] Figure 3 This is a graph comparing the network performance of MFCR-SleepNet on three datasets in an embodiment of the present invention;

[0049] Figure 4 This is a graph showing the ablation experiment results of MFCR-SleepNet on the SleepEDF-78 dataset in an embodiment of the present invention;

[0050] Figure 5 This is a graph comparing the BER of three datasets in an embodiment of the present invention. Detailed implementation manner

[0051] The present invention provides a sleep stage classification method based on multi-scale feature contrast learning, including:

[0052] Obtain the electroencephalogram (EEG) signal to be classified; preprocess the EEG signal to obtain an enhanced EEG signal;

[0053] Input the EEG signal into a trained multi-scale feature contrast representation learning sleep network model, and use the model to output the final classification result;

[0054] The multi-scale feature contrast representation learning sleep network model includes a contrastive representation learning module, a pyramid time context learning module, and a classification module connected in sequence, where:

[0055] The contrastive representation learning module is used to optimize the feature representation of the EEG signal, enabling the model to automatically extract discriminative features through a contrastive learning strategy, and obtaining a contrastive learning embedding space; the pyramid time context learning module is used to model the dynamic change relationship of the EEG signal at different time scales based on the embedding space, improving the ability to capture sleep stage transitions, and obtaining temporally consistent fused features; the classification module accurately classifies the sleep stage based on the fused features and finally outputs the corresponding sleep stage classification result;

[0056] A hard sample mining unit is set in the contrastive representation learning module to mine boundary samples and samples with low classification accuracy from the training dataset, introduce a sample weighting mechanism in the supervised contrastive loss function, set the weights of the samples, and limit the range of the weights through a maximum weight limit mechanism; finally, the network parameters in the contrastive representation learning module are optimized through the supervised contrastive loss function.

[0057] 1. Construct a training dataset.

[0058] The process of constructing the dataset for training the multi-scale feature contrast representation learning sleep network model in the present invention is as follows:

[0059] Three publicly available sleep datasets, SleepEDF-78, MASS, and SHHS, are used, and the data is preprocessed by standardization to ensure the consistency of the training and test data of the model.

[0060] For the selected sleep datasets, the following standard preprocessing is adopted:

[0061] In each sleep dataset, only the single-channel EEG signal is retained from the electroencephalogram signals, and redundant signals such as EOG and EMG are removed; all EEG signals are resampled to 100 Hz to maintain consistency; band-pass filtering (0.3 Hz - 35 Hz) is used to remove low-frequency drift and high-frequency noise in the EEG signals; an artifact removal method is adopted to reduce the interference of eye movement and muscle artifacts on the EEG signals; according to the AASM standard, stages N3 and N4 are combined, and the relevant EEG signals with incorrect labels are removed.

[0062] After standard preprocessing, a training dataset is obtained. Each sample in the training dataset is a 30-s EEG signal, and the label of the EEG signal is the manually annotated sleep stage (W, N1, N2, N3, REM); the length of the EEG signal depends on the sampling rate of the sleep dataset. For example, it is 3000 points for SleepEDF-78, 7680 points for MASS, and 3000 points for SHHS.

[0063] 2. Multi-scale Feature Contrast Representation Learning Sleep Network Model (MFCR-SleepNet).

[0064] See Appendix Figure 1 The multi-scale feature contrast representation learning sleep network model provided by the present invention includes a contrastive representation learning (CRL) module, a pyramid time context learning (PTCL) module, and a classification module, where:

[0065] 2.1 Contrastive Representation Learning (CRL) Module.

[0066] The described contrastive representation learning module is used to optimize the feature representation of EEG signals, enabling the model to automatically extract discriminative features through a contrastive learning strategy to obtain the sleep stage feature embedding representation.

[0067] Among them, the contrastive representation learning module includes a data augmentation unit, a momentum encoder, a projection network, and a hard sample mining unit, where:

[0068] (1) Data Augmentation Unit.

[0069] In the training stage, for a sample s in the training dataset, the sample is augmented using a data augmentation strategy to generate a positive sample s p , and a negative sample s is randomly selected q , and the augmented sample is represented as The number of positive samples and negative samples is the same, both are N b , that is, p, q = 1, 2, ..., N b .

[0070] Among them, the specific data augmentation strategy is as follows:

[0071] Positive samples: Positive samples are generated by applying six semantic-invariant data augmentation operations (amplitude scaling, time translation, amplitude shift, zero masking, additive Gaussian noise, and band-stop filtering) to the sample s; the positive samples have the same label as the current sample.

[0072] Negative sample pairs: Samples with different labels from the current sample are randomly selected from the training dataset as negative samples.

[0073] (2) Momentum encoder.

[0074] The momentum encoder receives the augmented samples and successively passes through a convolutional layer (Conv), a batch normalization layer (BN), an average pooling layer (Avg), and an activation function (ReLU θt ) to learn complex and abstract feature representations h t . The role of the momentum encoder is to reduce the feature drift problem during training, make the feature representations more consistent and robust, and ensure the stability of the features.

[0075] The processing process of the momentum encoder is shown as follows:

[0076]

[0077] The activation function in the momentum encoder uses ReLU, denoted as θ t is the learnable parameter of the momentum encoder (including the weights and biases of the convolutional layer, batch normalization layer, and average pooling layer), and is specifically expressed as follows:

[0078] θ t+1 = αθ t + (1 - α)θ′ t

[0079] θ t+1 is the parameter of the momentum encoder at the next moment, θ t is the parameter at the current moment, θ t ' is the parameter updated by network learning, and all are trainable parameters; α is the momentum coefficient, which is used to control the fusion ratio of historical parameters and new parameters, and is set to 0.9 in this embodiment.

[0080] (3) Projection network.

[0081] The projection network is used to project the feature representation h extracted by the momentum encodert Map to the embedding space for contrast learning to obtain corresponding embedding features for calculating the supervised contrast loss, so as to better implement the strategy of contrastive representation learning.

[0082] Use a multi-layer perceptron MLP with a single hidden layer dimension d z = 256 as the projection network; the embedding feature z is expressed as:

[0083]

[0084] where z' = P(h t ; w P ; b P ), P(·) is the projection network, w P , b P are the weights and biases of the projection network respectively, is the real number space.

[0085] The contrastive representation learning (CRL) module is trained using a supervised contrast loss function to optimize the network parameters (weights, biases, etc. of each functional layer) in the momentum encoder and the projection network. The specific design is as follows:

[0086] (4) Hard sample mining unit.

[0087] After obtaining the embedding features from the projection network, first perform hard sample mining on them, and then perform weighted optimization on the hard samples. The selected hard samples mainly include two categories: boundary samples and samples with low classification accuracy.

[0088] The hard sample mining module is used to identify and optimize difficult-to-classify samples, especially sleep stage boundary samples and minority class samples, to improve the model's learning ability for these classes. Its input is the embedding features output by the projection network, and the output is the weight w′ i required for weighting. The process is as follows:

[0089] Identify boundary samples at the sleep stage transition points through the labels of the samples, thereby forming a boundary sample set, which is defined as:

[0090] Β = {i|y true [i] ≠ y true [i + 1]}

[0091] where Β is the boundary sample set, and y true [i] is the label of the sleep stage corresponding to the i-th sample; the i-th sample that satisfies the condition y true [i] ≠ y true [i + 1] is used as a boundary sample.

[0092] Below the threshold θ accuracySamples with low classification accuracy for a certain category, which are defined as:

[0093] Α = {i|y true [i] = c, p o,c <θ accuracy}

[0094] Among them, A is the set of samples with low classification accuracy, c represents the category of a specific sleep stage, and p o,c represents the probability that the i-th sample belongs to category c predicted by the multi-scale feature comparison representation learning sleep network model; in this embodiment, θ accuracy is taken as 0.7; the samples that satisfy y true [i] = c, p o,c <θ accuracy are used as samples with low accuracy.

[0095] All samples in the set A of samples with low classification accuracy and the set B of boundary samples are used as difficult samples to construct a difficult sample mining target set Η = Β ∪ Α.

[0096] To enhance the model's attention to difficult samples, a sample weighting mechanism is introduced in the supervised contrast loss function, and the weight w of each sample i is defined as:

[0097]

[0098] Among them, the adjustable parameters λ a and λ b are set with initial values of 1. As the training progresses, they are dynamically adjusted according to the number of difficult samples and boundary samples. When the contribution of difficult samples to the overall loss is small, λ a and λ b will gradually increase to enhance the model's attention to these samples.

[0099] To prevent the gradient from being unstable due to excessive weights during model training, a maximum weight limit mechanism is introduced to limit the weight value within a reasonable range:

[0100] w′ i = min(w i , w max )

[0101] Among them, w′ i represents the weight of sample i after the maximum weight limit, and w max is a hyperparameter used to limit the maximum value of the weight. In this embodiment, w max is set to 3, and this value can effectively balance the impact of difficult samples on the loss and the overall training stability of the model.

[0102] (5) Supervised contrastive loss function.

[0103] The embedded feature z obtained by passing the sample through the projection network is combined with the weight w' of the sample i , and input into the supervised contrastive loss function L sc , specifically:

[0104]

[0105] where k and a represent the k-th and a-th samples, i ≠ k and i ≠ a, and the i-th sample and the k-th sample have the same label, z i , z k , z a are the embedded features of the i-th, k-th, and a-th samples respectively; N b represents the number of positive / negative samples, N(i) is the index set of all positive samples of sample i, and |N(i)| is its number; since the differences between sleep categories are small, in order to strengthen the influence of the small differences between embedded vectors, the temperature parameter τ is set to 0.07.

[0106] During the training of the network model, the above supervised contrastive loss function is used to optimize the network parameters in the momentum encoder and the projection network; in this way, the model can fully focus on difficult samples in the early stage of training and gradually stabilize in the later stage of training, avoiding over-focusing on certain samples and affecting the overall performance. In addition, regularization is also performed on the updates of λ a and λ b to make them have a certain smoothness during the optimization process, thus avoiding the violent fluctuations caused by too fast updates. This dynamic adjustment mechanism combined with the maximum weight limit significantly improves the adaptability of the model at different training stages, making its classification performance on low classification accuracy samples and boundary samples significantly improved, while maintaining the learning ability for mainstream samples and the stability of the training process.

[0107] 2.2 Pyramid temporal context learning module.

[0108] The PTCL module is used to model the dynamic change relationship of samples at different time scales, improve the ability to capture sleep stage transitions, and obtain more temporally consistent classification features.

[0109] As Figure 2 shown, the pyramid temporal context learning module includes a multi-scale pyramid pooling network (XSPP) and a convolutional network (CNN); the embedded feature z of the sample is first input into the convolutional network to obtain the output feature z out , expressed as follows:

[0110] z out = PReLU(w c*z + b c )

[0111] Among them, PReLU is the activation function of the convolutional network, enabling the network to better capture subtle differences in the data; * represents convolutional calculation, w c and b c are learnable weights and biases.

[0112] Then, the output feature z out is input into the multi-scale pyramid pooling network:

[0113] r s = XSPP(z out )

[0114] Among them, XSPP represents the processing process of the multi-scale pyramid pooling network, and r s represents the temporal feature obtained after pooling at the s-th layer of the multi-scale pyramid pooling network; the temporal features r1, r2,..., r s output by all layers of the multi-scale pyramid pooling network are concatenated to respectively capture features containing short-term, medium-term, and long-term temporal scale information, thereby forming a fused feature x concat :

[0115] x concat = concat(r1, r2,..., r s )

[0116] Among them, concat represents the concatenation operation.

[0117] 2.3 Classification module.

[0118] After obtaining the fused feature x containing short-term, medium-term, and long-term temporal scale information output by the pyramid time context learning module, the fused feature x concat is input into the self-attention mechanism of the classification module; among them, since the lengths of the three types of features contained in the fused feature x concat are inconsistent, padding operations can be used to make their lengths consistent. concat The self-attention mechanism can calculate the long-range dependencies of the data. By calculating the attention scores between different time points, the model can highlight the most crucial information and discard unimportant information. The specific implementation is as follows:

[0119]

[0120]

[0121] Among them, Q, K, and V are the queries, keys, and values respectively generated from the fused feature x concat according to three different learnable matrices; d k kis the dimension of the key vector, which helps to adjust the dot product size and prevent the large inner product from affecting the gradient of the softmax function; the superscript T represents the transpose.

[0122] Input the output A of the self-attention mechanism into the fully connected layer to obtain the sleep stage classification result corresponding to the sample, that is, the probability p that the sample belongs to class c predicted by the multi-scale feature contrast representation learning sleep network model. o,c 。

[0123] 4. Training of the network model.

[0124] During the training of the network model, the cross-entropy loss function is used to calculate and evaluate the classification accuracy, thereby optimizing the network parameters of the entire network model; the cross-entropy loss function is expressed as follows:

[0125]

[0126] Among them, C represents the number of categories of sleep stages, and y o,c is a binary indicator indicating whether the sample o belongs to class c; the gradient descent algorithm is used to train the network model, and the trained network model is saved.

[0127] In the actual application of the present invention, after obtaining an unknown EEG signal, it is preprocessed to obtain an EEG signal with a length of 30 s, and then it is input into the trained multi-scale feature contrast representation learning sleep network model, and finally the classification result of the EEG signal is output through the network model.

[0128] Example:

[0129] 1. Model training

[0130] 10-fold cross-validation is adopted, the optimizer uses Adam, the initial learning rate is 0.0005, the batch size is 512, and the early stopping strategy is adopted to prevent overfitting. All training processes are carried out on a Nvidia GeForce GTX3090 GPU in the Python 3.10.2 environment.

[0131] 2. Evaluation metrics

[0132] Include accuracy (ACC), macro-average F1 score (MF1), Cohen's Kappa, and boundary error rate (BER). Among them, BER is a new index proposed by the present invention, which is used to comprehensively evaluate the prediction accuracy of the model at the sleep stage transition point.

[0133] In the present invention, several common evaluation metrics are adopted to comprehensively evaluate the performance of the model. Accuracy (ACC) is used to measure the ratio of the number of samples correctly predicted by the model to the total number of samples, and the formula is as follows:

[0134]

[0135] Among them, TP, TN, FP, and FN represent True Positive, True Negative, False Positive, and False Negative respectively, referring to the number of samples in the correct or incorrect categories during recognition.

[0136] The macro-average F1 score (MF1) combines precision and recall, and is particularly suitable for multi-classification tasks in imbalanced datasets. It can comprehensively evaluate the performance of the model. Its formula is:

[0137]

[0138] Among them, C represents the number of sleep stage categories.

[0139] Cohen's Kappa value is used to measure the consistency between the model's prediction results and the true labels, and corrects the influence brought by random guessing. The F1 score for each category reflects the model's ability to distinguish each category. The boundary error rate (BER) is a new evaluation index proposed for the characteristics of sleep EEG data, used to measure the error between the predicted sleep stage boundary (switching point) and the true label. This index consists of two parts, including the time error (TE) of the boundary and the post-boundary stage consistency (PC). The formula for TE is as follows:

[0140]

[0141] Among them, t true,i is the position of the i-th true boundary, t pre d ,i is the position of the predicted boundary closest to this true boundary, and N represents the total number of true boundaries. The smaller the boundary error, the more accurate the model is in locating the boundary. Normalize TE TE max is the observed value of the maximum error. The formula for PC is as follows

[0142]

[0143] Among them, y pred,t is the stage label predicted by the model after the boundary t, and y true,t is the true stage label after the boundary t. Then calculate the mean of PC

[0144] Weight-multiply TE and PC avg to obtain the final comprehensive index BER. The formula for BER is:

[0145] BER = α·TE norm + β(1 - PC avg )

[0146] where α and β are the corresponding weighting factors, which are taken as 0.5 and 0.5 respectively here. (1 - PC avg ) is the reverse index of stage consistency because we hope that better stage consistency can reduce the value of BER. Specifically, the lower the BER, the better the performance of the model in boundary classification.

[0147] 3. Experimental Results

[0148] The experimental results show that MFCR - SleepNet outperforms the existing mainstream models on the SleepEDF - 78, MASS, and SHHS datasets. Especially, the classification accuracy in the N1 stage is significantly improved. t - SNE visualization shows that the model can better distinguish different sleep stages in the feature space, and the Boundary Error Rate (BER) index further verifies the high accuracy of the model at the stage transition points.

[0149] Figure 3 Shows the classification performance of MFCR - SleepNet on the three datasets of SleepEDF - 78, MASS, and SHHS. Compared with the existing models (DeepSleep, SeqSleepNet, XSleepNet, etc.), MFCR - SleepNet achieves the best results in terms of ACC (accuracy), MF1 (macro - average F1 - score), and Kappa score, indicating its stronger generalization ability. In addition, the classification performance in the N1 stage is significantly improved (e.g., the F1 - score reaches 61.8% on the MASS dataset), which benefits from the CRL module improving the feature discrimination ability of minority - class samples. At the same time, the F1 - score in the REM stage is also at the leading level, indicating that the model can effectively identify the key features in the sleep stages.

[0150] Figure 4The impacts of different module combinations on the model performance were analyzed. Among them, the results of EN+AP (using the encoding network and average pooling), EN+MSP (using the encoding network and pyramid pooling), ME+AP (using the momentum encoder and average pooling), ME+MSP (using the momentum encoder and pyramid pooling), and ME+MSP+HSM (the method of the present invention, combining the momentum encoder, pyramid pooling, and hard sample mining) show that the complete MFCR-SleepNet scheme (ME+MSP+HSM) reaches the optimal level in terms of ACC, MF1, and Kappa scores. Especially in the N1 stage, the F1-score is improved from 42.5% (EN+AP) to 55.1% (ME+MSP+HSM), proving that the CRL module and the PTCL module effectively improve the recognition ability of difficult-to-classify categories. In addition, MSP (multi-scale pooling) enhances the classification performance in the N3 and REM stages, while HSM (hard sample mining) optimizes the learning of boundary samples, further improving the classification stability.

[0151] Figure 5 The boundary error rates (BER) of different methods on the SleepEDF-78, MASS, and SHHS datasets are shown. The results show that MFCR-SleepNet achieves the lowest BER on all datasets (for example, the BER on the MASS dataset drops to 0.1932, lower than methods such as XSleepNet and SeqSleepNet), indicating that it has a lower misclassification rate at the sleep stage transition points and improves the recognition ability of stage transition samples. The low BER is attributed to the fact that the CRL supervised contrast learning optimizes the feature expression of boundary samples, while the PTCL time context modeling enhances the model's capture of long-term and short-term dependencies, making the classification in the transition region more stable and accurate.

[0152] Features and advantages of this solution:

[0153] Higher overall classification performance: MFCR-SleepNet has better ACC, MF1, and Kappa scores than existing mainstream models on the SleepEDF-78, MASS, and SHHS datasets, proving that it has stronger generalization ability and can maintain high performance on different datasets.

[0154] Significantly improved classification performance in the N1 stage: Traditional models have poor classification performance in the N1 stage, while in this solution, the F1-score is increased to 61.8% on the MASS dataset and also leads in the SleepEDF-78 and SHHS datasets. The CRL module improves the recognition ability in the N1 stage by optimizing the feature representation of minority class samples.

[0155] Lower boundary error rate (BER) and improved stage transition recognition ability: Experimental results show that the BER of this solution on all datasets is significantly lower than that of existing methods. For example, on the MASS dataset, the BER is only 0.1932, showing a significant decrease compared to XSleepNet (0.2413) and SeqSleepNet (0.2789). This is due to the fact that CRL supervised contrastive learning optimizes the feature representation of boundary samples, while PTCL time context modeling improves the stability of sleep stage transitions.

[0156] Effectively utilize time context information and improve classification stability: The PTCL module combines convolutional pyramid pooling (XSPP-CNN) and self-attention mechanism, which can capture short-term, medium-term, and long-term dependencies simultaneously, enabling the model to more accurately distinguish adjacent sleep stages and reduce misclassification.

[0157] Applicable to single-channel EEG and enhance the applicability of wearable devices: This solution mainly uses single-channel EEG, but its classification performance exceeds that of many multi-channel EEG methods, indicating its strong robustness and practical application value, especially suitable for low-cost EEG acquisition scenarios such as wearable devices and mobile monitoring systems.

[0158] This solution combines supervised contrastive learning (CRL) + time context modeling (PTCL), and demonstrates significant advantages in overall classification performance, minority class sample recognition, stage transition point classification, time series modeling, and single-channel EEG adaptability, providing an efficient and stable deep learning solution for the sleep stage classification task.

[0159] The above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A sleep stage classification method based on multi-scale feature contrast learning, characterized in that: include: Obtaining an EEG signal to be classified; Preprocess the EEG signal to obtain an enhanced EEG signal; Input the EEG signal into the trained multi-scale feature contrast representation learning sleep network model, and use the model to output the final classification result; The multi-scale feature contrast representation learning sleep network model includes a contrast representation learning module, a pyramid time context learning module and a classification module connected in sequence, wherein: The contrastive representation learning module is used to optimize the feature representation of the EEG signal and obtain the embedded features of the contrastive learning through the contrastive learning strategy; the pyramid time context learning module is used to model the dynamic change relationship of the EEG signal on different time scales according to the embedded features to obtain the fusion features with temporal consistency; the classification module classifies the sleep stages based on the fusion features and finally outputs the corresponding sleep stage classification results; The contrastive representation learning module is provided with a difficult sample mining unit for mining boundary samples and samples with low classification accuracy from the training data set, and a sample weighting mechanism is introduced into the supervised contrastive loss function to set the sample weight and limit the weight range through the maximum weight limit mechanism; finally, the network parameters in the contrastive representation learning module are optimized through the supervised contrastive loss function.

2. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: When constructing the training data set, only single-channel EEG signals are retained for the acquired EEG signals, all EEG signals are resampled to the preset frequency, and bandpass filtering is used to remove low-frequency drift and high-frequency noise in the EEG signals; artifact removal methods are used to reduce the interference of eye movement and muscle artifacts on EEG signals; According to the AASM standard, the N3 and N4 stages are merged, and the related EEG signals with incorrect labels are eliminated; a preset length is cut off for each EEG signal as a sample, and the label of the EEG signal is the corresponding sleep stage.

3. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The contrastive representation learning module includes a data enhancement unit, a momentum encoder, and a projection network; In the data enhancement unit, the input samples are enhanced using the data enhancement strategy to generate corresponding positive and negative samples; The sample and the corresponding positive sample and negative sample are taken as the enhanced sample; The enhanced samples are input into the momentum encoder and sequentially go through the convolution layer, batch normalization layer, average pooling layer, and activation function to learn complex and abstract feature representations; The projection network is used to map the feature representation extracted by the momentum encoder to the embedding space of contrastive learning to obtain the corresponding embedded features; A single hidden layer multi-layer perceptron MLP is used as the projection network.

4. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The embedding feature z is expressed as: Where z'=P(h t ;w P ; b P ), h t is to represent the features extracted by the momentum encoder, P(·) is the projection network, and w P 、b P are the weights and biases of the projection network, respectively. is the real number space.

5. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The contrast representation learning module is provided with a difficult sample mining unit for mining boundary samples and samples with low classification accuracy from the training data set, specifically: The boundary samples at the transition points of the sleep stages are identified by the labels of the samples, thereby forming a boundary sample set, which is defined as: Β={i|y true [i]≠y true [i+1]} Among them, Β is the boundary sample set, y true [i] is the label of the sleep stage corresponding to the i-th sample; it will satisfy y true [i]≠y true The i-th sample of the [i+1] condition is taken as the boundary sample; Below the threshold value θ accuracy The low classification accuracy samples of the category are defined as: A={i|y true [i]=c,p o,c <θ accuracy } Among them, A is a set of samples with low classification accuracy, c represents a category of sleep stage, and p o,c represents the probability of the i-th sample belonging to category c after multi-scale feature comparison and learning sleep network model prediction; true [i]=c,p o,c <θ accuracy The samples under the condition are regarded as low accuracy samples.

6. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The sample weighting mechanism is introduced into the supervised contrast loss function, the sample weight is set, and the range of the weight is limited by the maximum weight limit mechanism, including: A sample weighting mechanism is introduced in the supervised contrast loss function, and the weight w of each sample is i Defined as: The adjustable parameter λ a and λ b Set the initial value to 1; A maximum weight limit mechanism is introduced to limit the weight value to a reasonable range: In i ′=min(in i ,In max ) Among them, w i ′ represents the weight of sample i after the maximum weight limit, w max Is a hyperparameter used to limit the maximum value of the weight; Supervised contrast loss function L sc Specifically: Where k and a represent the kth and ath samples, i≠k and i≠a, and the ith and kth samples have the same label, z i 、z k 、z a are the embedded features of the i-th, k-th, and a-th samples respectively; N b represents the number of positive samples / negative samples, N(i) is the index set of all positive samples of sample i, and |N(i)| is its number; τ is the preset temperature parameter.

7. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The pyramid time context learning module includes a multi-scale pyramid pooling network and a convolutional network; the embedded features of the sample are first input into the convolutional network to obtain the output features; then the output features are input into the multi-scale pyramid pooling network, and the temporal features output by all layers of the multi-scale pyramid pooling network are cascaded to capture the features containing short-term, medium-term and long-term time scale information respectively, thereby forming a fused feature with temporal consistency.

8. The sleep stage classification method based on multi-scale feature contrast learning according to claim 1, characterized in that: The fused features are input into the self-attention mechanism of the classification module, and the output of the self-attention mechanism is input into the fully connected layer to obtain the sleep stage classification result corresponding to the sample, that is, the probability of the sample belonging to a certain category obtained by the multi-scale feature comparison representation learning sleep network model.

9. A terminal device comprising a processor, a memory and a computer program stored in the memory; characterized in that: When the processor executes the computer program, it implements the sleep stage classification method based on multi-scale feature contrast learning according to any one of claims 1-8.

10. A computer-readable storage medium, wherein a computer program is stored in the medium; characterized in that: When the computer program is executed by a processor, the sleep stage classification method based on multi-scale feature contrast learning according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Motor imagery electroencephalogram signal classification method based on frequency domain convolutional neural network

    CN113476056A

  • Automatic sleep staging method and system based on ResCNN-BiGRU

    CN119073924A

  • Automatic sleep staging method and system

    CN119312143A

  • Sleep detection and analysis system

    US11382534B1

  • Apparatus for automatically determining sleep disorder using deep running and operation method of the apparatus

    US20210045676A1

Cited By

  • Balanced contrast learning and multi-scale characterization-based electroencephalogram processing device and method

    CN121845614A

  • Single-channel electroencephalogram sleep analysis method and related equipment

    CN121971042A