A method and system for monitoring and evaluating sleep quality
Through the KAN-BiLSTM-Transformer model, the problems of insufficient multimodal information fusion and poor generalization performance are solved, efficient and accurate sleep staging prediction is achieved, and it is suitable for automated analysis of multimodal sleep signals.
Patent Information
- Application Number
- CN202411513301.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing sleep staging methods are insufficient in multimodal information fusion and poor generalization performance, resulting in insufficient accuracy and reliability, especially in cross-individual and cross-device scenarios.
The KAN-BiLSTM-Transformer model is adopted, combined with the Kolmogorov-Arnold network (KAN), bidirectional long and short-term memory network (BiLSTM) and Transformer, an adaptive spatiotemporal feature learning network is designed, and the model optimization is used to optimize the weighted Dice loss and timing smoothing loss through adaptive feature extraction, timing dependence modeling of multimodal signals.
It significantly improves the accuracy and generalization ability of sleep staging, can automatically learn the discriminant characteristics of multimodal sleep signals, captures space-time dependencies, and provides an efficient and interpretable sleep staging prediction model.
Smart Images

Figure BDA0005105992390000021 
Figure BDA0005105992390000031 
Figure BDA0005105992390000035
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sleep quality monitoring and evaluation, and in particular, a method and system for sleep quality monitoring and evaluation are proposed. Background Art
[0002] Sleep quality has a profound impact on human health and daily performance. Long-term sleep problems can not only lead to daytime fatigue and decreased concentration, but also increase the risk of various chronic diseases, such as obesity, diabetes, cardiovascular diseases, etc. The key to objectively evaluating sleep quality is to accurately judge sleep stages. Traditional sleep staging methods rely on manual interpretation of polysomnogram (PSG) data, which is time-consuming and easily affected by subjective factors. With the popularization of wearable devices and the development of artificial intelligence technology, it has become possible to automate sleep staging using machine learning and deep learning algorithms, and it has attracted increasing attention.
[0003] Existing automatic sleep staging methods can be roughly divided into the following categories:
[0004] 1. Traditional machine learning methods based on handcrafted features, such as support vector machine (SVM) and random forest (RF). These methods need to extract various statistical features and frequency domain features from the original PSG signals, and then input them into the classifier for training and prediction. Their limitation is that feature engineering requires a lot of domain knowledge and manpower, and it is difficult to fully utilize the temporal information in the original signals.
[0005] 2. End-to-end methods based on deep learning, such as convolutional neural network (CNN) and long short-term memory network (LSTM). These methods can directly input the original PSG signals, extract local features through convolutional layers, and then model the temporal dependencies using recurrent neural networks. It overcomes the limitations of handcrafted features, but has low computational efficiency and insufficient expressive ability when dealing with long sequences.
[0006] Although the above methods have made some progress in the sleep staging task, they still face the following problems:
[0007] 1. Insufficient multi-modal information fusion. In addition to core PSG signals such as EEG, EOG, and EMG, physiological indicators such as respiration, heart rate, body movement, and environmental factors such as temperature and noise also contain important sleep state information. Most existing methods only utilize a single signal source and lack the comprehensive utilization of multi-modal information.
[0008] 2. Poor generalization performance. Due to the large differences in the distributions of PSG data collected by different individuals and different devices, the generalization performance of existing models in cross-individual and cross-device scenarios often decreases significantly, and their reliability is insufficient in practical applications. Summary of the Invention
[0009] In view of the above problems, the present invention proposes a method and system for detecting and evaluating sleep quality,
[0010] A sleep stage prediction method based on KAN-BiLSTM-Transformer shows superiority over the prior art in multiple aspects:
[0011] 1. For the first time, KAN (Kolmogorov-Arnold Network) is introduced into the field of sleep staging. KAN is a novel neural network structure, and its core idea is to approximate any continuous function with a combination of unary functions. Compared with traditional neural networks, KAN has stronger representation ability. Replace CNN with KAN as the feature extractor to learn richer feature representations from multi-modal physiological signals.
[0012] 2. Integrate the respective advantages of BiLSTM and Transformer. BiLSTM is good at modeling long-term temporal dependencies, while Transformer is good at capturing key local patterns. First, input the features extracted by KAN into BiLSTM to learn temporal context information, and then input them into Transformer to enhance key features. The advantages of the two complement each other, enabling a more comprehensive exploration of the spatio-temporal features of multi-modal sleep signals.
[0013] 3. Design a novel learning framework. In addition to the main sleep staging task, introduce the prediction of physiological indicators such as heart rate and respiration as auxiliary means, so as to better utilize multi-modal information and understand the sleep state from different perspectives, thereby improving the prediction staging performance and generalization ability.
[0014] The significance of the present invention lies in: 1) exploring the application of KAN in the analysis of sleep quality physiological signals; 2) innovatively integrating the advantages of KAN, BiLBTM, and Transformer to design an efficient, interpretable, and robust end-to-end model for the sleep staging task; 3) making full use of multi-modal sleep information to lay a foundation for the development of a comprehensive and accurate sleep monitoring system.
[0015] On the one hand, the present invention proposes a method for detecting and evaluating sleep quality, which is based on the sleep stage prediction method of KAN-BiLSTM-Transformer, and includes the following steps:
[0016] 1. Multi-modal sleep signal acquisition and quality enhancement: Perform quality enhancement processing such as denoising, filtering, and normalization on signals such as EEG, EOG, EMG, ECG, respiration, SpO2, body movement, and environment. For EEG and EOG signals, use an adaptive filtering algorithm based on empirical mode decomposition (EMD) to remove noise, and the formula is:
[0017]
[0018] Among them, is the denoised signal, and c i (t) is the i-th intrinsic mode function (IMF) component, and β i is the weight coefficient of the i-th IMF component, and n is the total number of IMF components.
[0019] Then, an adaptive threshold denoising algorithm based on wavelet packet decomposition (WPD) is adopted to further improve the signal-to-noise ratio. The wavelet coefficients after soft threshold processing are:
[0020]
[0021] Among them, is the i-th processed wavelet coefficient, and w i is the original i-th wavelet coefficient, is the threshold of the i-th subband, and sign(·) is the sign function.
[0022] Finally, data augmentation is performed through random cropping, circular shifting, and additive Gaussian noise perturbation to balance the class distribution and improve the generalization of the model.
[0023] 2. Design of an Adaptive Spatiotemporal Feature Learning Network Based on KAN-BiLSTM-Transformer: Adaptive Feature Extraction Module: Use KAN to extract the adaptive feature representation of each modal signal. The formula is:
[0024]
[0025] Among them, represents the value of the d-th adaptive feature at the t-th time step; P is the expansion term number of KAN; is the combination coefficient of the d-th adaptive feature on the p-th item; is the unary function of the d-th adaptive feature on the p-th item. In the present invention, the Swish activation function is adopted:
[0026]
[0027] Among them, is the shape parameter of the d-th adaptive feature on the p-th item; is the projection vector of the p-th item; is the original feature vector at the t-th time step.
[0028] Bidirectional Long Short-Term Memory Network: Use BiLSTM to learn the temporal dependence relationship of the feature sequence. The update formulas for the forward and backward hidden states are:
[0029]
[0030] Among them, and are the forward and backward hidden states at the t-th time step respectively, is the adaptive feature vector at the t-th time step.
[0031] Multi-modal information interaction and fusion module: The interaction and fusion of different modal features are realized using the multi-head self-attention mechanism. The formula is:
[0032]
[0033] Among them, A h is the self-attention matrix of the h-th attention head; Q h , K h are the query matrix and key matrix of the h-th attention head respectively, and are obtained from the multi-modal feature matrix X through linear transformation:
[0034]
[0035] Among them, and are the projection matrices of the h-th attention head; D h is the subspace dimension of the h-th attention head, which is used to scale the dot product result.
[0036] Finally, high-level features are further extracted through a feed-forward neural network (FFN):
[0037] F = ReLU(Y′W1 + b1)W2 + b2;
[0038] Among them, Y′ is the multi-modal interaction feature matrix that integrates body movement and environmental information, and W1, b1, W2, b2 are the parameters of the FFN.
[0039] At the end of the model, a fully connected layer is used to map F to the sleep stage label space:
[0040] P = softmax(W cls F + b cls );
[0041] Among them, W cls and b cls are the parameters of the classification layer, and F is the feature matrix output by the multi-modal information interaction and fusion module.
[0042] 3. Construction of a multi-task self-supervised pre-training paradigm based on KAN-BiLSTM-Transformer:
[0043] Design of the loss function: Use the weighted Dice loss and the temporal smoothing loss The weighted sum is used as the loss function. The formula for the weighted Dice loss is:
[0044]
[0045] where is the weighted Dice loss for the \(c\)-th class, \(w\) c is the weight coefficient for the \(c\)-th class, \(p\) ct is the probability that the model predicts the \(c\)-th class at the \(t\)-th time step, \(y\) ct is the \(c\)-th element of the one-hot vector of the true label at the \(t\)-th time step, and \(\epsilon\) is the smoothing term. The formula for the temporal smoothing loss is:
[0046]
[0047] where \(N\) c is the total number of classes, and \(T\) is the total number of time steps.
[0048] Optimization algorithm and hyperparameter selection: The RMSprop optimization algorithm is adopted, and the learning rate is adjusted using the Cosine Annealing strategy. The formula is:
[0049]
[0050] where \(\eta\) t is the learning rate at the \(t\)-th step, \(\eta\) min and \(\eta\) max are the minimum and maximum learning rates respectively, \(T\) cur is the current training step, and \(T\) max is the total number of training steps.
[0051] Training process and evaluation: K-fold cross-validation is adopted, and the model performance is evaluated using accuracy, macro-average F1 score, and Cohen's Kappa coefficient. The formula for the macro-average F1 score is:
[0052]
[0053] where \(precision\) c and \(recall\) c are the precision and recall of the \(c\)-th class respectively. The formula for Cohen's Kappa coefficient is:
[0054]
[0055] where \(p\) o is the observed agreement, and \(p\) e is the chance agreement.
[0056] 4. Application of intelligent prediction of sleep staging: Perform quality enhancement and feature extraction on new sleep data samples, then use the trained model to predict the sleep staging probability at each time step, and make a decision on the final staging label based on the probability value to generate a sleep staging prediction result graph. The calculation formula for the prediction probability matrix is:
[0057] P = softmax(W cls F + b cls );
[0058] where W cls and b cls are the parameters of the classification layer, and F is the feature matrix output by the multi-modal information interaction and fusion module.
[0059] This model can automatically learn the discriminative features of multi-modal sleep signals, capture spatio-temporal dependence relationships, and perform cross-modal information fusion, significantly improving the performance of sleep staging. At the same time, it has good computational efficiency and interpretability, providing important value for clinical applications. In the future, this model can also be integrated into sleep monitoring devices or systems to achieve real-time online sleep staging analysis functions.
[0060] On the other hand, based on the above method, we also provide a sleep staging detection and evaluation system. This system consists of a data acquisition module, a data quality enhancement module, a feature extraction module, a sequence learning module, a feature fusion module, a staging prediction module, and a data visualization module, etc., and is equipped with necessary storage and computing resources to support the training and inference of the model.
[0061] The data acquisition module is responsible for collecting polysomnogram (PSG) data, including signals such as EEG, EOG, EMG, ECG, respiration, SpO2, body movement, and environment. This module consists of various physiological signal sensors and acquisition devices, such as EEG caps, EOG electrodes, EMG electrodes, ECG electrodes, respiratory belts, pulse oximeters, and accelerometers, etc. The collected raw signals are transmitted to the central data processing unit in a wired or wireless manner.
[0062] The data quality enhancement module performs a series of data quality enhancement operations on the collected raw signals, such as denoising, filtering, and segmentation, to improve the signal quality and reduce the influence of noise and artifacts. This module is mainly implemented by dedicated signal processing chips or high-performance CPUs / GPUs, and is equipped with sufficient memory and storage space for caching raw data and intermediate results. The signal data after data quality enhancement will be formatted and stored in the system's database for subsequent feature extraction and model training.
[0063] The feature extraction module uses a pre-trained KAN model to extract adaptive features from each signal after data quality enhancement. The parameters of the KAN model are stored in the system's memory or cache for quick loading and execution. The extracted adaptive features will be written to the memory or cache and passed to the sequence learning module.
[0064] The sequence learning module uses a BiLSTM model to encode the adaptive feature sequences of each signal to capture long-range dependencies in the time dimension. The parameters of the BiLSTM model are also stored in the memory or cache. The forward and backward hidden state vectors will be written to the memory or cache and passed to the feature fusion module.
[0065] The feature fusion module uses a Transformer model to perform cross-modal fusion of the sequence features of different signals. The parameters of the Transformer model are likewise stored in the memory or cache. The fused feature vectors will be written to the memory or cache and passed to the staging prediction module.
[0066] The staging prediction module uses a fully connected classifier to predict the sleep state at each time step based on the fused feature vectors. The parameters of the classifier are stored in the memory or cache, and the prediction results will be written to the system's database or external storage device for subsequent analysis and applications.
[0067] To support the operation of each of the above modules, the system is also equipped with a high-performance processor, a large-capacity memory, and storage devices. The processor can be a multi-core CPU, GPU, or a dedicated artificial intelligence chip such as a TPU, etc., to provide sufficient computing power. The memory can be built-in DRAM, cache, or external memory modules to meet the needs of data access and model operation. The storage device can be a hard disk, solid-state drive, or network storage for permanently storing the original data, intermediate results, and final prediction results.
[0068] In addition, the system also provides a friendly user interface and interaction methods to facilitate people's operations and result interpretation. Users can upload new sleep data, start an automatic staging task, view staging results and statistical reports, etc., through a graphical interface or a Web interface. The system also supports the import and export of multiple data formats, such as EDF, XML, JSON, etc., to achieve data exchange and sharing with other medical systems.
[0069] In summary, the system realizes the full-process automation from data acquisition, data quality enhancement, feature extraction, sequence learning, feature fusion, staging prediction to data visualization, and provides flexible user interaction and data management functions. Detailed implementation
[0070] The present invention will be further explained and illustrated below in conjunction with specific embodiments, but it is not a restrictive constraint on the present invention.
[0071] A method for detecting and evaluating sleep quality, which is based on the sleep stage prediction method of KAN-BiLSTM-Transformer, includes the following steps:
[0072] 1. Multimodal sleep signal acquisition and quality enhancement:
[0073] The data used in this study came from polysomnogram (PSG) records of a sleep center. A total of 100 people (50 males and 50 females; age range 18 - 65 years, average age μ age = 42.5, standard deviation σ age = 12.1) of overnight PSG data were included. All subjects signed an informed consent form before the experiment.
[0074] The original PSG data included the following signals:
[0075] (1) Electroencephalogram (EEG): Recorded EEG signals of multiple channels such as F3, F4, C3, C4, O1, O2, etc., with a sampling frequency The original EEG signal can be represented as a matrix X EEG .
[0076] (2) Electrooculogram (EOG): Recorded EOG signals of two channels, left and right, with a sampling frequency The original EOG signal can be represented as a matrix X EOG .
[0077] (3) Electromyogram (EMG): Recorded EMG signals of the mandibular muscles, with a sampling frequency The original EMG signal can be represented as a vector x FMG .
[0078] (4) Electrocardiogram (ECG): Recorded single-lead ECG signals, with a sampling frequency The original ECG signal can be represented as a vector x ECG .
[0079] (5) Respiratory signal: Recorded the respiratory signal reflected by the chest and abdominal movements, with a sampling frequency The original respiratory signal can be represented as a vector x RESP .
[0080] (6) Blood oxygen saturation (SpO2): Recorded the SpO2 time series at a frequency of using a pulse oximeter, which can be represented as a vector x SpO2 .
[0081] (7) Body movement: The hand movement signal was recorded using a wrist accelerometer at a frequency and can be represented as matrix X MOT .
[0082] (8) Environment: The temperature x of the laboratory, humidity x TEMP , and light intensity x HUMI were recorded using a multi-modal sensor at a LUMI frequency.
[0083] For different types of raw signals, the following data quality enhancement methods were adopted:
[0084] (1) EEG and EOG signals: First, an adaptive filtering algorithm based on empirical mode decomposition (EMD) was used to remove power frequency interference and baseline drift. This algorithm decomposes the signal into several intrinsic mode functions (IMFs) and adaptively identifies and removes the noise components. Let the original EEG signal be x(t), and the IMF components obtained by EMD decomposition be {c1(t), c2(t), …, c n (t)}, then the denoised signal is:
[0085]
[0086] where β i is the weight coefficient of the i-th IMF component and can be obtained by optimizing the following objective function:
[0087]
[0088] In the above formula, the first term is the signal reconstruction error, and the second term is the L1 regularization term, which is used to encourage sparsity, that is, automatically reduce the weight of the noise IMF component to 0. λ is a hyperparameter that balances the two terms.
[0089] Secondly, in order to further improve the signal-to-noise ratio, an adaptive threshold denoising algorithm based on wavelet packet decomposition (WPD) was adopted. This algorithm performs multi-scale wavelet packet transform on the signal, adaptively sets thresholds in each sub-band, and performs soft threshold processing on the wavelet coefficients to remove high-frequency noise. Let the EEG signal after the first step of denoising be and its wavelet packet coefficients be w{w1, w2, …, w N}, then the coefficients after soft threshold processing are:
[0090]
[0091] where is the threshold of the i-th sub-band and can be determined by the following adaptive strategy:
[0092]
[0093]
[0094] In the above formula, σ i is the standard deviation estimate of the i-th subband coefficient, and N is the signal length. Through wavelet packet reconstruction, the finally denoised EEG signal can be obtained
[0095] (2) EMG signal: First, use a Butterworth high-pass filter to remove low-frequency interference. Second, adopt an envelope extraction algorithm based on Hilbert transform to obtain the envelope of the EMG signal. Let the EMG signal after band-pass filtering be x(t), and its Hilbert transform be Then the analytic signal z(t) and the envelope signal e(t) are respectively:[[]]
[0096]
[0097] Finally, perform moving average smoothing on the envelope signal e(t) to obtain the finally data quality-enhanced EMG signal .
[0098] (3) ECG signal: First, use a median filter to remove baseline drift and electromyographic noise. Second, adopt the Pan-Tompkins algorithm to detect the R-wave peaks and calculate the sequence of adjacent R-wave peak intervals (RRI). This algorithm is implemented through the following steps:
[0099] Band-pass filtering: Use a band-pass filter to extract the signal in the QRS complex wave frequency band (5 - 15 Hz).
[0100] Differentiation: Differentiate the signal after band-pass filtering to enhance the slope characteristics of the QRS complex wave.[[]]
[0101] Squaring: Square the differentiated signal to further enhance the energy of the high-frequency QRS complex wave.[[]]
[0102] Moving window integration: Use a moving window to integrate and smooth the squared signal to obtain the QRS complex wave energy envelope.[[]]
[0103] Adaptive threshold detection: Set the threshold adaptively according to the signal mean and noise level, and perform threshold detection on the QRS complex wave energy envelope to obtain the R-wave peak positions.[[]]
[0104] Finally, perform 3σ principle outlier rejection and cubic spline interpolation on the RRI sequence to obtain the finally data quality-enhanced RRI time series x RRJ (t).
[0105] (4) Respiratory signal: First, use a FIR low-pass filter to remove high-frequency noise. Second, adopt an instantaneous phase extraction algorithm based on the Hilbert transform to obtain the respiratory phase signal. Let the filtered respiratory signal be x(t), and its Hilbert transform be Then the instantaneous phase signal φ(t) is:
[0106]
[0107] Based on φ(t), features such as respiratory frequency and respiratory depth can be further calculated. Finally, perform z-score normalization on the original respiratory signal x(t) and various respiratory features to obtain the respiratory feature vector x with enhanced data quality RESP .
[0108] (5) SpO2 signal: First, use the LOESS regression algorithm to smooth the SpO2 time series and remove outliers caused by probe displacement, etc. Second, extract the statistical features of SpO2, including the mean μ SpO2 , standard deviation σ SpO2 , maximum value max(SpO2), minimum value min((SpO2), etc., to obtain the SpO2 feature vector x SpO2 .
[0109] (6) Body movement signal: First, use a Butterworth low-pass filter to remove high-frequency noise. Second, perform vector synthesis on the three-axis acceleration data to obtain the synthetic acceleration time series:
[0110]
[0111] Third, adopt a double-threshold method to detect body movement events. Let the synthetic acceleration sequence be a = [a1, a2,..., a T , and the two thresholds be θ1 and θ2, then the detection rule for the body movement event m = [m1, m2,..., m T | is:
[0112]
[0113] where k is the time window length, and θ1 and θ2 can be set according to experience. Finally, perform statistics on the detected body movement events to obtain the body movement feature vector x MOT , including the number of body movements, total body movement duration, average body movement duration, etc.
[0114] (7) Environmental signal: Perform min-max normalization on temperature, humidity, and light intensity respectively to obtain the normalized environmental feature vector x ENV .
[0115] After the above data quality enhancement is completed, the time series signals such as EEG, EOG, and EMG are segmented into 30-second time windows, and each segment corresponds to a unique sleep stage label (Wake, N1, N2, N3, or REM). At the same time, the ECG, respiration, SpO2, body movement, and environmental feature vectors are aligned with each time window to form a complete multi-modal sleep data sample.
[0116] In addition, considering the problem of class imbalance in sleep data (such as relatively few N1 and REM segments), the following data augmentation strategies are adopted:
[0117] (1) EEG and EOG signals: Data augmentation methods in the time and frequency domains are used. In the time domain, the following operations are sequentially performed on the original signal segment to obtain the enhanced signal :
[0118]
[0119] where S L (·) represents randomly intercepting a segment of length L from the original signal, represents circularly shifting the signal points to the right, ε is additive Gaussian noise, and σ 2 is the noise variance.
[0120] In the frequency domain, first perform a fast Fourier transform (FFT) on the signal to obtain the spectrum X(f). Then, add a random phase shift φ r to the spectrum, and then perform an inverse FFT to obtain the enhanced signal
[0121]
[0122]
[0123] In the above formula, U(0, 2π) represents the uniform distribution on [0, 2π]. The phase shift in the frequency domain is equivalent to introducing a random time delay in the time domain, which can increase data diversity.
[0124] (2) ECG signal: Random resampling is performed on the RRI time series x RRI . Specifically, a subsequence s RRI is randomly selected from x RRI , and then linear interpolation is performed on s RRI to obtain a new sample of the same length as the original sequence
[0125]
[0126] Among them, Interpolate(·) represents the linear interpolation function, and T is the length of the original sequence. Resampling can change the rhythm characteristics of the RRI sequence and simulate different heart rate variabilities.
[0127] (3) Respiratory signal: Randomly perturb the respiratory frequency and respiratory depth characteristics. Let the original respiratory frequency be f0 and the respiratory depth be d0. The perturbed characteristics and are:
[0128]
[0129] where and are the variances of frequency and depth perturbations respectively, which can be set according to prior knowledge. This kind of perturbation can simulate the natural variation of respiratory patterns.
[0130] (4) SpO2 signal: Randomly shift each feature in the SpO2 feature vector x SpO2 . For example, the shift of the mean value μ SpO2 is:
[0131]
[0132] where is the variance of the shift amount. The shift methods for other features are similar. This kind of shift can simulate the individual differences and measurement errors of SpO2.
[0133] (5) Body movement and environmental signals: Use the SMOTE algorithm to oversample the body movement and environmental feature vectors x MOT and x ENV to synthesize new minority class samples. The basic idea of the SMOTE algorithm is that for each minority class sample x i , randomly select a sample x j from its k nearest neighbors, and then interpolate between x i and x j to generate a new sample
[0134]
[0135] where λ is the interpolation coefficient. Through SMOTE, the class distribution of body movement and environmental data can be balanced, and the robustness of the model can be improved.
[0136] In summary, through data quality enhancement and data augmentation of the original PSG signal, a standardized and diversified multi-modal sleep dataset is obtained where is the multi-modal input of the i-th sample, and y i∈ {0, 1, 2, 3, 4} is the corresponding sleep stage label, and M is the total number of samples. This dataset can be used for subsequent model training and evaluation.
[0137] 2. Design of an Adaptive Spatiotemporal Feature Learning Network Based on KAN-BiLSTM-Transformer:
[0138] The present invention proposes an innovative adaptive spatiotemporal feature learning network, which consists of three key components: an Adaptive Feature Extraction Module, a Bidirectional Long Short-Term Memory (BiLSTM), and a Multi-Modal Information Interaction Fusion Module. Among them, the Adaptive Feature Extraction Module is based on the Kolmogorov-Arnold network (KAN) and is responsible for extracting discriminative and adaptive high-level feature representations from multi-modal sleep signals. BiLSTM is used to mine the long-term temporal dependencies within each modal feature sequence. The Multi-Modal Information Interaction Fusion Module takes the self-attention mechanism as the core to achieve dynamic interaction and deep fusion between different modal features. The three components complement each other to jointly construct an efficient and accurate deep learning framework for sleep stage prediction from the input end to the output end.
[0139] 2.1 Adaptive Feature Extraction Module:
[0140] The input of the Adaptive Feature Extraction Module is the multi-modal sleep signals obtained in the multi-modal sleep signal acquisition and quality enhancement stage, including EEG, EOG, EMG, ECG, respiration, SpO2, body movement, and environment, etc. For each type of signal, an adaptive feature extractor based on KAN is designed to automatically learn the optimal feature representation of the signal. Taking the EEG signal as an example, let the EEG segment after data quality enhancement be where C is the number of channels and T is the number of time steps. The core idea of KAN is to use the combination of unary functions to approximate complex multi-variable signals. Specifically, the KAN feature extraction process of the EEG signal can be expressed as:
[0141]
[0142] where is the t-th column of X EEG representing the original EEG features at the t-th time step. is the d-th adaptive feature extracted from the EEG signal at the t-th time step. P is the number of expansion terms of KAN, is the projection vector of the p-th term, is the unary function of the d-th adaptive feature on the p-th term is the corresponding combination coefficient. is the extracted EEG adaptive feature matrix. In actual implementation, the Swish activation function is adopted:
[0143]
[0144] where is the shape parameter of the d-th adaptive feature on the p-th term, which can be obtained through training and learning from both the input end to the output end. Swish can adaptively adjust the non-linearity degree of the feature, thereby enhancing the representation ability of the feature.
[0145] For EOG signals, a KAN feature extractor similar to that of EEG is adopted to convert the EOG signal after data quality enhancement into an adaptive feature matrix
[0146]
[0147]
[0148] For EMG signals (1D) and ECG signals (1D) with lower dimensions, simplified versions of the KAN feature extractor are adopted respectively. Taking EMG as an example:
[0149]
[0150] where is the t-th element of x EMG and is the scalar projection coefficient. The processing method of ECG signals is the same as that of EMG, and finally the adaptive feature matrix
[0151] For respiratory signals, first use KAN to extract feature sequences such as respiratory frequency and respiratory depth, and then splice them together to form an adaptive feature matrix
[0152] H RESP =[H RESP_RATE ; H RESP_DEPTH ; …];
[0153] where H RESP_RATE and H RESP_DEPTH are the respiratory frequency and respiratory depth feature sequences extracted by KAN respectively.
[0154] For the SpO2 signal, together with other slowly varying physiological indicators (such as heart rate, body temperature, etc.), it is mapped to the same time step as signals such as EEG using a fully connected layer to obtain an adaptive feature matrix.
[0155]
[0156] where x SpO2 , x HR , x TEMP etc. are the original feature sequences after data quality enhancement of SpO2, heart rate, body temperature, etc., and FC(·) represents the fully connected layer.
[0157] For the body movement signal, first encode the original three-axis acceleration data using a Gated Recurrent Unit (GRU) layer, and then take the hidden state output of the last time step as the adaptive feature vector of the body movement.
[0158]
[0159]
[0160] where is the original feature of the three-axis acceleration at the t-th time step.
[0161] For the environmental signal, splice the features after data quality enhancement of temperature, humidity, light intensity, etc., and then obtain the adaptive feature vector of the environment through a fully connected layer.
[0162] h ENV = ReLU(W ENV [x TEMP ; x HUMI ; x LUMI ; …]+b ENV );
[0163] where W ENV and b ENV are the parameters of the fully connected layer.
[0164] In summary, the adaptive feature extraction module extracts a series of adaptive feature matrices {H EEG , H EOG , H EMG , H ECG , H RESP H SpO2} and feature vectors h MOT and h ENV, which will be used as the input for the subsequent BiLSTM and multi-modal information interaction and fusion module.
[0165] 2.2 Bidirectional Long Short-Term Memory Network:
[0166] BiLSTM is used to mine the long-term temporal dependencies in the adaptive feature sequences of each modality. For each feature matrix it is first sliced into a series of feature vectors by time step and then they are sequentially input into the BiLSTM:
[0167]
[0168] where and are the forward and backward LSTM hidden state outputs at the t-th time step respectively, and K is the dimension of the hidden state. is the BiLSTM output at the t-th time step, which is composed of the concatenation of the forward and backward hidden states. The internal structure of LSTM includes the input gate i t , forget gate f t , output gate o t and memory cell c t , and its forward propagation process is:
[0169] i t = σ(W ii h t + b ii + W hi z t-1 + b hi );
[0170] f t = σ(W if h t + b if + W hf z t-1 + b hf );
[0171] o t = σ(W io h t + b io + W ho z t-1 + b ho );
[0172] g t = tanh(W ig h t + b ig + W hg z t-1 + b hg );
[0173] c t = f t ⊙c t-1 + i t ⊙g t ;
[0174] z t = o t ⊙tanh(c t );
[0175] where W * and b * are the parameter matrix and bias vector of the LSTM, σ(·) is the Sigmoid activation function, tanh(·) is the hyperbolic tangent activation function, and ⊙ is the element-wise multiplication operation.
[0176] After inputting the feature matrices of all modalities into the BiLSTM, a set of high-level temporal feature matrices {Z EEG , Z EOG , Z EMG , Z ECG , Z RESP , Z SpO2} is obtained, where these matrices characterize the long-range dependencies and evolution laws of the signals of each modality in the time dimension, providing important prior knowledge for subsequent cross-modal information interaction and fusion.
[0177] 2.3 Multivariate Information Interaction and Fusion Module:
[0178] The multivariate information interaction and fusion module takes the self-attention mechanism as the core to achieve dynamic interaction and deep fusion of different modality features in both the time and space dimensions. The input of this module is the high-level temporal feature matrices extracted by the BiLSTM and the body movement and environmental feature vectors output by the adaptive feature extraction module. First, all the feature matrices are concatenated in the time dimension to obtain a multi-modal temporal feature matrix
[0179] X = [Z EEG ; Z EOG ; Z EMG ; Z ECG ; Z RESP ; Z SpO2 ;
[0180] where D - 2K×6 is the dimension of the concatenated features.
[0181] Then, the multi-head self-attention mechanism is used for adaptive cross-modal interaction and fusion of X. Specifically, the multi-head self-attention mechanism first projects X into H different subspaces through a linear transformation to obtain H groups of query matrices key matrices Sum matrix
[0182]
[0183] where is the projection matrix of the h-th attention head, and D h = D / H is the subspace dimension of each head.
[0184] Next, for each attention head, calculate the normalized dot product of the query matrix and the key matrix to obtain a self-attention matrix
[0185]
[0186] where softmax(·) is the Softmax function used to normalize the dot product result into attention weights. is the scaling factor used to control the variance of the dot product.
[0187] Then, multiply the self-attention matrix by the value matrix to obtain the output matrix of the h-th attention head :
[0188] U h = V h A h , h = 1, 2,..., H;
[0189] Finally, concatenate the output matrices of all attention heads in the feature dimension and obtain the final output matrix of the multi-head self-attention through a linear transformation
[0190] Y = W O [U1; U2;...; U H + b O ;
[0191] where and are the parameters of the output linear transformation.
[0192] Next, expand the body movement feature vector h MOT and the environmental feature vector h ENV to the same number of time steps as Y and concatenate them in the feature dimension:
[0193] Y′ = [Y; h MOT 1 T ; h ENN 1 T ;
[0194] where is a vector of all 1s. This step introduces the global information of body movement and environment into the multi-modal interaction fusion features.
[0195] Finally, Y′ is further processed by a Feed-Forward Network (FFN) to extract higher-level cross-modal fusion features:
[0196] F = ReLU(Y′W1 + b1)W2 + b2;
[0197] where are the parameters of the FFN, and D FFN is the dimension of the hidden layer of the FFN.
[0198] So far, the multi-source information interaction fusion module has output a fusion feature matrix that combines the spatio-temporal features of multi-modal sleep signals which contains both the long-range dependencies of different channel signals in the time dimension and the cross-complementary information of different modal signals in the feature space. This provides a rich and comprehensive feature representation for the subsequent sleep stage prediction task.
[0199] At the end of the model, a fully connected layer is used to map F to the sleep stage label space:
[0200] P = softmax(W cls F + b cls );
[0201] where and are the parameters of the classification layer, and N c is the number of sleep stage categories (usually 5, including Wake, N1, N2, N3, and REM). is the predicted probability matrix, where P ct represents the probability that the t-th time step belongs to the c-th sleep stage.
[0202] In summary, the adaptive spatio-temporal feature learning network based on KAN-BiLSTM-Transformer proposed by the present invention makes full use of the complementary information of multi-modal sleep signals. Through a series of designs such as adaptive feature extraction, temporal modeling, and cross-modal interaction fusion, a high-efficient and accurate deep learning framework from the input end to the output end is constructed. This model can automatically mine the discriminative features in different physiological signals, capture their spatio-temporal evolution laws, and perform comprehensive information fusion, thus significantly improving the performance of sleep stage classification. At the same time, due to the parallel computing structure of Transformer and the interpretability of the attention mechanism, this model also has good computational efficiency and understandability, which provides important advantages for its practical clinical applications.
[0203] 3. Construction of a multi-task self-supervised pre-training paradigm based on KAN-BiLSTM-Transformer:
[0204] After defining the structure of the sleep staging prediction model, the next key step is to design an effective training strategy to optimize the model parameters and improve the prediction performance. The following will introduce the training process of the model in detail, including the design of the loss function, the selection of the optimization algorithm, the adjustment of hyperparameters, and the evaluation method during training.
[0205] 3.1 Loss function design:
[0206] For the multi-classification task of sleep staging, a composite loss function is designed to guide the model training. This loss function consists of two parts: weighted Dice loss and temporal smoothing loss
[0207] The weighted Dice loss is a loss function based on the Dice coefficient, which is especially suitable for the class imbalance problem. For the c-th class, its weighted Dice loss is defined as:
[0208]
[0209] where y ct is the c-th element of the one-hot vector of the true label at the t-th time step, p ct is the probability that the model predicts as the c-th class at the t-th time step, w c is the weight coefficient of the c-th class, and ∈ is a smoothing term used to avoid the denominator being zero. The weight coefficient w c can be adjusted according to the number of samples of each class in the training set to alleviate the class imbalance problem, and its calculation formula is:
[0210]
[0211] where N c is the total number of classes, and N i is the number of samples of the i-th class.
[0212] The final weighted Dice loss is the sum of the losses of all classes:
[0213]
[0214] The definition of the temporal smoothing loss is the same as before, aiming to encourage the model to make smooth and consistent predictions between adjacent time steps:
[0215]
[0216] The final composite loss function is the weighted sum of the weighted Dice loss and the temporal smoothing loss:
[0217]
[0218] where λ WD and λ TS is the weight coefficient to balance the two losses.
[0219] 3.2 Optimization algorithm and hyperparameter selection:
[0220] To optimize the composite loss function, the RMSprop optimization algorithm is used to update the model parameters. RMSprop is a gradient descent method with an adaptive learning rate. It adjusts the learning rate of each parameter by taking an exponentially weighted moving average of the squared gradients. Compared to traditional SGD algorithms, RMSprop can automatically adjust the learning rate, accelerating convergence.
[0221] The update rule of the RMSprop algorithm is:
[0222]
[0223] where υ i is the squared gradient The exponentially weighted moving average of , β is the decay rate, η is the learning rate, ∈ is the smoothing term, θ t are the model parameters at step t.
[0224] In the experiment, β is set to 0.9 and ∈ is set to 10 -6 The initial value of the learning rate η is set to 10 -3 , and use the CosineAnnealing strategy for dynamic adjustment:
[0225]
[0226] where η min and η max are the minimum and maximum learning rates, T cur is the current training step number, T max is the total number of training steps. Through the Cosine Annealing strategy, the learning rate changes periodically, which can help the model escape the local optimum in the later stages of training.
[0227] In addition, the following hyperparameters and training strategies were adopted:
[0228] Batch size: To strike a balance between training efficiency and memory usage, the batch size is set to 128.
[0229] L2 regularization: To prevent overfitting, L2 regularization is applied to the model parameters, and the regularization coefficient is set to 10 -1。
[0230] Dropout: After modules such as KAN, BiLSTM, and FFN, a Dropout layer is added to randomly discard a certain proportion (0.5) of neurons to improve the generalization of the model.
[0231] Gradient clipping: To prevent the problem of gradient explosion, the gradients in the backpropagation process are clipped, and the gradient norm is restricted within 1.0.
[0232] 3.3 Training process and evaluation:
[0233] The model is trained and evaluated using K-fold cross-validation. Specifically, the dataset is randomly divided into K mutually exclusive subsets. Each time, one of the subsets is selected as the validation set, and the remaining K - 1 subsets are used as the training set. This is repeated K times, and the average of the K results is taken as the final performance metric. In the experiment, K = 5.
[0234] For each fold, the model is trained on the training set and the model performance is evaluated on the validation set. The following metrics are used to evaluate the sleep staging performance:
[0235] Accuracy: The proportion of correctly predicted time steps to the total number of time steps.
[0236] Macro-average F1 score: First, calculate the F1 score for each class separately, and then take the average. The macro-average F1 score can measure the average performance of the model across different classes and is more robust to class imbalance problems.
[0237] Cohen's Kappa coefficient: Measures the consistency between the predicted results and the true labels, with a value range of [-1, 1]. The closer it is to 1, the better the consistency. This metric takes into account the impact of random guessing and can better reflect the actual performance of the model.
[0238] When there is no performance improvement for 10 consecutive epochs on the validation set, the early stopping mechanism is triggered to stop training and save the model parameters with the best performance.
[0239] The complete training process is as shown in the following algorithm.
[0240] Training process of the KAN - BiLSTM - Transformer model
[0241] Input: Training dataset Validation dataset Test dataset
[0242] Output: Trained model parameters 0
[0243] 1: Randomly initialize model parameters 0
[0244] 2: For epoch - 1 to N cpochs do
[0245] 3:
[0246] 4: P i = KAN - BiLSTM - Transformer(X i ; θ)
[0247] 5:
[0248] 6:
[0249] 7: g = clip(g, 1.0)
[0250] 8: θ = θ - η t ·RMSprop(g, β, ∈)
[0251] 9: end for
[0252] 10:
[0253] 11: if score > best_score then
[0254] 12: best_score = score
[0255] 13: θ best = θ
[0256] 14: no improve - 0
[0257] 15: else
[0258] 16: no improve + 1
[0259] 17: if no_improve >= 10 then
[0260] 18: break
[0261] 19: end if
[0262] 20: end if
[0263] 21: end for
[0264] 22:
[0265] 23: return θ best
[0266] Among them, N cpochs is the maximum number of training epochs, η t is the learning rate at the t-th step, which is dynamically adjusted by the Cosine Annealing strategy. clip(·, 1.0) represents the gradient clipping operation, and θ best are the model parameters with the optimal performance.
[0267] 3.4 Application of intelligent sleep stage prediction:
[0268] After obtaining the trained KAN - BiLSTM - Transformer model, it can be used to automatically predict the sleep stage of new sleep data. There is a new sleep data sample, which contains the time series of various physiological signals (such as EEG, EOG, EMG, etc.). It is hoped to use the trained model to predict the sleep stage for each time step (such as a 30 - second time window) of this new sample.
[0269] The specific prediction process is as follows:
[0270] (1) Data quality enhancement: First, perform the same data quality enhancement operations on the new sample as on the training data, including denoising, filtering, normalization, etc., to obtain the sample with enhanced data quality.
[0271] (2) Feature extraction: Next, input the sample with enhanced data quality into the trained KAN - BiLSTM - Transformer model. The model will automatically extract the adaptive feature representations of each modality signal and learn the temporal dependencies through BiLSTM, and finally obtain a fused feature representation matrix.
[0272] (3) Sleep stage probability prediction: Input the feature representation matrix into the last fully - connected network layer of the model to obtain the probability values of each time step belonging to each sleep stage, forming a probability matrix.
[0273] (4) Sleep stage label decision: According to the probability matrix, different decision strategies can be adopted to obtain the final sleep stage label. The strategy is to select the sleep stage category with the highest probability as the prediction label for each time step.
[0274] (5) Result visualization: Finally, the predicted sleep stage labels can be aligned with the original physiological signal waveform diagram to generate an intuitive sleep stage prediction result diagram.
[0275] Sleep is an indispensable part of human life activities and is closely related to our physical and mental health. For a long time, how to objectively, accurately and comprehensively evaluate sleep quality has been a key issue of great concern. Traditional sleep monitoring and staging methods mainly rely on manual interpretation of polysomnogram (PSG) data, which is not only time-consuming and laborious, but also easily affected by subjective factors.
[0276] Related artificial intelligence technologies have been applied to the monitoring and evaluation of sleep quality. The present invention introduces KAN-BiLSTM-Transformer into the field of sleep staging. By carefully designing the acquisition scheme, preprocessing process, network structure, training paradigm and application process, a complete intelligent auxiliary diagnosis system for sleep staging is constructed. The system can automatically learn the discriminative features contained in multi-modal sleep signals, reveal the internal relationship between different physiological indicators, and output objective and accurate staging prediction results in an end-to-end manner.
Claims
1. A method for detecting and evaluating sleep quality, characterized in that, The method is based on model. The method includes the following steps: Step 1: Multimodal sleep signal acquisition and quality enhancement; Step 2: Design of an adaptive spatio-temporal feature learning network based on to use to extract the adaptive feature representations of each modal signal: ; Among them, represents the value of the -th adaptive feature at the -th time step; is the number of expansion terms of ; is the combination coefficient of the -th adaptive feature on the -th term; is the unary function of the -th adaptive feature on the -th term, using the ; Among them, is the shape parameter of the th adaptive feature on the th item; is the projection vector of the th item; is the original feature vector at the th time step; Bidirectional Long Short-Term Memory Network: Using to learn the temporal dependence of the feature sequence, the forward and backward hidden state update formulas are as follows: ; ; Among them, and are the forward and backward hidden states at the -th time step respectively, and is the adaptive feature vector at the -th time step; Multi - information Interaction and Fusion Module: Implement the interaction and fusion of different modality features using the multi - head self - attention mechanism: ; Among them, is the self-attention matrix of the th attention head; and are the query matrix and key matrix of the th attention head respectively, and are obtained from the multimodal feature matrix through linear transformation: ; Among them, and are the projection matrices of the -th attention head; is the subspace dimension of the -th attention head, which is used to scale the dot product result; Finally, further extract high - level features through a feed - forward neural network (FFN): ; Among them, is a multi-modal interaction feature matrix that fuses body movement and environmental information, , , , are the parameters of the FFN; At the end of the model, a fully connected layer is used to map to the sleep stage label space: ; Among them, and are the parameters of the classification layer, is the feature matrix output by the multi-source information interaction and fusion module; Step 3: Based on the construction of the multi-task self-supervised pre-training paradigm, Using weighted loss and the temporal smoothing loss as the weighted sum as the loss function: ; Among them, is the weighted loss of the th class, is the weight coefficient of the th class, is the probability that the model predicts as the th class at the th time step, is the th element of the one-hot vector of the true label at the th time step, is the smoothing term; the formula for the temporal smoothing loss is: ; Among them, is the total number of categories, is the total number of time steps; Optimization Algorithm and Hyperparameter Selection: Adopt the RMSprop optimization algorithm and use the CosineAnnealing strategy to adjust the learning rate. The formula is: ; Among them, is the learning rate of the step, and are the minimum and maximum learning rates respectively, is the current training step number, is the total number of training steps; Training Process and Evaluation: Adopt K - fold cross - validation and use accuracy, macro - average F1 - score, and Cohen's Kappa coefficient to evaluate the model performance; Step 4: Application of intelligent prediction of sleep staging.
2. A sleep quality detection and evaluation method according to claim 1, wherein Said step 1 specifically includes: Perform quality enhancement processing on EEG, EOG, EMG, ECG, respiration, SpO2, body movement, and environmental signals; For EEG and EOG signals, use an adaptive filtering algorithm based on empirical mode decomposition (EMD) to remove noise. The formula is: ; Among them, is the denoised signal, is the th intrinsic mode function (IMF) component, is the weight coefficient of the th IMF component, is the total number of IMF components; Then, adopt an adaptive threshold denoising algorithm based on wavelet packet decomposition (WPD) to further improve the signal - to - noise ratio. The wavelet coefficients after soft - threshold processing are: ; Among them, is the th wavelet coefficient after processing, is the original th wavelet coefficient, is the threshold of the th sub-band, is the sign function; Finally, perform data augmentation through random cropping, circular shifting, and additive Gaussian noise perturbation to balance the class distribution and improve the model generalization ability.
3. The sleep quality detection and evaluation method according to claim 2, characterized in that Said step 4 specifically includes: Perform quality enhancement and feature extraction on new sleep data samples, then use the trained model to predict the sleep staging probability at each time step, and make a decision on the final staging label according to the probability value to generate a sleep staging prediction result graph. The calculation formula for the prediction probability matrix is: ; Among them, and are the parameters of the classification layer, is the feature moment output by the multi-source information interaction and fusion module.
4. A sleep quality detection and evaluation system, said system comprises a data acquisition module, a data quality enhancement module, a feature extraction module, a sequence learning module, a feature fusion module, a staging prediction module, and a data visualization module, and it further comprises storage resources and computing resources, used to implement the method in any one of claims 1 - 3.
Citation Information
Patent Citations
Method and device for measuring resting heart rate as well as wearable equipment comprising device
CN106343979A
Transform model and comparative learning-based sleep staging method and system
CN114881105A