ECG-EEG sleep staging method and system based on single-path vq-vaes

CN122767786APending Publication Date: 2026-09-18HUZHOU ENMEIDI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610939293.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0014]针对现有技术所存在的上述缺点,本发明提供了基于单路VQ-VAE的ECG-EEG睡眠分期方法及系统,能够有效克服现有技术所存在的分期精度较低、结构冗余、训练流程复杂和跨场景泛化能力差的缺陷

Benefits of technology

1)显式引入睡眠通用语义先验:将通过EEG预训练的睡眠通用语义码本作为跨模态对齐的“先验锚点”,克服了传统知识蒸馏/微调迁移/分布对齐忽略睡眠阶段语义先验、关键判别信息丢失的缺陷;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122767786A_ABST
    Figure CN122767786A_ABST
Patent Text Reader

Abstract

The present application relates to sleep staging, in particular to an ECG-EEG sleep staging method and system based on single-path VQ-VAE, in the first stage, the VQ-VAE is pre-trained through the EEG signal, and a set of discrete, compact and interpretable sleep general semantic codebook is obtained as the "prior anchor point" of cross-modal alignment; in the second stage, based on the frozen sleep general semantic codebook, the distance between the ECG feature and the sleep general semantic codebook in the distance space is constrained by the contrastive learning loss and the mean square error loss, so as to realize the cross-modal migration of the sleep staging knowledge of the EEG gold standard to the ECG domain; the technical scheme provided by the present application can effectively overcome the defects of low staging accuracy, structural redundancy, complex training process and poor cross-scene generalization ability existing in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to sleep staging, specifically to an ECG-EEG sleep staging method and system based on a single-channel VQ-VAE. Background Technology

[0002] The American Academy of Sleep Medicine (AASM) clinical sleep staging standards primarily rely on electroencephalogram (EEG) signals, supplemented by electrooculogram (EOG) and electromyogram (EMG) signals to interpret sleep stages in 30-second frames. EEG signals are the direct physiological anchors for sleep staging: stable light sleep N2 corresponds to sleep spindles and K-complexes; deep sleep N3 corresponds to delta waves; and REM sleep corresponds to low-amplitude desynchronized EEG. However, EEG signal acquisition equipment is expensive and cumbersome to wear, requiring electrodes to be in contact with the scalp for extended periods, making it unsuitable for long-term home monitoring. In contrast, single-lead electrocardiogram (ECG) signals can be conveniently acquired using wearable devices such as ECG patches and offer 24-hour dynamic monitoring capabilities, making them more suitable for practical applications.

[0003] Existing automatic ECG sleep staging methods mainly include the following categories:

[0004] 1) Direct end-to-end deep learning method: CNN, CNN-Transformer and other models are used to extract features such as HRV and QRS complex from ECG signals to achieve sleep staging. However, since ECG signals can only indirectly reflect autonomic nervous activity, the staging accuracy is generally lower than the EEG gold standard, and the ability to identify the sleep transition period N1, deep sleep N3 and REM sleep is insufficient.

[0005] 2) Knowledge distillation transfer method: Distilling the well-trained EEG sleep staging teacher model onto the ECG sleep staging student model can improve the staging performance based on ECG signals to some extent, but it requires high adaptability between the teacher and student models, and the accuracy of the student model has a clear upper limit.

[0006] 3) Fine-tuning the transfer paradigm: Freeze the bottom feature extraction layer of the EEG pre-trained model and only fine-tune the top classifier to adapt to the ECG signal. The training is simple, but the effect of feature alignment between EEG and ECG signals with huge rhythm differences is limited.

[0007] 4) Cross-modal feature fusion methods (such as SleepECGFusion): Deeply fuse EEG pre-trained feature knowledge with ECG native features, but usually adopt redundant structures such as dual-path VQ-VAE, which have a large number of parameters, complex training, and poor adaptability to small sample datasets.

[0008] 5) Domain adaptive transfer methods (such as Holter-to-Sleep): Align the distribution of the EEG / ECG domains through adversarial training and other means, but ignore the inherent general semantic priors of the sleep stages. Key discriminative information is lost during the transfer process, and the results are not good for identifying niche stages such as the sleep transition period N1 and the rapid eye movement sleep period REM.

[0009] The aforementioned existing technologies share the following common drawbacks:

[0010] 1) Lack of general semantic priors for sleep: The transfer method of direct alignment distribution ignores the "semantic anchors" corresponding to different sleep stages, resulting in the loss of key discriminative information during the transfer;

[0011] 2) The accuracy of the staging, especially in core stages such as the sleep transition period (N1), deep sleep (N3), and REM sleep, is far lower than the EEG gold standard;

[0012] 3) Existing dual-path transfer schemes have redundant structures, complex training processes, insufficient stability in small sample and cross-dataset scenarios, and poor cross-scenario generalization ability. Summary of the Invention

[0013] (a) Technical problems to be solved

[0014] In view of the above-mentioned shortcomings of the existing technology, the present invention provides an ECG-EEG sleep staging method and system based on single-channel VQ-VAE, which can effectively overcome the defects of low staging accuracy, structural redundancy, complex training process and poor cross-scenario generalization ability of the existing technology.

[0015] (II) Technical Solution

[0016] To achieve the above objectives, the present invention provides the following technical solution:

[0017] The ECG-EEG sleep staging method based on single-channel VQ-VAE includes the following two stages:

[0018] Phase 1: Pre-training of the EEG VQ-VAE sleep general semantic codebook, which includes the following steps:

[0019] S1. Acquire tagged EEG signals and perform preprocessing;

[0020] S2. Input the EEG signal into the first LAMBA-Net intra-frame encoder to obtain continuous hidden features;

[0021] S3. Construct a learnable discrete codebook, perform hard coding and quantization on continuous latent features based on the discrete codebook to obtain quantized feature maps, input the quantized feature maps into the decoder, reconstruct the EEG signal, construct a joint optimization loss function for joint training and optimization, and freeze the sleep general semantic codebook obtained from the training as prior knowledge for cross-modal alignment.

[0022] Phase Two: ECG-EEG cross-modal feature alignment and sleep staging, specifically including the following steps:

[0023] S4. Acquire single-lead ECG signals and perform preprocessing;

[0024] S5. Input the ECG signal into the second LAMBA-Net intra-frame encoder to obtain the ECG intra-frame features;

[0025] S6. Calculate the distance matrix between continuous latent features, ECG intra-frame features and general sleep semantic codebook respectively, and calculate the contrastive learning target matrix. Construct a joint alignment loss function to align ECG intra-frame features, reduce the gap between ECG intra-frame features and general sleep semantic codebook in the distance space, and transfer the sleep staging knowledge of the EEG gold standard across modalities to the ECG domain.

[0026] S7. Input the aligned ECG intra-frame features into the LAMBA-Net inter-frame encoder to model long-range dependencies across sleep cycles and obtain global feature representations.

[0027] S8. Project the global feature representation onto the classification space and output the sleep staging results.

[0028] Preferably, in S1, tagged EEG signals are acquired and preprocessed, including:

[0029] Acquire tagged single-channel / multi-channel EEG signals, uniformly downsample, remove invalid frames, and then divide them according to a preset time interval.

[0030] Preferably, in S3, a learnable discrete codebook is constructed, and continuous latent features are hard-coded and quantized based on the discrete codebook to obtain quantized feature maps. The quantized feature maps are input into the decoder to reconstruct the EEG signal. A joint optimization loss function is constructed for joint training and optimization. The sleep general semantic codebook obtained from the training is frozen as prior knowledge for cross-modal alignment, including: S31. Construct a learnable discrete codebook; S32. Calculate the Euclidean distance between each feature vector in the continuous latent features and all code vectors in the discrete codebook. Select the nearest neighbor code vector to perform hard coding quantization on the continuous latent features to obtain the quantized feature map. S33. Input the quantized feature map into the decoder to reconstruct the EEG signal; S34. Construct a joint optimization loss function based on reconstruction loss, codebook learning loss, and commitment loss, and perform joint training optimization. S35. Freeze the general semantic codebook of sleep obtained during training as prior knowledge for cross-modal alignment.

[0031] Preferably, in S34, a joint optimization loss function is constructed based on reconstruction loss, codebook learning loss, and commitment loss, including: Joint optimization loss function L total Expressed as follows: ; Among them, L rec The reconstruction loss is given by x, which represents the original input EEG signal. The reconstructed EEG signal, This represents the calculation of the square of the L2 Euclidean distance; L codebook For codebook learning loss, z e e is a single feature vector in a continuous latent feature set. k A single code vector in a discrete codebook. This represents the gradient truncation operator; L commit β represents the commitment loss, and β is the commitment loss weighting coefficient.

[0032] Preferably, in S4, a single-lead ECG signal is acquired and preprocessed, including: S41. Acquire single-lead ECG signals using a wearable single-lead ECG signal acquisition device; S42. Anti-aliasing filtering is performed through an FIR low-pass filter, and then unified resampling is performed through linear interpolation to remove power frequency interference, abnormal frames with baseline drift, and invalid frames labeled as motion period / unknown period. S43. Using a certain number of consecutive frames as the window size, perform sliding segmentation to obtain temporal context samples. Use Z-Score standardization to normalize the entire dataset using the mean and standard deviation of the training set, and save it to ensure deployment consistency.

[0033] Preferably, in S6, the distance matrices between continuous latent features, ECG intra-frame features, and the general sleep semantic codebook are calculated respectively, and the contrastive learning target matrix is ​​calculated. A joint alignment loss function is constructed to align ECG intra-frame features, reducing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space, and transferring the EEG gold standard sleep staging knowledge across modalities to the ECG domain, including: S61. Calculate the first distance matrix between continuous latent features and the general semantic codebook for sleep, and the second distance matrix between ECG intra-frame features and the general semantic codebook for sleep. S62. Perform L2 normalization on the first distance matrix and the second distance matrix respectively, and then calculate the cosine similarity according to the temperature coefficient to obtain the contrast learning target matrix. S63. Construct a joint alignment loss function to align ECG intra-frame features, reduce the gap between ECG intra-frame features and the general semantic codebook of sleep in the distance space, and transfer the sleep staging knowledge of the EEG gold standard across modalities to the ECG domain.

[0034] Preferably, in S63, constructing a joint alignment loss function for ECG intra-frame feature alignment includes: Joint alignment loss function L align Expressed as follows: ; Among them, L align For the joint alignment loss function, L InfoNCE To compare learning loss, L MSE α represents the mean error loss, and α is the weighting coefficient of the contrastive learning loss.

[0035] Preferably, in S8, the global feature representation is projected onto the classification space, and the sleep staging results are output, including: The global feature representation is projected onto a five-class classification space through a linear mapping to output the sleep staging results.

[0036] The ECG-EEG sleep staging system based on a single-channel VQ-VAE includes the following functional modules: The EEG signal acquisition and preprocessing module acquires tagged EEG signals and performs preprocessing. The continuous latent feature coding module uses the first LAMBA-Net intra-frame encoder to extract features from the EEG signal to obtain continuous latent features; The VQ-VAE pre-training module constructs a learnable discrete codebook, hard-codes and quantizes continuous latent features based on the discrete codebook to obtain quantized feature maps, inputs the quantized feature maps into the decoder, reconstructs the EEG signal, constructs a joint optimization loss function for joint training and optimization, and freezes the sleep general semantic codebook obtained during training as prior knowledge for cross-modal alignment. The ECG signal acquisition and preprocessing module acquires single-lead ECG signals and performs preprocessing. The ECG intra-frame feature coding module uses the second LAMBA-Net intra-frame encoder to extract features from the ECG signal and obtain ECG intra-frame features.

[0037] The cross-modal alignment module calculates the distance matrix between continuous latent features, ECG intra-frame features and the general sleep semantic codebook, respectively, and calculates the contrastive learning target matrix. It constructs a joint alignment loss function to align ECG intra-frame features, narrowing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space, and transferring the sleep staging knowledge of the EEG gold standard to the ECG domain across modalities.

[0038] The inter-frame temporal modeling module uses the LAMBA-Net inter-frame encoder to model the long-range dependency of aligned ECG intra-frame features across sleep cycles, and obtains global feature representations.

[0039] The classification output module projects the global feature representation onto the classification space and outputs the sleep staging results.

[0040] The visualization module plots the sleep staging results as an overnight sleep staging curve for reference.

[0041] Preferably, the first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder share parameters but have independent weights. Both the first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder include the following structure:

[0042] Non-stationary branch, extract non-stationary transient frequency domain features, and complete the association expression between multi-scale local context modeling and non-stationary transient frequency domain features;

[0043] Stationary branching is used to extract time-domain features of stationary rhythms.

[0044] The feature fusion module based on the channel attention mechanism takes non-stationary transient frequency domain features and stationary rhythmic time domain features as input by adding them element by element. First, it compresses the temporal information of each channel through adaptive average pooling, then generates channel weights through two fully connected layers and the Sigmoid activation function. The channel weights are multiplied by the input features channel by channel, and finally the channel dimension is restored through one-dimensional convolution.

[0045] Intra-frame BiLSTM branch, through bidirectional gating structure, finely captures local short-term temporal correlations and nonlinear transient details between adjacent subframes within an epoch;

[0046] The intra-frame Mamba branch, relying on the selective state-space model Mamba, dynamically selects the input-related state matrix and step size to efficiently model the long-range dependency of the sub-frame sequence with linear time complexity. It is then concatenated with the output features of the intra-frame BiLSTM branch along the channel dimension, and then fused through one-dimensional convolution to obtain continuous latent features.

[0047] The LAMBA-Net inter-frame encoder includes the following structure:

[0048] Inter-frame BiLSTM branch captures short-term transition patterns during sleep phases;

[0049] The inter-frame Mamba branch models long-range dependencies across sleep cycles and fuses them with the output features of the inter-frame BiLSTM branch through one-dimensional convolution, resulting in a global feature representation that simultaneously integrates fine intra-frame time-frequency features and long-range inter-frame temporal information.

[0050] (III) Beneficial Effects Compared with the prior art, the ECG-EEG sleep staging method and system based on single-channel VQ-VAE provided by the present invention has the following beneficial effects: 1) Explicitly introduce general semantic priors for sleep: Use the general semantic codebook for sleep pre-trained by EEG as the "prior anchor" for cross-modal alignment, which overcomes the shortcomings of traditional knowledge distillation / fine-tuning transfer / distribution alignment that ignore the semantic priors of the sleep stage and lose key discriminative information. 2) The single-path structure is lighter and more stable than the dual-path structure, and is better suited for the extraction of temporal features of sleep signals and cross-scene transfer. 3) Contrastive learning + MSE joint alignment loss: While constraining similarity structure and numerical distance at the metric learning level, it takes into account both coarse-grained class separability and fine-grained feature consistency. 4) Significantly improves staging accuracy: Breakthrough improvement in the identification effect of core clinical stages such as sleep transition N1, deep sleep N3, and rapid eye movement sleep REM, solving the pain points of sample imbalance and difficulty in identifying niche stages; 5) Strong cross-dataset generalization: It consistently achieves optimal staging results on two datasets with different collection devices and population distributions, namely MASS and SHHS; 6) Deployment-friendly: Clinical-grade sleep staging can be completed with only a wearable single-lead ECG signal acquisition device, greatly reducing the threshold for home sleep monitoring; 7) Reusable: The pre-trained general semantic codebook for sleep can be further transferred to sleep staging tasks involving other physiological signals such as PPG, SCG, and respiration. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0052] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0054] This invention provides a cross-modal transfer learning method for sleep staging based on single-channel VQ-VAE in ECG-EEG. This method adopts a two-stage training strategy: In the first stage, VQ-VAE is pre-trained using EEG signals to obtain a set of discrete, compact, and interpretable general semantic codebooks for sleep as "prior anchors" for cross-modal alignment; In the second stage, based on the frozen general semantic codebooks for sleep, the difference between ECG features and general semantic codebooks in the distance space is constrained by comparing the learning loss and mean squared error loss, thereby realizing the cross-modal transfer of the gold standard sleep staging knowledge of EEG to the ECG domain.

[0055] The following describes the specific process of the ECG-EEG sleep staging method based on a single-channel VQ-VAE (vector quantization variational autoencoder) provided by this invention, using specific examples (e.g.) Figure 1 (as shown) and technical effects.

[0056] This method includes the following two stages: Phase 1: Pre-training of the EEG VQ-VAE sleep general semantic codebook, which includes the following steps: S1. Acquire tagged EEG signals and perform preprocessing, including: Acquire tagged single-channel / multi-channel EEG signals, and uniformly downsample (to 100Hz), remove invalid frames, and then divide them according to a preset time interval (30s).

[0057] S2. Input the EEG signal into the first LAMBA-Net intra-frame encoder to obtain continuous hidden features.

[0058] S3. Construct a learnable discrete codebook, hard-encode and quantize continuous latent features based on the discrete codebook to obtain quantized feature maps, input the quantized feature maps into the decoder, reconstruct the EEG signal, construct a joint optimization loss function for joint training and optimization, and freeze the sleep general semantic codebook obtained during training as prior knowledge for cross-modal alignment, including: S31. Construct a learnable discrete codebook; S32. Calculate the Euclidean distance between each feature vector in the continuous latent features and all code vectors in the discrete codebook. Select the nearest neighbor code vector to perform hard coding quantization on the continuous latent features to obtain the quantized feature map. S33. Input the quantized feature map into the decoder to reconstruct the EEG signal; S34. Construct a joint optimization loss function based on reconstruction loss, codebook learning loss, and commitment loss, and perform joint training optimization. S35. Freeze the general semantic codebook of sleep obtained during training as prior knowledge for cross-modal alignment.

[0059] Specifically, S34 constructs a joint optimization loss function based on reconstruction loss, codebook learning loss, and commitment loss, including: Joint optimization loss function L total Expressed as follows: ; Among them, L rec The reconstruction loss is given by x, which represents the original input EEG signal. The reconstructed EEG signal, This represents the calculation of the square of the L2 Euclidean distance; L codebook For codebook learning loss, z e e is a single feature vector in a continuous latent feature set. k A single code vector in a discrete codebook. This represents the gradient truncation operator; L commit The commitment loss is represented by β, which is the weighting coefficient for the commitment loss (β is preferably 0.25).

[0060] Phase Two: ECG-EEG cross-modal feature alignment and sleep staging, specifically including the following steps: S4. Acquire single-lead ECG signals and perform preprocessing, including: S41. Acquire single-lead ECG signals using wearable single-lead ECG signal acquisition devices (such as flexible electrode ECG patches); S42. Anti-aliasing filtering is performed using an FIR low-pass filter (1024th order Hanning window), and then linear interpolation is used for unified resampling (up to 100Hz, corresponding to a fixed sequence length of 3000 points for a single 30s frame) to remove power frequency interference, abnormal baseline drift frames, and invalid frames labeled as motion period / unknown period. S43. Using a certain number (100) consecutive frames as the window size, perform sliding segmentation (50% overlap rate) to obtain temporal context samples (in the form of [number of windows, number of channels, sequence length]). Use Z-Score standardization to normalize the entire dataset using the mean and standard deviation of the training set, and save it to ensure deployment consistency.

[0061] S5. Input the ECG signal into the second LAMBA-Net intra-frame encoder to obtain the ECG intra-frame features.

[0062] S6. Calculate the distance matrices between continuous latent features, ECG intra-frame features, and the general sleep semantic codebook, respectively. Also calculate the contrastive learning target matrix, construct a joint alignment loss function to align ECG intra-frame features, narrowing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space. This allows for cross-modal transfer of EEG gold standard sleep staging knowledge to the ECG domain, including: S61. Calculate the first distance matrix between continuous latent features and the general semantic codebook for sleep, and the second distance matrix between ECG intra-frame features and the general semantic codebook for sleep. S62. Perform L2 normalization on the first distance matrix and the second distance matrix respectively, and then calculate the cosine similarity according to the temperature coefficient to obtain the contrast learning target matrix. S63. Construct a joint alignment loss function to align ECG intra-frame features, reduce the gap between ECG intra-frame features and the general semantic codebook of sleep in the distance space, and transfer the sleep staging knowledge of the EEG gold standard across modalities to the ECG domain.

[0063] Specifically, S63 constructs a joint alignment loss function for intra-frame ECG feature alignment, including: Joint alignment loss function L align Expressed as follows: ; Among them, L align For the joint alignment loss function, L InfoNCE To compare learning loss, L MSE The mean error loss is α, and the contrastive learning loss weight coefficient is α (preferably 0.7).

[0064] S7. Input the aligned ECG intra-frame features into the LAMBA-Net inter-frame encoder to model long-range dependencies across sleep cycles and obtain global feature representations.

[0065] S8. Project the global feature representation onto the classification space and output the sleep staging results, including: The global feature representation is projected onto a five-class classification space through a linear mapping to output the sleep staging results.

[0066] Based on the above-disclosed ECG-EEG sleep staging method based on single-channel VQ-VAE, this invention also discloses an ECG-EEG sleep staging system based on single-channel VQ-VAE, comprising the following functional modules: The EEG signal acquisition and preprocessing module acquires tagged EEG signals and performs preprocessing. The continuous latent feature coding module uses the first LAMBA-Net intra-frame encoder to extract features from the EEG signal to obtain continuous latent features; The VQ-VAE pre-training module constructs a learnable discrete codebook, hard-codes and quantizes continuous latent features based on the discrete codebook to obtain quantized feature maps, inputs the quantized feature maps into the decoder, reconstructs the EEG signal, constructs a joint optimization loss function for joint training and optimization, and freezes the sleep general semantic codebook obtained during training as prior knowledge for cross-modal alignment. The ECG signal acquisition and preprocessing module acquires single-lead ECG signals and performs preprocessing. The ECG intra-frame feature coding module uses the second LAMBA-Net intra-frame encoder to extract features from the ECG signal and obtain ECG intra-frame features. The cross-modal alignment module calculates the distance matrix between continuous latent features, ECG intra-frame features and the general sleep semantic codebook, respectively, and calculates the contrastive learning target matrix. It constructs a joint alignment loss function to align ECG intra-frame features, narrowing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space, and transferring the sleep staging knowledge of the EEG gold standard to the ECG domain across modalities. The inter-frame temporal modeling module uses the LAMBA-Net inter-frame encoder to model the long-range dependency of aligned ECG intra-frame features across sleep cycles, and obtains global feature representations. The classification output module projects the global feature representation onto the classification space and outputs the sleep staging results. The visualization module plots the sleep staging results as an overnight sleep staging curve for reference.

[0067] In the technical solution of this application, the first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder share parameters but have independent weights. Both the first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder include the following structure: Non-stationary branch, extract non-stationary transient frequency domain features, and complete the association expression between multi-scale local context modeling and non-stationary transient frequency domain features; Stationary branching is used to extract time-domain features of stationary rhythms. The feature fusion module based on the channel attention mechanism takes non-stationary transient frequency domain features and stationary rhythmic time domain features as input by adding them element by element. First, it compresses the temporal information of each channel through adaptive average pooling, then generates channel weights through two fully connected layers and the Sigmoid activation function. The channel weights are multiplied by the input features channel by channel, and finally the channel dimension is restored through one-dimensional convolution. Intra-frame BiLSTM branch, through bidirectional gating structure, finely captures local short-term temporal correlations and nonlinear transient details between adjacent subframes within an epoch; The intra-frame Mamba branch, relying on the selective state-space model Mamba, dynamically selects the input-related state matrix and stride to efficiently model the long-range dependency of the sub-frame sequence with linear time complexity. It is then concatenated with the output features of the intra-frame BiLSTM branch along the channel dimension, and finally fused by one-dimensional convolution to obtain continuous latent features.

[0068] In the technical solution of this application, the LAMBA-Net inter-frame encoder includes the following structure: Inter-frame BiLSTM branch captures short-term transition patterns during sleep phases; The inter-frame Mamba branch models long-range dependencies across sleep cycles and fuses them with the output features of the inter-frame BiLSTM branch through one-dimensional convolution, resulting in a global feature representation that simultaneously integrates fine intra-frame time-frequency features and long-range inter-frame temporal information.

[0069] To better illustrate the technical solution of this application, the following explanation is based on subject-level 5-fold cross-validation conducted on the “ECGI” channel of the MASS dataset and the “ECG” channel of the SHHS dataset (the random seed is set to 42, and 200 subjects are divided into 5 mutually exclusive subsets).

[0070] Experimental results show that the present invention achieves the highest global Kappa coefficient, overall MF1, and MF1 for all five sleep stages across both datasets, significantly outperforming the knowledge distillation model, Holter-to-Sleep model, dual-path VQ-VAE scheme, and the ECG baseline model that directly uses LAMBA-Net (Long-term Associated Mamba Integrated Network). The present invention also exhibits the smallest standard deviation across all metrics on the SHHS dataset, demonstrating optimal cross-scenario stability.

[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An ECG-EEG sleep staging method based on single-channel VQ-VAE, characterized in that: It includes the following two stages: Phase 1: Pre-training of the EEG VQ-VAE sleep general semantic codebook, which includes the following steps: S1. Acquire tagged EEG signals and perform preprocessing; S2. Input the EEG signal into the first LAMBA-Net intra-frame encoder to obtain continuous hidden features; S3. Construct a learnable discrete codebook, perform hard coding and quantization on continuous latent features based on the discrete codebook to obtain quantized feature maps, input the quantized feature maps into the decoder, reconstruct the EEG signal, construct a joint optimization loss function for joint training and optimization, and freeze the sleep general semantic codebook obtained from the training as prior knowledge for cross-modal alignment. Phase Two: ECG-EEG cross-modal feature alignment and sleep staging, specifically including the following steps: S4. Acquire single-lead ECG signals and perform preprocessing; S5. Input the ECG signal into the second LAMBA-Net intra-frame encoder to obtain the ECG intra-frame features; S6. Calculate the distance matrix between continuous latent features, ECG intra-frame features and general sleep semantic codebook respectively, and calculate the contrastive learning target matrix. Construct a joint alignment loss function to align ECG intra-frame features, reduce the gap between ECG intra-frame features and general sleep semantic codebook in the distance space, and transfer the sleep staging knowledge of the EEG gold standard across modalities to the ECG domain. S7. Input the aligned ECG intra-frame features into the LAMBA-Net inter-frame encoder to model long-range dependencies across sleep cycles and obtain global feature representations. S8. Project the global feature representation onto the classification space and output the sleep staging results.

2. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 1, characterized in that: S1 acquires tagged EEG signals and performs preprocessing, including: Acquire tagged single-channel / multi-channel EEG signals, uniformly downsample, remove invalid frames, and then divide them according to a preset time interval.

3. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 2, characterized in that: In S3, a learnable discrete codebook is constructed. Based on the discrete codebook, continuous latent features are hard-coded and quantized to obtain quantized feature maps. These quantized feature maps are then input into the decoder to reconstruct the EEG signal. A joint optimization loss function is constructed and jointly trained for optimization. The sleep general semantic codebook obtained during training is frozen as prior knowledge for cross-modal alignment, including: S31. Construct a learnable discrete codebook; S32. Calculate the Euclidean distance between each feature vector in the continuous latent features and all code vectors in the discrete codebook. Select the nearest neighbor code vector to perform hard coding quantization on the continuous latent features to obtain the quantized feature map. S33. Input the quantized feature map into the decoder to reconstruct the EEG signal; S34. Construct a joint optimization loss function based on reconstruction loss, codebook learning loss, and commitment loss, and perform joint training optimization. S35. Freeze the general semantic codebook of sleep obtained during training as prior knowledge for cross-modal alignment.

4. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 3, characterized in that: S34 constructs a joint optimization loss function based on reconstruction loss, codebook learning loss, and commitment loss, including: Joint optimization loss function L total Expressed as follows: ; Among them, L rec The reconstruction loss is given by x, which represents the original input EEG signal. The reconstructed EEG signal, This represents the calculation of the square of the L2 Euclidean distance; L codebook For codebook learning loss, z e e is a single feature vector in a continuous latent feature set. k A single code vector in a discrete codebook. This represents the gradient truncation operator; L commit β represents the commitment loss, and β is the commitment loss weighting coefficient.

5. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 3, characterized in that: S4 acquires a single-lead ECG signal and performs preprocessing, including: S41. Acquire single-lead ECG signals using a wearable single-lead ECG signal acquisition device; S42. Anti-aliasing filtering is performed through an FIR low-pass filter, and then unified resampling is performed through linear interpolation to remove power frequency interference, abnormal frames with baseline drift, and invalid frames labeled as motion period / unknown period. S43. Using a certain number of consecutive frames as the window size, perform sliding segmentation to obtain temporal context samples. Use Z-Score standardization to normalize the entire dataset using the mean and standard deviation of the training set, and save it to ensure deployment consistency.

6. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 5, characterized in that: In S6, the distance matrices between continuous latent features, ECG intra-frame features, and the general sleep semantic codebook are calculated respectively. The contrastive learning target matrix is ​​also calculated, and a joint alignment loss function is constructed to align ECG intra-frame features, reducing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space. This allows for cross-modal transfer of EEG gold standard sleep staging knowledge to the ECG domain, including: S61. Calculate the first distance matrix between continuous latent features and the general semantic codebook for sleep, and the second distance matrix between ECG intra-frame features and the general semantic codebook for sleep. S62. Perform L2 normalization on the first distance matrix and the second distance matrix respectively, and then calculate the cosine similarity according to the temperature coefficient to obtain the contrast learning target matrix. S63. Construct a joint alignment loss function to align ECG intra-frame features, reduce the gap between ECG intra-frame features and the general semantic codebook of sleep in the distance space, and transfer the sleep staging knowledge of the EEG gold standard across modalities to the ECG domain.

7. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 6, characterized in that: S63 constructs a joint alignment loss function for ECG intra-frame feature alignment, including: Joint alignment loss function L align Expressed as follows: ; Among them, L align For the joint alignment loss function, L InfoNCE To compare learning loss, L MSE α represents the mean error loss, and α is the weighting coefficient of the contrastive learning loss.

8. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 6, characterized in that: In S8, the global feature representation is projected onto the classification space, and the sleep staging results are output, including: The global feature representation is projected onto a five-class classification space through a linear mapping to output the sleep staging results.

9. An ECG-EEG sleep staging system based on a single-channel VQ-VAE, applicable to the ECG-EEG sleep staging method based on a single-channel VQ-VAE as described in claim 1, characterized in that: Includes the following functional modules: The EEG signal acquisition and preprocessing module acquires tagged EEG signals and performs preprocessing. The continuous latent feature coding module uses the first LAMBA-Net intra-frame encoder to extract features from the EEG signal to obtain continuous latent features; The VQ-VAE pre-training module constructs a learnable discrete codebook, hard-codes and quantizes continuous latent features based on the discrete codebook to obtain quantized feature maps, inputs the quantized feature maps into the decoder, reconstructs the EEG signal, constructs a joint optimization loss function for joint training and optimization, and freezes the sleep general semantic codebook obtained during training as prior knowledge for cross-modal alignment. The ECG signal acquisition and preprocessing module acquires single-lead ECG signals and performs preprocessing. The ECG intra-frame feature coding module uses the second LAMBA-Net intra-frame encoder to extract features from the ECG signal and obtain ECG intra-frame features. The cross-modal alignment module calculates the distance matrix between continuous latent features, ECG intra-frame features and the general sleep semantic codebook, respectively, and calculates the contrastive learning target matrix. It constructs a joint alignment loss function to align ECG intra-frame features, narrowing the gap between ECG intra-frame features and the general sleep semantic codebook in the distance space, and transferring the sleep staging knowledge of the EEG gold standard to the ECG domain across modalities. The inter-frame temporal modeling module uses the LAMBA-Net inter-frame encoder to model the long-range dependency of aligned ECG intra-frame features across sleep cycles, and obtains global feature representations. The classification output module projects the global feature representation onto the classification space and outputs the sleep staging results. The visualization module plots the sleep staging results as an overnight sleep staging curve for reference.

10. The ECG-EEG sleep staging method based on single-channel VQ-VAE according to claim 9, characterized in that: The first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder share parameters but have independent weights. Both the first LAMBA-Net intra encoder and the second LAMBA-Net intra encoder include the following structure: Non-stationary branch, extract non-stationary transient frequency domain features, and complete the association expression between multi-scale local context modeling and non-stationary transient frequency domain features; Stationary branching is used to extract time-domain features of stationary rhythms. The feature fusion module based on the channel attention mechanism takes non-stationary transient frequency domain features and stationary rhythmic time domain features as input by adding them element by element. First, it compresses the temporal information of each channel through adaptive average pooling, then generates channel weights through two fully connected layers and the Sigmoid activation function. The channel weights are multiplied by the input features channel by channel, and finally the channel dimension is restored through one-dimensional convolution. Intra-frame BiLSTM branch, through bidirectional gating structure, finely captures local short-term temporal correlations and nonlinear transient details between adjacent subframes within an epoch; The intra-frame Mamba branch, relying on the selective state-space model Mamba, dynamically selects the input-related state matrix and step size to efficiently model the long-range dependency of the sub-frame sequence with linear time complexity. It is then concatenated with the output features of the intra-frame BiLSTM branch along the channel dimension, and then fused through one-dimensional convolution to obtain continuous latent features. The LAMBA-Net inter-frame encoder includes the following structure: Inter-frame BiLSTM branch captures short-term transition patterns during sleep phases; The inter-frame Mamba branch models long-range dependencies across sleep cycles and fuses them with the output features of the inter-frame BiLSTM branch through one-dimensional convolution, resulting in a global feature representation that simultaneously integrates fine intra-frame time-frequency features and long-range inter-frame temporal information.