A non-invasive sleep staging method and system based on space-time transformer

CN122767779APending Publication Date: 2026-09-18HUNAN VENTMED MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610747936.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明旨在提供一种基于时空Transformer的非侵入式睡眠分期方法及系统,以解决现有非侵入式睡眠分期技术在家用场景下,难以有效感知电极接触质量的动态变化,难以实现微观波形与宏观拓扑特征的深度融合,导致的睡眠分期准确率不佳的问题

Benefits of technology

(1)本发明针对家庭监测场景下信号质量易波动、大脑空间拓扑挖掘不足以及微观与宏观特征难以兼顾的痛点,提供了一种全流程的睡眠分期方案;首先,通过引入阻抗感知与调制机制,感知各通道电极接触质量的动态变化,并以此对脑电特征贡献度进行实时调适,有效抑制了因受试者翻身或出汗导致的接触噪声污染,为后续分析奠定了可靠的信号基础;在此基础上,通过挖掘多频段的动态空间连接关系并实现稀疏化处理,使得模型能够排除伪连接干扰,捕获不同睡眠深度对应的脑区协同模式;进一步地,利用时空融合策略解决了微观精细波形与宏观脑网络拓扑在时间尺度上的不匹配问题,实现了特征在多尺度上的深度协同,使模型能够根据当前的全局脑功能状态灵活调配对瞬态生理事件(如纺锤波)的解读权重;最后,结合睡眠生理连续性约束进行序列级决策,修正了不合理的阶段跳变,确保输出的分期序列既符合微观特征表现又满足宏观生理演变规律,提升了非侵入式设备在居家场景下的分期准确性与结果稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122767779A_ABST
    Figure CN122767779A_ABST
Patent Text Reader

Abstract

This invention discloses a non-invasive sleep staging method and system based on spatiotemporal Transformer, comprising: A1: acquiring multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocessing them respectively; A2: sequentially extracting basic temporal features of EEG, impedance state encoding, impedance-modulated EEG features, and self-attention enhancement features, and outputting multi-channel EEG data after deep noise reduction; A3: calculating dynamic spatial connectivity graph features; A4: sequentially extracting local microscopic waveform features, aligned and fused spatiotemporal features, and spatial enhancement features, and outputting spatiotemporal hybrid features; A5: calculating the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages, and constructing a two-level loss function to provide feedback correction to the model parameters, thereby completing sleep staging. This invention can solve the problem of poor sleep staging accuracy caused by the inability of traditional methods to effectively perceive dynamic changes in electrode contact quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sleep monitoring technology, and in particular to a non-invasive sleep staging method and system based on spatiotemporal Transformer. Background Technology

[0002] Sleep staging, based on international standards established by the American Academy of Sleep Medicine (AASM), divides the sleep process into wakefulness (W), non-rapid eye movement (NREM) sleep (including three sub-stages: N1, N2, and N3) and rapid eye movement (REM) sleep through characteristic analysis of electroencephalogram (EEG) signals. As a crucial foundation for sleep quality assessment, sleep disorder screening, and brain function analysis, the accuracy, continuity, and stability of sleep staging results are of great significance to the effectiveness of clinical interventions. With the development of brain-computer interface (BCI) technology, non-invasive, long-term home-based dry electrode EEG monitoring devices, with their advantages of being non-invasive, portable, and easy to operate, are gradually becoming an important development direction to replace the traditional hospital polysomnography (PSG) gold standard.

[0003] However, in the practical application of home-based non-invasive BCI devices, automatic sleep staging technology faces multiple real-world challenges. First, in a home environment, dry electrodes are used for data collection, and signal quality is affected by environmental interference and hardware limitations, resulting in non-stationary fluctuations. During sleep, subjects inevitably experience physiological behaviors such as turning over, head movement, or sweating, causing the contact state between the electrodes and the scalp to dynamically change. At any given moment, the frontal lobe electrodes may deviate significantly from normal impedance due to displacement, leading to severe signal distortion, while the central area electrodes maintain good contact. Second, sleep staging is highly dependent on the spatial topological differences in brain functional connectivity at different stages. For example, in N3 stage, the frontal and parietal lobes often show high synchronous enhancement in the low-frequency delta band; this inter-regional coordination pattern reflects a specific sleep depth. Furthermore, sleep staging needs to simultaneously consider two types of heterogeneous features that are severely mismatched in observation scale: one is microscopic waveform features lasting only 0.5 to 3 seconds (such as spindle waves and K-complexes); the other is macroscopic brain network topological features based on a 30-second window. Existing methods struggle to achieve deep synergy between these two aspects: focusing on micro-waveform extraction risks losing the discriminative value of macro-features such as synchronized slow waves across the entire brain; focusing on macro-topology results in the loss of crucial transient micro-events that distinguish between N1 and N2 phases. Furthermore, the same micro-waveform often carries different physiological meanings in different whole-brain topological contexts. Current approaches lack effective spatiotemporal alignment and hybrid modeling mechanisms, making it difficult to dynamically adjust the interpretation weights of micro-waveforms based on the current brain network state. Summary of the Invention

[0004] In view of this, the present invention aims to provide a non-invasive sleep staging method and system based on spatiotemporal Transformer to solve the problem that existing non-invasive sleep staging technologies are difficult to effectively sense dynamic changes in electrode contact quality and difficult to achieve deep integration of microscopic waveforms and macroscopic topological features in home settings, resulting in poor sleep staging accuracy.

[0005] A non-invasive sleep staging method based on spatiotemporal Transformer includes: A1: Collect multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocess them to obtain preprocessed multi-channel EEG data and preprocessed multi-channel impedance data; A2: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, extract the basic temporal features of EEG and impedance state encoding, impedance modulation EEG features and self-attention enhancement features in sequence, and output the multi-channel EEG data after deep noise reduction. A3: Based on the multi-channel EEG data after deep noise reduction, extract multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix and features of each channel node in sequence, and output dynamic spatial connectivity graph features. A4: Based on the multi-channel EEG data after deep noise reduction and the dynamic spatial connectivity map features, local micro waveform features, aligned and fused spatiotemporal features, and spatial enhancement features are extracted sequentially to output spatiotemporal hybrid features. A5: Based on the spatiotemporal hybrid characteristics, the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages are calculated sequentially, and a two-level loss function is constructed to provide feedback correction to the model parameters, thus completing the sleep staging.

[0006] Furthermore, step A1 also includes: A11: Multi-channel EEG signals during the subject's sleep process are collected using a home-use non-invasive dry electrode brain-computer interface device. The data type is multi-channel time-series voltage data, including frontal lobe lead EEG signals, central area lead EEG signals, and reference lead EEG signals. The frontal lobe lead EEG signals include Fp1 or Fp2 channel signals, the central area lead EEG signals include C3 or C4 channel signals, and the reference lead EEG signals include A1 and A2 earlobe reference signals. Simultaneously, electrode-scalp contact impedance data output by the device's built-in sensors are collected. The data type is time-series resistance data, including the real-time contact impedance values ​​of each lead. A12: The multi-channel EEG signal is detected and corrected using the AP0 jump zeroing method to obtain the baseline-corrected multi-channel EEG signal; the electrode-scalp contact impedance data is processed by removing abnormal impedance values ​​using the three-times-standard-deviation method, then filling in missing values ​​using linear interpolation, and finally using sliding window mean filtering to smooth high-frequency fluctuations to obtain the pre-processed electrode-scalp contact impedance data. A13: Based on the spatial topological coordinates of each lead electrode on the scalp, the baseline-corrected multi-channel EEG signal is spatially denoised by using a common average reference method combined with an adaptive artifact subspace reconstruction method to obtain a spatially denoised multi-channel EEG signal. A14: For the EEG data of each channel in the spatially denoised multi-channel EEG signal and the electrode-scalp contact impedance data after preliminary preprocessing, an alignment method based on the system global timestamp is used to achieve temporal synchronization of the multi-dimensional data, and the synchronized data is linearly normalized to unify the amplitude dimensions, so as to obtain the preprocessed multi-channel EEG data and the preprocessed multi-channel impedance data.

[0007] Furthermore, step A2 also includes: A21: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, the initial feature space is mapped through a one-dimensional convolutional network to extract the basic temporal features of EEG and impedance state encoding. A22: Based on the basic temporal features of EEG and impedance state encoding, dynamic impedance gating weights are calculated through feature concatenation and convolution operations, and the basic temporal features of EEG are modulated to obtain impedance-modulated EEG features. A23: Based on impedance-modulated EEG characteristics, a physiological transient response mask is calculated using second-order temporal difference. Then, the evolution pattern of long-term sleep is extracted, and the physiological transient response mask is injected as a residual supplement term to obtain self-attention enhancement features. A24: Based on the self-attention enhancement features, the sequence length is decoded and restored through a one-dimensional deconvolution layer, and finally, it is connected to the preprocessed multi-channel EEG data to output the deep noise-reduced multi-channel EEG data.

[0008] Furthermore, step A2 also includes: The preprocessed multichannel EEG data were processed by a one-dimensional convolutional layer to obtain the basic temporal features of EEG. The preprocessed multichannel impedance data were processed by a one-dimensional convolutional layer and a modified linear unit with leakage to obtain the impedance state code. The basic EEG temporal features and impedance state codes are processed sequentially through dimension concatenation, a one-dimensional convolutional layer, and a Sigmoid function to obtain dynamic impedance gating weights. Then, the dynamic impedance gating weights are combined with the basic EEG temporal features to obtain impedance-modulated EEG features. For the impedance modulation EEG features at time t, the impedance modulation EEG features at time t+1 are subtracted by twice the impedance modulation EEG features at time t, and then the impedance modulation EEG features at time t-1 are added to obtain the characteristic second-order temporal difference component at time t. The absolute value of the characteristic second-order temporal difference component at time t is taken and processed by the ReLU function to obtain the physiological transient response mask at time t. Then, the impedance modulation EEG features are processed by the multi-head self-attention mechanism and added element-wise with the Hadamard product of the impedance modulation EEG features at time t and the physiological transient response mask at time t to obtain the self-attention enhancement features at time t. The self-attention enhancement features are processed sequentially through a one-dimensional deconvolution layer and a linear layer to obtain the decoded and reconstructed features. The product of the decoded and reconstructed features and the scaling coefficient vector is then processed by layer normalization and added element-wise to the preprocessed multi-channel EEG data to obtain the deep noise-reduced multi-channel EEG data.

[0009] It is important to further clarify that in long-term sleep EEG data collection in a home setting, the contact state between the dry electrodes and the scalp is not constant. Subjects inevitably experience physiological and behavioral changes throughout the night, such as turning over, head movement, and sweating. These factors lead to varying electrode-scalp contact quality across different channels at different times. Some channels may capture distorted data dominated by contact noise at specific times, rather than true EEG activity. Furthermore, this degradation in contact quality is not a globally synchronous phenomenon but exhibits channel heterogeneity and time-varying characteristics. For example, at one moment, the frontal lobe electrodes may experience a sudden increase in impedance and severe signal distortion due to sweat accumulation, while the central area electrodes maintain good contact. The situation may reverse in the next time period. If all channels and all times are treated equally without distinction, severely distorted channels will enter the subsequent feature extraction process with the same weight as normal channels, contaminating the entire sleep staging analysis process. To address the aforementioned challenges, this invention designs a dynamic signal reliability adjustment mechanism based on impedance sensing. The implementation logic of its technical effect is as follows: First, one-dimensional convolution mapping is performed on the preprocessed multi-channel impedance data and the preprocessed multi-channel EEG data, respectively, transforming the original impedance time series into impedance state codes. The impedance data itself is a sequence of resistance values ​​that continuously changes over time, and its level directly reflects the physical contact tightness between the electrode and the scalp at the current moment—when the subject turns over, causing a certain electrode to loosen, or when sweating in the frontal lobe causes an abnormal increase in the skin-electrode interface impedance, the impedance value of that channel at the corresponding moment will significantly deviate from the normal level. Through one-dimensional convolution, local pattern extraction is performed on this time-series impedance signal, allowing the model to… Different types of contact anomaly patterns, such as "impedance consistently high within a certain time window" and "abrupt impedance jumps," are transformed into impedance state codes with semantic discriminative capabilities. This allows the network to perceive the contact quality state of each channel at each time point, rather than relying solely on the absolute impedance value at a single moment, thus robustly capturing the dynamic process of contact degradation. The second step involves concatenating these impedance state codes with basic EEG temporal features, followed by convolution to generate dynamic impedance gating weights. These weights are then applied to the basic EEG temporal features element-wise using a Hadamard product to obtain impedance-modulated EEG features. The gating weights are not manually set fixed thresholds but are adaptively learned by the network based on the joint distribution of the impedance state codes and EEG features. When the impedance state encoding of a channel indicates severe degradation in contact quality at a certain moment, the corresponding gating weight approaches zero, significantly reducing the contribution of the EEG features of that channel at that moment. Conversely, when the impedance state encoding indicates good contact, the corresponding gating weight approaches one, and the EEG features are better preserved. This soft modulation method allows the dynamic changes in signal quality to be expressed relatively correspondingly in the feature space, avoiding information loss caused by the removal of the entire channel to a certain extent, and also avoiding noise pollution caused by equal weighting to a certain extent. In the third step, for the impedance-modulated EEG features, the physiological transient response mask is calculated using second-order temporal difference, and the long-term sleep evolution pattern is extracted by combining the multi-head self-attention mechanism. The state response mask is injected as a residual complement to obtain self-attention enhancement features. Among them, the sleep EEG signal contains both slowly evolving macroscopic sleep rhythm structure, such as the periodic alternation from light sleep to deep sleep and then to REM sleep, and transient physiological events within local time periods, such as the waveform features of K complex waves and sleep spindle waves that suddenly appear within short time windows. The multi-head self-attention mechanism is good at capturing long-distance sequence dependencies and can perceive the rhythmic evolution pattern across time periods throughout the night. The second-order temporal difference is highly sensitive to the local acceleration changes of features on the time axis, and can effectively locate the positions of transient physiological events that fluctuate significantly in a short period of time, and mark and enhance these positions in the form of physiological transient response mask.The two are fused by residual superposition, so that the self-attention enhancement features take into account both macroscopic sleep structure and local transient response, which helps the subsequent stage model to make more accurate judgments at key nodes such as sleep transition boundaries. In the fourth step, the self-attention enhancement features are decoded and restored by a one-dimensional deconvolution layer, and then skip connections are made with the preprocessed multi-channel EEG data to output deeply denoised multi-channel EEG data. The role of skip connections is that after multi-layer feature transformation, the network learns the residual correction amount of the original signal, rather than the re-estimation of the signal. This means that when the electrode contact quality is good and the original EEG signal itself is relatively clean, the correction amplitude of the network will naturally tend to be smaller, and the effective components of the original signal can be better preserved. When the contact noise is severe, the network can apply a larger residual correction based on the quality perception information transmitted by the preceding impedance gating, so that the output deeply denoised multi-channel EEG data can maintain relatively reliable signal quality throughout the night's sleep. Compared to existing methods, traditional home EEG acquisition solutions typically input all acquisition channels into the analysis model with equal weight, or rely solely on fixed thresholds to remove channels that exceed the acceptable range. This approach fails to adequately address the dynamic drift of impedance throughout the night and struggles to distinguish the reliability differences of the same channel at different times. This invention introduces impedance data as an auxiliary information stream parallel to EEG data into the network and achieves deep fusion of the two in the feature space through a gating modulation mechanism. This allows the model to dynamically reweight the contribution of each channel to EEG features at each time point based on physically measurable contact quality indicators. This helps alleviate the problem of dynamically changing signal reliability being treated with equal weight in long-term dry electrode acquisition in home settings, providing a more reliable signal quality foundation for subsequent sleep feature extraction.

[0010] Furthermore, step A3 also includes: A31: Based on the multi-channel EEG data after deep noise reduction, it is divided into time windows according to the time dimension, and the signals in each window are decomposed into time and frequency to obtain the multi-channel EEG data and multi-band time and frequency energy characteristics after the time window is divided. A32: Based on the multi-band time-frequency energy characteristics, the dynamic functional connection matrix between each channel is calculated through adaptive time-frequency coherence; A33: Based on the dynamic functional connection matrix, adaptive filtering is performed to achieve sparsity, resulting in a sparse functional connection graph matrix and the characteristics of each channel node; A34: Based on the characteristics of each channel node and the sparse functional connection matrix, spatial topology encoding is performed through a graph convolutional network, and then the dynamic spatial connection graph features are fused and output.

[0011] Furthermore, step A3 also includes: For the deep noise-reduced multi-channel EEG data, the time window is divided into multiple channels according to the time window size and window step size. Then, the time window divided multi-channel EEG data is processed by short-time Fourier transform and frequency band extraction to obtain multi-frequency time-frequency energy features. The components of the multi-band time-frequency energy characteristics in the b-th frequency band are processed sequentially by global average pooling, linear layer, and Softmax function to obtain the adaptive frequency band weighting coefficients of the b-th frequency band. The coherence coefficient between channel i and channel j in the b-th frequency band is obtained by dividing the cross power spectrum of channel i and channel j in the b-th frequency band by the square root of the product of the self power spectrum of channel i and the self power spectrum of channel j in the b-th frequency band. The product of the adaptive frequency band weighting coefficients of each frequency band and the coherence matrix of the corresponding frequency band is summed over all frequency bands and then processed by the Sigmoid function to obtain the dynamic functional connection matrix. An adaptive sparsification threshold is obtained by processing the dynamic functional connection matrix using a fixed percentile threshold sparsification method. A sparsification mask is generated at the position in the dynamic functional connection matrix that is greater than the adaptive sparsification threshold. The sparsification mask is then multiplied by the dynamic functional connection matrix to obtain a sparse functional connection graph matrix. At the same time, the multi-band time-frequency energy features of all frequency bands are processed by dimension splicing and linear layers to obtain the features of each channel node. The node features of each channel and the sparse functional connectivity graph matrix are processed by a graph convolutional network to obtain graph convolutional network node features. Then, the graph convolutional network node features and the sparse functional connectivity graph matrix are processed by graph pooling to obtain graph pooling features. Finally, the graph convolutional network node features and the graph pooling features are processed by dimension concatenation to obtain dynamic spatial connectivity graph features.

[0012] It should be further explained that the basis for sleep stage differentiation lies in the significant differences in the topological structure of brain functional connectivity at different sleep stages. For example, in N3 stage deep non-rapid eye movement sleep, the frontal and parietal lobes show highly synchronized enhancement in the low-frequency delta band. This inter-regional co-activation directly reflects the unique pattern of brain-to-brain communication in deep sleep. Existing multi-channel EEG acquisition schemes generally face two limitations when dealing with this type of problem. First, many methods adopt a strategy of simply splicing multi-channel signals and then performing independent temporal modeling, extracting features only in the time dimension, but failing to fully explore the functional connectivity relationships between different brain regions across the spatial dimension. This results in the neglect of important stage differentiation information such as brain-to-brain co-connectivity patterns during the modeling process. Second, when calculating brain-to-brain connectivity, many methods tend to construct fully connected systems. The graph, which calculates the connection strength between any two channels, assigns connection weights to channel pairs that have no substantial connection in brain function. These "pseudo-connections" are mixed in with the real discriminative connection patterns, polluting the network's learning of sleep stage-related topological features. In addition, in real sleep, different sleep stages often have their own dominant frequency bands—the typical feature of N2 sleep is the appearance of σ-band spindle waves, while N3 sleep is dominated by δ-band slow waves, and REM sleep exhibits higher frequency characteristics. However, most existing traditional methods assign the same weight to all frequency bands or ignore the differences in frequency bands. They cannot adaptively adjust the emphasis on the connection of brain regions in different frequency bands according to the real-time changes in the current sleep depth. Therefore, they cannot effectively capture the brain region synergistic relationships that truly play a discriminative role in specific sleep stages. To address the aforementioned challenges, this invention designs a multi-band sleep-specific spatial connectivity topology mining mechanism. The implementation logic of its technical effect is as follows: First, the multi-channel EEG data after deep noise reduction is uniformly divided into time windows along the time dimension. Then, a short-time Fourier transform is performed on the signals within each window to obtain a time-frequency representation, thereby separating the sleep-related frequency band energy from the time-frequency matrix. The significance lies in the fact that the original EEG signal is in time-series form, containing mixed frequency components that cannot directly distinguish the EEG activities corresponding to different frequency bands. Through time window division, the model achieves a balance between time accuracy and frequency resolution. Through short-time Fourier transform, the signal is converted to the time-frequency domain, enabling the network to retain temporal location information. Simultaneously, the energy distribution of each frequency band at different times is obtained. The resulting multi-band time-frequency energy features can clearly present the EEG activity level of each channel in each frequency band at each time, laying the foundation for subsequent frequency band-specific functional connectivity analysis. The second step involves calculating the coherence coefficient between channels in each frequency band based on the multi-band time-frequency energy features to construct a frequency band-level functional connectivity matrix. Then, an adaptive weighting mechanism is used to dynamically assign weights to different frequency bands. Finally, the connectivity information of the multi-bands is fused to obtain the overall dynamic functional connectivity matrix. When calculating the coherence coefficient, the network utilizes the ratio of cross-power spectrum to self-power spectrum. This index reflects the similarity and synchronization degree of the EEG signals of two channels in the frequency domain. A higher σ-band energy level indicates stronger functional synergy between two brain regions in that frequency band. By calculating this index separately for each frequency band, the model can perceive that the same channel pair may exhibit different connection strengths in different frequency bands. Subsequently, the network adaptively learns weighting coefficients for each frequency band based on the global distribution characteristics of the multi-band time-frequency energy features. This results in higher weights for frequency bands that contribute more energy and are richer in information within the current time window, while lower weights are given to frequency bands that contribute less energy and have poorer discriminative power within the same time window. This frequency band-level adaptive weighting mechanism enables the model to dynamically adapt to different dominant frequency bands at different sleep stages. For example, when a significant increase in the σ-band energy is detected in a given time window, the importance of brain region connectivity information in the σ-band is automatically increased during the fusion process. This allows for better utilization of the typical characteristics of sleep stage N2 during stage segmentation. The third step involves adaptively sparsifying the dynamic functional connectivity matrix to generate a sparse functional connectivity graph. The principle of sparsification is that while a fully connected graph contains all connections between pairs of channels, most of these connections have low weights and lack discriminative significance. These low-weight edges introduce a lot of noise in subsequent graph convolution modeling. By adopting a fixed percentile threshold strategy, the network retains only those strong connections that are at the upper end of the distribution and are the most representative, while setting the remaining weak connections to zero. This helps to filter out pseudo-connections, allowing the sparse graph to more clearly present the true brain functional organization of the current sleep stage and reduce the interference of low-weight pseudo-connections on subsequent classification.Simultaneously, the model fuses the time-frequency energy features of each frequency band through linear transformation to generate node features for each channel. These node features carry comprehensive EEG activity information across multiple frequency bands and will be propagated and aggregated as attribute information of graph nodes in subsequent graph convolutions. In the fourth step, spatial topological encoding is performed using the node features of each channel and the sparse functional connectivity graph as input. The graph convolution network performs message passing according to the graph's topological structure—for a certain brain region channel, the graph convolution aggregates its neighboring channel information on the functional connectivity graph, enabling the representation of the brain region to encode the spatial dimension information of "the cooperative relationship between surrounding nodes and this node." Furthermore, the node features after graph convolution are globally aggregated through graph pooling operations, allowing the network to capture the macroscopic topological characteristics of the brain functional network within the current time window from an overall perspective. Finally, the node-level features and the global pooled features are concatenated, so that the output dynamic spatial connectivity graph features simultaneously possess spatial topological information of both local brain regions and the global network, providing rich, multi-scale spatial dimension discriminative features for subsequent sleep period classification. Compared to existing methods, traditional multi-channel EEG sleep staging schemes typically rely solely on independent modeling of temporal features or employ fully connected graphs with fixed frequency band weights. This makes it difficult to fully leverage the discriminative role of multi-channel spatial location information and adapt to the dynamic changes in dominant frequency bands at different sleep stages. This invention, through a progressive architecture that gradually constructs time window partitioning, multi-band coherent computation, adaptive frequency band weighting, sparsification processing, and graph convolutional spatial encoding, enables the model to adaptively mine the unique brain functional connectivity topology of each sleep stage based on multi-band data. This effectively filters out interference from pseudo-connections and fully utilizes the synergistic relationships between brain regions to encode spatial dimension discriminative information for sleep staging, providing more realistic and reliable brain functional topological features for subsequent sleep stage classification.

[0013] Furthermore, step A4 also includes: After deep noise reduction, the multi-channel EEG data are processed by a one-dimensional convolutional layer of the first receptive field and a one-dimensional convolutional layer of the second receptive field, respectively. Then, a dimensional splicing operation is performed to obtain local microscopic waveform features. At the same time, the dynamic spatial connectivity graph features are processed by copying and expanding along the time dimension and then processed by a linear layer. The result is added element-wise with the result of the local microscopic waveform features processed by the linear layer and then processed by the Sigmoid function to obtain the spatiotemporal alignment coupling operator. The dynamic spatial connectivity graph features are then processed by copying and expanding along the time dimension and then performed with the spatiotemporal alignment coupling operator to form a Hadamard product. Finally, the result is added element-wise with the local microscopic waveform features to obtain the aligned and fused spatiotemporal features. The dynamic spatial connectivity graph features are processed sequentially by a linear layer and a Sigmoid function to obtain the topological manifold modulation factor. Then, the aligned and fused spatiotemporal features are processed by a spatial multi-head self-attention mechanism and added element-wise with the aligned and fused spatiotemporal features. After layer normalization, the Hadamard product is performed with the topological manifold modulation factor to obtain the spatial enhancement features. The spatial augmentation features are processed by a temporal multi-head self-attention mechanism to obtain temporal evolution features. The temporal evolution features are then processed sequentially by a gated convolutional layer and a Sigmoid function to obtain temporal evolution gate weights. The temporal evolution features and the temporal evolution gate weights are then multiplied by a Hadamard product. This product is then added element-wise to the spatial augmentation features and the Hadamard product obtained by subtracting the temporal evolution gate weights from 1 to obtain the spatiotemporal hybrid features.

[0014] It needs further explanation that the diagnostic criteria for sleep staging involve two types of essentially complementary but significantly mismatched feature information in terms of observation scale and time span. The first is microscopic waveform features, referring to the microscopic EEG structures presented at the time-by-time level, such as typical sleep waveforms like spindle waves and K-complexes. These waveforms last only a short time range of 0.5 to 3 seconds, appearing as localized, transient waveform changes on the EEG. These microscopic events play a crucial role in distinguishing between N2 and N1 stages and in identifying the characteristics within N2. The second is brain network topological features, referring to the spatial coordination patterns formed by the functional connections between different brain regions over a longer time scale, such as a 30-second time window. For example, the synchronous slow wave phenomenon in the low-frequency range of the whole brain during N3 deep sleep. This macroscopic brain network organization is of great significance for determining sleep depth. Existing sleep staging schemes struggle to balance these two aspects simultaneously. Two types of information: If only microscopic waveform features are relied upon for frame-by-frame extraction and modeling, such as using convolutional neural networks to independently analyze the signals of each channel at the time level, the model will obtain fine waveform details but will inevitably lose the important discriminative value of macroscopic brain network features such as N3 stage whole-brain synchronous slow waves for staged sleep, resulting in a decline in the ability to identify deep sleep states; conversely, if only brain network topological features are relied upon, although the overall brain functional organization can be captured well, the key discriminative role of transient microscopic events such as N2 stage spindle waves and K complex waves for staged sleep will be lost, especially when it is necessary to accurately distinguish between the two easily confused sleep stages N1 and N2. Insufficient discriminative power of these microscopic waveforms will lead to a decline in classification accuracy; in addition, the diagnostic significance of the same spindle wave waveform for sleep staged sleep often varies under different brain network backgrounds and brain functional states, requiring dynamic adjustment of the interpretation and weighting of microscopic waveforms according to the current brain network state; To address the aforementioned challenges, this invention designs a spatiotemporal heterogeneity fusion mechanism for microscopic waveforms and brain network topology. The implementation logic of this mechanism is as follows: First, based on multi-channel EEG data after deep noise reduction, local microscopic waveform features are extracted using two convolutional layers with different receptive fields. The short receptive field convolutional layer focuses on capturing local high-frequency waveform changes occurring within seconds, such as spindle waves and K-complex waves. Its kernel size ensures that each output position aggregates only a few adjacent time-series information, thus preserving the waveform morphology of microscopic events. The long receptive field convolutional layer, on the other hand, targets a slightly broader time range, capturing the context of microscopic events. The outputs of the two convolutional layers are then concatenated to obtain a fusion mechanism that includes both precise and detailed microscopic waveform features. The first step involves designing a spatiotemporal alignment coupling operator to fuse microscopic waveform features with macroscopic spatial features derived from brain network topology. The core idea of ​​this coupling operator is that there is a mismatch in dimensionality and time scale between microscopic waveform features and brain network topology features—the former is generated point-by-point at the time-level, while the latter is a global representation within a 30-second time window. Linear transformations are used to process both microscopic features and brain network topology features, and then a spatiotemporal alignment coupling operator is generated using a sigmoid gating function to learn when and where to strengthen the expression of microscopic features and when and where to weaken them so that the brain network topology information can play a better role. Afterwards... The model first replicates and expands the topological features of the brain network in the temporal dimension to align them with the microscopic features. Then, it performs weighted modulation using a coupling operator, and finally fuses these features with the microscopic features themselves. During periods when the brain network exhibits strong synchronous characteristics, this modulation tends to strengthen the discriminative patterns associated with the brain network; during periods when the brain network exhibits local asynchrony, it retains more microscopic waveform details. The third step involves generating a topological manifold modulation factor based on the aligned and fused spatiotemporal features, and performing nonlinear weight correction on the multi-head self-attention output in the spatial dimension. Specifically, the model first reweights and interacts with the spatiotemporal features in the spatial dimension through a multi-head self-attention mechanism, enabling features from different brain regions to interact according to each other. The correlations are fused; then, based on the topological features of the brain network, a topological manifold modulation factor is generated through linear transformation and Sigmoid activation. This factor is essentially a general description of the current brain functional topological state, reflecting the overall characteristics of the brain network topology at that time. This factor is used as the weight multiplied by the output of spatial self-attention, so that spatial interaction patterns that conform to the current brain network topological features are enhanced, while interaction patterns that do not conform to the current brain network state are suppressed accordingly, so that the feature encoding of the spatial dimension can better adapt to the current brain functional state; in the fourth step, for the spatial enhancement features, a multi-head self-attention mechanism is used in the temporal dimension to capture long-term dependencies spanning multiple time windows.Sleep is not an isolated, window-by-window process, but a coherent evolutionary sequence. Strong temporal correlations exist between adjacent windows, and distant windows may exhibit similar patterns due to sleep cycle repetition. Multi-head self-attention in the temporal dimension enables the model to automatically learn which historical moments are most relevant to the current moment, thus obtaining features that reflect the temporal evolution of sleep. Finally, temporal evolution gating weights are designed, and convolution and sigmoid activation are used to process the temporal evolution features—higher gating weights are used during periods of relatively stable sleep and slow evolution, allowing temporal evolution features to dominate feature generation; lower gating weights are used during periods of rapid sleep transition and more sensitive responses, increasing the contribution of spatial enhancement features. Adaptive mixing of the two based on gating weights results in a spatiotemporal hybrid feature that flexibly balances the contributions of microscopic waveform details and brain network topology information under different sleep scenarios. This invention employs a multi-layered, progressive fusion mechanism, including spatiotemporal alignment coupling, topological manifold modulation, and temporal evolution gating. This mechanism ensures the full preservation of microscopic waveform details while organically integrating macroscopic topological information of the brain network. Furthermore, it dynamically adjusts the interpretation weights and patterns of microscopic waveforms based on the current state of the brain network. This enables the model to achieve deep synergy between microscopic and macroscopic features in both temporal and spatial dimensions, providing more comprehensive and multi-scale discriminative features for sleep period classification.

[0015] Furthermore, step A5 also includes: The spatiotemporal hybrid features at time t are processed sequentially by linear layer and layer normalization operations to obtain the classification embedding features at time t. The classification embedding features at time t are then divided by the temperature adjustment coefficient and processed by the Softmax function to obtain the initial probability distribution of the sleep stage at time t. The spatiotemporal hybrid features are processed by a bidirectional gated recurrent unit to obtain the bidirectional gated recurrent unit features at time t. The bidirectional gated recurrent unit features at time t are then processed by a linear layer and a Sigmoid function to obtain the temporal transition gating weights at time t. The initial probability distribution of the sleep stage at time t is multiplied by the result of "1 minus the temporal evolution gating weights" at time t. The result of the Hadamard product of the bidirectional gated recurrent unit features at time t processed by a linear layer and the temporal transition gating weights at time t is added element-wise and then processed by a Softmax function to obtain the corrected probability distribution of the sleep stage at time t. The logarithm of the corrected sleep stage probability distribution for the c-th sleep stage at time t is multiplied by the one-hot encoded label of the c-th sleep stage at time t. This multiplication is then accumulated over the total number of sleep stage categories and the total number of times, and multiplied by the reciprocal of the negative total number of times to obtain the classification cross-entropy loss. The square of the difference between the corrected sleep stage probability distribution at time t and the corrected sleep stage probability distribution at time t-1 is calculated and accumulated over the time to obtain the temporal evolution smoothing loss. The balance coefficient of the smoothing term is multiplied by the temporal evolution smoothing loss and then added to the classification cross-entropy loss to obtain the total loss function. The model parameters are corrected by minimizing the total loss function.

[0016] It should be further explained that if a frame-by-frame independent classification strategy is adopted, physiologically unreasonable jump phenomena may occur in the sleep stage discrimination; for example, a single frame discrimination of wakefulness W may suddenly appear in deep sleep N3, or a micro-awakening misjudgment may occur in N2→W→N2; sleep is essentially a continuous physiological evolution process, and the stages at adjacent moments should have a strong correlation. Therefore, the physiological continuity constraint of sleep should be introduced to achieve sequence-level global modeling. This invention designs a stage sequence decision-making mechanism under the constraint of sleep physiological continuity, and the implementation logic is as follows: The first step involves calculating the initial probability distribution of sleep stages at each moment through linear transformation and normalization based on spatiotemporal hybrid features, representing local discrimination at the feature level. The second step utilizes bidirectional gated recurrent units to capture long-range information across the entire sleep sequence's temporal dimension. The bidirectional structure simultaneously perceives historical and future information, enabling the network to gain a global understanding of the sleep sequence and effectively capture the physiological patterns of sleep stage transitions—for example, the transition from N2 to N3 is usually gradual rather than abrupt. The third step designs temporal transition gating weights to adaptively correct the initial probability distribution. These weights dynamically change between zero and one; when the weights are close to zero, they retain more feature information, while when they are close to one, they tend to favor sequence constraints. This adaptive mixing of the two allows the model to retain feature accuracy while correcting abnormal discriminations through physiological continuity constraints when necessary. The fourth step constructs a two-level loss function for constraint optimization. The classification cross-entropy loss measures the deviation between the prediction and the true label; the temporal evolution smoothing loss directly constrains the difference in probability distribution between adjacent moments, penalizing drastic stage changes and encouraging physiologically coherent transitions. The two loss terms are weighted and fused through a balance coefficient, allowing for a flexible trade-off between accuracy and smoothness. Compared with existing methods, this invention introduces physiological continuity soft constraints throughout the training and inference process by organically combining initial probability distribution, bidirectional gated loop, temporal migration gated hybrid and two-level loss. This maintains the accuracy of feature discrimination and effectively filters out stage jumps and fragmented misjudgments, outputting a smooth and reliable sequence that conforms to the physiological laws of sleep.

[0017] This invention also discloses a non-invasive sleep staging system based on spatiotemporal Transformer, comprising: EEG and impedance data acquisition and preprocessing module: Acquires multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocesses them respectively to obtain preprocessed multi-channel EEG data and preprocessed multi-channel impedance data; Deep noise reduction module: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, it sequentially extracts the basic temporal features of EEG and impedance state encoding, impedance modulation EEG features, and self-attention enhancement features, and outputs the deep noise-reduced multi-channel EEG data. Spatial feature extraction module: Based on the multi-channel EEG data after deep noise reduction, the module sequentially extracts multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix and features of each channel node, and outputs dynamic spatial connectivity graph features. Spatiotemporal fusion module: Based on the multi-channel EEG data after deep noise reduction and the dynamic spatial connectivity map features, it sequentially extracts local micro waveform features, aligned and fused spatiotemporal features, and spatial enhancement features, and outputs spatiotemporal fusion features; Sleep staging module: Based on the spatiotemporal hybrid characteristics, the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages are calculated sequentially, and a two-level loss function is constructed to provide feedback correction to the model parameters, thus completing the sleep staging.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This invention addresses the pain points of fluctuating signal quality, insufficient brain spatial topology mining, and difficulty in simultaneously considering microscopic and macroscopic features in home monitoring scenarios by providing a full-process sleep staging scheme. First, by introducing an impedance sensing and modulation mechanism, the dynamic changes in the contact quality of electrodes in each channel are sensed, and the contribution of EEG features is adjusted in real time accordingly. This effectively suppresses contact noise pollution caused by the subject turning over or sweating, laying a reliable signal foundation for subsequent analysis. On this basis, by mining the dynamic spatial connectivity of multiple frequency bands and implementing sparsity processing, the model can eliminate pseudo-connection interference and capture non-speculative signals. The study identifies brain region collaboration patterns corresponding to sleep depth. Furthermore, it utilizes a spatiotemporal fusion strategy to address the time-scale mismatch between microscopic fine waveforms and macroscopic brain network topology, achieving deep collaboration of features across multiple scales. This allows the model to flexibly adjust the interpretation weights of transient physiological events (such as spindle waves) based on the current global brain functional state. Finally, it incorporates sequence-level decision-making based on sleep physiological continuity constraints to correct unreasonable stage jumps, ensuring that the output staging sequence conforms to both microscopic feature representation and macroscopic physiological evolution patterns. This enhances the staging accuracy and result stability of non-invasive devices in home settings.

[0019] (2) To address the signal distortion caused by dynamic drift of contact impedance between dry electrodes and scalp in home settings and channel heterogeneity, this invention designs a dynamic adjustment mechanism for signal credibility based on impedance sensing. First, one-dimensional convolution is used to transform the original temporal impedance into an impedance state code with semantic discriminative ability, thereby robustly capturing the dynamic process of contact degradation. Then, the code is combined with the basic features of EEG to generate dynamic impedance gating weights, and the EEG features are softly modulated using the Hadamard product to obtain impedance-modulated EEG features, thereby realizing dynamic reweighting of the contribution of each channel at each time. Furthermore, second-order temporal difference is used to calculate the physiological transient response mask, and the long-term evolution law is extracted by combining the multi-head self-attention mechanism to construct self-attention enhancement features. Finally, the residual correction of the original signal is achieved through skip connections, and the multi-channel EEG data after deep noise reduction is output, which effectively alleviates the problem of dynamic fluctuation of signal credibility but equal weighting, and can automatically suppress noise pollution based on physical contact indicators, providing a reliable signal basis for subsequent feature extraction.

[0020] (3) To address the problems of insufficient spatial topology attribute mining and unclear functional connectivity patterns caused by pseudo-connection interference in multi-channel EEG processing, this invention designs a multi-band sleep-specific spatial connectivity topology mining mechanism. First, multi-band time-frequency energy features are extracted from the deep noise-reduced signal through short-time Fourier transform, and coherence coefficients are calculated on each frequency band to construct a functional connectivity matrix. Then, an adaptive weighting mechanism is introduced to dynamically assign weights to different frequency bands to adapt to the dominant frequency band changes in different sleep stages (such as N2 or N3). A sparse functional connectivity graph is generated through adaptive sparsification processing with a fixed percentile threshold, effectively filtering out weak connections that lack discriminative significance. Graph convolutional networks are used to aggregate neighborhood channel information and node attributes, outputting dynamic spatial connectivity graph features that contain local brain region and global network characteristics. This mechanism enables the model to adaptively mine brain functional connectivity structures unique to each sleep stage, fully utilize the synergistic relationship between brain regions to encode spatial dimension discriminative information, thereby providing more realistic and reliable brain functional topology features.

[0021] (4) To address the problem of the mismatch between micro-waveform details and macro-brain network topology in terms of observation scale and the difficulty in co-modeling, this invention designs a spatiotemporal heterogeneity fusion mechanism for micro-waveforms and brain network topology. First, micro-waveform features containing fine waveform details and local context are extracted through convolutional layers with different receptive fields. Then, a spatiotemporal alignment coupling operator is designed to align micro-features and macro-spatial features on the time scale, and a topological manifold modulation factor generated by the brain network topology is introduced to perform nonlinear weight correction on the spatial dimension attention output, so that the feature encoding can better adapt to the current brain functional state. Furthermore, long-term dependencies are captured through multi-head self-attention in the time dimension, and temporal evolution gating weights are used to adaptively mix spatial enhancement features and temporal evolution features in different scenarios of stable and rapid sleep state transitions, finally obtaining spatiotemporal hybrid features. While ensuring that the micro-waveform morphology is fully preserved, macro-topological information is organically integrated, realizing deep collaboration between micro- and macro-features in both spatiotemporal dimensions, and providing a comprehensive and multi-scale discrimination basis for sleep staging. Attached Figure Description

[0022] Figure 1 A flowchart illustrating a non-invasive sleep staging method based on spatiotemporal Transformer provided by this invention; Figure 2 The graph shows the changes in the two-stage loss function and learning rate provided by this invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention in any way. Any modifications or substitutions made based on the teachings of the present invention shall fall within the protection scope of the present invention.

[0024] Example 1: A non-invasive sleep staging method based on spatiotemporal Transformer, such as Figure 1 As shown, it includes the following steps: A1: Multichannel EEG signals and electrode-scalp contact impedance data were collected from the subjects during sleep and preprocessed to obtain preprocessed multichannel EEG data and preprocessed multichannel impedance data, including: A11: Multi-channel EEG signals during the subject's sleep process are collected using a home-use non-invasive dry electrode brain-computer interface device. The data type is multi-channel time-series voltage data, including frontal lobe lead EEG signals, central area lead EEG signals, and reference lead EEG signals. The frontal lobe lead EEG signals include Fp1 or Fp2 channel signals, the central area lead EEG signals include C3 or C4 channel signals, and the reference lead EEG signals include A1 and A2 earlobe reference signals. Simultaneously, electrode-scalp contact impedance data output by the device's built-in sensors are collected. The data type is time-series resistance data, including the real-time contact impedance values ​​of each lead. A12: The multi-channel EEG signal is detected and corrected using the AP0 jump zeroing method to obtain the baseline-corrected multi-channel EEG signal; the electrode-scalp contact impedance data is processed by removing abnormal impedance values ​​using the three-times-standard-deviation method, then filling in missing values ​​using linear interpolation, and finally using sliding window mean filtering to smooth high-frequency fluctuations to obtain the pre-processed electrode-scalp contact impedance data. A13: Based on the spatial topological coordinates of each lead electrode on the scalp, the baseline-corrected multi-channel EEG signal is spatially denoised by using a common average reference method combined with an adaptive artifact subspace reconstruction method to obtain a spatially denoised multi-channel EEG signal. A14: For the EEG data of each channel in the spatially denoised multi-channel EEG signal and the electrode-scalp contact impedance data after preliminary preprocessing, an alignment method based on the system global timestamp is used to achieve temporal synchronization of the multi-dimensional data, and the synchronized data is linearly normalized to unify the amplitude dimensions, so as to obtain the preprocessed multi-channel EEG data and the preprocessed multi-channel impedance data.

[0025] A2: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, the basic temporal features of EEG, impedance state encoding, impedance-modulated EEG features, and self-attention enhancement features are extracted sequentially. The deep noise-reduced multi-channel EEG data is then output, including: A21: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, initial feature space mapping is performed using a one-dimensional convolutional network to extract basic EEG temporal features and impedance state codes. The calculation method is as follows: ; ; in, This refers to the basic temporal characteristics of electroencephalography (EEG). It is a one-dimensional convolutional layer. For preprocessed multichannel EEG data, For impedance state encoding, For a modified linear unit with leakage, This is the preprocessed multi-channel impedance data; A22: Based on the basic temporal features of EEG and impedance state encoding, dynamic impedance gating weights are calculated through feature concatenation and convolution operations. The basic temporal features of EEG are then modulated to obtain impedance-modulated EEG features. The calculation method is as follows: ; ; in, For dynamic impedance-gated weights, For the Sigmoid function, This is a dimension splicing operation. This is an impedance-modulated EEG characteristic. For Hadamah accumulation; A23: Based on impedance-modulated EEG characteristics, a physiological transient response mask is calculated using second-order temporal difference. Subsequently, the evolutionary pattern of long-term sleep is extracted, and the physiological transient response mask is injected as a residual supplement term to obtain self-attention enhancement features. The calculation method is as follows: ; ; ; in, Let be the characteristic second-order temporal difference component at time t, where t is the time index. The impedance-modulated EEG characteristics at time t+1. The impedance-modulated EEG characteristics at time t. The impedance-modulated EEG characteristics at t-1 This is a mask for the physiological transient response at time t. For ReLU function, To take the absolute value, The self-attention enhancement features at time t, This is a multi-head self-attention mechanism. This is an impedance-modulated EEG characteristic. This is an element-wise addition operation; A24: Based on the self-attention enhancement features, sequence length is decoded and restored using a one-dimensional deconvolutional layer. Finally, skip connections are made with the preprocessed multi-channel EEG data to output deep noise-reduced multi-channel EEG data. The calculation method is as follows: ; ; in, To decode and reconstruct features, For linear layers, It is a one-dimensional deconvolution layer. This is a self-attention enhancement feature. This is multi-channel EEG data after deep noise reduction. For layer normalization operation, This is a vector of scaling factors.

[0026] Specifically, for scenarios where impedance and EEG signals exhibit time-varying coupling during long-term sleep data acquisition, this invention also provides an impedance-adaptive gating initial feature mapping calculation method to replace step A21. The calculation method is as follows: ; ; ; ; ; in, These are the initial convolutional features of the EEG. Encoding the original impedance state, It is an impedance-adaptive coupling gate factor. This method is based on the fundamental temporal features of EEG. By learning the coupling relationship between impedance and EEG, it enhances the expression of EEG features when the impedance quality is good, and simultaneously suppresses the interference of impedance state encoding when the impedance quality is poor, so that the two types of features can achieve closer coupling and complementarity in the initial mapping stage.

[0027] A3: Based on the multi-channel EEG data after deep noise reduction, multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix, and features of each channel node are extracted sequentially, and dynamic spatial connectivity graph features are output, including: A31: Based on the deep noise-reduced multi-channel EEG data, it is divided into time windows along the time dimension, and the signals within each window are decomposed into time-frequency components to obtain the multi-channel EEG data and multi-band time-frequency energy features after time window division. The calculation method is as follows: ; ; ; in, Multichannel EEG data after time window division. Divide the time window into operations. The size of the time window. For window step size, It is a time-frequency characteristic. This is a short-time Fourier transform operation. Multi-band time-frequency energy characteristics, For frequency band extraction operation; A32: Based on the multi-band time-frequency energy characteristics, the dynamic functional connectivity matrix between each channel is calculated using adaptive time-frequency coherence. The calculation method is as follows: ; ; ; in, For the first Channel within the frequency band With channel The coherence coefficients between channels, where i is the index of the first channel, j is the index of the second channel, and b is the frequency band index. For the first Channel within the frequency band and cross power spectrum, and The first Channel within the frequency band and The self-power spectrum, For the first Adaptive band weighting coefficients for the frequency band. For global average pooling, For the multi-band time-frequency energy characteristics in the first Frequency band components, For the first Apply the Softmax function to the frequency band. For dynamic function connection matrix, For the first Coherence matrix within the frequency band; A33: Based on the dynamic functional connectivity matrix, adaptive filtering is performed to achieve sparsity, resulting in a sparse functional connectivity graph matrix and the features of each channel node. The calculation method is as follows: ; ; ; ; in, To adapt the sparsity threshold, For fixed percentile threshold sparsity method, To preserve the percentile threshold of edge weights, For sparsity masking, For a sparse functional connectivity graph matrix, Features of each channel node These are the time-frequency energy characteristics of the 1st, ..., Bth frequency bands, respectively, where B is the total number of frequency bands. A34: Based on the features of each channel node and the sparse functional connectivity matrix, spatial topological encoding is performed through a graph convolutional network, and then the dynamic spatial connectivity graph features are fused and output. The calculation method is as follows: ; ; ; in, For graph convolutional network node features, For graph convolutional networks, For graph pooling features, For graph pooling operations, This is a feature of dynamic spatial connectivity graphs.

[0028] A4: Based on the multi-channel EEG data after deep noise reduction and the dynamic spatial connectivity map features, local microscopic waveform features, aligned and fused spatiotemporal features, and spatial enhancement features are extracted sequentially to output spatiotemporal hybrid features, including: A41: Based on the multi-channel EEG data after deep noise reduction, local micro-waveform features are extracted, and the aligned and fused spatiotemporal features are calculated using a spatiotemporal alignment coupling operator. The calculation method is as follows: ; ; ; in, For local microscopic waveform characteristics, For the first receptive field, a one-dimensional convolutional layer, For the second receptive field, a one-dimensional convolutional layer, For spatiotemporal alignment coupling operators, For copying and expanding operations along the time dimension, To align and fuse the spatiotemporal features; A42: Based on the aligned and fused spatiotemporal features, a topological manifold modulation factor is generated, and nonlinear weight correction is applied to the multi-head self-attention output in the spatial dimension to obtain spatial enhancement features. The calculation method is as follows: ; ; in, The topological manifold modulation factor. For spatial enhancement features, For spatial multi-head self-attention mechanism, To align and fuse the spatiotemporal features; A43: Based on the aforementioned spatial enhancement features, a multi-head self-attention mechanism in the temporal dimension is used to capture long-term dependencies across windows. Temporal evolution gating weights are designed and fused with temporal patterns to output spatiotemporal hybrid features. The calculation method is as follows: ; ; ; in, Features of temporal evolution, For time-multi-head self-attention mechanism, For temporal evolution gating weights, It is a gated convolutional layer. It is a spatiotemporal hybrid feature.

[0029] A5: Based on the spatiotemporal hybrid characteristics, the initial probability distribution and the corrected probability distribution of sleep stages are calculated sequentially. A two-level loss function is then constructed to provide feedback correction to the model parameters, completing the sleep staging process, including: A51: Based on the spatiotemporal mixing characteristics, the initial probability distribution of the sleep stage is calculated as follows: ; ; in, Let be the classification embedding features at time t. The spatiotemporal mixing characteristics at time t, Let be the initial probability distribution of the sleep stage at time t. For the Softmax function, This is the temperature regulation coefficient; A52: Based on the initial probability distribution and spatiotemporal mixing characteristics of sleep stages, a bidirectional gated recurrent unit is used to capture long-range temporal information, and the corrected sleep stage probability distribution is calculated. The calculation method is as follows: ; ; ; in, The bidirectional gated cyclic cell characteristic at time t. It is a bidirectional gated loop unit. Let be the temporal transition gating weight at time t. Let be the corrected probability distribution of the sleep stage at time t; A53: Construct a two-stage loss function consisting of classification cross-entropy loss and temporal evolution smoothing loss. By minimizing the deviation between the predicted results and the true labels, feedback correction of the model parameters is achieved. The calculation method is as follows: ; in, For classification cross-entropy loss, The total number of moments. The total number of sleep stage categories. Index for sleep categories, This is the one-hot encoded label for the c-th sleep stage at time t. For time-series evolution smoothing loss, Let be the corrected probability distribution of the sleep stage at time t-1. For the total loss function, This is the balance coefficient for the smoothing term.

[0030] The specific parameter settings for each neural network module involved in this invention are as follows: In step A2, a one-dimensional convolutional layer, Conv1D, is used for the initial mapping of basic EEG temporal features to impedance state encoding. The kernel size is set to 15, the stride is 1, and the number of output channels is 64, using the ReLU activation function. In the calculation of dynamic impedance gating weights, the kernel size is 11 and the number of output channels is 32 to extract the local dependencies after feature concatenation. The multi-head self-attention mechanism uses 8 attention heads, with a hidden layer dimension of 512 and 2048 neurons in the intermediate layers of the feedforward neural network. The kernel size of the one-dimensional deconvolutional layer is 15, the stride is 1, and the number of output channels is consistent with the number of input channels. The output dimension of the linear layer is 128. In step A3, the short-time Fourier transform uses the Hanning window function with a window size of 4000 sampling points (corresponding to 16 seconds at a sampling rate of 250Hz) and a window overlap rate of 75%. The frequency band extraction operation covers five sleep-related frequency bands: δ band (0.5-4Hz), θ band (4-8Hz), α band (8-13Hz), σ band (13-16Hz), and β band (16-30Hz). The graph convolutional network uses a two-layer GCN structure, with 128 output channels in both the first and second layers, and uses the ReLU activation function. The graph pooling operation uses global average pooling. In step A4, the kernel size of the short receptive field one-dimensional convolutional layer is 5, and the kernel size of the long receptive field one-dimensional convolutional layer is 25. Both have 32 output channels, and after concatenation, a 64-dimensional microscopic waveform feature is obtained. The output dimension of the linear layer in the spatiotemporal alignment coupling operator is 32. The spatial multi-head self-attention mechanism uses 8 attention heads and the hidden layer dimension is 512. The temporal multi-head self-attention mechanism uses 8 attention heads and the hidden layer dimension is 512. The kernel size of the gated convolutional layer is 7, and the number of output channels is 128. In step A5, the output dimension of the linear layer for classifying embedded features is set to the number of sleep stage categories (usually 5 categories, including W, N1, N2, N3, and REM), and the temperature adjustment coefficient τ' is set to 1.5; the number of hidden layer neurons in the bidirectional gated recurrent unit is 256; the output dimension of the linear layer in the temporal transfer gating weight calculation is the number of sleep stage categories. For model training optimization, the Adam optimizer was used, with an initial learning rate of 0.0005. The learning rate decay strategy employed a cosine annealing scheduler with a minimum learning rate of 0.00001. The batch size was set to 32, and the number of training epochs was 100. The L2 regularization coefficient was set to 0.0001 to prevent overfitting. The smoothing term balancing coefficient γ was set to 0.5 to achieve a balance between classification accuracy and temporal continuity. A gradient pruning strategy was adopted, with the upper limit of the gradient norm set to 2.0.

[0031] Example 2: This invention also discloses a non-invasive sleep staging system based on spatiotemporal Transformer, comprising: EEG and impedance data acquisition and preprocessing module: Acquires multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocesses them respectively to obtain preprocessed multi-channel EEG data and preprocessed multi-channel impedance data; Deep noise reduction module: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, it sequentially extracts the basic temporal features of EEG and impedance state encoding, impedance modulation EEG features, and self-attention enhancement features, and outputs the deep noise-reduced multi-channel EEG data. Spatial feature extraction module: Based on the multi-channel EEG data after deep noise reduction, the module sequentially extracts multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix and features of each channel node, and outputs dynamic spatial connectivity graph features. Spatiotemporal fusion module: Based on the multi-channel EEG data after deep noise reduction and the dynamic spatial connectivity map features, it sequentially extracts local micro waveform features, aligned and fused spatiotemporal features, and spatial enhancement features, and outputs spatiotemporal fusion features; Sleep staging module: Based on the spatiotemporal hybrid characteristics, the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages are calculated sequentially, and a two-level loss function is constructed to provide feedback correction to the model parameters, thus completing the sleep staging.

[0032] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0033] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0034] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A spatio-temporal Transformer-based non-invasive sleep staging method, characterized in that, Includes the following steps: A1: Collect multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocess them to obtain preprocessed multi-channel EEG data and preprocessed multi-channel impedance data; A2: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, extract the basic temporal features of EEG and impedance state encoding, impedance modulation EEG features and self-attention enhancement features in sequence, and output the multi-channel EEG data after deep noise reduction. A3: Based on the multi-channel EEG data after deep noise reduction, extract multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix and features of each channel node in sequence, and output dynamic spatial connectivity graph features. A4: Based on the multi-channel EEG data after deep noise reduction and the dynamic spatial connectivity map features, local micro waveform features, aligned and fused spatiotemporal features, and spatial enhancement features are extracted sequentially to output spatiotemporal hybrid features. A5: Based on the spatiotemporal hybrid characteristics, the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages are calculated sequentially, and a two-level loss function is constructed to provide feedback correction to the model parameters, thus completing the sleep staging.

2. The Spatio-Temporal Transformer-based non-invasive sleep staging method of claim 1, wherein, Step A1 includes: A11: Multi-channel EEG signals during the subject's sleep process are collected using a home-use non-invasive dry electrode brain-computer interface device. The data type is multi-channel time-series voltage data, including frontal lobe lead EEG signals, central area lead EEG signals, and reference lead EEG signals. The frontal lobe lead EEG signals include Fp1 or Fp2 channel signals, the central area lead EEG signals include C3 or C4 channel signals, and the reference lead EEG signals include A1 and A2 earlobe reference signals. Simultaneously, electrode-scalp contact impedance data output by the device's built-in sensors are collected. The data type is time-series resistance data, including the real-time contact impedance values ​​of each lead. A12: The multi-channel EEG signal is detected and corrected using the AP0 jump zeroing method to obtain the baseline-corrected multi-channel EEG signal; the electrode-scalp contact impedance data is processed by removing abnormal impedance values ​​using the three-times-standard-deviation method, then filling in missing values ​​using linear interpolation, and finally using sliding window mean filtering to smooth high-frequency fluctuations to obtain the pre-processed electrode-scalp contact impedance data. A13: Based on the spatial topological coordinates of each lead electrode on the scalp, the baseline-corrected multi-channel EEG signal is spatially denoised by using a common average reference method combined with an adaptive artifact subspace reconstruction method to obtain a spatially denoised multi-channel EEG signal. A14: For the EEG data of each channel in the spatially denoised multi-channel EEG signal and the electrode-scalp contact impedance data after preliminary preprocessing, an alignment method based on the system global timestamp is used to achieve temporal synchronization of the multi-dimensional data, and the synchronized data is linearly normalized to unify the amplitude dimensions, so as to obtain the preprocessed multi-channel EEG data and the preprocessed multi-channel impedance data.

3. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 1, characterized in that, Step A2 includes: A21: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, the initial feature space is mapped through a one-dimensional convolutional network to extract the basic temporal features of EEG and impedance state encoding. A22: Based on the basic temporal features of EEG and impedance state encoding, dynamic impedance gating weights are calculated through feature concatenation and convolution operations, and the basic temporal features of EEG are modulated to obtain impedance-modulated EEG features. A23: Based on impedance-modulated EEG characteristics, a physiological transient response mask is calculated using second-order temporal difference. Then, the evolution pattern of long-term sleep is extracted, and the physiological transient response mask is injected as a residual supplement term to obtain self-attention enhancement features. A24: Based on the self-attention enhancement features, the sequence length is decoded and restored through a one-dimensional deconvolution layer, and finally, it is connected to the preprocessed multi-channel EEG data to output the deep noise-reduced multi-channel EEG data.

4. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 3, characterized in that, Step A2 includes: The preprocessed multichannel EEG data were processed by a one-dimensional convolutional layer to obtain the basic temporal features of EEG. The preprocessed multichannel impedance data were processed by a one-dimensional convolutional layer and a modified linear unit with leakage to obtain the impedance state code. The basic EEG temporal features and impedance state codes are processed sequentially through dimension concatenation, a one-dimensional convolutional layer, and a Sigmoid function to obtain dynamic impedance gating weights. Then, the dynamic impedance gating weights are combined with the basic EEG temporal features to obtain impedance-modulated EEG features. For the impedance modulation EEG features at time t, the impedance modulation EEG features at time t+1 are subtracted by twice the impedance modulation EEG features at time t, and then the impedance modulation EEG features at time t-1 are added to obtain the characteristic second-order temporal difference component at time t. The absolute value of the characteristic second-order temporal difference component at time t is taken and processed by the ReLU function to obtain the physiological transient response mask at time t. Then, the impedance modulation EEG features are processed by the multi-head self-attention mechanism and added element-wise with the Hadamard product of the impedance modulation EEG features at time t and the physiological transient response mask at time t to obtain the self-attention enhancement features at time t. The self-attention enhancement features are processed sequentially through a one-dimensional deconvolution layer and a linear layer to obtain the decoded and reconstructed features. The product of the decoded and reconstructed features and the scaling coefficient vector is then processed by layer normalization and added element-wise to the preprocessed multi-channel EEG data to obtain the deep noise-reduced multi-channel EEG data.

5. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 3, characterized in that, Step A3 includes: A31: Based on the multi-channel EEG data after deep noise reduction, it is divided into time windows according to the time dimension, and the signals in each window are decomposed into time and frequency to obtain the multi-channel EEG data and multi-band time and frequency energy characteristics after the time window is divided. A32: Based on the multi-band time-frequency energy characteristics, the dynamic functional connection matrix between each channel is calculated through adaptive time-frequency coherence; A33: Based on the dynamic functional connection matrix, adaptive filtering is performed to achieve sparsity, resulting in a sparse functional connection graph matrix and the characteristics of each channel node; A34: Based on the characteristics of each channel node and the sparse functional connection matrix, spatial topology encoding is performed through a graph convolutional network, and then the dynamic spatial connection graph features are fused and output.

6. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 5, characterized in that, Step A3 includes: For the deep noise-reduced multi-channel EEG data, the time window is divided into multiple channels according to the time window size and window step size. Then, the time window divided multi-channel EEG data is processed by short-time Fourier transform and frequency band extraction to obtain multi-frequency time-frequency energy features. The components of the multi-band time-frequency energy characteristics in the b-th frequency band are processed sequentially by global average pooling, linear layer, and Softmax function to obtain the adaptive frequency band weighting coefficients of the b-th frequency band. The coherence coefficient between channel i and channel j in the b-th frequency band is obtained by dividing the cross power spectrum of channel i and channel j in the b-th frequency band by the square root of the product of the self power spectrum of channel i and the self power spectrum of channel j in the b-th frequency band. The product of the adaptive frequency band weighting coefficients of each frequency band and the coherence matrix of the corresponding frequency band is summed over all frequency bands and then processed by the Sigmoid function to obtain the dynamic functional connection matrix. An adaptive sparsification threshold is obtained by processing the dynamic functional connection matrix using a fixed percentile threshold sparsification method. A sparsification mask is generated at the position in the dynamic functional connection matrix that is greater than the adaptive sparsification threshold. The sparsification mask is then multiplied by the dynamic functional connection matrix to obtain a sparse functional connection graph matrix. At the same time, the multi-band time-frequency energy features of all frequency bands are processed by dimension splicing and linear layers to obtain the features of each channel node. The node features of each channel and the sparse functional connectivity graph matrix are processed by a graph convolutional network to obtain graph convolutional network node features. Then, the graph convolutional network node features and the sparse functional connectivity graph matrix are processed by graph pooling to obtain graph pooling features. Finally, the graph convolutional network node features and the graph pooling features are processed by dimension concatenation to obtain dynamic spatial connectivity graph features.

7. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 5, characterized in that, The A4 step includes: After deep noise reduction, the multi-channel EEG data are processed by a one-dimensional convolutional layer of the first receptive field and a one-dimensional convolutional layer of the second receptive field, respectively. Then, a dimensional splicing operation is performed to obtain local microscopic waveform features. At the same time, the dynamic spatial connectivity graph features are processed by copying and expanding along the time dimension and then processed by a linear layer. The result is added element-wise with the result of the local microscopic waveform features processed by the linear layer and then processed by the Sigmoid function to obtain the spatiotemporal alignment coupling operator. The dynamic spatial connectivity graph features are then processed by copying and expanding along the time dimension and then performed with the spatiotemporal alignment coupling operator to form a Hadamard product. Finally, the result is added element-wise with the local microscopic waveform features to obtain the aligned and fused spatiotemporal features. The dynamic spatial connectivity graph features are processed sequentially by a linear layer and a Sigmoid function to obtain the topological manifold modulation factor. Then, the aligned and fused spatiotemporal features are processed by a spatial multi-head self-attention mechanism and added element-wise with the aligned and fused spatiotemporal features. After layer normalization, the Hadamard product is performed with the topological manifold modulation factor to obtain the spatial enhancement features. The spatial augmentation features are processed by a temporal multi-head self-attention mechanism to obtain temporal evolution features. The temporal evolution features are then processed sequentially by a gated convolutional layer and a Sigmoid function to obtain temporal evolution gate weights. The temporal evolution features and the temporal evolution gate weights are then multiplied by a Hadamard product. This product is then added element-wise to the spatial augmentation features and the Hadamard product obtained by subtracting the temporal evolution gate weights from 1 to obtain the spatiotemporal hybrid features.

8. The non-invasive sleep staging method based on spatiotemporal Transformer according to claim 7, characterized in that, Step A5 includes: The spatiotemporal hybrid features at time t are processed sequentially by linear layer and layer normalization operations to obtain the classification embedding features at time t. The classification embedding features at time t are then divided by the temperature adjustment coefficient and processed by the Softmax function to obtain the initial probability distribution of the sleep stage at time t. The spatiotemporal hybrid features are processed by a bidirectional gated recurrent unit to obtain the bidirectional gated recurrent unit features at time t. The bidirectional gated recurrent unit features at time t are then processed by a linear layer and a Sigmoid function to obtain the temporal transition gating weights at time t. The initial probability distribution of the sleep stage at time t is multiplied by the result of "1 minus the temporal evolution gating weights" at time t. The result of the Hadamard product of the bidirectional gated recurrent unit features at time t processed by a linear layer and the temporal transition gating weights at time t is added element-wise and then processed by a Softmax function to obtain the corrected probability distribution of the sleep stage at time t. The logarithm of the corrected sleep stage probability distribution for the c-th sleep stage at time t is multiplied by the one-hot encoded label of the c-th sleep stage at time t. This multiplication is then accumulated over the total number of sleep stage categories and the total number of times, and multiplied by the reciprocal of the negative total number of times to obtain the classification cross-entropy loss. The square of the difference between the corrected sleep stage probability distribution at time t and the corrected sleep stage probability distribution at time t-1 is calculated and accumulated over the time to obtain the temporal evolution smoothing loss. The balance coefficient of the smoothing term is multiplied by the temporal evolution smoothing loss and then added to the classification cross-entropy loss to obtain the total loss function. The model parameters are corrected by minimizing the total loss function.

9. A non-invasive sleep staging system based on spatiotemporal Transformer, characterized in that, include: EEG and impedance data acquisition and preprocessing module: Acquires multi-channel EEG signals and electrode-scalp contact impedance data during the subject's sleep process and preprocesses them to obtain preprocessed multi-channel EEG data and preprocessed multi-channel impedance data; Deep noise reduction module: Based on the preprocessed multi-channel EEG data and preprocessed multi-channel impedance data, it sequentially extracts the basic temporal features of EEG and impedance state encoding, impedance modulation EEG features, and self-attention enhancement features, and outputs the deep noise-reduced multi-channel EEG data. Spatial feature extraction module: Based on the multi-channel EEG data after deep noise reduction, the module sequentially extracts multi-band time-frequency energy features, dynamic functional connectivity matrix, sparse functional connectivity graph matrix and features of each channel node, and outputs dynamic spatial connectivity graph features. Spatiotemporal fusion module: Based on the multi-channel EEG data after deep noise reduction and the features of dynamic spatial connectivity map, it sequentially extracts local micro waveform features, aligned and fused spatiotemporal features, and spatial enhancement features, and outputs spatiotemporal fusion features; Sleep staging module: Based on the spatiotemporal hybrid characteristics, the initial probability distribution of sleep stages and the corrected probability distribution of sleep stages are calculated sequentially, and a two-level loss function is constructed to provide feedback correction to the model parameters, thereby completing sleep staging; so as to realize the non-invasive sleep staging method based on spatiotemporal Transformer as described in any one of claims 1-8.