A multi-modal physiological signal sleep staging method and system based on multi-scale cross-view alignment

CN122767784APending Publication Date: 2026-09-18SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610861620.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]针对现有技术中提到的问题,本发明提出一种基于多尺度跨视图对齐的多模态生理信号睡眠分期方法及系统,解决了现有技术多视图睡眠分期方法中视图间特征难以精确对齐、单一尺度特征提取不充分以及模型计算量过大的问题

Benefits of technology

本发明采用多个堆叠的残差块,每个残差块内集成空间通道多头注意力模块,能够从时频图像视图中提取从浅层到深层的多尺度时频特征;同时,拓扑图特征提取采用阶数分别为2、4、9的三个切比雪夫图卷积分支并行提取多邻域邻居信息,通过上述两个多尺度提取机制解决了现有技术中单一尺度特征提取不充分的问题,使得融合后的特征既保留了局部精细信息,又融合了全局拓扑信息,显著提升了特征表达的丰富性和判别力;另外,本发明通过将时频特征与拓扑图特征进行跨视图对齐,解决了现有技术中两类视图特征因排列方式不同(时频特征按通道组织、图特征按节点组织)而导致的对齐困难问题,通过特征重组织实现精确位置对齐,使得后续融合真正实现了跨视图信息的互补,而非简单拼接,消融实验表明,单独添加跨视图对齐使准确率从82.1%提升至83.6%,验证了该操作的有效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122767784A_ABST
    Figure CN122767784A_ABST
Patent Text Reader

Abstract

The present application relates to physiological signal processing and sleep medicine field, specifically to a kind of multi-modal physiological signal sleep staging method and system based on multi-scale cross-view alignment, the present application first obtains multi-modal physiological signals such as electroencephalogram, electrooculogram, electromyogram and electrocardiogram, respectively constructs time-frequency image view and topological graph view;Multi-scale residual network and spatial channel multi-head attention module are used to extract the time-frequency characteristics of time-frequency image, while Chebyshev graph convolution of order 2, 4 and 9 is used to extract the topological graph characteristics of topological graph view in parallel;Based on the corresponding relationship of the two types of views in time dimension, cross-view accurate alignment and fusion are realized by feature reorganization;And introduce common feature learning module to enhance cross-view semantic consistency.The present application effectively solves the semantic mismatch problem in multi-view fusion, and significantly improves the recognition performance of difficult-to-classify sleep stages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of physiological signal processing and sleep medicine, specifically to a method and system for sleep staging of multimodal physiological signals based on multi-scale cross-view alignment. Background Technology

[0002] Sleep staging is a key technique for assessing sleep quality and diagnosing sleep disorders. Polysomnography (PSG) records multimodal signals, including electroencephalography (EEG), electrooculography (EOG), electromyography (EMG), and electrocardiography (ECG), providing rich information for sleep staging.

[0003] Existing deep learning-based automatic sleep staging methods mostly employ single-view inputs, such as the original time-series signal or time-frequency graph, making it difficult to simultaneously capture the signal's frequency domain distribution, spatial topology, and temporal dynamic features. In recent years, multi-view methods have achieved superior performance compared to single-view methods by fusing information from different views, such as time-frequency graphs and topology graphs. However, existing methods still suffer from the following technical limitations: First, the feature extraction scale is singular, making it difficult to simultaneously capture both fine local features and global topological information. For example, using only a single-order graph convolution cannot simultaneously capture the details of a node's first-order neighbors and the global structure of multi-hop neighbors. Second, there is a semantic mismatch between features from different views. The time-frequency graph is a two-dimensional representation generated by short-time Fourier transform, and its time dimension corresponds to the time window of the original signal; the topology graph is a graph structure independently constructed based on each sleep segment, and its node features also correspond to the same time window. Although the two have a natural temporal correspondence, direct splicing due to different feature arrangements leads to misalignment, limiting the effectiveness of cross-view interaction. Third, stacking multiple branches in pursuit of performance significantly increases the number of model parameters and computational complexity, making it difficult to meet the requirements for lightweight deployment. Summary of the Invention

[0004] To address the problems mentioned in the prior art, this invention proposes a multimodal physiological signal sleep staging method and system based on multi-scale cross-view alignment, which solves the problems of difficult accurate alignment of features between views, insufficient extraction of single-scale features, and excessive computational load in the existing multi-view sleep staging methods.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: This invention discloses a multimodal physiological signal sleep staging method based on multi-scale cross-view alignment, comprising the following steps: S1. Acquire multimodal physiological signals, wherein the multimodal physiological signals include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals; S2. Construct time-frequency image views and topological graph views for multimodal physiological signals respectively; wherein, the time-frequency image views are obtained by short-time Fourier transform, and the topological graph views are constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges; S3. Input the time-frequency image view into the multi-scale time-frequency feature extractor to extract multi-scale features and obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; S4. Input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches with orders of 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. S5. Align the time-frequency features and topology graph features across the view to obtain aligned time-frequency features and topology graph features; including reorganizing the time-frequency features according to time position based on the correspondence between the time-frequency features and topology graph features in the time dimension, so that the time-frequency features and topology graph features correspond one-to-one at each time position. S6. Fuse the aligned time-frequency features with the topology map features to obtain cross-view fusion features; input the cross-view fusion features into the classifier and output the corresponding sleep classification results.

[0006] As a further improvement of the present invention, the time-frequency image view is obtained by short-time Fourier transform, including: Short-time Fourier transform is performed on each signal using a Hanning window with a window length of 100 sampling points, a step size of 50 sampling points, and 100 FFT points. The time-frequency plot size is expanded to 100×100 by zero-filling, the amplitude spectrum is extracted and logarithmically compressed; Stack the time-frequency images of all channels along the channel dimension to form a three-dimensional time-frequency image tensor.

[0007] As a further improvement of the present invention, the topological graph view is constructed through adaptive graph learning, with signal channels as nodes and inter-channel feature correlations as edges, including: Each signal is treated as a node, and 64-dimensional initial features of each node are extracted through a one-dimensional convolutional network. An adaptive graph learning algorithm with a regularization parameter of 0.001 is used to generate an adjacency matrix between nodes, thus obtaining a topological graph view.

[0008] As a further improvement of the present invention, the multi-scale time-frequency feature extractor is composed of four residual blocks stacked sequentially. Each residual block uses a depthwise separable convolution instead of a standard convolution, with a kernel size of 3×3 and zero padding. The number of convolutional kernels increases by a factor of 2 from 64 in the first residual block to 128, 256, and 512 in each subsequent block. The feature maps output by each residual block are concatenated after global average pooling and then fed into a fully connected layer containing 128 hidden units to output time-frequency features.

[0009] As a further improvement of the present invention, the spatial channel multi-head attention module is composed of a spatial attention unit, a channel attention unit, and a multi-head attention unit connected in series.

[0010] As a further improvement of the present invention, in the multi-neighborhood Chebyshev diagram convolutional feature extractor, the number of convolution kernels in each Chebyshev diagram convolutional branch is 64. A Chebyshev graph convolution branch of order 2 is used to capture local topological features; A Chebyshev plot convolution branch of order 4 is used to capture mesoscale semantic structure features; A Chebyshev graph convolutional branch of order 9 is used to capture global topological features; The output features of the three branches are concatenated along the channel dimension and then sequentially modeled using a temporal convolutional layer and a temporal attention mechanism to obtain topological graph features.

[0011] As a further improvement of the present invention, the reorganization of time-frequency features according to time position in step S5 includes: Features are extracted from the set of two-dimensional feature maps output by the first residual block of the multi-scale time-frequency feature extractor; For each time position, extract all frequency dimension data of the corresponding time position from all channel feature maps and combine them to form a subtensor; By concatenating all subtensors in chronological order, the arrangement of time-frequency features is changed from channel-based organization to time-position-based organization.

[0012] As a further improvement to this invention, a common feature learning module is also included: 128-dimensional common semantic features were extracted from a multi-scale time-frequency feature extractor and a multi-neighborhood Chebyshev diagram convolutional feature extractor, respectively. The cosine similarity between two common semantic features is calculated as the similarity loss, and cross-view similarity is enhanced by maximizing the cosine similarity. Figure 1 To the point of being responsive.

[0013] This invention proposes a multimodal physiological signal sleep staging system based on multi-scale cross-view alignment, comprising: The acquisition module is used to acquire multimodal physiological signals, which include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals. The module is used to construct a time-frequency image view and a topological graph view for multimodal physiological signals, respectively. The time-frequency image view is obtained by short-time Fourier transform, and the topological graph view is constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges. The first extraction module is used to input the time-frequency image view into the multi-scale time-frequency feature extractor for multi-scale feature extraction to obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; The second extraction module is used to input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches of order 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. The cross-view alignment module is used to align time-frequency features with topological graph features across views, resulting in aligned time-frequency features and topological graph features. This includes reorganizing the time-frequency features according to their time positions based on the correspondence between the time-frequency features and topological graph features in the time dimension, so that the time-frequency features and topological graph features correspond one-to-one at each time position. The fusion classification module is used to fuse the aligned time-frequency features with the topological map features to obtain cross-view fusion features; the cross-view fusion features are input into the classifier to output the corresponding sleep classification results.

[0014] This invention proposes a multimodal physiological signal sleep staging device based on multi-scale cross-view alignment, comprising a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multimodal physiological signal sleep staging method based on multi-scale cross-view alignment as described above.

[0015] Compared with the prior art, the present invention achieves the following technical effects: This invention employs multiple stacked residual blocks, each integrating a spatial channel multi-head attention module, enabling the extraction of multi-scale time-frequency features from shallow to deep layers from the time-frequency image view. Simultaneously, topological graph feature extraction utilizes three Chebyshev graph convolution branches of orders 2, 4, and 9 to extract multi-neighbor information in parallel. These two multi-scale extraction mechanisms address the problem of insufficient single-scale feature extraction in existing technologies, ensuring that the fused features retain both local fine-grained information and global topological information, significantly enhancing the richness and discriminative power of feature representation. Furthermore, this invention solves the alignment difficulties caused by the different arrangement of the two types of view features (time-frequency features organized by channel, graph features organized by node) by performing cross-view alignment between time-frequency features and topological graph features. Precise position alignment is achieved through feature reorganization, enabling subsequent fusion to truly achieve complementary cross-view information rather than simple splicing. Ablation experiments show that adding cross-view alignment alone improves the accuracy from 82.1% to 83.6%, verifying the effectiveness of this operation.

[0016] The model of this invention has approximately 18M parameters, which is about 35% less than traditional three-view models. The training time per fold is approximately 1.8 hours, achieving lightweight deployment while maintaining high performance. Furthermore, on the ISRUC-S3 public dataset, this method achieves an accuracy of 85.3%, a macro-average F1 score of 0.843, and a Kappa coefficient of 0.810. Notably, it achieves an F1 score of 0.674 in the most difficult N1 classification stage, all of which outperform several existing comparative methods. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a schematic diagram of the model structure of the present invention; Figure 3 This is a schematic diagram of the spatial channel multi-head attention module of the present invention; Figure 4 This is a schematic diagram of cross-view alignment according to the present invention; Figure 5 This is a schematic diagram showing the comparison results between the method of the present invention and the benchmark model on the ISRUC-S3 dataset; Figure 6 This is a schematic diagram of the ablation experiment using the method of the present invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0019] See Figure 1 This embodiment proposes a multimodal physiological signal sleep staging method based on multi-scale cross-view alignment, including the following steps: S1. Acquire multimodal physiological signals, wherein the multimodal physiological signals include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals; S2. Construct time-frequency image views and topological graph views for multimodal physiological signals respectively; wherein, the time-frequency image views are obtained by short-time Fourier transform, and the topological graph views are constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges; S3. Input the time-frequency image view into the multi-scale time-frequency feature extractor to extract multi-scale features and obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; S4. Input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches with orders of 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. S5. Align the time-frequency features and topology graph features across the view to obtain aligned time-frequency features and topology graph features; including reorganizing the time-frequency features according to time position based on the correspondence between the time-frequency features and topology graph features in the time dimension, so that the time-frequency features and topology graph features correspond one-to-one at each time position. S6. Fuse the aligned time-frequency features with the topology map features to obtain cross-view fusion features; input the cross-view fusion features into the classifier and output the corresponding sleep classification results.

[0020] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments: This embodiment is applied to the publicly available ISRUC-S3 sleep dataset. Those skilled in the art can apply this method to other polysomnography datasets or clinical acquisition devices based on the description in this embodiment.

[0021] The ISRUC-S3 dataset in this embodiment contains overnight polysomnography recordings of 10 healthy subjects, including 9 men and 1 woman, aged 30 to 58 years, with a mean age of 44.2 years. The recording time for each subject was approximately 8 hours, with a raw sampling frequency of 200 Hz. Electrodes were placed according to the international 10-20 system during acquisition, with reference electrodes A1 and A2 (left and right earlobes). The data simultaneously recorded EEG, EMG, EMG, and ECG signals. To ensure the uniformity and validity of the input signals, this invention selected 10 channels from the raw recordings, specifically including 6 EEG channels (C3-A2, C4-A1, F3-A2, F4-A1, O1-A2, O2-A1), 2 EMG channels (LOC-A2 and ROC-A1), 1 EMG channel (X1, mandibular EMG), and 1 ECG channel (X2). All channels used either A1 or A2 as the reference electrode.

[0022] First, in this embodiment, each signal is processed sequentially as follows: For EEG and EOG signals, a 4th-order Butterworth filter is used for bandpass filtering, with cutoff frequencies set to 0.3 Hz and 35 Hz, respectively. For EMG signals, a 10–70 Hz bandpass filter is used; for ECG signals, no bandpass filtering is performed. All signals are filtered using a 50 Hz notch filter.

[0023] In this embodiment, independent component analysis is preferably used to remove eye movement artifacts. Then, all signals are downsampled from the original 200Hz to 100Hz. After downsampling, each 30-second sleep segment contains 3000 sampling points. The original labels are mapped from REM sleep (numbered 5) to 4 according to the AASM standard, so that the five categories (awake, N1, N2, N3, REM) are consecutively numbered from 0 to 4, and converted to one-hot encoding. After removing the first and last 30 invalid segments from each subject's record, a total of 8549 valid segments are obtained. To evaluate the model's generalization ability on subjects without prior observation, this invention uses leave-one-out cross-validation. All data from one subject are selected sequentially as the test set, and the data from the remaining nine subjects are used as the training set, for a total of 10 fold experiments.

[0024] like Figure 2 As shown, this embodiment reduces the number of model parameters and computational complexity by constructing a time-frequency image view and a topology graph view; Figure 2 The following is a definition of the English and Chinese terms: Multi-Scale Time-Frequency Feature Extractor; STFT (Short Time Fourier Transform); TF Image (Time-Frequency Image View); SCM-Attn (Spatial Channel Multi-Head Attention); C (Synthesis); FC (Fully Connected Layer).

[0025] Common features learning; Lc contrastive loss; Multi-Neighborhood Chebyshev Graph Feature Extractor; AGL adaptive graph learning; Chebyshev Graph Conv; CVA cross-view alignment; Sleep Staging classifier. Predicted represents the predicted result; Ground Truth represents the true label.

[0026] First, the time-frequency image view is constructed as follows: For each signal in each 30-second segment, a short-time Fourier transform is performed using a Hanning window with a length of 100 sampling points, a step size of 50, and 100 FFT points. The time-frequency image size is expanded to 100×100 by zero-padding. The amplitude spectrum is taken and logarithmically compressed. Then, the time-frequency images of the 10 channels are stacked along the channel dimension to form a tensor with a size of 100×100×10. This tensor preserves the energy distribution of each channel on the time-frequency plane and can be regarded as a multi-channel time-frequency image view.

[0027] The construction method of the topology graph view in this embodiment is as follows: Each signal is regarded as a node, with a total of 10 nodes. A 64-dimensional feature vector is extracted from each node using a one-dimensional convolutional network. An adaptive graph learning method is used to learn the adjacency matrix between nodes. Specifically, the similarity between the feature vectors of any two nodes is calculated, and regularization (with a regularization parameter set to 0.001) is applied to give nodes with similar features a larger connection weight, while nodes with dissimilar features have weights close to zero. The graph structure constructed in this way is not predefined, but is jointly optimized with the classification loss during training, thereby adaptively learning the functional connectivity relationships between channels to obtain the topology graph view; this topology graph view reflects the correlation structure between different physiological signals within the same sleep segment.

[0028] like Figure 2 As shown, this embodiment inputs a time-frequency image view into a multi-scale time-frequency feature extractor, which consists of four residual blocks stacked sequentially. Each residual block integrates a spatial channel multi-head attention module (SCM-Attn). To reduce computational cost while maintaining performance, each residual block uses depthwise separable convolution instead of standard convolution. The kernel size is uniformly 3×3, and zero padding is used to ensure that the spatial size of the output feature map is not reduced. The number of convolution kernels starts at 64 in the first residual block and increases by a factor of 2 in each subsequent block, namely 64, 128, 256, and 512.

[0029] Each residual block outputs a feature map with a different spatial size. Specifically, the first residual block outputs a feature map with a size of 64×50×50, the second residual block outputs a feature map with a size of 128×25×25, the third residual block outputs a feature map with a size of 256×13×13, and the fourth residual block outputs a feature map with a size of 512×7×7. It can be seen that as the network deepens, the time and frequency dimensions gradually decrease, while the number of feature channels continuously increases, thereby extracting increasingly abstract multi-scale spectral features.

[0030] Within each residual block, the SCM-Attn module performs spatial attention, channel attention, and multi-head attention sequentially in a cascaded manner, such as... Figure 3 As shown, the spatial attention module first averages and maximizes the input feature maps along the channel dimension, resulting in two two-dimensional attention maps. These two maps are then fused through a convolutional layer to generate spatial weights, enabling the network to focus on the importance of different time-frequency locations in the time-frequency map. The channel attention module performs global average pooling and global max pooling on the feature maps along the spatial dimension, resulting in two one-dimensional vectors. These vectors are then used to generate channel weights through two fully connected layers, thereby calibrating the responses of each frequency channel. Finally, the multi-head attention module uses the channel attention output as the query, the spatial attention output as the key, and the input data stream of the original residual block as the value. A weighted sum is calculated by multiple parallel attention heads, allowing different attention heads to focus on different subspace interactions while preserving the original information of the residual connections. The overall effect of the SCM-Attn module is to enhance the network's ability to focus on discriminative time-frequency regions while suppressing irrelevant noise.

[0031] The multi-scale feature maps output from the four residual blocks are each subjected to global average pooling to obtain a feature vector with lengths of 64, 128, 256, and 512, respectively. These vectors are concatenated to form a 960-dimensional vector, which is then fed into a fully connected layer with 128 hidden units to output 128-dimensional time-frequency features.

[0032] In this embodiment, the topological graph view is input into a multi-neighborhood Chebyshev graph convolutional feature extractor, such as... Figure 2 As shown, this extractor employs three parallel Chebyshev graph convolution branches with orders k set to 2, 4, and 9, respectively. The number of graph convolution kernels in each branch is set to 64.

[0033] When k=2, each node aggregates the features of its first- and second-order neighbors, mainly capturing fine-grained local topological information, such as short-term correlations between adjacent channels. When k=4, the node aggregates the features of its four-hop neighbors, covering a wider range of semantic structures. When k=9, the node aggregates the features of its nine-hop neighbors, effectively covering almost all nodes in the entire graph, thus integrating global topological information. The output feature dimension of each of the three branches is 64. Concatenating these three feature vectors along the channel dimension yields a 192-dimensional feature.

[0034] However, sleep staging relies not only on the graph structure of individual segments but also on the temporal evolution of sleep segments. To model this dynamic characteristic, the current sleep segment and the two segments before and after it (a total of 5 consecutive segments) are taken, and the aforementioned multi-branch graph convolution is performed independently on each segment. Each segment then generates a 192-dimensional feature vector, and the 5 segments are stacked into a 5×192 sequence. Next, a temporal convolutional layer is introduced, with a kernel size of 3, a stride of 1, and an output channel count of 192, to capture the temporal dependencies between adjacent segments. Then, a temporal attention module is passed, which learns an importance weight for each of the 5 time positions, and the weighted sum yields a 192-dimensional global graph feature vector. Temporal attention allows the model to automatically focus on the most discriminative neighboring segments. Finally, a fully connected layer reduces the dimensionality from 192 to 128, outputting the topological graph features.

[0035] After the above processing, 128-dimensional time-frequency features and 128-dimensional topological graph features were obtained. In this embodiment, in order to effectively fuse the two, as follows: Figure 4 As shown, this embodiment achieves this through cross-view alignment.

[0036] Cross-view alignment leverages the temporal correspondence between time-frequency features and topology features. The time axis of the time-frequency plot corresponds to the time window of the original signal, while the topology features, although constructed independently for each segment, also correspond to the same time window. Therefore, by rearranging the time-frequency features from channel-based organization to time-position-based organization, a one-to-one temporal correspondence with the topology features can be achieved.

[0037] The specific implementation involves obtaining a set of two-dimensional feature maps from the first residual block (output size 64×50×50) of the multi-scale time-frequency feature extractor. For each time position t (from 1 to 50), the data point in column t is extracted from the feature maps of all 10 channels, i.e., all frequency dimension information (64 values) corresponding to time position t in each channel's feature map. These column vectors are combined along the frequency and channel dimensions to form a sub-tensor with a size of 64×10. This process is repeated to obtain a total of 50 sub-tensors. These 50 sub-tensors are then concatenated in chronological order to form a new feature tensor with a size of 50×640. After this reorganization, the original feature arrangement changes from being organized by channel and then by time within each channel to being organized by time position and then by channel and frequency within each time position. The reorganized time-frequency features now have a clear time axis (the first dimension is 50 time positions), with each time position corresponding to a 640-dimensional feature vector. The topological graph features were originally 128-dimensional global vectors. To align them, they were copied 50 times to obtain a 50×128 sequence.

[0038] At this point, the two feature sequences have corresponding representations at each time position t, thus achieving accurate time-point alignment and solving the semantic misalignment problem in multi-view fusion.

[0039] After alignment, the reorganized time-frequency features (50×640) and the expanded topology map features (50×128) are concatenated along the feature dimension to obtain a 50×768 cross-view fusion feature.

[0040] This embodiment also introduces a common feature learning module. Specifically, two common semantic feature vectors are extracted from the penultimate layer of the multi-scale time-frequency feature extractor (i.e., the 128-dimensional features before the fully connected layer) and the penultimate layer of the multi-neighborhood Chebyshev diagram convolutional feature extractor (also 128-dimensional features). Then, the cosine similarity between these two vectors is calculated, which is their dot product divided by the product of their respective magnitudes, with a value between -1 and 1. This cosine similarity is used as a similarity loss; by maximizing this similarity, the features extracted by the two views are forced to converge in the semantic space. The similarity loss, as an auxiliary regularization term, is jointly optimized with the main classification loss using a weighted average, with a weight set to 0.1.

[0041] The classifier employs a fully connected layer, taking 768-dimensional cross-view fusion features as input and outputting 5-dimensional features. This output is then processed through a Softmax activation function to obtain the probability distributions for five sleep stages (awake, N1, N2, N3, and REM). The total loss function is a weighted sum of cross-entropy loss and similarity loss. Training utilizes the Adam optimizer with an initial learning rate of 5e-5, a batch size of 32, a maximum training epoch count of 120, and an early stopping mechanism to prevent overfitting. The entire model training process employs ten-fold cross-validation, with each fold using one subject as the test set and the remaining nine as the training set.

[0042] like Figure 5 As shown, on the ISRUC-S3 dataset, the model of this invention achieved the following average performance: accuracy of 85.28%, macro-average F1 score of 0.8429, and Kappa coefficient of 0.8102. The F1 scores for each category were: W stage 0.9103, N1 stage 0.6741, N2 stage 0.8438, N3 stage 0.9001, and REM stage 0.8861. Compared with existing methods, the accuracy of this invention (85.3%) is significantly better than MSTGCN (82.1%), MSF-SleepNet (83.6%), MVF-SleepNet (84.1%), and STAGN (84.4%). Especially in the most difficult N1 stage, the F1 score of this invention reached 0.674, higher than the comparative methods.

[0043] like Figure 6 As shown, three variant models were designed for ablation experiments: baseline (without SCM-Attn, without cross-view alignment), with only SCM-Attn, and with only cross-view alignment. Results showed that adding SCM-Attn alone improved accuracy from 82.1% to 82.9%, adding cross-view alignment alone improved it to 83.6%, and combining both resulted in an accuracy of 85.3%, validating the effectiveness of each module. In terms of computational efficiency, the model has approximately 18M parameters, a reduction of about 35% compared to the 28M in the first patent, with a training time of approximately 1.8 hours per fold.

[0044] Based on the same inventive concept, this invention also provides a multimodal physiological signal sleep staging system based on multi-scale cross-view alignment. Since the principle of solving the problem by this multimodal physiological signal sleep staging system based on multi-scale cross-view alignment is similar to that of the aforementioned multimodal physiological signal sleep staging method based on multi-scale cross-view alignment, the implementation of this multimodal physiological signal sleep staging system based on multi-scale cross-view alignment can refer to the implementation of the multimodal physiological signal sleep staging method based on multi-scale cross-view alignment, and the repeated parts will not be described again.

[0045] In specific implementation, the multimodal physiological signal sleep staging system based on multi-scale cross-view alignment provided in this embodiment of the invention specifically includes: The acquisition module is used to acquire multimodal physiological signals, which include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals. The module is used to construct a time-frequency image view and a topological graph view for multimodal physiological signals, respectively. The time-frequency image view is obtained by short-time Fourier transform, and the topological graph view is constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges. The first extraction module is used to input the time-frequency image view into the multi-scale time-frequency feature extractor for multi-scale feature extraction to obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; The second extraction module is used to input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches of order 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. The cross-view alignment module is used to align time-frequency features with topological graph features across views, resulting in aligned time-frequency features and topological graph features. This includes reorganizing the time-frequency features according to their time positions based on the correspondence between the time-frequency features and topological graph features in the time dimension, so that the time-frequency features and topological graph features correspond one-to-one at each time position. The fusion classification module fuses the aligned time-frequency features with the topological map features to obtain cross-view fusion features. These cross-view fusion features are then input into the classifier, which outputs the corresponding sleep classification result. Accordingly, embodiments of the present invention also provide a multimodal physiological signal sleep staging device based on multi-scale cross-view alignment, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multimodal physiological signal sleep staging method based on multi-scale cross-view alignment provided in embodiments of the present invention.

[0046] For a more detailed explanation of the above method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0047] Accordingly, embodiments of the present invention also provide a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described multimodal physiological signal sleep staging method based on multi-scale cross-view alignment provided in embodiments of the present invention.

[0048] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0049] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0050] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0051] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0052] The above provides a detailed description of the sleep staging method, system, device, and storage medium based on multi-scale cross-view alignment of multimodal physiological signals provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for sleep staging based on multi-scale cross-view alignment of multimodal physiological signals, characterized in that, Includes the following steps: S1. Acquire multimodal physiological signals, wherein the multimodal physiological signals include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals; S2. Construct time-frequency image views and topological graph views for multimodal physiological signals respectively; wherein, the time-frequency image views are obtained by short-time Fourier transform, and the topological graph views are constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges; S3. Input the time-frequency image view into the multi-scale time-frequency feature extractor to extract multi-scale features and obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; S4. Input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches with orders of 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. S5. Align the time-frequency features and topology graph features across the view to obtain aligned time-frequency features and topology graph features; including reorganizing the time-frequency features according to time position based on the correspondence between the time-frequency features and topology graph features in the time dimension, so that the time-frequency features and topology graph features correspond one-to-one at each time position. S6. Fuse the aligned time-frequency features with the topology map features to obtain cross-view fusion features; input the cross-view fusion features into the classifier and output the corresponding sleep classification results.

2. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, The time-frequency image view is obtained through short-time Fourier transform, including: Short-time Fourier transform is performed on each signal using a Hanning window with a window length of 100 sampling points, a step size of 50 sampling points, and 100 FFT points. The time-frequency plot size is expanded to 100×100 by zero-filling, the amplitude spectrum is extracted and logarithmically compressed; Stack the time-frequency images of all channels along the channel dimension to form a three-dimensional time-frequency image tensor.

3. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, The topological graph view, constructed through adaptive graph learning, uses signal channels as nodes and inter-channel feature correlations as edges, and includes: Each signal is treated as a node, and 64-dimensional initial features of each node are extracted through a one-dimensional convolutional network. An adaptive graph learning algorithm with a regularization parameter of 0.001 is used to generate an adjacency matrix between nodes, thus obtaining a topological graph view.

4. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, The multi-scale time-frequency feature extractor is composed of four residual blocks stacked sequentially. Each residual block uses a depthwise separable convolution instead of a standard convolution, with a kernel size of 3×3 and zero padding. The number of convolutional kernels increases by a factor of 2 from 64 in the first residual block to 128, 256, and 512 in each subsequent block. The feature maps output by each residual block are concatenated after global average pooling and then fed into a fully connected layer containing 128 hidden units to output time-frequency features.

5. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 4, characterized in that, The spatial channel multi-head attention module is composed of a spatial attention unit, a channel attention unit, and a multi-head attention unit connected in series.

6. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, In the multi-neighborhood Chebyshev graph convolutional feature extractor, the number of convolution kernels in each Chebyshev graph convolutional branch is 64; A Chebyshev graph convolution branch of order 2 is used to capture local topological features; A Chebyshev plot convolution branch of order 4 is used to capture mesoscale semantic structure features; A Chebyshev graph convolutional branch of order 9 is used to capture global topological features; The output features of the three branches are concatenated along the channel dimension and then sequentially modeled using a temporal convolutional layer and a temporal attention mechanism to obtain topological graph features.

7. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, The reorganization of time-frequency features according to time position in S5 includes: Features are extracted from the set of two-dimensional feature maps output by the first residual block of the multi-scale time-frequency feature extractor; For each time position, extract all frequency dimension data of the corresponding time position from all channel feature maps and combine them to form a subtensor; By concatenating all subtensors in chronological order, the arrangement of time-frequency features is changed from channel-based organization to time-position-based organization.

8. The sleep staging method based on multi-scale cross-view alignment of multimodal physiological signals according to claim 1, characterized in that, It also includes a common feature learning module: 128-dimensional common semantic features were extracted from a multi-scale time-frequency feature extractor and a multi-neighborhood Chebyshev diagram convolutional feature extractor, respectively. The cosine similarity between two common semantic features is calculated as the similarity loss, and cross-view consistency is enhanced by maximizing the cosine similarity.

9. A multimodal physiological signal sleep staging system based on multi-scale cross-view alignment, characterized in that, include: The acquisition module is used to acquire multimodal physiological signals, which include at least electroencephalogram (EEG) signals, electrooculogram (EOG) signals, electromyogram (EMG) signals, and electrocardiogram (ECG) signals. The module is used to construct a time-frequency image view and a topological graph view for multimodal physiological signals, respectively. The time-frequency image view is obtained by short-time Fourier transform, and the topological graph view is constructed by adaptive graph learning with signal channels as nodes and inter-channel feature correlations as edges. The first extraction module is used to input the time-frequency image view into the multi-scale time-frequency feature extractor for multi-scale feature extraction to obtain time-frequency features; the multi-scale time-frequency feature extractor is composed of multiple residual blocks stacked together, and each residual block integrates a spatial channel multi-head attention module; The second extraction module is used to input the topology graph view into the multi-neighbor Chebyshev graph convolutional feature extractor for multi-neighbor feature extraction. The multi-neighbor Chebyshev graph convolutional feature extractor uses three Chebyshev graph convolutional branches of order 2, 4 and 9 to extract multi-scale neighbor information in parallel, and combines temporal convolutional layers and temporal attention modules to perform temporal modeling to obtain topology graph features. The cross-view alignment module is used to align time-frequency features with topological graph features across views, resulting in aligned time-frequency features and topological graph features. This includes reorganizing the time-frequency features according to their time positions based on the correspondence between the time-frequency features and topological graph features in the time dimension, so that the time-frequency features and topological graph features correspond one-to-one at each time position. The fusion classification module is used to fuse the aligned time-frequency features with the topological map features to obtain cross-view fusion features; the cross-view fusion features are input into the classifier to output the corresponding sleep classification results.

10. A multimodal physiological signal sleep staging device based on multi-scale cross-view alignment, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multimodal physiological signal sleep staging method based on multi-scale cross-view alignment as described in any one of claims 1 to 8.