A sleep staging method and system based on multimodal physiological signal fusion

By performing multimodal signal fusion processing on the multi-somnography monitoring signal, multi-view features are extracted and fused, the problem of subjectivity of traditional sleep staging methods and insufficient single-modal signals is solved, and the accuracy of sleep staging is improved.

CN115349821BActive Publication Date: 2025-05-16SHENZHEN TECH UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210675112.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-16
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The traditional sleep staging method relies on the naked eye observation of medical staff, is time-consuming and labor-intensive and prone to errors due to subjectivity. Most automated sleep staging models are based on single-modal physiological signals and cannot capture the communication information between polysomnography signals, resulting in room for improvement in the accuracy of sleep staging.

Method used

The sleep staging method based on multimodal physiological signal fusion is adopted. By downsampling, short-time Fourier transform, lead feature extraction and machine self-learning graph generation of multi-view feature fusion is carried out to improve the accuracy of sleep staging by combining the time-frequency graph and the time-domain features of the self-learning graph.

Benefits of technology

It improves the accuracy of sleep staging, reduces human subjective errors, and can more effectively capture the communication information between polysomnography signals, providing more accurate sleep state recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115349821B_ABST
    Figure CN115349821B_ABST
Patent Text Reader

Abstract

The present invention discloses a sleep staging method and system based on multimodal physiological signal fusion, the method comprising: performing downsampling processing on the collected polysomnography signal to obtain a multi-lead signal; performing short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency diagram, and using one-dimensional convolution to extract the lead features in the multi-lead signal, and then generating a machine self-learning diagram according to the lead features; respectively combining the time-domain features of adjacent sleep stages on the time-frequency diagram and the machine self-learning diagram to obtain frequency-domain-time-domain fusion features and space-domain-time-domain fusion features; performing multi-view feature fusion on the frequency-domain-time-domain fusion features and space-domain-time-domain fusion features to obtain a sleep staging result. The present invention can improve the accuracy of sleep staging and can be widely used in the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a sleep staging method and system based on multimodal physiological signal fusion. Background Art

[0002] Sleep is a basic physiological function of human beings. Its manifestations are characterized by a series of changes in the activities of the brain, muscles, eyes, heart and breathing, and it plays an important role in restoring various human functions. Through sleep, the human body's fatigue of the day is eliminated and energy is restored, which can ensure that the brain is clear-minded and responsive in the awake state of the next day. Sleep is closely related to human mental illness. Lack of sleep can lead to mental depression. Depression patients are often accompanied by symptoms such as insomnia and parasomnias. Sleep state monitoring has been a research focus in the intersection of physiological signals and artificial intelligence in recent years. It is of great significance to human health. Some sleep-related diseases, such as insomnia, schizophrenia and autism, can be distinguished by analyzing sleep quality. Sleep staging is an important method to identify sleep states, which helps to better locate the onset period of different abnormalities.

[0003] Traditional sleep staging methods mostly rely on medical staff to observe polysomnography (PSG) with the naked eye, record patients at night, and use sensors attached to the body to measure a variety of physiological signals, such as electroencephalogram (EEG), electromyogram (EMG), electrocardiogram (ECG) and electrooculogram (EOG), which are used to monitor respiratory changes and other physiological changes in patients. However, this method is time-consuming and labor-intensive, and is prone to errors due to human subjectivity. Therefore, it is necessary to adopt an automated sleep staging model to solve this problem. Most automated sleep staging models are based on single-modal physiological signals and cannot capture the connectivity information between polysomnographic signals to provide users with multi-angle physiological information. Models based on multimodal physiological signals mostly extract a single feature type and cannot fuse the complementary information of multiple views. The two ultimately lead to the accuracy of sleep staging still has room for improvement. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a sleep staging method and system based on multimodal physiological signal fusion, which can improve the accuracy of sleep staging.

[0005] An aspect of an embodiment of the present invention provides a sleep staging method based on multimodal physiological signal fusion, comprising:

[0006] Downsampling the collected polysomnography signal to obtain a multi-lead signal;

[0007] Performing short-time Fourier transform processing on the multi-lead signal to obtain a time-frequency diagram, and extracting lead features from the multi-lead signal using one-dimensional convolution, and generating a machine self-learning diagram according to the lead features;

[0008] Combining the time-domain features of adjacent sleep stages on the time-frequency graph and the machine self-learning graph respectively to obtain frequency-domain-time-domain fusion features and space-domain-time-domain fusion features;

[0009] The frequency-time domain fusion features and the space-time domain fusion features are fused by multi-view features to obtain a sleep staging result.

[0010] Optionally, performing short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency diagram includes:

[0011] Configure sampling frequency and STFT window width;

[0012] According to the configured sampling frequency and STFT window width, a short-time Fourier transform is performed on the signal of each 30-second segment to obtain an STFT image corresponding to each lead signal;

[0013] After stacking the STFT images of all lead signals, the time-frequency diagram is obtained;

[0014] When the window width intercepted during the short-time Fourier transform process is less than 100, both ends of the time-frequency diagram are filled with zeros.

[0015] Optionally, generating a machine self-learning graph according to lead features includes:

[0016] Each lead is regarded as each node in the graph; wherein the feature of the node is composed of a feature matrix extracted from the original signal by two one-dimensional convolution kernels;

[0017] A self-learning graph is generated according to the extracted feature matrix, thereby forming a physiological structure relationship graph between different signal leads.

[0018] Optionally, combining the time domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency domain-time domain fusion features and space domain-time domain fusion features includes:

[0019] Performing feature extraction on the time-frequency graph using frequency domain convolution to extract frequency domain features;

[0020] Using spatial convolution to perform feature extraction on the self-learning graph to extract spatial features between human physiological structures;

[0021] Extract the time domain features of the time-frequency graph and the self-learning graph respectively;

[0022] Multi-view feature fusion is performed based on the extracted frequency domain features, spatial features and time domain features to obtain frequency domain-time domain fusion features and space domain-time domain fusion features.

[0023] Optionally, the extracting features of the time-frequency graph by using frequency domain convolution to extract frequency domain features includes:

[0024] Perform feature extraction on the time-frequency graph through a VGG-16 network to obtain frequency domain features;

[0025] The VGG-16 network includes 5 convolutional layers, 3 fully connected layers and 1 SoftMax output layer; maximum pooling is used between each layer, and the activation function Relu is used to activate all hidden layers, and the final features are input into the 128-dimensional fully connected layer.

[0026] Optionally, the using of spatial domain convolution to perform feature extraction on the self-learning graph to extract spatial features between human physiological structures includes:

[0027] A spatial attention mechanism is added to the self-learning graph, and Chebyshev graph convolution is used to capture the topological structure in the self-learning graph and extract spatial features;

[0028] Among them, the adjacency matrix and spatial attention matrix learned in the Chebyshev graph convolution process are used to dynamically adjust the update of the nodes.

[0029] Optionally, respectively extracting time domain features of the time-frequency graph and the self-learning graph includes:

[0030] Combine the currently extracted frequency feature and the two segments before and after the frequency feature into a layer of GRU network;

[0031] The output of each sequence is input into an attention network to learn the weight of each sequence, and then the features of the five sequences are fused into a 256-dimensional temporal feature;

[0032] After combining the fused features with the features of the current sleep stage itself, the features are input into a 128-dimensional fully connected layer;

[0033] Add a temporal attention mechanism to the construction graph of the current record and the two previous and next records, and use time for convolution to obtain an attention matrix;

[0034] The attention matrix is ​​normalized using the Softmax operation to obtain the time domain features.

[0035] Another aspect of the embodiment of the present invention further provides a sleep staging system based on multimodal physiological signal fusion, comprising:

[0036] The first module is used to downsample the collected polysomnography signal to obtain a multi-lead signal;

[0037] The second module is used to perform short-time Fourier transform processing on the multi-lead signal to obtain a time-frequency diagram, and to extract lead features in the multi-lead signal using one-dimensional convolution, and then generate a machine self-learning diagram according to the lead features;

[0038] The third module is used to combine the time domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency domain-time domain fusion features and space domain-time domain fusion features;

[0039] The fourth module is used to perform multi-view feature fusion on the frequency-time domain fusion features and the space-time domain fusion features to obtain a sleep staging result.

[0040] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0041] The memory is used to store programs;

[0042] The processor executes the program to implement the method described above.

[0043] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the method described above.

[0044] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0045] The embodiment of the present invention performs downsampling processing on the collected polysomnography signal to obtain a multi-lead signal; performs short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency diagram, and uses one-dimensional convolution to extract the lead features in the multi-lead signal, and then generates a machine self-learning diagram according to the lead features; respectively combines the time-frequency diagram and the machine self-learning diagram with the time domain features of adjacent sleep stages to obtain frequency domain-time domain fusion features and space domain-time domain fusion features; performs multi-view feature fusion on the frequency domain-time domain fusion features and space domain-time domain fusion features to obtain a sleep staging result. The present invention can improve the accuracy of sleep staging. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0047] Figure 1 An overall step flow chart provided for an embodiment of the present invention;

[0048] Figure 2 A flowchart of the MVF-SleepNet provided by an embodiment of the present invention;

[0049] Figure 3 A schematic diagram of generating a self-learning graph provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] In view of the problems existing in the prior art, an embodiment of the present invention provides a sleep staging method based on multimodal physiological signal fusion, such as Figure 1 As shown, the method of the present invention comprises the following steps:

[0052] Downsampling the collected polysomnography signal to obtain a multi-lead signal;

[0053] Performing short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency graph, and using one-dimensional convolution to extract lead features in the multi-lead signal, and then generating a machine self-learning graph according to the lead features;

[0054] Combining the time-domain features of adjacent sleep stages on the time-frequency graph and the machine self-learning graph respectively to obtain frequency-domain-time-domain fusion features and space-domain-time-domain fusion features;

[0055] The frequency-time domain fusion features and the space-time domain fusion features are fused by multi-view features to obtain a sleep staging result.

[0056] Optionally, performing short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency diagram includes:

[0057] Configure sampling frequency and STFT window width;

[0058] According to the configured sampling frequency and STFT window width, a short-time Fourier transform is performed on the signal of each 30-second segment to obtain an STFT image corresponding to each lead signal;

[0059] After stacking the STFT images of all lead signals, the time-frequency diagram is obtained;

[0060] When the window width intercepted during the short-time Fourier transform process is less than 100, both ends of the time-frequency diagram are filled with zeros.

[0061] Optionally, generating a machine self-learning graph according to lead features includes:

[0062] Each lead is regarded as each node in the graph; wherein the feature of the node is composed of a feature matrix extracted from the original signal by two one-dimensional convolution kernels;

[0063] A self-learning graph is generated according to the extracted feature matrix, thereby forming a physiological structure relationship graph between different signal leads.

[0064] Optionally, combining the time domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency domain-time domain fusion features and space domain-time domain fusion features includes:

[0065] Performing feature extraction on the time-frequency graph using frequency domain convolution to extract frequency domain features;

[0066] Using spatial convolution to perform feature extraction on the self-learning graph to extract spatial features between human physiological structures;

[0067] Extract the time domain features of the time-frequency graph and the self-learning graph respectively;

[0068] Multi-view feature fusion is performed based on the extracted frequency domain features, spatial features and time domain features to obtain frequency domain-time domain fusion features and space domain-time domain fusion features.

[0069] Optionally, the extracting features of the time-frequency graph by using frequency domain convolution to extract frequency domain features includes:

[0070] Perform feature extraction on the time-frequency graph through a VGG-16 network to obtain frequency domain features;

[0071] The VGG-16 network includes 5 convolutional layers, 3 fully connected layers and 1 SoftMax output layer; maximum pooling is used between each layer, and the activation function Relu is used to activate all hidden layers, and the final features are input into the 128-dimensional fully connected layer.

[0072] Optionally, the using of spatial domain convolution to perform feature extraction on the self-learning graph to extract spatial features between human physiological structures includes:

[0073] A spatial attention mechanism is added to the self-learning graph, and Chebyshev graph convolution is used to capture the topological structure in the self-learning graph and extract spatial features;

[0074] Among them, the adjacency matrix and spatial attention matrix learned in the Chebyshev graph convolution process are used to dynamically adjust the update of the nodes.

[0075] Optionally, respectively extracting time domain features of the time-frequency graph and the self-learning graph includes:

[0076] Combine the currently extracted frequency feature and the two segments before and after the frequency feature into a layer of GRU network;

[0077] The output of each sequence is input into an attention network to learn the weight of each sequence, and then the features of the five sequences are fused into a 256-dimensional temporal feature;

[0078] After combining the fused features with the features of the current sleep stage itself, the features are input into a 128-dimensional fully connected layer;

[0079] Add a temporal attention mechanism to the construction graph of the current record and the two previous and next records, and use time for convolution to obtain an attention matrix;

[0080] The attention matrix is ​​normalized using the Softmax operation to obtain the time domain features.

[0081] Another aspect of the embodiment of the present invention further provides a sleep staging system based on multimodal physiological signal fusion, comprising:

[0082] The first module is used to downsample the collected polysomnography signal to obtain a multi-lead signal;

[0083] The second module is used to perform short-time Fourier transform processing on the multi-lead signal to obtain a time-frequency diagram, and to extract lead features in the multi-lead signal using one-dimensional convolution, and then generate a machine self-learning diagram according to the lead features;

[0084] The third module is used to combine the time domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency domain-time domain fusion features and space domain-time domain fusion features;

[0085] The fourth module is used to perform multi-view feature fusion on the frequency-time domain fusion features and the space-time domain fusion features to obtain a sleep staging result.

[0086] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0087] The memory is used to store programs;

[0088] The processor executes the program to implement the method described above.

[0089] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the method described above.

[0090] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0091] The specific implementation principle of the present invention is described in detail below in conjunction with the accompanying drawings:

[0092] In order to use polysomnography for automated sleep staging, some researchers have proposed various sleep staging models using traditional machine learning methods. However, most of their models are based on signal processing or data mining methods, extracting frequency domain or time domain features from physiological signals, and then constructing models after feature selection. The feature extraction methods of these models rely on human prior knowledge, and the feature selection methods rely on the experience of researchers. In today's world where the amount of medical data is growing, the overall performance is inferior to that of deep learning models. In recent years, some researchers have extracted features from polysomnography signals and constructed automated sleep staging models based on two classic methods in deep learning methods: convolutional neural networks (CNN) and recurrent neural networks (RNN). However, their research has the following problems: some models only use one or two leads in the physiological signal, rather than multiple leads, which ignores the connectivity information between different lead channels; some models simply extract features from the time domain, frequency domain or spatial domain, without choosing to organically integrate the features of multiple domains; and most of the above studies only consider regular grid-like input data types, such as time series data and images, but do not consider non-Euclidean input data types, such as graph networks, which can well combine different physiological structures of the human body.

[0093] In recent years, complex networks and graph neural networks (GNNs) have become hot topics in machine learning research and have also been widely used in sleep staging methods based on physiological signals. Graph structures can be used to ideally represent human physiological structures, and the connectivity between different human physiological structures can be captured by graph data types.

[0094] Different from the methods in the prior art, the present invention proposes a model based on multimodal physiological signal fusion. (1) The input of the model includes EEG, EOG, ECG and EMG signals in polysomnography. (2) The present invention constructs a time-frequency graph (TFImage) and a self-learning graph (GL Graph) to represent the relationship between regularized lead signals and the relationship between irregular physiological structures, respectively, and uses a deep learning model to extract multi-view complementary features from the time domain, frequency domain, and spatial domain of physiological signals, and fuses the features to improve the accuracy of sleep staging. (3) The present invention conducted experiments on the public sleep signal dataset ISRUC-Sleep S3 to verify the superiority of this method.

[0095] In the present invention, the data set ISRUC-Sleep S3 data set used in the present invention is publicly provided by the University of Coimbra, Portugal. The sampling frequency of the polysomnographic signal is 200 Hz. The data set segments the polysomnographic records of the subjects throughout the night, each segment is 30 seconds long, and the data segments are marked as awake period, N1 period, N2 period, N3 period or REM period. The total number of polysomnographic records used in the experiment is 10, and the total number of segments is 8549.

[0096] The workflow of the multimodal fusion model MVF-SleepNet proposed in this invention is as follows: Figure 2 As shown. First, after the signal is downsampled to 100 Hz, the short-time Fourier transform (STFT) is performed on the multi-lead signal to convert the signal into a time-frequency diagram, the regularization information of the signal is integrated, and the lead features are extracted using a one-dimensional convolution (1-D CNN). A machine self-learning diagram is generated based on the lead features to integrate the physiological structure information of the human body. For the time-frequency diagram of the current marking segment, the present invention first uses the VGG-16 network to extract the frequency domain features, and then the extracted features are combined with the frequency domain features of the two adjacent segments on the left and right, and input into the GRU network to extract the time domain features, and then combined with the frequency domain features of the original marking segment to obtain the frequency domain-time domain fusion features; for the self-learning diagram of the current marking segment, the present invention first uses the Chebyshev graph convolution to extract the spatial domain features, and then the extracted features are combined with the frequency domain features of the two adjacent segments on the left and right, and the time domain features are extracted using the time domain convolution to obtain the spatial domain-time domain fusion features. Finally, the extracted multi-view features are fused for sleep staging.

[0097] The technical solution of the present invention is described in detail below:

[0098] (1) Multi-lead relationship representation:

[0099] In order to better represent the relationship between multi-channel data and extract rich information from Euclidean data and non-Euclidean data, the present invention proposes to construct a time-frequency diagram for representing regular data in multi-channel physiological signals and to construct a machine self-learning diagram for representing irregular data in multi-channel physiological signals.

[0100] a. Time-frequency diagram construction

[0101] In signal analysis, Fourier transform can be used to analyze the components of a signal or to synthesize a signal. It is mainly used to process stationary signals. Although the frequency components of a signal can usually be obtained through Fourier transform, the time of each component is unknown. For non-stationary signals, in order to know the time when each frequency occurs, a short-time Fourier transform is required. The essence of the short-time Fourier transform is windowing, which decomposes the time domain process into infinitely small processes of equal length, each of which is approximately stationary, and then Fourier transforms them. The STFT formula is defined as follows:

[0102]

[0103] Wherein w(t) is the window function, and x(t) is the signal to be converted. In this embodiment, the present invention performs STFT conversion on the sleep signal of each 30-second period, and the sampling frequency of the conversion is set to 1 Hz. The STFT window width is set to 100. If the width of some intercepted windows is less than 100, the two ends of the signal will be filled with 0. After the signal of each lead is converted, the corresponding STFT image will be obtained. Finally, the present invention stacks all the images together to obtain a time-frequency diagram with a resolution of 100×100×10.

[0104] b. Self-learning graph construction

[0105] The self-learning graph is generated as follows Figure 3 As shown: the graph structure can capture the relationship between different lead signals, however, previous studies mostly rely on manually pre-defining the graph structure, which relies on human prior knowledge, and there is currently no universal graph structure that can be applied to different data sets. Different data often require the construction of different graph structures, which is not conducive to the generalization ability of the model. The use of self-learning graphs, that is, allowing the graph to automatically learn to generate graph structures according to the characteristics of different lead node features according to the characteristics of the data set, can effectively improve the generalization of the model. In the present invention, the present invention regards each lead as each node in the graph, and the node features are composed of feature matrices extracted from the original signal by two one-dimensional convolution kernels (sizes of 32 and 64, respectively). Then, a self-learning graph is generated based on the feature matrix to form a physiological structure relationship diagram between different signal leads.

[0106] (2) Multi-view feature extraction

[0107] After constructing the time-frequency graph and the self-learning graph, the present invention needs to extract different features from them. Previous studies have shown that multi-channel signals have rich information in the frequency, spatial and time domains that can be used to identify different sleep states. In order to extract these multi-view features simultaneously, the present invention uses frequency domain convolution to extract frequency domain features from the time-frequency graph, and uses spatial domain convolution to extract spatial features between human physiological structures from the self-learning graph. In addition, in real life, sleep experts often use neighborhood information of adjacent sleep stages to help identify the current sleep stage. Inspired by this, the present invention extracts time domain features from the two graphs respectively. Multi-view feature extraction consists of 4 modules, namely: frequency domain feature extraction module, spatial domain feature extraction module, time domain feature extraction module and multi-view feature fusion module.

[0108] Frequency domain feature extraction module. After constructing the time-frequency graph, the present invention uses the VGG-16 network widely used in the field of computer vision to extract frequency domain features. VGG-16 consists of 5 convolutional layers, 3 fully connected layers and 1 SoftMax output layer. In addition, Maxpooling is used between each layer, and the activation function Relu is used to activate all hidden layers. The final features are input into the 128-dimensional fully connected layer.

[0109] Spatial feature extraction module. After constructing the self-learning graph, the present invention adds a spatial attention mechanism to the graph and uses Chebyshev graph convolution (Cheb GCN) to capture the topological structure in the graph and extract spatial features. The definition of spatial attention is as follows:

[0110]

[0111] In the formula is the input of the lth layer. p ,b p , Z1, Z2, Z3 are learnable parameters, and σ is the Sigmoid activation function. P represents the spatial attention matrix, which is dynamically calculated by the input of the current layer. The calculated attention matrix P is then normalized by the Softmax operation. In the model of the present invention, when performing graph convolution, the learned adjacency matrix and spatial attention matrix P can dynamically adjust the update of nodes. The Chebyshev graph convolution formula is defined as follows:

[0112]

[0113] L=DA

[0114]

[0115] Among them, g represents the convolution kernel, * GRepresents the graph convolution operation, θ is the vector of Chebyshev coefficients, and x is the input data. max is the largest eigenvalue of the Lass matrix, and I N is the identity matrix. T k is a recursive Chebyshev polynomial. After graph convolution, each channel lead integrates the features of other channel leads.

[0116] Time domain feature extraction module. Inspired by the fact that sleep experts often use the features of adjacent sleep periods to judge the current sleep stage, the present invention combines the currently extracted frequency features and the two segments before and after them into a layer of GRU network. The update expression of GRU is defined as follows:

[0117] h t =(1-z)⊙h t-1 +z⊙h′

[0118] Among them, z is the gate signal, h t is the information of the current signal, h t-1 The information sent by the previous unit. Then, the present invention inputs the output of each sequence into an attention network, learns the weight of each sequence, and then fuses the features of these five sequences into a 256-dimensional temporal feature. Finally, the present invention combines the fused dimensional features with the features of the current sleep stage itself, and then inputs them into a 128-dimensional fully connected layer. Similarly, in order to understand the characteristics of adjacent sleep phases, the present invention adds a temporal attention mechanism to the construction graph of the current record and its previous and next two segments, and uses temporal convolution. The definition of temporal attention is as follows:

[0119]

[0120] Among them, V q ,b q ,M1,M2,M3 are learnable parameters. Q u,v Represents the sleeping brain network G u and G v Finally, the attention matrix Q is normalized using the Softmax operation, and the input of the spatiotemporal stream is adjusted by the temporal attention to make it pay more attention to the rich temporal information. The definition of the temporal convolution is as follows:

[0121]

[0122] Wherein, ReLU is the activation function, Φ is the parameter of the convolution kernel, and * is the standard convolution operation. After the temporal convolution, the present invention flattens the feature matrix and inputs it into a 128-dimensional fully connected layer.

[0123] Multi-view feature fusion module. After the model obtains frequency-domain-time domain features from TF Image and spatial-domain-time domain features from GL Graph, it performs splicing operations on these feature matrices to perform multi-view feature fusion. The splicing operation is defined as follows:

[0124]

[0125] In the formula, X FT ,X S T represents the features extracted from the frequency domain-time domain and the spatial domain-time and space, respectively, and || represents the concatenation operation. Finally, the present invention inputs the fused features into a 128-dimensional fully connected layer, and after activation by the Softmax function, outputs a 5-category sleep staging result.

[0126] (3) Experimental verification

[0127] In order to prove its effectiveness, the proposed algorithm is compared with existing machine learning methods (SVM, RF, etc.), methods based on classical deep learning (MLP+LSTM, CNN and CNN+BiLSTM, etc.), and methods based on graph neural networks (STGCN and MSTGCN, etc.). The experimental results are the average of 10-fold cross validation. Table 1 is a comparison of the sleep stage recognition performance of each segment on the public dataset ISRUC-Sleep S3.

[0128] In the present invention, a new multimodal fusion algorithm MVF-SleepNet based on deep learning is developed for sleep staging. The MVF-SleepNet proposed by the present invention uses short-time Fourier transform and self-learning graph to represent the relationship between multi-lead signals respectively, extracts frequency domain features through VGG-16 network, extracts spatial domain features through ChebGCN, and extracts time domain features through GRU and time domain convolution. On ISRUC-Sleep S3, the sleep staging accuracy of the present invention is 83.3%.

[0129] Table 1 below shows the comparison results of the sleep staging performance of each model.

[0130] Table 1

[0131]

[0132]

[0133] Among them, the first method in Table 1 refers to the method adopted in the prior art literature “E.Alickovic and A.Subasi, “Ensemble SVMmethod for automatic sleep stage classification,” IEEE Instrum.Meas., vol.67, no.6, pp.1258–1265, Jun.2018”.

[0134] The second method refers to the method adopted in the prior art literature “P.Memar and F.Faradji, “A novel multi-class EEG-based sleep stage classification system,” IEEE Trans.Neural Syst.Rehabil.Eng., vol.26, no.1, pp.84–95, Jan.2018”.

[0135] The third method refers to the method adopted in the prior art literature “H. Dong, A. Supratak, W. Pan, C. Wu, PM Matthews, and Y. Guo, “Mixed neural network approach for temporal sleep stage classification,” IEEE Trans. Neural Syst. Rehabil. Eng., vol. 26, no. 2, pp. 324–333, Feb. 2018”.

[0136] The fourth method refers to the method adopted in the prior art document “A.Supratak, H.Dong, C.Wu, and Y.Guo, “DeepSleepNet: A modelfor automatic sleep stage scoring based on raw single-channel EEG,” IEEE Trans.Neural Syst.Rehabil.Eng., vol.25, no.11, pp.1998–2008, Nov.2017”.

[0137] The fifth method refers to the method adopted in the prior art literature “S. Chambon, MN Galtier, PJ Arnal, G. Wainrib, and A. Gramfort, “A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series,” IEEE Trans. Neural Syst. Rehabil. Eng., vol. 26, no. 4, pp. 758–769, Apr. 2018”.

[0138] The sixth method refers to the method adopted in the prior art literature “H. Phan, F. Andreotti, N. Cooray, OY Chén, and M. De Vos, “SeqSleepNet: End-to-end hierarchical recurrent neural network for sequence-to-sequence automatic sleep staging,” IEEE Trans. Neural Syst. Rehabil. Eng., vol. 27, no. 3, pp. 400–410, Mar. 2019”.

[0139] The seventh method refers to the method adopted in the prior art literature “Z. Jia et al., “GraphSleepNet: Adaptive spatial-temporal graph convolutional networks for sleep stage classification,” in Proc. 29th Int. Joint Conf. Artif. Intell. (IJCAI), Jul. 2020, pp. 1324–1330”.

[0140] The eighth method refers to the method adopted in the prior art literature “Z. Jia, Y. Lin, J. Wang, X. Ning, Y. He, R. Zhou, Y. Zhou, and H. L. Li-wei, “Multi-view spatial-temporal graph convolutional networks with domain generalization for sleep stage classification,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 29, pp. 1977–1986, 2021”.

[0141] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented by the present invention. The optional embodiment is expected, wherein the order of various operations is changed and the sub-operation of a part of the larger operation is wherein described is executed independently.

[0142] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed in the present invention, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0143] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0144] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0145] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0146] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0147] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0148] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0149] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A sleep staging method based on multimodal physiological signal fusion, characterized in that: include: Downsampling the collected polysomnography signal to obtain a multi-lead signal; Performing short-time Fourier transform processing on the multi-lead signal to convert it into a time-frequency graph, and using one-dimensional convolution to extract lead features in the multi-lead signal, and then generating a machine self-learning graph according to the lead features; Combining the time-domain features of adjacent sleep stages on the time-frequency graph and the machine self-learning graph respectively to obtain frequency-domain-time-domain fusion features and space-domain-time-domain fusion features; Perform multi-view feature fusion on the frequency-domain-time-domain fusion features and the space-domain-time-domain fusion features to obtain a sleep staging result; The generating of a machine self-learning graph according to lead features comprises: Each lead is regarded as each node in the graph; wherein the feature of the node is composed of a feature matrix extracted from the original signal by two one-dimensional convolution kernels; Generate a self-learning graph according to the extracted feature matrix, and then form a physiological structure relationship graph between different signal leads; The step of combining the time-domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency-domain-time-domain fusion features and space-domain-time-domain fusion features includes: Performing feature extraction on the time-frequency graph using frequency domain convolution to extract frequency domain features; Using spatial convolution to perform feature extraction on the self-learning graph to extract spatial features between human physiological structures; Extract the time domain features of the time-frequency graph and the self-learning graph respectively; Multi-view feature fusion is performed based on the extracted frequency domain features, spatial features, and time domain features to obtain frequency domain-time domain fusion features and space domain-time domain fusion features; The method of using spatial domain convolution to extract features from the self-learning graph to extract spatial features between human physiological structures includes: A spatial attention mechanism is added to the self-learning graph, and Chebyshev graph convolution is used to capture the topological structure in the self-learning graph and extract spatial features; Among them, the adjacency matrix and spatial attention matrix learned in the Chebyshev graph convolution process are used to dynamically adjust the update of the nodes.

2. The sleep staging method based on multimodal physiological signal fusion according to claim 1, characterized in that: The step of performing short-time Fourier transform processing on the multi-lead signal to convert the multi-lead signal into a time-frequency diagram includes: Configure sampling frequency and STFT window width; According to the configured sampling frequency and STFT window width, a short-time Fourier transform is performed on the signal of each 30-second segment to obtain an STFT image corresponding to each lead signal; After stacking the STFT images of all lead signals, the time-frequency diagram is obtained; When the window width intercepted during the short-time Fourier transform process is less than 100, both ends of the time-frequency diagram are filled with zeros.

3. The sleep staging method based on multimodal physiological signal fusion according to claim 1, characterized in that: The step of extracting features from the time-frequency graph using frequency domain convolution to extract frequency domain features includes: Perform feature extraction on the time-frequency graph through a VGG-16 network to obtain frequency domain features; The VGG-16 network includes 5 convolutional layers, 3 fully connected layers and 1 SoftMax output layer; maximum pooling is used between each layer, and the activation function Relu is used to activate all hidden layers, and the final features are input into the 128-dimensional fully connected layer.

4. The sleep staging method based on multimodal physiological signal fusion according to claim 1, characterized in that: The extracting of time domain features of the time-frequency graph and the self-learning graph respectively includes: Combine the currently extracted frequency feature and the two segments before and after the frequency feature into a layer of GRU network; The output of each sequence is input into an attention network to learn the weight of each sequence, and then the features of the five sequences are fused into a 256-dimensional temporal feature; After combining the fused features with the features of the current sleep stage itself, the features are input into a 128-dimensional fully connected layer; Add a temporal attention mechanism to the construction graph of the current record and the two previous and next records, and use time for convolution to obtain an attention matrix; The attention matrix is ​​normalized using the Softmax operation to obtain the time domain features.

5. A system using the sleep staging method based on multimodal physiological signal fusion as described in any one of claims 1 to 4, characterized in that: include: The first module is used to downsample the collected polysomnography signal to obtain a multi-lead signal; The second module is used to perform short-time Fourier transform processing on the multi-lead signal to obtain a time-frequency diagram, and to extract lead features in the multi-lead signal using one-dimensional convolution, and then generate a machine self-learning diagram according to the lead features; The third module is used to combine the time domain features of adjacent sleep stages of the time-frequency graph and the machine self-learning graph to obtain frequency domain-time domain fusion features and space domain-time domain fusion features; The fourth module is used to perform multi-view feature fusion on the frequency-time domain fusion features and the space-time domain fusion features to obtain a sleep staging result.

6. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-lead physiological signal analysis method and device

    CN110008790A

  • Sleep awakening analysis method based on deep learning

    CN110811558A