Sleep staging method based on spatial-temporal feature coding and multi-source fusion

By adopting spatiotemporal feature coding and multi-source fusion neural network model in sleep staging, the problem of difficult to utilize the synergistic effect of multi-channel spatial information and multi-signal sources in multi-somnography data is solved, and higher sleep staging accuracy and individualized help are achieved.

CN120217200AActive Publication Date: 2025-06-27PEKING UNIV

Patent Information

Application Number
CN202510287610.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the synergy between multi-channel spatial information and multi-signal sources in polysomnography data, resulting in the impact of the accuracy of sleep staging results.

Method used

An automatic sleep staging neural network model based on spatiotemporal feature encoding and multi-source fusion is adopted, including a graph space encoder, a multi-signal source fusion module and a timing Transformer encoder, to process the spatial and temporal features in multisomnography data.

Benefits of technology

By effectively integrating information from multiple channels and multiple signal sources, the accuracy of sleep staging is improved, providing practical help to individuals, providing convenient tools for doctors, and providing heuristic auxiliary guidance for precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217200A_ABST
    Figure CN120217200A_ABST
Patent Text Reader

Abstract

The invention provides a sleep staging method based on spatial-temporal feature coding and multi-source fusion, and belongs to the technical field of big data analysis. According to the method, collected multi-channel multi-source physiological data including electroencephalogram, myoelectricity and electro-oculogram signals are utilized to construct a sleep staging neural network model fusing the multi-channel multi-source physiological data; establishing a graph space encoder for modeling multi-channel spatial features, and encoding interaction and position information among different channels; establishing a multi-signal source fusion module for fusing multi-source signals, and identifying different contributions of multi-signal sources to stages based on an attention mechanism; a time sequence Transform encoder used for modeling time features is established, and local and global time features are encoded at the same time. According to the method, automatic staging of sleep can be realized, the staging accuracy is improved, personal help can be provided for individuals, a convenient tool is provided for doctors, and heuristic auxiliary guidance is provided for precision medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a sleep analysis technology based on the electrophysiological data of a polysomnograph (PSG), and particularly relates to an automatic sleep staging method based on spatio-temporal feature encoding and multi-source fusion, belonging to the technical field of big data analysis. Background Art

[0002] Advances in communication and computing technologies have made their applications in collecting and analyzing physiological data widespread. Due to their highly flexible, non-invasive, and portable characteristics, physiological data analysis has become an effective tool for evaluating physiological states. Physiological data contains a large amount of potentially valuable information that is extremely important for life and health monitoring. Among them, sleep staging is a key application because sleep has a great impact on a person's health, performance, and well-being. Through sleep staging technology, various psychological and physical health problems can be diagnosed and prevented. Sleep is divided into several stages, including three non-rapid eye movement (NREM) sleep stages (N1, N2, N3), rapid eye movement (REM) sleep stage (R), and wakefulness state (W). The correct identification of each stage is crucial for supporting specific physiological and psychological functions and overall sleep quality. However, manually performing sleep staging is a highly demanding, time-consuming, and labor-intensive task for doctors, and the results may be affected by doctors' subjective judgments. Therefore, developing automatic sleep staging technology will greatly assist doctors and patients and has high practical value.

[0003] Physiological data for sleep staging is usually collected by a polysomnograph. These physiological data are highly complex. Specifically, they have multiple electrode channels and each channel strictly corresponds to a specific part of the human body; they come from multiple physiological signal sources, including electroencephalogram (EEG), electromyogram (EMG), and electrooculogram (EOG); in addition, these signals contain some characteristic waveforms, and the types, quantities, and durations of the waveforms will affect the results of sleep staging. Doctors use one frame as the basic time unit for sleep staging, usually one frame being 30 seconds. In recent years, academic and industrial people have developed many methods for identifying and analyzing physiological data. Existing technologies focus more on the analysis of single channels or very few channels and lack the utilization of spatial information between multiple channels. Moreover, existing technologies mainly use a single physiological signal source for analysis and ignore the synergistic effect of multiple signal sources. Among the few methods that consider the influence of multi-channel spatial information or multiple signal sources, existing technologies usually simply fuse this information by addition or splicing methods, resulting in the neglect of the time characteristics of the unique waveforms of polysomnography data and making it difficult to effectively utilize multi-source information in terms of time and space.

[0004] Currently, the identification and analysis of polysomnography data still face the following challenges: First, polysomnography data has multiple channels with complex spatial characteristics and interactions between channels, which are difficult to capture. Second, polysomnography data comes from multiple physiological signal sources, and each signal source has a unique impact on the staging results. For example, doctors mainly rely on the characteristics of EEG to distinguish N1, N2, and N3 stages, and mainly rely on the characteristics of EOG to judge the W stage. How to effectively fuse information from multiple sources is a difficult point. Third, doctors judge the staging results based on the characteristic waveforms in a frame of polysomnography data, including the types, quantities, positions, durations, etc. of these waveforms. How to simultaneously process the global and local time characteristics of characteristic waveforms requires in-depth research. Summary of the Invention

[0005] In view of the deficiencies of the above-mentioned existing technologies, the present invention provides a sleep staging method based on spatio-temporal feature encoding and multi-source fusion. In view of the characteristics of multi-channel and multi-source data, it fully considers the spatial and temporal characteristics of polysomnography data, and based on the waveform characteristics, electrode spatial sites and other medical prior knowledge referred to by doctors when manually staging sleep, an automatic sleep staging neural network model based on spatio-temporal feature encoding and multi-source fusion is established to realize the identification and analysis of new polysomnography data.

[0006] The technical solution provided by the present invention is as follows:

[0007] A sleep staging method based on spatio-temporal feature encoding and multi-source fusion, which uses the collected polysomnography data of an individual to establish an automatic sleep staging neural network model based on spatio-temporal feature encoding and multi-source fusion to realize the staging and identification of the subsequent sleep stages of the individual. The method includes the following steps:

[0008] The sleep staging neural network model includes three modules: 1) a graph space encoder, 2) a multi-signal source fusion module, and 3) a temporal Transformer encoder. Among them, the graph space encoder is used to encode the spatial interaction information in polysomnography data; the multi-signal source fusion module fuses the information of EEG, EMG, and EOG signals in a learnable manner; the temporal Transformer encoder captures local and global time characteristics. For each frame of polysomnography data, sleep staging is performed. Let C represent the number of channels of polysomnography data, L e represent the number of sampling points in a frame, X e represent an initially input frame, that is To better utilize the characteristic waveforms, first divide a frame of polysomnography data into multiple segments. Let L p represent the number of sampling points in a segment, and X p represent a segment, that is The input of the graph space encoder and the multi-signal source fusion module is a single segment, while the temporal Transformer encoder takes all the segments that have undergone graph space encoding and multi-signal source fusion as input.

[0009] 1) Construct a graph space encoder for spatial features; including steps 1a) - 1b):

[0010] 1a) Construct a multi-layer perceptron encoder for electromyogram and electrooculogram signals:

[0011] The channels of electromyogram and electrooculogram signals are only one or two, and their spatial information is much less than that of electroencephalogram signals with relatively more channels. For electromyogram signals and electrooculogram signals, a multi-layer perceptron (MLP) composed of linear layers and residual connections is used to encode them initially. The multi-layer perceptron encoders for electromyogram and electrooculogram signals are respectively expressed as Equation 1 and Equation 2:

[0012] H′ EMG = MLP(H EMG ) (Equation 1)

[0013] H′ EOG = MLP(H EOG ) (Equation 2)

[0014] Where, are the encoded electromyogram and electrooculogram signals respectively, are the electromyogram channel and electrooculogram channel signals selected from a single segment X p respectively. C EMa and C EOG are the channel numbers of electromyogram signals and electrooculogram signals in the segment respectively.

[0015] 1b) Construct a sparse graph neural network for electroencephalogram signals:

[0016] The electroencephalogram signal has relatively more channels and a strict position is defined on the head, and its multi-channel spatial connection information is significant. Use one-dimensional convolution to convert the electroencephalogram signal into the embedding of graph nodes, and the conversion process is expressed as Equation 3:

[0017]

[0018] Where, is the initial embedding of graph nodes converted from electroencephalogram signals, is the electroencephalogram channel signal selected from a single segment X p respectively, and C EEG is the channel number of electroencephalogram signals in the segment.

[0019] Based on the graph nodes transformed from EEG signals, a sparse connection graph is constructed, and the multi-head attention mechanism is used to explicitly learn the spatial connections between channels. Compared with the fully connected graph, the sparse connection graph adopted by this method reduces the number of edges and thus reduces the computational complexity. The steps to construct the sparse connection graph are as follows:

[0020] Random regular graph: Randomly permute the graph nodes transformed from EEG signals, and connect the points in the rearranged order so that each point is connected to a constant number of other points to construct a random regular graph.

[0021] Virtual connection graph: Establish a constant number of virtual nodes. The virtual nodes have learnable node embeddings, and the virtual nodes are connected to each real graph node to construct a virtual connection graph.

[0022] The random regular graph and the virtual connection graph jointly construct a sparse connection on the graph nodes. The multi-head attention mechanism is used to calculate the weights of the edges connecting the nodes. Taking a single edge as an example, the weight calculation is expressed as Equation 4:

[0023] e i:j =(W K h i ) T ·W Q h j (Equation 4)

[0024] where i, j = 1, 2, 3, …, N, and N is the total number of segments into which the polysomnography data is divided. W K , represents the learnable parameter, represents the weight of the edge connecting node i and node j. is the embedding of node i in

[0025] The node embeddings of the graph are updated using the message propagation mechanism as shown in Equation 5:

[0026]

[0027] where σ is the activation function. b is the bias, is the set of nodes connected to node i. When performing message propagation, the attention weight e j:i of each edge connected to the node calculated by Equation 4 is applied. is the updated embedding of node i, and all the node embeddings transformed from EEG signals jointly form the encoded EEG signals

[0028] 2) Construct a multi-signal source fusion module:

[0029] PSG data includes multiple signal sources such as EEG, EMG, and EOG. This method uses the multi-head attention mechanism to fuse multi-source signals in the channel dimension C. First, multiple signal sources encoded by the graph space encoder are concatenated, as shown in Equation 6:

[0030]

[0031] Among them, H′ EEG , H' EOG , H′ EMG are the EEG, EOG, and EMG signals that have passed through the graph space encoder respectively. After concatenating them, a transpose is performed to obtain After that, multi-head attention is used to focus on the contributions of different signal sources. The attention formula is as shown in Equation 7:

[0032]

[0033] Among them, Q, K, and V correspond to the query, key, and value in the Transformer architecture respectively. The softmax function maps the attention to the interval [0, 1]. This method uses self-attention, and the values of Q, K, and V in Equation 7 are the same and equal to H′. After that, a linear layer is used to fuse the multi-signal sources in the channel dimension C. The output of the multi-signal source fusion module is shown in Equation 8:

[0034] x′ p = Linear(Attn(H′)) (Equation 8)

[0035] Among them, is a single segment that has undergone graph space encoding and multi-source signal fusion.

[0036] 3) Construct a temporal Transformer encoder for temporal features

[0037] Physiological signals have diverse local waveforms, which belong to the local temporal features of physiological signals. The number and occurrence position of the waveforms belong to the global temporal features of physiological signals. Both local and global temporal features are important bases for judging the staging. The temporal Transformer encoder designed in this method takes into account both the local and global temporal features of the signals. First, all segments that have undergone graph space encoding and multi-source signal fusion are concatenated as the input to the temporal Transformer encoder, as shown in Equation 9.

[0038] z [0] = [x′ p1 , x′ p2 , x′ p3 ,..., x′ pN T (Equation 9)​

[0039] is the input of the sequential Transformer encoder, which is composed of the spliced segments encoded by 1) and 2). N is the total number of segments into which the polysomnography data is divided, and the subscript number of x' p is the sequential number of the segment. The sequential Transformer encoder uses the Transformer architecture to capture important local time features. The encoder includes a total of L layers of Transformer layers. A single Transformer layer includes a multi-layer perceptron (MLP) composed of multiple connected linear layers, layer normalization (LN), and multi-head attention (Attn). The formula of a single Transformer layer is represented by Equation 10 and Equation 11.

[0040]

[0041] are the input and intermediate values of the i-th layer of the Transformer layer respectively. To retain the global information, the encoded local time features are subjected to residual connection with the initial global time features. Finally, the concatenated time features are flattened and passed through a linear layer to obtain the sleep staging result of this frame, as shown in Equation 12:

[0042]

[0043] is the automatic staging result of this frame obtained by this method, and its values correspond to the confidence levels of the 5 sleep stages (N1, N2, N3, R, W) respectively. z [L] is the output of the last layer of the Transformer layer, are the flattened local time features and global time features.

[0044] 4) Training and solution of the sleep staging neural network model

[0045] The present invention uses categorical cross-entropy (CCE) as the loss function and uses the stochastic gradient descent method for training. There are S samples and M categories (in this method, M = 5, corresponding to 5 sleep stages). The categorical cross-entropy is as shown in Equation 13:

[0046]

[0047] Among them, represents the corresponding confidence level of the m-th category of the s-th sample, and y sm is the label of the s-th sample. If the sample belongs to the m-th category, the value is 1; if it does not belong to the m-th category, the value is 0.

[0048] Through the above steps, the physiological state recognition and analysis based on multi-source information fusion are realized, and the sleep staging result is obtained.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] The present invention provides a sleep staging method based on spatio-temporal feature encoding and multi-source fusion. By using the collected multi-channel and multi-source physiological data, including electroencephalogram (EEG), electromyogram (EMG), and electrooculogram (EOG), a sleep staging neural network model that fuses multi-channel and multi-source physiological data is constructed: a graph space encoder for modeling multi-channel spatial features is established to encode the interaction and position information between different channels; a multi-signal source fusion module for fusing multi-source signals is established to identify the different contributions of multi-signal sources to staging based on the multi-head attention mechanism; a temporal Transformer encoder for modeling temporal features is established to encode both local and global temporal features simultaneously. By adopting the technical solution provided by the present invention, it helps to realize the automatic staging of sleep, improve the accuracy of staging, can provide practical help for individuals, provide a convenient tool for doctors, and provide heuristic auxiliary guidance for precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flowchart of a sleep staging method based on spatio-temporal feature encoding and multi-source fusion according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present invention will be further described below with reference to the accompanying drawings through embodiments, but it is not limited to the scope of the present invention in any way.

[0053] The present invention provides a sleep staging method based on spatio-temporal feature encoding and multi-source fusion, Figure 1 The flow of the method of the present invention is shown, including: using the collected multi-channel and multi-source physiological data (polysomnography data), including EEG, EMG, and EOG, to construct a neural network model that fuses multi-channel and multi-source physiological data: establishing a graph space encoder for modeling multi-channel spatial features to encode the interaction and position information between different channels; establishing a multi-signal source fusion module for fusing multi-source signals to identify the different contributions of multi-signal sources to staging based on the multi-head attention mechanism; establishing a temporal Transformer encoder for modeling temporal features to encode both local and global temporal features simultaneously. By adopting the technical solution provided by the present invention, it helps to realize the automatic staging of sleep, improve the accuracy of staging, can provide practical help for individuals, provide a convenient tool for doctors, and provide heuristic auxiliary guidance for precision medicine.

[0054] The following is an example of polysomnographic data of multiple individuals actually collected. The number of individuals is 26, all of whom are healthy people without sleep disorders. The data was collected by sampling at a rate of 256 Hz for one night of polysomnography for each individual and has been carefully revised. The average recording time for an individual is 30,765 seconds. The physiological signal recording includes different sleep stages (such as N1, N2, N3, R, W). The method of the present invention is used to implement sleep analysis based on spatiotemporal feature coding and multi-source fusion, which provides assistance for the identification of various sleep stages. Specifically, the experimental data is annotated by a professional sleep physician, and sleep stages are divided according to one frame every 30 seconds. The signals collected by the polysomnographic monitoring equipment include EEG signals, EOG signals, and EMG signals. The signal acquisition electrodes are installed at the corresponding sites in strict accordance with the "Manual of Interpretation of Sleep and Related Events of the American Academy of Sleep Medicine". Number of EEG signal channels C EEG =6, the electrode sites are F3-M2, F4-M1, C3-M2, C4-M1, O1-M2, O2-M1; the number of electrooculogram signal channels is C EOG =2, the electrode sites are E1-M2 and E2-M2 respectively; the number of myoelectric signal channels is C EMG =1, and the electrode sites are Chin 1-Chin 2. When the polysomnography equipment used in the experiment collected signals, the EEG and EOG signals were filtered using 0.3Hz low-pass and 35Hz high-pass filters, and the EMG signals were processed using 10Hz low-pass and 100Hz high-pass filters. All channels were then normalized to the range of [0, 1].

[0055] The polysomnographic data contains C = 9 channels in total. Each frame belongs to and only belongs to one sleep stage. This method divides each frame into N = 30 segments, each segment lasting 1 second.

[0056] The specific implementation method of identifying the subsequent physiological state data of an individual based on the collected physiological state data of the individual using the method of the present invention includes the following steps:

[0057] 1) Constructing a graph space encoder for spatial features; including steps 1a) to 1b):

[0058] 1a) Construct a multi-layer perceptron encoder for myoelectric and electrooculographic signals:

[0059] The channels of myoelectric and electroocular signals are only one or two, and their spatial information is far less than that of EEG signals with relatively more channels. For myoelectric and electroocular signals, a multi-layer perceptron (MLP) consisting of linear layers and residual connections is used to encode them for the first time. The multi-layer perceptron encoders of myoelectric and electroocular signals are expressed as Equation 1 and Equation 2 respectively:

[0060] H′ EMG =MLP(H EMG)(Formula 1)

[0061] H′ EOG = MLP(H EOG )(Formula 2)

[0062] where are the encoded EMG and EOG signals respectively, are the EMG channel and EOG channel signals selected from a single segment X p respectively. C EMG and C EOG are the number of channels of the EMG signal and the EOG signal in the segment respectively.

[0063] 1b) Construct a sparse graph neural network for EEG signals:

[0064] EEG signals have relatively more channels and strict positions are specified on the head, and their multi-channel spatial connection information is significant. Use one-dimensional convolution to convert EEG signals into the embedding of graph nodes, and the conversion process is expressed as Formula 3:

[0065]

[0066] where is the initial embedding of graph nodes converted from EEG signals, is the EEG channel signal selected from a single segment X p respectively. C EEG is the number of channels of the EEG signal in the segment.

[0067] Based on the graph nodes converted from EEG signals, construct a sparse connection graph and use the multi-head attention mechanism to explicitly learn the spatial connection between channels. Compared with the fully connected graph, the sparse connection graph adopted by this method reduces the number of edges and thus reduces the computational complexity. The steps to construct the sparse connection graph are as follows:

[0068] Random regular graph: Randomly permute the graph nodes converted from EEG signals, and connect the points in the rearranged order so that each point is connected to a constant number of other points to construct a random regular graph.

[0069] Virtual connection graph: Establish a constant number of virtual nodes. The virtual nodes have learnable node embeddings, and the virtual nodes are connected to each real graph node to construct a virtual connection graph.

[0070] The random regular graph and the virtual connection graph jointly construct a sparse connection on the graph nodes. Use the multi-head attention mechanism to calculate the weights of the edges connecting the nodes. Taking a single edge as an example, the weight calculation is expressed as Formula 4:

[0071] e i:j =(W K h i) T ·W Q h j (Equation 4)

[0072] where \(i, j = 1, 2, 3, \ldots, N\), \(W\) K , represents the learnable parameter, represents the weight of the edge connecting node \(i\) and node \(j\). is the embedding of node \(i\) in \(H\) EEG . The superscript \(T\) is the transpose symbol.

[0073] Use the message propagation mechanism and update the node embeddings of the graph, as shown in Equation 5:

[0074]

[0075] where \(\sigma\) is the activation function. \(b\) is the bias, is the set of nodes connected to node \(i\). \(h_i'\) is the updated embedding of node \(i\). When performing message propagation, the attention weight \(e\) of each edge connected to the node calculated by Equation 4 is applied j:i . All the node embeddings transformed from EEG signals together constitute the encoded EEG signals

[0076] 2) Construct a multi-signal source fusion module:

[0077] The PSG data includes multiple signal sources such as EEG, EMG, and EOG. This method uses the multi-head attention mechanism to fuse multi-source signals in the channel dimension \(C\). First, splice multiple signal sources encoded by the graph space encoder, as shown in Equation 6:

[0078]

[0079] where \(H'\) EEG , \(H'\) EOG , \(H'\) EMG are the EEG, EOG, and EMG signals that have passed through the graph space encoder respectively. After splicing them and performing transposition, we get Then use multi-head attention to focus on the contributions of different signal sources. The attention formula is shown in Equation 7:

[0080]

[0081] Among them, Q, K, and V respectively correspond to the query, key, and value in the Transformer architecture. The softmax function maps the attention to the interval [0, 1]. This method uses self-attention, where Q, K, and V in Equation 7 are the same and equal to H'. Subsequently, a linear layer is used on the channel dimension C to fuse multiple signal sources. The output of the multi-signal source fusion module is shown in Equation 8:

[0082] x′ p = Linear(Attn(H')) (Equation 8)

[0083] Where is a single segment that has undergone graph space encoding and multi-source signal fusion.

[0084] 3) Construct a temporal Transformer encoder for temporal features

[0085] Physiological signals have diverse local waveforms, which belong to the local temporal features of physiological signals. The number and occurrence positions of the waveforms belong to the global temporal features of physiological signals. Both local and global temporal features are important bases for interpreting the staging. The designed temporal Transformer encoder in this method considers both the local and global temporal features of the signals. First, all segments that have undergone graph space encoding and multi-source signal fusion are concatenated and input into the temporal Transformer encoder, as shown in Equation 9.

[0086] z [0] = [x′ p1 , x′ p2 , x′ p3 ,..., x′ pN T (Equation 9)

[0087] is the input of the temporal Transformer encoder, which is concatenated by the segments encoded in 1) and 2). N is the total number of segments into which the polysomnography data is divided, and the subscript number of x′ p is the sequential number of the segment. The temporal Transformer encoder uses the Transformer architecture to capture important local temporal features. The encoder includes a total of L Transformer layers. A single Transformer layer includes a multi-layer perceptron (MLP) composed of multiple connected linear layers, layer normalization (LN), and multi-head attention (Attn). The formula for a single Transformer layer is represented by Equation 10 and Equation 11.

[0088]

[0089] ​ They are the input and intermediate values of the i-th layer Transformer layer respectively. To preserve the global time features, the encoded local time features are subjected to residual connection with the initial global time features. Finally, the concatenated time features are flattened and passed through a linear layer to obtain the sleep staging result of this frame, as shown in Equation 12:

[0090]

[0091] is the automatic staging result of this frame obtained by this method, and its values correspond to the confidence levels of the 5 sleep stages (N1, N2, N3, R, W) respectively, z L is the output of the last layer Transformer layer, are the flattened local time features and global time features.

[0092] 4) Training and solution of the sleep staging neural network model

[0093] The present invention uses categorical cross-entropy (CCE) as the loss function and uses the stochastic gradient descent method for training. There are S samples and M categories (M = 5 in this method, corresponding to 5 sleep stages), and the categorical cross-entropy is as shown in Equation 13:

[0094]

[0095] where, represents the confidence level of the s-th sample for the corresponding category m, and y sm is the label of the s-th sample. If the sample belongs to category m, the value is 1; if it does not belong to category m, the value is 0.

[0096] The present invention is compared with three advanced deep learning methods in the current sleep analysis field, namely DeepSleepNet, AttnSleepNet, and CareSleepNet. The results are shown in Tables 1 and 2. It can be seen from Table 1 that the method proposed by the present invention is generally superior to the three advanced deep learning methods in terms of precision, recall, and F1-score in each sleep stage. On average for each stage, compared with the best results of other methods, the precision is increased by 20.8%, the recall is increased by 13.3%, and the F1-score is increased by 2.4%. This shows the superiority of the present invention.

[0097] Table 1. Comparison of the method of the present invention with other methods in each sleep stage

[0098]

[0099] It can be seen from Table 2 that the accuracy of the method of the present invention is increased by 9.2% compared with the best results of other methods,

[0100] Table 2. Comparison of the method of the present invention with other methods in terms of overall indicators

[0101]

[0102] Meanwhile, ablation experiments were conducted on three parts of the graph space encoder, multi-source fusion module, and temporal Transformer encoder in the method proposed by the present invention respectively. The results are shown in Tables 3, 4, and 5. It can be seen from Table 3 that the graph space encoder for spatial features proposed by the present invention (i.e., the above-mentioned step 1)) is meaningful. Compared with the graph neural network without attention mechanism and the one-dimensional convolutional neural network, the accuracy of analyzing the sleep period using the method proposed by the present invention has increased by 4.26%, and the Kappa has increased by 9.7%. It can be seen from Table 4 that the multi-source fusion module proposed by the present invention (i.e., the above-mentioned step 2)) has better effects than other fusion methods. Compared with only using a linear layer or taking the average value of all channels for multi-source fusion, the accuracy of the method proposed by the present invention has increased by 10.2%, and the Kappa value has increased by 14.5%. It can be seen from Table 5 that the temporal Transformer encoder for temporal features proposed by the present invention is effective. Compared with removing the temporal Transformer encoder, the accuracy of the method proposed by the present invention has increased by 13.2%, and the Kappa has increased by 17.9%

[0103] Table 3. Results of ablation experiment on graph space encoder

[0104]

[0105] Table 4. Results of ablation experiment on multi-source fusion module

[0106]

[0107] Table 5. Results of ablation experiment on temporal Transformer encoder

[0108]

[0109] The disclosed embodiments are intended to help better understand the present invention. However, professionals should understand that various substitutions and modifications can be made without departing from the essence of the invention and the appended claims. Therefore, the present invention should not be limited only to the content shown in these embodiments. The protection scope of the present invention should be determined according to the scope defined in the claims

Claims

1. A sleep staging method based on spatiotemporal feature coding and multi-source fusion, using the collected individual polysomnography data to establish an automatic sleep staging neural network model based on spatiotemporal feature coding and multi-source fusion, to achieve the staging and identification of the individual's subsequent sleep stages; characterized in that, The sleep staging neural network model includes three modules: 1) a graph-space encoder, 2) a multi-signal source fusion module, and 3) a temporal Transformer encoder, wherein the graph-space encoder is used to encode spatial interaction information in polysomnographic data, the multi-signal source fusion module fuses information of EEG, EMG and EOG signals in a learnable manner, and the temporal Transformer encoder captures local and global temporal features; sleep staging is performed on each frame of polysomnographic data, and C is set to represent the number of channels of the polysomnographic data, and L e Indicates the number of sampling points in a frame, X e Represents a frame of initial input, that is A frame of polysomnography data is divided into multiple segments, L p Indicates the number of sampling points in a fragment, X p Represents a fragment, i.e. The input of the graph space encoder and the multi-signal source fusion module is a single segment, and the temporal Transformer encoder takes all the segments that have undergone graph space encoding and multi-signal source fusion as input; the method comprises the following steps: (1) Construct a graph space encoder for spatial features to encode the spatial features of EEG, EOG and EMG respectively; (2) Construct a multi-signal source fusion module, by splicing the EEG, EOG, and EMG signals encoded in step 1), and fuse them on the channel dimension C using a multi-head attention mechanism; (3) Construct a temporal Transformer encoder for temporal features, concatenate the fused segments and input them into an L-layer temporal Transformer encoder to combine local and global temporal features; (4) The sleep staging neural network model is trained through the classification cross entropy loss function, and the sleep staging results are output.

2. The sleep staging method based on spatiotemporal feature coding and multi-source fusion as claimed in claim 1, characterized in that: The graph space encoder is constructed in step (1), specifically: 1a) Construct a multi-layer perceptron encoder for EMG and EOG. For EMG and EOG, a multi-layer perceptron MLP consisting of linear layers and residual connections is used for initial encoding to meet the following requirements: H′ EMG =MLP(H EMG ),H′ EOG =MLP(H EOG ) in, and are the encoded myoelectric and oculoscopic signals, and Before encoding, they are p The myoelectric channel and eye electrooculogram channel signals selected from C EMG and C EOG are the number of channels of myoelectric signal and electrooculographic signal in the segment, L p Indicates the number of sampling points in a fragment; 1b) Construct a sparse graph neural network for EEG signals. Convert EEG into graph node embeddings through one-dimensional convolution. Based on the graph nodes converted from EEG signals, construct a sparse connection graph containing random regular graphs and virtual connection graphs. Use the multi-head attention mechanism to calculate the edge weights of the connection nodes, and use the message propagation mechanism to update the node embeddings of the graph. All node embeddings converted from EEG signals together constitute the encoded EEG signal C EEG is the number of channels of the EEG signal in the clip.

3. The sleep staging method based on spatiotemporal feature coding and multi-source fusion as claimed in claim 2, characterized in that: In step 1b), the EEG signal is converted into the embedding of the graph node, and the conversion formula is: in, is the initial embedding of the graph nodes converted from EEG signals, For a single fragment X p The EEG channel signal selected from C EEG is the number of channels of EEG signals in the clip; The edge weight calculation formula is: have been i:j =(W K h i ) T ·W Q h j Where i, j = 1, 2, 3, ..., N, N is the total number of segments into which the polysomnographic data is divided, represents the learnable parameters, Represents the weight of the edge connecting node i and node j. yes The embedding of node i; The formula for updating the node embedding of the graph is: Among them, σ is the activation function, b is the bias, is the set of nodes connected to node i, is the updated embedding of node i. All node embeddings converted from EEG signals together constitute the encoded EEG signal C EEG is the number of channels of the EEG signal in the clip.

4. The sleep staging method based on spatiotemporal feature coding and multi-source fusion as claimed in claim 2, characterized in that: The construction of the sparse connection graph specifically includes: Random regular graph: Randomly arrange the graph nodes converted from EEG signals, connect the nodes in the rearranged order, make each node connected to other constant number of nodes, and construct a random regular graph; Virtual connection graph: A constant number of virtual nodes are established. The virtual nodes have learnable node embeddings. The virtual nodes are connected to each real graph node to build a virtual connection graph.

5. The sleep staging method based on spatiotemporal feature coding and multi-source fusion as claimed in claim 1, characterized in that: The multi-signal fusion in step (2) specifically includes: The EEG, EOG, and EMG signals encoded by the graph space encoder are concatenated and transposed to obtain Calculated through multi-head self-attention mechanism Among them, Q, K, and V correspond to the query, key, and value in the Transformer architecture respectively. The softmax function maps attention to the interval [0, 1]. The values ​​of Q, K, and V are the same and equal to H′. And use linear layer fusion on channel dimension C to get x′ p =Linear(Attn(H′)), It is a single segment that has undergone graph space encoding and multi-source signal fusion.

6. The sleep staging method based on spatiotemporal feature coding and multi-source fusion as claimed in claim 1, characterized in that: The construction of the temporal Transformer encoder in step (3) specifically includes: Splice all segments that have been coded in graph space and fused with multi-source signals z [0] =[x′ p1 , x′ p2 , x′ p3 , ..., x′ pN ] T , is the input of the temporal Transformer encoder, N is the total number of segments the polysomnography data is divided into, and x′ p The subscript numbers are the sequence numbers of the fragments; The temporal Transformer encoder consists of L layers of Transformer layers. A single Transformer layer consists of a multi-layer perceptron (MLP) consisting of multiple linear layers connected together, layer normalization (LN), and multi-head attention (Attn). The formula is: and They are the input and intermediate values ​​of the,th Transformer layer respectively; The encoded local time features are residually connected with the initial global time features, the connected time features are flattened and passed through a linear layer to obtain the sleep stage result of the frame. The formula is: z [L] is the output of the last Transformer layer, The values ​​correspond to the confidence levels of the five sleep periods N1, N2, N3, R, and W respectively.

7. The method according to claim 1, characterized in that In step (4), the sleep staging neural network model training uses classification cross entropy as the loss function and is trained using the stochastic gradient descent method. The classification cross entropy is defined as: Among them, S is the number of samples, M is the number of types, M = 5, corresponding to 5 sleep periods, represents the sth sample The confidence level corresponding to type m, y sm is the label of the sth sample.

Citation Information

Patent Citations

  • Sleep staging detection system based on graph attention mechanism and space-time graph convolution

    CN117158912A

  • Sleep staging method and device based on graph structure, electronic equipment and storage medium

    CN118986286A

  • Sleep stage detection method based on multi-mode and visual transformation network

    CN119548095A

  • Apparatus for automatically determining sleep disorder using deep running and operation method of the apparatus

    US20210045676A1

Cited By

  • Explanatable sleep staging method and device based on training after visual language model supervision fine tuning

    CN121542800A

  • EEG-EMG fusion rehabilitation control and evaluation method based on multi-band time-space cross attention

    CN121647700A