A Multimodal Sleep Staging Method Based on a Bidirectional Interactive RWKV Network
The multimodal sleep staging method using a bidirectional interactive RWKV network solves the problem of inconsistent contributions between modalities, achieving more efficient and accurate sleep staging, especially with significant improvements in long-term data.
Patent Information
- Application Number
- CN202511386065.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing multimodal physiological signal fusion methods struggle to dynamically and specifically utilize the strengths of each modality when dealing with inconsistencies in contributions between different modalities, resulting in insufficient staging efficiency and accuracy, especially evident in long-term sleep data.
A multimodal sleep staging method based on a bidirectional interactive RWKV network is adopted. The first and second RWKV modules are set in parallel to extract deep features from EEG, EOS and EMG signals respectively. The bidirectional self-context encoding module, cross-modal dynamic interaction module and bidirectional channel hybrid module are used to perform contextual modeling and fusion of intramodal and extramodal information, and finally generate the posterior probability distribution of sleep stages.
It improves the processing efficiency and accuracy of sleep staging, can more accurately aggregate complementary information from different modalities, and enhances the comprehensive analysis capability of complex multimodal physiological data, especially showing excellent performance in long-term sleep data.
Smart Images

Figure CN120873825B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a multimodal sleep staging method based on a bidirectional interactive RWKV network. Background Technology
[0002] Automated sleep staging is a core component in assessing individual sleep quality and diagnosing sleep disorders. Multimodal fusion analysis utilizing various physiological signals from polysomnography (PSG), such as electroencephalography (EEG), electrooculography (EOG), and electromyography (EMG), has become a mainstream research trend. Different modalities of physiological signals originate from different physiological processes and each carries unique, complementary information characterizing sleep states. Effectively fusing this information is crucial for constructing high-performance automated sleep staging methods.
[0003] However, a key challenge in multimodal physiological signal fusion is effectively managing the inconsistencies in contributions between different modalities. Specifically, EEG signals, because they directly reflect cortical activity, are generally considered the core modality dominating sleep stages; EOG signals have unique value in identifying REM sleep, while EMG signals can effectively distinguish muscle tone states. These differences in modality characteristics mean that their discriminative abilities and importance in representing different sleep stages vary. Using simple feature splicing or fixed weighted fusion strategies often makes it difficult to dynamically and specifically utilize the strengths of each modality, and may even suppress the contribution of key modalities by introducing redundant or secondary information.
[0004] To address this issue, we have drawn on advanced attention mechanisms in deep learning, with a particular emphasis on multimodal sleep staging methods based on the Transformer architecture and cross-modal attention. The core idea behind these sleep staging methods is to allow dynamic, context-dependent interactions between feature representations of different modalities. However, there are still issues with the efficiency and accuracy of staging, especially for long-term sleep data, where these shortcomings are even more pronounced. Summary of the Invention
[0005] Therefore, it is necessary to provide a multimodal sleep staging method based on a bidirectional interactive RWKV network to address the problems of low staging efficiency and low accuracy.
[0006] To solve the above problems, the present disclosure adopts the following technical solution:
[0007] This disclosure provides a multimodal sleep staging method based on a bidirectional interactive RWKV network, including the following steps:
[0008] Step 1: Obtain sleep data to be processed, including electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electromyogram (EMG) signals;
[0009] Step 2: Extract depth features from the EEG, EOS, and EMG signals, obtaining the depth features in a one-to-one correspondence. Depth features Depth features ;
[0010] Step 3, using depth features As principal modality features and depth features As auxiliary modal features, the first RWKV module is used to realize context modeling of intramodal information and fusion of intermodal information; with deep features As principal modality features and depth features As an auxiliary modal feature, a second RWKV module, which runs parallel to the first RWKV module, is used to realize context modeling of intramodal information and fusion of intermodal information; the outputs of the first RWKV module and the second RWKV module are integrated by element-wise addition to form a fused feature;
[0011] Step 4: Generate the posterior probability distribution for each sleep stage based on the fusion features, and take the category with the highest probability as the sleep stage prediction result.
[0012] In a preferred embodiment, step 2 includes: extracting depth features from the electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) signals one-to-one using three modal signal feature extraction modules to obtain depth features in a one-to-one correspondence. Depth features Depth features The modal signal feature extraction module is composed of a dual-branch convolution module and a residual attention module connected in series.
[0013] In a preferred embodiment, the dual-branch convolution module includes a first convolutional sub-unit as a small kernel branch and a second convolutional sub-unit as a large kernel branch; the first and second convolutional sub-units have the same structure, each including a first one-dimensional convolutional layer, a first maximum pooling layer, a first dropout layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, a fourth one-dimensional convolutional layer, and a second maximum pooling layer, which are sequentially connected; the residual attention module includes a residual layer composed of a fifth one-dimensional convolutional layer, a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function, which are sequentially arranged.
[0014] In a preferred embodiment, step 2 includes:
[0015] Step 2.1: The first and second convolutional sub-units obtain the input from the modal signal feature extraction module and perform convolution and pooling operations;
[0016] Step 2.2: Concatenate the feature maps output by the first and second convolutional sub-units along the feature dimension to form a combined feature map that integrates multi-scale information. ;
[0017] Step 2.3: Combining Feature Maps The first intermediate feature is obtained after processing with the residual layer. ;
[0018] Step 2.4: Use a global average pooling layer to process the first intermediate features. Global average pooling is performed along the time dimension to obtain the second intermediate feature. ;
[0019] Step 2.5: Transfer the second intermediate feature After learning the channel weights through the first fully connected layer, ReLU activation function, second fully connected layer, and Sigmoid activation function, the channel attention weight vector in the channel dimension is output.
[0020] Step 2.6: The channel attention weight vector and the intermediate features The attention-enhanced features are obtained by multiplying each channel element-wise. The first intermediate feature is connected via residual link. Features after attention enhancement The elements are added together to obtain the final output of the residual attention module.
[0021] In a preferred embodiment, the first RWKV module and the second RWKV module have the same structure, both including a first bidirectional self-context encoding module, a second bidirectional self-context encoding module, a cross-modal dynamic interaction module, and a bidirectional channel hybrid module; the first bidirectional self-context encoding module is used to implement context-aware enhancement of the main modality features and output context-aware enhanced main modality features. The second bidirectional self-context encoding module is used to enhance the context awareness of auxiliary modal features and output context-aware enhanced auxiliary modal features. The cross-modal dynamic interaction module is used to implement the context-aware enhanced main modality features. and the context-aware enhanced auxiliary modal features The dynamic interaction fusion yields the fused sequence features. The bidirectional channel mixing module is used to perform nonlinear transformation on the fused sequence features in the channel dimension, and after parallel modeling in the channel dimension, the features are fused. The fused sequence features are then restored to the dimension of the fused sequence features through linear mapping to form a semantically enhanced modal representation. The semantically enhanced modal representation and the fused sequence features are connected through residuals to form the final output of the RWKV module.
[0022] In a preferred embodiment, the first and second bidirectional self-context encoding modules have the same structure, including parallel forward RWKV-7 units and backward RWKV-7 units, and a linear fusion layer. The forward RWKV-7 units are used to process deep features to obtain a forward context representation. The input to the backward RWKV-7 unit is the time-reversed depth features, which, after processing by the backward RWKV-7 unit, yield the backward context representation. The linear fusion layer is used to project the first stitching result back to the dimension of its corresponding depth feature to form a final bidirectional context representation, wherein the first stitching result is a forward context representation. and backward context representation It is obtained by splicing along the feature dimension.
[0023] In a preferred embodiment, the forward RWKV-7 unit and the backward RWKV-7 unit have the same structure, both including a third normalization layer, a time mixing layer, a fourth normalization layer, and a first MLP layer containing ReLU².
[0024] In a preferred embodiment, the cross-modal dynamic interaction module includes a cross-modal temporal mixing layer and a fifth normalization layer. The cross-modal temporal mixing layer is used to adjust and update the context state matrix of the auxiliary modality features, and to generate a receptance vector based on the updated context state matrix of the auxiliary modality features and the main modality features. The interaction features are obtained by multiplication, and the fifth normalization layer is used to normalize the interaction features generated by the cross-modal temporal hybrid layer.
[0025] In a preferred embodiment, step 3 includes:
[0026] Step 3.3.1, the aforementioned The receptance vector is generated after Token Shift and linear transformation. Time decay vector Context learning rate and output gating Meanwhile, the aforementioned Generate normalized removal keys Normalized replacement key and normalized value vector ;
[0027] Step 3.3.2, at each time step Context state matrix of auxiliary modal features It evolves dynamically according to the following formula:
[0028] ;
[0029] in, express At time step The constructed context state matrix, express At time step The constructed context state matrix, This represents the Diag function. Indicates by At the current time step The generated time decay vector, Indicates by At the current time step The generated context learning rate, Indicates by At the current time step The generated normalized removal key, Indicates by At the current time step The generated normalized value vector, Indicates by At the current time step The generated normalized replacement key, This represents the XOR operation;
[0030] Step 3.3.3: The updated version in Step 3.3.2 Receptance vector generated from main modality features Multiplication yields interactive features ,in, This indicates that the dominant modality features at the current time step... The generated receptance vector, Indicates the current time step Interactive features This represents element-wise multiplication.
[0031] Step 3.3.4: Normalize the interaction features generated in Step 3.3.3, and then... At the current time step Generated output gating The output control of the cross-modal dynamic interaction module is completed under the adjustment, and the fused sequence features are formed through linear mapping;
[0032] Step 3.4.1: The fused sequence features generated in step 3.3.4 enter the bidirectional channel mixing module, and undergo nonlinear transformations in the forward and backward channels of the bidirectional channel mixing module respectively. After parallel modeling in the channel dimension, they are fused and restored to the dimension of the fused sequence features through linear mapping to form a semantically enhanced modal representation.
[0033] Step 3.4.2: Perform residual connection between the fused sequence features to obtain the final output of the RWKV module.
[0034] In a preferred embodiment, step 4 includes:
[0035] Step 4.1: Input the fused features into the first linear layer and embed them onto the target classification dimension;
[0036] The outputs of steps 4.2 and 4.1 are normalized by the Softmax layer to generate the posterior probability distribution for each sleep stage;
[0037] Step 4.3: Select the category with the highest probability in the output of the Softmax layer in Step 4.2 as the prediction result for sleep stage.
[0038] The aforementioned multimodal sleep staging method based on a bidirectional interactive RWKV network extracts deep features from EEG, EEG, and EMG signals. It then integrates the results of a first RWKV module with a parallel second RWKV module, generating a posterior probability distribution for each sleep stage. The category with the highest probability is then selected as the predicted sleep stage. This design improves the efficiency of sleep staging; it enables more precise and dynamic aggregation of complementary information from different modalities, effectively addressing the problem of inconsistent modal contributions; it enhances the comprehensive analysis capability of complex multimodal physiological data, improving the accuracy of sleep staging; and it maintains high efficiency and accuracy even for long-term sleep data. Attached Figure Description
[0039] Figure 1 This is a flowchart illustrating a method in one embodiment of the present disclosure;
[0040] Figure 2 This is a schematic diagram of one embodiment of the present disclosure;
[0041] Figure 3 This is a schematic diagram of a modal signal feature extraction module in one embodiment of the present disclosure;
[0042] Figure 4 This is a schematic diagram of the RWKV-7 unit in one embodiment of this disclosure. Detailed Implementation
[0043] The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and preferred embodiments.
[0044] Terminology Explanation:
[0045] Sleep staging: Sleep staging refers to the process of dividing the sleep process into different stages based on changes in physiological signals such as electroencephalography (EEG), electrooculography (EOG), and electromyography (EMG). Typically, sleep staging is divided into five stages: wakefulness (W), rapid eye movement (REM) sleep, and three non-rapid eye movement (N1, N2, and N3) sleep stages.
[0046] See Figure 1 This embodiment provides a multimodal sleep staging method for bidirectional interactive RWKV networks, the method comprising:
[0047] Step 1: Obtain sleep data to be processed, including electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electromyogram (EMG) signals;
[0048] Step 2: Extract depth features from the EEG, EOS, and EMG signals respectively, obtaining the depth features in a one-to-one correspondence. Depth features Depth features ;
[0049] Step 3, using the depth features As the primary modality feature, the depth feature As auxiliary modal features, the first RWKV module is used to realize context modeling of intramodal information and fusion of intermodal information; with the aforementioned deep features As the primary modality feature, the depth feature As an auxiliary modal feature, a second RWKV module, which runs in parallel with the first RWKV module, is used to realize context modeling of intramodal information and fusion of intermodal information; the outputs of the first RWKV module and the second RWKV module are integrated by element-wise addition to form a fused feature;
[0050] Step 4: Generate the posterior probability distribution for each sleep stage based on the fusion features, and take the category with the highest probability as the sleep stage prediction result.
[0051] See Figure 2 The following examples will describe the method described above in detail.
[0052] Step 1: Obtain sleep data to be processed, including electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electromyogram (EMG) signals;
[0053] Understandably, in step 1, after obtaining the electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) signals, the EEG, EOG, and EMG signals are preprocessed, and the preprocessed signals are used in step 2.
[0054] Step 2: The multimodal feature extraction module extracts depth features from the EEG, EOS, and EMG signals, obtaining the depth features one-to-one. Depth features Depth features .
[0055] In step 2, the electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) signals serve as input signals to the multimodal feature extraction module. This module comprises three modal signal feature extraction modules, each corresponding one-to-one with the EEG, EOG, and EMG signals for feature extraction. These three modules are respectively referred to as the EEG modal signal feature extraction module, the EOG modal signal feature extraction module, and the EMG modal signal feature extraction module. In other words, the deep features are obtained by performing deep feature extraction on the EEG signal through the EEG modal signal feature extraction module. The depth features are obtained by performing depth feature extraction on the electrooculogram (EOG) signal using the EOG modal signal feature extraction module. The depth features are obtained by performing depth feature extraction on the electromyography signal through the electromyography modal signal feature extraction module. .
[0056] The EEG modal signal feature extraction module, EOS modal signal feature extraction module, and EMG modal signal feature extraction module have the same structural architecture. They are all composed of a dual-branch convolutional module and a residual attention module (Residual SE Block) connected in series. That is, the modal signal feature extraction module includes a first convolutional sub-unit, a second convolutional sub-unit, and a residual attention module.
[0057] The first and second convolutional subunits constitute the dual-branch convolutional module, with the first convolutional subunit serving as the small kernel branch and the second convolutional subunit serving as the large kernel branch. The input terminals of both the first and second convolutional subunits are used to input EEG signals, EEG signals, or EMG signals, and their output terminals correspond to the input terminals of the residual attention module.
[0058] Since the three modal signal feature extraction modules have the same structure, the following detailed explanation will focus on the EEG modal signal feature extraction module as an example. For the specific structure of the modal signal feature extraction module, please refer to [link / reference]. Figure 3 .
[0059] The first and second convolutional subunits have the same architecture, each consisting of four one-dimensional convolutional layers, two max-pooling layers, and one dropout layer, in the following order: first one-dimensional convolutional layer, first max-pooling layer, first dropout layer, second one-dimensional convolutional layer, third one-dimensional convolutional layer, fourth one-dimensional convolutional layer, and second max-pooling layer, sequentially connected in series. Each one-dimensional convolutional layer in the dual-branch convolutional module is followed by a first batch normalization layer (first BN layer) using LeakyReLU as the activation function, i.e., following the order of one-dimensional convolutional layer, first BN layer, and activation function.
[0060] The residual attention module includes, in sequence, a residual layer consisting of a fifth one-dimensional convolutional layer, a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function.
[0061] Specifically, it includes the following sub-steps:
[0062] Step 2.1: Use the preprocessed 30-second EEG signal segment (this is just an example, and the 30 seconds is also just an example) as the input to the EEG modality signal feature extraction module. That is, the input signal is simultaneously sent to two parallel convolution branches, the first convolution sub-unit and the second convolution sub-unit.
[0063] Step 2.2: After the first and second convolutional sub-units each complete the convolution and pooling operations, their output feature maps are concatenated along the feature dimension to form a combined feature map that integrates multi-scale information. Understandable, preferred, combined feature maps Before entering the residual attention module, it passes through a second Dropout layer to prevent overfitting.
[0064] Step 2.3: Combine the feature maps obtained in Step 2.2. The feature maps are fed into the residual attention module and combined. First, the first intermediate feature is obtained through a residual layer consisting of a fifth one-dimensional convolutional layer. .
[0065] Step 2.4: Take the first intermediate feature from Step 2.3. Global average pooling is performed along the time dimension to obtain the second intermediate feature. This is the "Squeeze" operation (which compresses the spatial dimension of each channel through global average pooling to generate a global descriptor).
[0066] Step 2.5: The second intermediate feature obtained in step 2.4 The channel attention weight vector is output after performing an "Excitation" operation through two fully connected layers (learning channel weights through two fully connected layers and the activation function [ReLU+Sigmoid]).
[0067] The channel attention weight vectors obtained in steps 2.6 and 2.5 are used to process the intermediate features obtained in step 2.3. The attention-enhanced features are obtained by multiplying each channel element-wise (weighted). Subsequently, the first intermediate feature is connected via residual link. Features after attention enhancement Adding elements one by one yields the final output of the residual attention module. = + Output This refers to the depth feature representation obtained after the modality has passed through a modality signal feature extraction module (EEG modality signal feature extraction module, EOL modality signal feature extraction module, or EMG modality signal feature extraction module), which serves as the output of the modality signal feature extraction module. As the output feature of the EEG modal signal feature extraction module As the output feature of the electrooculogram modal signal feature extraction module As the output feature of the electromyographic modal signal feature extraction module.
[0068] In step 3, the RWKV module can be understood as an RWKV model or RWKV layer. The RWKV module achieves real-time fusion of multimodal data through dynamic state evolution and cross-modal adaptation. The first RWKV module and the second RWKV module are in parallel structure. The first RWKV module is called the first bidirectional RWKV cross-modal fusion module (abbreviated as the first Bi-IFM), and the second RWKV module is called the second bidirectional RWKV cross-modal fusion module (abbreviated as the second Bi-IFM). The bidirectional RWKV-7 cross-modal fusion module is a Bidirectional RWKV-7 Interaction Fusion Module, where RWKV stands for ReceptanceWeighted Key Value. EEG corresponds to the primary modality, while EOG and EMG serve as auxiliary modalities, using deep features. As the primary modality feature, with depth features and depth features As an auxiliary modal feature, the first Bi-IFM is responsible for the main modal feature and an auxiliary modal feature ( The interaction between the primary modality feature and another auxiliary modality feature ( ) is handled by the second Bi-IFM. The interaction between the first and second Bi-IFMs is as follows. The first Bi-IFM and the second Bi-IFM have the same structure, and Bi-IFM represents either the first or the second Bi-IFM. The first Bi-IFM is denoted as... Module, the second Bi-IFM is denoted as Module.
[0069] The Bi-IFM (Bi-Self-RWKV) cross-modal fusion module in Bi-IFM is used to achieve deep contextual modeling of intramodal information and fusion of intermodal information (i.e., dynamic interaction and fusion). Key components of Bi-IFM include the Bi-Self-RWKV module, the cross-modal dynamic interaction module, and the bi-channel hybridization module. Through the synergistic effect of these three modules, Bi-IFM can achieve in-depth mining of intramodal semantics and intermodal complementarity. There are two Bi-Self-RWKV modules: one called the first Bi-Self-RWKV module and the other called the second Bi-Self-RWKV module. The cross-modal dynamic interaction module and the bi-channel hybridization module use independent, non-shared parameters to capture temporal dependency features in different directions. The bidirectional self-context encoding module is used to perform context-aware enhancement processing on deep features to obtain context-aware enhanced main modality features / auxiliary modality features. The cross-modal dynamic interaction module is used to utilize the dynamic state update mechanism of RWKV-7 to treat the main modality features as the "query party" and generate control signals in real time to modulate the context state matrix constructed by the auxiliary modalities, and selectively extract key complementary information from it, thereby effectively addressing the problem of inconsistent modal contributions.
[0070] The Bi-IFM includes a first bidirectional self-context encoding module, a second bidirectional self-context encoding module, a cross-modal dynamic interaction module, and a bidirectional channel mixing module. Specifically, Bi-IFM also includes a first normalization layer and a second normalization layer. The first normalization layer is used to obtain the main modality features (deep features). The first bidirectional self-context encoding module is used to normalize the main modality features (output of the first normalization layer) and output context-aware enhanced main modality features. The second normalization layer is used to obtain auxiliary modality features (deep features). or depth features The first module performs normalization processing on the auxiliary master modality features (output of the second normalization layer) and outputs context-aware enhanced auxiliary modality features. The second bidirectional self-context encoding module is used to obtain the output results of the two bidirectional self-context encoding modules and also to realize the dynamic interaction fusion of context-aware enhanced master modality features and context-aware enhanced auxiliary modality features to obtain the fused sequence features. The second bidirectional channel mixing module is used to obtain the output of the cross-modal dynamic interaction module. As a bidirectional feedforward structure, the second bidirectional channel mixing module is used to perform nonlinear transformation on the fused sequence features in the channel dimension, perform parallel modeling in the channel dimension and then fuse them, and form the final output of Bi-IFM through residual connection.
[0071] The bidirectional self-context coding module in the Bi-IFM includes a forward RWKV-7 unit, a backward RWKV-7 unit, and a linear fusion layer. The forward RWKV-7 unit is referred to as... The backward RWKV-7 unit is called Parallel forward and backward RWKV-7 units are used to achieve context-aware enhancement. Both forward and backward RWKV-7 units are parameter-independent. A bidirectional self-context encoding module performs deep modeling of a single modality sequence using parameter-independent forward and backward RWKV-7 units, generating a feature representation that fuses complete temporal information from both ends, providing a robust foundation for subsequent intermodal interactions. A linear fusion layer projects the first concatenation result back to the dimension of its corresponding deep feature, forming the final bidirectional context representation. The forward and backward RWKV-7 units have identical structures; an RWKV-7 unit represents any one of the forward or backward RWKV-7 units. See also... Figure 4 The RWKV-7 unit includes: a third normalization layer, a time mixing layer, a fourth normalization layer, and a first MLP layer (Multi-Layer Perceptron, MLP) containing ReLU². In the bidirectional self-context encoding module, the input to the forward RWKV-7 unit is... After processing by the forward RWKV-7 unit, the following is obtained: The input to the backward RWKV-7 unit is the time-dimension inverted value. After processing by the RWKV-7 unit, the following is obtained: Representing the forward context and backward context representation The first concatenation result is obtained by concatenating the features along their dimensions. Then, a linear fusion layer projects the first concatenation result back to the dimension of its corresponding deep features, forming the final bidirectional context representation. For the first bidirectional self-context encoding module, the input to the forward RWKV-7 unit is... The result of the first normalization layer is then input to the RWKV-7 unit. The result of the first normalization layer normalization, followed by time-dimension inversion; for the second bidirectional self-context encoding module, its forward RWKV-7 unit input is... or The result after normalization by the second normalization layer is then input to the RWKV-7 unit. or The result is normalized by the second normalization layer and then reversed by the time dimension.
[0072] The cross-modal dynamic interaction module includes a cross-modal temporal mixing layer (Cross-TimeMix) and a fifth normalization layer. The cross-modal temporal mixing layer is used to adjust and update the context state matrix of the auxiliary modality features, and is used to generate a receptance vector based on the updated context state matrix of the auxiliary modality features and the main modality features. The interaction features are obtained by multiplication, and the fifth normalization layer is used to normalize the interaction features generated by the cross-modal temporal mixing layer.
[0073] The bidirectional channel mixing module includes a bidirectional channel mixing layer (Bi-ChannelMix), which comprises a forward channel and a backward channel. The bidirectional channel mixing module performs a nonlinear transformation on the fused sequence features along the channel dimension. It models the forward and backward channels in parallel along their respective channel dimensions, fuses the outputs of the forward and backward channels, and restores the fused sequence features to their original dimension through a linear mapping, forming a semantically enhanced modal representation. The semantically enhanced modal representation and the fused sequence features are then joined via a residual connection to form the final output of Bi-IFM.
[0074] The parallel structure of the bidirectional RWKV-7 cross-modal fusion module also includes an integration module, which is used to integrate the output results of the first bidirectional RWKV-7 cross-modal fusion module and the second bidirectional RWKV-7 cross-modal fusion module through element-wise addition to form a unified deep fusion multimodal representation.
[0075] Step 3 specifically includes:
[0076] Step 3.1 (First or Second Normalization Layer): Normalize the deep features input to the bidirectional RWKV cross-modal fusion module. Then, feed the normalized deep features into the first bidirectional self-context encoding module in the original sequence order. Obtain the forward context representation The normalized deep features are inverted in the time dimension and then fed into the second bidirectional self-context encoding module. For the second bidirectional self-context encoding module After reversing the time dimension of the output again, a backward context representation with the same order as the original sequence is obtained. .
[0077] Step 3.2: Represent the forward context obtained in Step 3.1 and backward context representation The features are concatenated along their respective dimensions, and the concatenated result is projected back to the dimensions of the deep features through a linear fusion layer to form the final bidirectional context representation. The bidirectional context representation of the dominant modality obtained through the bidirectional self-context encoding module is denoted as the context-aware enhanced dominant modality feature. The bidirectional context representation obtained by the auxiliary modality through the bidirectional self-context encoding module is denoted as the context-aware enhanced auxiliary modality feature. .
[0078] Step 3.3: The cross-modal dynamic interaction module obtains the output results (main modality features) of the bidirectional self-context encoding module. and auxiliary modal features It also enables dynamic interactive fusion of primary modal features and secondary modal features.
[0079] Step 3.3.1: Context-aware enhanced main modality features The key control signal, the receptance vector, is generated after Token Shift (an optimization technique for sequence modeling) and linear transformation. (Used to receive information from the past), time decay vector Context learning rate and output gating Simultaneously, context-aware enhanced auxiliary modal features Generate normalized removal keys Normalized replacement key and normalized value vector This forms a key-value pair structure for subsequent dynamic interactions.
[0080] Step 3.3.2, at each time step Context state matrix constructed from auxiliary modal features It evolves dynamically according to the following formula:
[0081]
[0082] in, express At time step The constructed context state matrix, express At time step The constructed context state matrix, This represents the Diag function. Indicates by At the current time step The generated time decay vector, Indicates by At the current time step The generated context learning rate, Indicates by At the current time step The generated normalized removal key, Indicates by At the current time step The generated normalized value vector, Indicates by At the current time step The generated normalized replacement key, This indicates the XOR operation.
[0083] The update process in step 3.3.2 embodies a fine-grained control mechanism for cross-modal information: the old state first passes through... Retain some memories, and then use and Selectively erase specific content, and finally... new value vector With Replace key Introduce updated information. The entire process is underway. It is carried out under dynamic control to ensure that the state evolution has time consistency and modal sensitivity.
[0084] Step 3.3.3: The updated version in Step 3.3.2 and The generated Receptance vector Multiplication yields interactive features Reflecting the current moment Dynamic retrieval of auxiliary information. Among them, Indicates by At the current time step The generated receptance vector, Indicates the current time step Interactive features The symbol represents element-wise multiplication.
[0085] Step 3.3.4: (Fifth Normalization Layer) Normalize the interaction features generated in Step 3.3.3, and optionally combine them with... , , Residual enhancement is then performed. Finally, the gated vector (by...) At the current time step (Generated output gating) The output control of the cross-modal dynamic interaction module is completed under the regulation of the module, and the output result of the cross-modal dynamic interaction module is formed by linear mapping. This is called the semantically enhanced modal representation, and the fused sequence feature is called the fused sequence feature.
[0086] Step 3.4.1: The fused sequence features generated in Step 3.3.4 enter the bidirectional channel mixing module and undergo nonlinear transformations in the forward and backward channels of the bidirectional channel mixing module, i.e., ReLU² activated MLP. After parallel modeling in the channel dimensions, they are fused (parallel modeling is performed in the forward and backward channels in their respective channel dimensions, and the outputs of the forward and backward channels are fused). After fusion, the original embedding dimension (i.e., the dimension of the fused sequence features) is restored through linear mapping to form a semantically enhanced modal representation.
[0087] Step 3.4.2: Perform a residual connection between the final output of step 3.4.1 and the final output of step 3.3.4 to obtain the final output of Bi-IFM.
[0088] Step 3.5: The outputs of the two parallel Bi-IFMs are integrated through element-wise addition to form a unified deep fusion multimodal representation, called the fusion feature, which is used for subsequent sleep stage classification tasks.
[0089] Step 4: The linear classification module generates the posterior probability distribution of each sleep stage based on the fusion features, and the selection module takes the category with the highest probability as the sleep stage prediction result.
[0090] The linear classification module includes a first linear layer and a classification layer, wherein the classification layer is a Softmax layer.
[0091] The specific process of step 4 is as follows:
[0092] Step 4.1: Input the fused features output from Step 3.5 into the first linear layer and map them from the high-dimensional embedding to the target classification dimension.
[0093] The outputs of steps 4.2 and 4.1 are normalized by the Softmax function to generate the posterior probability distribution for each sleep stage.
[0094] Step 4.3: Select the category with the highest probability in the output of the Softmax layer in Step 4.2 as the prediction result of the sleep stage of the current 30-second signal segment.
[0095] If a longer period of sleep (such as a whole night's sleep) is to be staged, the whole night's data is divided into several segments, and each segment is processed using the method described above. That is, each segment of data is used as the sleep data to be processed in step 1. Through this multimodal sleep staged method, a complete sleep stage sequence of the whole night's sleep can be generated.
[0096] As an example, sleep is divided into five stages: sleep onset, light sleep, deep sleep, non-rapid eye movement (NREM) sleep, and REM sleep. This method allows for the segmentation of sleep data throughout the night. The results are typically: some data belong to the waking period, some to the REM period, some to the non-rapid eye movement (NREM) period (N1), some to the NREM period (N2), and some to the NREM period (N3).
[0097] Table 1 compares the performance of the multimodal sleep staging method disclosed herein with other existing methods. As can be seen from Table 1, the sleep staging method of this disclosure has high accuracy, effectively avoids neglecting minority classes, and exhibits high classification consistency.
[0098] Table 1
[0099]
[0100] In Table 1, ISRUC-S3 represents the dataset released by the ISRUC (Intelligent Systems Research Unit of Coimbra) team; DeepSleepNet represents an end-to-end sleep staging model based on deep learning; AttnSleep represents a sleep staging model based on attention mechanisms; BiT-MamSleep represents a sleep staging model that fuses bidirectional Transformer with multimodal attention; SalientSleepNet represents a lightweight sleep staging model based on saliency detection; MMASleepNet represents a multi-scale attention sleep staging network; MultiChannel represents a sleep staging method based on multi-channel feature engineering; MHFNet represents a sleep staging model that mixes high-frequency features with a lightweight network; SleepWaveNet represents a temporal sleep staging model based on WaveNet; Ours represents the method disclosed in this paper; ACC represents accuracy; and MF1 represents the macro average F1 score. This represents the Kappa coefficient, Cohen's Kappa.
[0101] The multimodal sleep staging method based on a bidirectional interactive RWKV network disclosed herein achieves the following results: Based on deep feature extraction from EEG, EEG, and EMG signals, and utilizing a parallel first RWKV module and a parallel second RWKV module, the results of the two RWKV modules are fused to generate a posterior probability distribution for each sleep stage. The category with the highest probability is then selected as the sleep staging prediction result. This design improves the accuracy and efficiency of sleep staging. Especially for long-term sleep data, the staging efficiency and accuracy of this disclosed method are significantly higher than existing sleep staging methods.
[0102] Specifically, this disclosure, leveraging the linear computational complexity of the RWKV architecture, significantly improves the efficiency of sleep staging calculations and reduces resource consumption when processing long-term sleep data. Simultaneously, RWKV's integrated RNN-like sequence modeling capabilities and bidirectional processing flow design make it more adaptable to the temporal characteristics of physiological signals, enabling more accurate and dynamic aggregation of complementary information from different modalities. This more effectively addresses the problem of inconsistent modal contributions and enhances the comprehensive analysis capabilities for complex multimodal physiological data, thereby improving the accuracy of sleep staging. Especially for long-term data, the staging efficiency and accuracy of the method disclosed in this disclosure are significantly higher than existing sleep staging methods.
[0103] This disclosure provides a multimodal sleep staging system based on a bidirectional interactive RWKV network, including:
[0104] The acquisition module is used to acquire sleep data to be processed, including electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electromyogram (EMG) signals.
[0105] The feature extraction module is used to extract deep features from electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) signals, obtaining deep features in a one-to-one correspondence. Depth features Depth features ;
[0106] The feature fusion module is used to fuse deep features. As principal modality features and depth features As auxiliary modal features, the first RWKV module is used to realize context modeling of intramodal information and fusion of intermodal information; used for deep features As principal modality features and depth features As an auxiliary modal feature, a second RWKV module, which runs parallel to the first RWKV module, is used to realize context modeling of intramodal information and fusion of intermodal information; it is used to integrate the outputs of the first RWKV module and the second RWKV module through element-wise addition to form a fused feature;
[0107] The prediction module is used to generate a posterior probability distribution for each sleep stage based on the fused features, and take the category with the highest probability as the sleep stage prediction result.
[0108] In specific implementation, the multimodal sleep staging system based on a bidirectional interactive RWKV network can be implemented by referring to the multimodal sleep staging method based on a bidirectional interactive RWKV network in any of the above embodiments. The specific implementation steps will not be repeated here.
[0109] An electronic device can be implemented according to the method of this disclosure, the electronic device comprising: a memory; one or more processors; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for executing the multimodal sleep staging method based on a bidirectional interactive RWKV network according to any of the above embodiments.
[0110] This disclosure also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the steps of the multimodal sleep staging method based on a bidirectional interactive RWKV network described in any of the above embodiments.
[0111] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0112] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A multimodal sleep staging method based on a bidirectional interactive RWKV network, characterized in that, Includes the following steps: Step 1: Obtain sleep data to be processed, including electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electromyogram (EMG) signals; Step 2: Extract depth features from the EEG, EOS, and EMG signals, obtaining the depth features in a one-to-one correspondence. Depth features Depth features ; Step 3, using depth features As principal modality features and depth features As auxiliary modal features, the first RWKV module is used to realize context modeling of intramodal information and fusion of intermodal information; with deep features As principal modality features and depth features As an auxiliary modal feature, a second RWKV module, which runs parallel to the first RWKV module, is used to realize context modeling of intramodal information and fusion of intermodal information; the outputs of the first RWKV module and the second RWKV module are integrated by element-wise addition to form a fused feature; Step 4: Generate the posterior probability distribution for each sleep stage based on the fusion features, and take the category with the highest probability as the sleep stage prediction result. The first and second RWKV modules have the same structure, both including a first bidirectional self-context encoding module, a second bidirectional self-context encoding module, a cross-modal dynamic interaction module, and a bidirectional channel mixing module; the first bidirectional self-context encoding module is used to implement context-aware enhancement of the main modality features and output context-aware enhanced main modality features. The second bidirectional self-context encoding module is used to enhance the context awareness of auxiliary modal features and output context-aware enhanced auxiliary modal features. The cross-modal dynamic interaction module is used to implement the context-aware enhanced main modality features. and the context-aware enhanced auxiliary modal features The dynamic interactive fusion yields the fused sequence features. The bidirectional channel mixing module is used to perform nonlinear transformation on the fused sequence features in the channel dimension, and after parallel modeling in the channel dimension, the features are fused. The fused sequence features are then restored to the dimension of the fused sequence features through linear mapping to form a semantically enhanced modal representation. The semantically enhanced modal representation and the fused sequence features are connected through residuals to form the final output of the RWKV module. The first and second bidirectional self-context encoding modules have the same structure, including parallel forward RWKV-7 units and backward RWKV-7 units, and a linear fusion layer. The forward RWKV-7 units are used to process deep features to obtain a forward context representation. The input to the backward RWKV-7 unit is the deep features after time-dimension reversal. After processing by the backward RWKV-7 unit, the backward context representation is obtained. The linear fusion layer is used to project the first stitching result back to the dimension of its corresponding depth feature to form a final bidirectional context representation, wherein the first stitching result is a forward context representation. and backward context representation It is obtained by concatenating along the feature dimension; The cross-modal dynamic interaction module includes a cross-modal temporal mixing layer and a fifth normalization layer. The cross-modal temporal mixing layer is used to adjust and update the context state matrix of the auxiliary modality features, and to generate a receptance vector based on the updated context state matrix of the auxiliary modality features and the main modality features. The interaction features are obtained by multiplication, and the fifth normalization layer is used to normalize the interaction features generated by the cross-modal temporal hybrid layer.
2. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to claim 1, characterized in that, Step 2 includes: extracting depth features from the electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG) signals one-to-one using three modal signal feature extraction modules, thereby obtaining the depth features in a one-to-one correspondence. Depth features Depth features The modal signal feature extraction module is composed of a dual-branch convolution module and a residual attention module connected in series.
3. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to claim 2, characterized in that, The dual-branch convolution module includes a first convolutional sub-unit as a small kernel branch and a second convolutional sub-unit as a large kernel branch; the first and second convolutional sub-units have the same structure, each including a first one-dimensional convolutional layer, a first maximum pooling layer, a first dropout layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, a fourth one-dimensional convolutional layer, and a second maximum pooling layer, which are sequentially connected; the residual attention module includes a residual layer composed of a fifth one-dimensional convolutional layer, a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function, which are sequentially arranged.
4. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to claim 3, characterized in that, Step 2 includes: Step 2.1: The first and second convolutional sub-units obtain the input from the modal signal feature extraction module and perform convolution and pooling operations; Step 2.2: Concatenate the feature maps output by the first and second convolutional sub-units along the feature dimension to form a combined feature map that integrates multi-scale information. ; Step 2.3: Combining Feature Maps The first intermediate feature is obtained after processing with the residual layer. ; Step 2.4: Use a global average pooling layer to process the first intermediate features. Global average pooling is performed along the time dimension to obtain the second intermediate feature. ; Step 2.5: Transfer the second intermediate feature After learning the channel weights through the first fully connected layer, ReLU activation function, second fully connected layer, and Sigmoid activation function, the channel attention weight vector in the channel dimension is output. Step 2.6: The channel attention weight vector and the intermediate features The attention-enhanced features are obtained by multiplying each channel element-wise. The first intermediate feature is connected via residual link. Features after attention enhancement The elements are added together to obtain the final output of the residual attention module.
5. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to claim 1, characterized in that, The forward RWKV-7 unit and the backward RWKV-7 unit have the same structure, both including a third normalization layer, a time mixing layer, a fourth normalization layer, and a first MLP layer containing ReLU².
6. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to any one of claims 1-5, characterized in that, Step 3 includes: Step 3.3.1, the aforementioned The receptance vector is generated after Token Shift and linear transformation. Time decay vector Context learning rate and output gating Meanwhile, the aforementioned Generate normalized removal keys Normalized replacement key and normalized value vector ; Step 3.3.2, at each time step Context state matrix of auxiliary modal features It evolves dynamically according to the following formula: ; in, express At time step The constructed context state matrix, express At time step The constructed context state matrix, This refers to the Diag function. Indicates by At the current time step The generated time decay vector, Indicates by At the current time step The generated context learning rate, Indicates by At the current time step The generated normalized removal key, Indicates by At the current time step The generated normalized value vector, Indicates by At the current time step The generated normalized replacement key, This represents the XOR operation; Step 3.3.3: The updated version in Step 3.3.2 Receptance vector generated from main modality features Multiplication yields interactive features ,in, This indicates that the dominant modality features at the current time step... The generated receptance vector, Indicates the current time step Interactive features This represents element-wise multiplication. Step 3.3.4: Normalize the interaction features generated in Step 3.3.3, and then... At the current time step Generated output gating The output control of the cross-modal dynamic interaction module is completed under the adjustment, and the fused sequence features are formed through linear mapping; Step 3.4.1: The fused sequence features generated in step 3.3.4 enter the bidirectional channel mixing module, and undergo nonlinear transformations in the forward and backward channels of the bidirectional channel mixing module respectively. After parallel modeling in the channel dimension, they are fused and restored to the dimension of the fused sequence features through linear mapping to form a semantically enhanced modal representation. Step 3.4.2: Perform residual connection between the fused sequence features to obtain the final output of the RWKV module.
7. The multimodal sleep staging method based on a bidirectional interactive RWKV network according to claim 1, characterized in that, Step 4 includes: Step 4.1: Input the fused features into the first linear layer and embed them onto the target classification dimension; The outputs of steps 4.2 and 4.1 are normalized by the Softmax layer to generate the posterior probability distribution for each sleep stage; Step 4.3: Select the category with the highest probability in the output of the Softmax layer in Step 4.2 as the prediction result for sleep stage.
Citation Information
Patent Citations
Automatic sleep staging method, device and equipment and storage medium
CN117860203A
KR20230095431A