Sleep Apnea Subtype Detection Method Based on Multimodal Information Fusion
By constructing sleep heterogeneous graphs and combining graph convolutional layers and multi-layer perceptrons, the problem of insufficient space-time dependence of deep learning models when dealing with heterogeneous physiological signals is solved, and more accurate sleep apnea subtype detection is achieved.
Patent Information
- Application Number
- CN202510522731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-24
AI Technical Summary
When existing deep learning models deal with heterogeneous physiological signals, it is difficult to effectively capture the spatiotemporal dependence and intrinsic links of sleep apnea signals, resulting in insufficient accuracy of sleep apnea subtype detection.
The multimodal information fusion method is adopted to construct sleep heterogeneous graphs, feature extraction is performed using residual convolution neural networks and graph convolution layers, and subtype classification is performed in combination with graph attention networks and multi-layer perceptrons to realize spatiotemporal correlation modeling of multimodal signals.
It significantly improves the complementary utilization efficiency of multimodal signals, can more accurately characterize the dynamic changes of sleep breathing events, provide clinically interpretable diagnostic basis, and improves the accuracy of sleep apnea subtype detection.
Smart Images

Figure CN120048499B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medical artificial intelligence and biosignal processing, and particularly to a method for detecting subtypes of sleep apnea based on multimodal information fusion. Background Art
[0002] Sleep Apnea is a quite common sleep disorder disease, which is widely affecting the health of many people around the world. According to the statistical data of relevant professional research, among the adult population, the incidence of sleep apnea ranges from 10% to 30%. Moreover, with the acceleration of the population aging process and the continuous increase in the number of obese people, its incidence shows a continuous upward trend. For clinical diagnosis and the formulation of precise treatment plans, the subtype classification of sleep apnea is of crucial significance. Specifically, different subtypes of sleep apnea, such as Obstructive Sleep Apnea (OSA), Central Sleep Apnea (CSA), and Mixed Sleep Apnea (MSA), have obvious differences in terms of pathogenesis, pathological characteristics, and the impact on the body, which also leads to significantly different corresponding treatment methods.
[0003] In recent years, deep learning models represented by Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM) have been extremely widely applied in the field of sleep apnea detection, and have effectively improved the detection performance in this field to a certain extent. Nevertheless, these deep learning models still have obvious limitations, which are prominently manifested as insufficient ability to process signal heterogeneity and lack of modeling of spatio-temporal dependence. Specifically, physiological signals of different modalities, such as nasal airflow, chest movement, and blood oxygen saturation, each have unique sampling frequencies, amplitude ranges, and noise characteristics. When models such as CNN and LSTM process these heterogeneous signals, due to the large differences between the signals, it is difficult for the models to fully explore the potential internal connections between these signals and the key features they contain. At the same time, the spatio-temporal dependence presented by sleep breathing signals is extremely complex, which is reflected in the dynamic changes of the signals in the time dimension and the mutual correlation in the space dimension. However, the existing deep learning models currently have many technical bottlenecks in capturing these intricate spatio-temporal relationships, and it is difficult to comprehensively and accurately depict the entire dynamic change process of sleep breathing events from occurrence to development, thereby affecting the accurate detection and judgment of sleep apnea conditions. Summary of the Invention
[0004] In view of the above situation, the main objective of the present invention is to propose a method for detecting subtypes of sleep apnea based on multi-modal information fusion to solve the above technical problems.
[0005] The present invention proposes a method for detecting subtypes of sleep apnea based on multi-modal information fusion, and the method includes the following steps:
[0006] Step 1: Obtain the original multi-modal physiological signals, and preprocess the original multi-modal physiological signals to obtain a preprocessed signal window;
[0007] Step 2: Perform node definition operations on the preprocessed signal window to obtain a node set. Based on the node set, construct homogeneous edges and heterogeneous edges respectively to obtain a homogeneous edge set and a heterogeneous edge set. The node set, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph;
[0008] Step 3: Based on the node set, use a residual convolutional neural network (ResCNN) to extract features from the signal window corresponding to the node to obtain node embedding features;
[0009] Step 4: Input the sleep heterogeneous graph into a graph convolutional layer for processing to obtain an updated node feature vector;
[0010] Step 5: Concatenate the updated node feature vectors to obtain a concatenated feature matrix. Construct a subtype classification module based on the multi-head self-attention mechanism and a multi-layer perceptron, and use the subtype classification module to process the concatenated feature matrix to obtain classification results for different subtypes of sleep apnea.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0012] 1. The present invention uses a heterogeneous graph structure to encode the spatio-temporal correlations of multi-modal physiological signals, constructs cross-modal connections through desaturation characteristic edges and mutual information edges, breaks through the limitations of traditional models that only rely on a single modality or simple concatenation, and significantly improves the utilization efficiency of multi-modal signal complementarity;
[0013] 2. The present invention designs a hybrid convolutional layer by combining a graph attention network convolutional layer and a graph sampling and aggregation network convolutional layer, which can process the intra-modal time continuity and the inter-modal spatial correlation respectively. Compared with traditional CNN / LSTM that can only capture time series or local spatial features, it realizes cross-modal and cross-level complex relationship modeling through multiple graph convolutional layers;
[0014] 3. The present invention designs graph edge rules based on physiological mechanisms, making the decision-making basis traceable to specific physiological features, with stronger clinical interpretability, providing an intuitive diagnostic basis for clinicians, and solving the pain point of poor interpretability of traditional black-box models.
[0015] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. Description of the Drawings
[0016] Figure 1 It is a flowchart of the steps of the sleep apnea subtype detection method based on multi-modal information fusion proposed by the present invention;
[0017] Figure 2 It is the overall structure diagram of the sleep apnea subtype detection method based on multi-modal information fusion proposed by the present invention. Detailed Embodiments
[0018] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0019] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention. However, it should be understood that the scope of the embodiments of the present invention is not limited by this.
[0020] Please refer to Figure 1 , this embodiment provides a sleep apnea subtype detection method based on multi-modal information fusion, and the method includes the following steps:
[0021] Step 1: Obtain the original multi-modal physiological signals, preprocess the original multi-modal physiological signals, and obtain the preprocessed signal window.
[0022] In Step 1, obtaining the original multi-modal physiological signals and preprocessing the original multi-modal physiological signals to obtain the preprocessed signal window specifically includes the following sub-steps:
[0023] Obtain the original multi-modal physiological signals, screen the original multi-modal physiological signals, and obtain the blood oxygen saturation signal, nasal airflow signal, and chest movement signal;
[0024] Use the interpolation technique to uniformly resample the blood oxygen saturation signal, nasal airflow signal, and chest movement signal to 20HZ to obtain the resampled blood oxygen saturation signal, resampled nasal airflow signal, and resampled chest movement signal;
[0025] It should be noted that after resampling the signals, it can ensure the alignment in the time dimension.
[0026] The resampled nasal airflow signal and the resampled chest movement signal are filtered using a fourth-order Butterworth band-pass filter to obtain the filtered nasal airflow signal and the filtered chest movement signal. The following relational expressions exist in the corresponding process:
[0027] ;
[0028] Among them, represents the transfer function of the fourth-order Butterworth band-pass filter, represents the upper corner frequency of the fourth-order band-pass filter, represents the lower corner frequency of the fourth-order band-pass filter, represents the center corner frequency of the fourth-order band-pass filter and , represents the bandwidth and , represents the complex frequency variable;
[0029] It should be noted that for the nasal airflow signal and the chest movement signal, a fourth-order Butterworth band-pass filter is used for filtering, and the cut-off frequency is set to 2 Hz. This filter can effectively suppress high-frequency noise and retain the characteristics of mid-low frequency signals related to respiratory events. High-frequency noise may be generated by factors such as environmental interference and equipment errors. After filtering, the signal can be made smoother and more stable, highlighting the effective information.
[0030] The resampled blood oxygen saturation signal, the filtered nasal airflow signal, and the filtered chest movement signal are standardized to obtain the standardized blood oxygen saturation signal, the standardized nasal airflow signal, and the standardized chest movement signal. The following relational expressions exist in the corresponding process:
[0031] ;
[0032] represents the standardized signal, represents the input signal, represents the data mean, represents the data standard deviation;
[0033] It should be noted that by standardizing the signal, the values of different signals are unified to the same scale range, eliminating the influence caused by differences in dimension and numerical distribution between signals.
[0034] The standardized blood oxygen saturation signal, the standardized nasal airflow signal, and the standardized chest movement signal are subjected to data screening and balancing to obtain the preprocessed signal window.
[0035] It should be noted that in actual physiological recordings, there may be some windows that contain a large number of null values. Such windows need to be discarded to ensure data quality. At the same time, the number of normal signal windows is usually much larger than that of respiratory event windows. This data imbalance will affect the detection effect of respiratory events. To solve this problem, under-sampling is performed on the normal signal windows, and the normal windows close to the respiratory event windows are preferentially selected. This can not only reduce the number of normal class data and balance the dataset, but also retain the context information related to normal signals and respiratory events, and improve the recognition ability of respiratory events.
[0036] Step 2: Perform node definition operations on the preprocessed signal windows to obtain a node set. Based on the node set, homogeneous edge construction and heterogeneous edge construction are respectively carried out to obtain a homogeneous edge set and a heterogeneous edge set. The node set, homogeneous edge set, and heterogeneous edge set constitute a sleep heterogeneous graph.
[0037] Please refer to Figure 2 , in Step 2, perform node definition operations on the preprocessed signal windows to obtain a node set. Based on the node set, homogeneous edge construction and heterogeneous edge construction are respectively carried out to obtain a homogeneous edge set and a heterogeneous edge set. The node set, homogeneous edge set, and heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps:
[0038] Based on the preprocessed signal windows, each 20-second signal window is defined as a node and the node types are divided according to the signal channels to obtain a node set;
[0039] It should be noted that the node types are divided into blood oxygen saturation nodes, nasal airflow nodes, and chest movement nodes according to the signal channels.
[0040] Based on the node set, the nodes within the same signal channel are fully connected to obtain a homogeneous edge set;
[0041] It should be noted that in order to capture the local time continuity of the signals within the same channel, the nodes within the same channel are connected in a fully connected manner. The adjacent window nodes within the blood oxygen saturation channel are fully connected, the adjacent window nodes within the nasal airflow channel are fully connected, and the adjacent window nodes within the chest movement channel are fully connected.
[0042] Based on the node set, the respiratory event nodes are connected to the nodes corresponding to the lowest blood oxygen saturation values to obtain a cross-channel edge set. The following relational expressions exist in the corresponding process:
[0043] ;
[0044] Among them, represents the cross-channel edge set, represents the node corresponding to the central event window characterizing the critical respiratory event moment in the nasal airflow or chest movement channel. Nodes corresponding to the nasal airflow channel Nodes corresponding to the chest movement channel Nodes corresponding to the lowest oxygen saturation value within the corresponding sample window in the blood oxygen saturation channel Represents the lowest blood sample saturation value
[0045] Based on the node set and a given mutual information threshold, calculate the mutual information value between the nasal airflow node and the chest movement node, and construct mutual information edges for the nodes corresponding to the mutual information values that exceed the mutual information threshold to obtain a set of mutual information edges. The following relational expressions exist in the corresponding process:
[0046] ;
[0047] Among them, Represents the set of mutual information edges Represents the mutual information between the nasal airflow and the chest movement nodes Represents the mutual information threshold;
[0048] It should be noted that through the above connection method, it can ensure that only those connections with significant information associations are retained in the graph structure, avoiding the introduction of too many meaningless edges. Through the edges constructed based on mutual information, the co-variation relationship between the nasal airflow and chest movement signals can be captured, further enriching the multi-modal relationships expressed by the graph structure.
[0049] The combination of the cross-channel edge set and the mutual information edge set constitutes a heterogeneous edge set.
[0050] It should be noted that in Appendix Figure 2 ,[[]] Represents blood oxygen saturation, GATConv represents a convolutional operation based on a graph attention network for processing information transmission of homogeneous edges; SAGEConv represents a convolutional operation based on a graph sampling and aggregation network for processing information aggregation of heterogeneous edges.
[0051] Step 3: Based on the node set, use a residual convolutional neural network to extract features from the signal windows corresponding to the nodes to obtain node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph.
[0052] In Step 3, based on the node set, use a residual convolutional neural network to extract features from the signal windows corresponding to the nodes to obtain node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps:
[0053] S101. Input the signal window corresponding to the node into the first residual block, and successively perform two - layer one - dimensional convolutional layer processing, batch normalization processing, and LeakyReLU function processing to obtain the output of the second layer of the first residual block. The following relational expressions exist in the corresponding process:
[0054] ;
[0055] Among them, represents the output of the first layer of the first residual block, represents the output of the second layer of the first residual block, represents the signal window corresponding to the node, represents the convolutional weight of the first layer in the first residual block, represents the bias of the first layer in the first residual block, represents the convolutional layer weight of the second layer in the first residual block, represents the bias of the second layer in the first residual block, represents after one - dimensional convolutional layer processing, represents after activation function processing, represents after batch normalization processing.
[0056] S102. Perform a residual skip connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block. The following relational expressions exist in the corresponding process:
[0057] ;
[0058] Among them, represents the output of the first residual block.
[0059] Iteratively repeat the steps of S101 and S102 for the output of the first residual block to obtain the output of the third residual block. Perform global average pooling operation on the output features of the third residual block to obtain the node embedding features. The following relational expressions exist in the corresponding process:
[0060] ;
[0061] Among them, represents the node embedding features, represents the number of time steps, represents the output features of the 3rd residual block at the th time step.
[0062] It should be noted that each node in the sleep heterogeneous graph corresponds to a node embedding feature respectively. The residual convolutional neural network adopts a hierarchical structure, including three residual blocks. Each residual block contains two layers, and each layer includes a one-dimensional convolutional layer, a batch normalization layer, and a LeakyReLU activation function. The kernel size of the one-dimensional convolutional layer is set to 3, and the stride is set to 1. The residual skip connection can alleviate the problem of gradient disappearance, enabling the network to learn multi-scale time patterns.
[0063] Step 4: Input the sleep heterogeneous graph into the graph convolutional layer for processing to obtain an updated node feature vector.
[0064] In step 4, inputting the sleep heterogeneous graph into the graph convolutional layer for processing to obtain an updated node feature vector specifically includes the following sub-steps:
[0065] Based on the graph attention network convolutional layer, use the LeakyReLU function to calculate the attention coefficients of the node embedding features of homogeneous edges, obtaining the attention coefficients of homogeneous edge nodes. There are the following relational expressions in the corresponding process:
[0066] ;
[0067] Among them, represents the attention coefficient between node and its neighbor node , represents the feature vector of node , represents the feature vector of the neighbor node of node , represents the attention mechanism parameter, represents the learnable weight matrix, represents the vector concatenation operation, represents the homogeneous edge node, represents node 's homogeneous edge neighbor node.
[0068] Use the softmax function to normalize the attention coefficients of homogeneous edge nodes, obtaining the normalized attention coefficients of homogeneous edge nodes. There are the following relational expressions in the corresponding process:
[0069] ;
[0070] Among them, represents the normalized attention coefficient between node and its neighbor node , represents being subjected to normalization processing.
[0071] The features of neighbor nodes are weighted and aggregated using the attention coefficients after homogeneous edge node normalization to obtain the updated homogeneous edge node features. The following relational expressions exist in the corresponding process:
[0072] ;
[0073] Among them, represents the feature vector of the updated node . represents the set of neighbor nodes . represents the activation function.
[0074] Based on the graph sampling and aggregation network convolutional layer, neighbor nodes of heterogeneous edge nodes are sampled to obtain the set of sampled neighbor nodes. Based on the set of sampled neighbor nodes, the mean aggregation function is used to aggregate the feature vectors of the sampled neighbor nodes to obtain the aggregated features of the neighbor nodes. The following relational expressions exist in the corresponding process:
[0075] ;
[0076] Among them, represents the aggregated features of the neighbor nodes, represents the feature vector of the sampled neighbor node . represents the sampled neighbor node, represents the set of sampled neighbor nodes, represents being processed by the mean aggregation function.
[0077] It should be noted that for each node associated with a heterogeneous edge, a fixed number of neighbor nodes are sampled separately from the sets of neighbor nodes of different types corresponding to this node.
[0078] The feature vectors of heterogeneous edge nodes and the aggregated features of neighbor nodes are linearly transformed respectively and input into the ReLU function for processing to obtain the updated heterogeneous edge node features. The following relational expressions exist in the corresponding process:
[0079] ;
[0080] Among them, represents the feature vector of the updated heterogeneous edge node , represents being processed by the rectified linear unit activation function, represents the feature vector of the heterogeneous edge node , represents the learnable weight matrix for linearly transforming the feature vector of the heterogeneous edge node, A learnable weight matrix for linearly transforming the aggregated features of neighbor nodes.
[0081] Perform a cross-edge type minimum aggregation operation on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain the updated node feature vector. The following relational expressions exist in the corresponding process:
[0082] ;
[0083] Among them, represents the updated node feature vector, represents the minimum value operation.
[0084] It should be noted that the graph convolutional layer contains 4 layers. By stacking these layers, messages can be passed and fused multiple times in the graph, gradually mining more complex and in-depth relationships in the graph. Each layer of the graph convolutional layer will further update the node's embedding representation on the basis of the previous layer, so as to capture richer information.
[0085] Step 5: Concatenate the updated node feature vectors to obtain the concatenated feature matrix. Build a subtype classification module based on the multi-head self-attention mechanism and the multi-layer perceptron, and use the subtype classification module to process the concatenated feature matrix to obtain the classification results of different sleep apnea subtypes.
[0086] In Step 5, concatenate the updated node feature vectors to obtain the concatenated feature matrix. Build a subtype classification module based on the multi-head self-attention mechanism and the multi-layer perceptron, and use the subtype classification module to process the concatenated feature matrix to obtain the classification results of different sleep apnea subtypes. Specifically, it includes the following sub-steps:
[0087] Concatenate the updated node feature vectors to obtain the concatenated feature matrix. The following relational expressions exist in the corresponding process:
[0088] ;
[0089] Among them, represents the concatenated feature matrix, both represent the updated node feature vectors, represents the number of all nodes in the sleep heterogeneous graph.
[0090] Based on the multi-head self-attention mechanism, perform a linear transformation on the concatenated feature matrix with each attention head respectively to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head. The following relational expressions exist in the corresponding process:
[0091] ;
[0092] Among them, represents the query matrix of the -th attention head, represents the key matrix of the -th attention head, represents the value matrix of the -th attention head, represents the weight matrix for generating the query matrix of the -th attention head, represents the weight matrix for generating the key matrix of the -th attention head, represents the weight matrix for generating the value matrix of the -th attention head, represents the index of the attention head, represents the number of attention heads.
[0093] The query matrix of the attention head and the key matrix of the attention head are processed using the Softmax function to obtain the attention score matrix. There is the following relational expression in the corresponding process:
[0094] ;
[0095] Among them, represents the -th attention score matrix, represents the dimension of the attention head.
[0096] The value matrix of the attention head and the attention score matrix are weighted and fused to obtain the output of the attention head. There is the following relational expression in the corresponding process:
[0097] ;
[0098] Among them, represents the output of the -th attention head;
[0099] The outputs of the attention heads are concatenated to obtain the concatenated matrix, and the concatenated matrix is linearly transformed to obtain the final graph feature. There is the following relational expression in the corresponding process:
[0100] ;
[0101] Among them, represents the concatenated matrix, both represent the outputs of the attention heads, represents the final graph feature, represents the learnable weight matrix for linearly transforming the concatenated matrix.
[0102] In the step of sequentially performing a fully connected layer transformation process and an activation process on the final graph features to obtain the output of the fully connected layer, the following relational expressions exist in the corresponding process:
[0103] ;
[0104] Among them, represents the output of the first fully connected layer, represents the weight matrix of the first layer of the fully connected layer, represents the bias vector of the first layer of the fully connected layer, represents the output of the th layer of the fully connected layer, represents the output of the th layer of the fully connected layer, represents the weight matrix of the th layer of the fully connected layer, represents the bias vector of the th layer of the fully connected layer, represents the total number of layers of the fully connected layer in the multi-layer perceptron.
[0105] Input the output of the last fully connected layer into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes. The following relational expressions exist in the corresponding process:
[0106] ;
[0107] Among them, represents the classification results of different sleep apnea subtypes.
[0108] It should be noted that the classification results of different sleep apnea subtypes are a probability distribution matrix. Each row in the matrix corresponds to the predicted probability of each category (normal, hypopnea, obstructive sleep apnea, central sleep apnea, and mixed sleep apnea). Select the category with the highest probability value as the final predicted sleep breathing event type, thereby completing the classification of sleep breathing events and providing an accurate judgment basis for the detection of sleep apnea subtypes.
[0109] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the sequence indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0110] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0111] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for detecting subtypes of sleep apnea based on multimodal information fusion, characterized in that, The method includes the following steps: Step 1: Obtain the original multimodal physiological signals, preprocess the original multimodal physiological signals to obtain a preprocessed signal window; Step 2: Perform node definition operations on the preprocessed signal window to obtain a set of nodes, and respectively construct homogeneous edges and heterogeneous edges based on the set of nodes to obtain a homogeneous edge set and a heterogeneous edge set; Step 3: Based on the set of nodes, use a residual convolutional neural network to extract features from the signal window corresponding to the nodes to obtain node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph; Step 4: Input the sleep heterogeneous graph into a graph convolutional layer for processing to obtain an updated node feature vector; Step 5: Concatenate the updated node feature vectors to obtain a concatenated feature matrix. Construct a subtype classification module based on the multi-head self-attention mechanism and a multi-layer perceptron, and use the subtype classification module to process the concatenated feature matrix to obtain classification results for different sleep apnea subtypes.
2. The method for detecting subtypes of sleep apnea based on multi-modal information fusion according to claim 1, wherein In the said Step 1, obtaining the original multimodal physiological signals, preprocessing the original multimodal physiological signals to obtain a preprocessed signal window specifically includes the following sub-steps: Obtain the original multimodal physiological signals, screen the original multimodal physiological signals to obtain a blood oxygen saturation signal, a nasal airflow signal, and a chest movement signal; Use interpolation technology to uniformly resample the blood oxygen saturation signal, the nasal airflow signal, and the chest movement signal to 20HZ to obtain a resampled blood oxygen saturation signal, a resampled nasal airflow signal, and a resampled chest movement signal; Use a fourth-order Butterworth band-pass filter to filter the resampled nasal airflow signal and the resampled chest movement signal to obtain a filtered nasal airflow signal and a filtered chest movement signal; Perform normalization processing on the resampled blood oxygen saturation signal, the filtered nasal airflow signal, and the filtered chest movement signal to obtain a normalized blood oxygen saturation signal, a normalized nasal airflow signal, and a normalized chest movement signal; Perform data screening and balancing on the normalized blood oxygen saturation signal, the normalized nasal airflow signal, and the normalized chest movement signal to obtain a preprocessed signal window.
3. The method for detecting subtypes of sleep apnea based on multimodal information fusion according to claim 2, wherein In the said Step 2, performing node definition operations on the preprocessed signal window to obtain a set of nodes, and respectively constructing homogeneous edges and heterogeneous edges based on the set of nodes to obtain a homogeneous edge set and a heterogeneous edge set. The set of nodes, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph, specifically including the following sub-steps: Based on the preprocessed signal window, define each 20-second signal window as a node and divide the node types according to the signal channels to obtain a set of nodes; Based on the set of nodes, fully connect the nodes within the same signal channel to obtain a homogeneous edge set; Based on the set of nodes, connect the respiratory event nodes to the nodes with the lowest blood oxygen saturation value to obtain a cross-channel edge set. The following relationship exists in the corresponding process: ; Among them, represents the cross-channel edge set, represents the node corresponding to the central event window characterizing the key respiratory event moment in the nasal airflow or chest movement channel, represents the node corresponding to the nasal airflow channel, represents the node corresponding to the chest movement channel, represents the node corresponding to the lowest oxygen saturation value within the sample window in the blood oxygen saturation channel, represents the lowest blood sample saturation value; Based on the node set and a given mutual information threshold, calculate the mutual information value between the nasal airflow node and the chest movement node, construct mutual information edges for the nodes corresponding to the mutual information values exceeding the mutual information threshold to obtain a set of mutual information edges. The following relational expressions exist in the corresponding process: ; Among them, represents the set of mutual information edges, represents the mutual information between the nasal airflow and the chest movement nodes, represents the mutual information threshold; The combination of the cross-channel edge set and the mutual information edge set constitutes the heterogeneous edge set.
4. The method for detecting subtypes of sleep apnea based on multimodal information fusion according to claim 3, wherein In step 3, based on the node set, use a residual convolutional neural network to extract features from the signal window corresponding to the node to obtain node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps: S101. Input the signal window corresponding to the node into the first residual block and perform two-layer one-dimensional convolutional layer processing, batch normalization processing, and LeakyReLU function processing in sequence to obtain the output of the second layer of the first residual block; S102. Perform a residual skip connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block; Iteratively repeat the steps of S101 and S102 for the output of the first residual block to obtain the output of the third residual block. Perform global average pooling on the output features of the third residual block to obtain node embedding features; The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph.
5. The method for detecting subtypes of sleep apnea based on multimodal information fusion according to claim 4, wherein Input the signal window corresponding to the node into the first residual block and perform two-layer one-dimensional convolutional layer processing, batch normalization processing, and LeakyReLU function processing in sequence to obtain the output of the second layer of the first residual block. The following relational expressions exist in the corresponding process: ; Among them, represents the output of the first layer of the first residual block, represents the output of the second layer of the first residual block, represents the signal window corresponding to the node, represents the convolution weight of the first layer in the first residual block, represents the bias of the first layer in the first residual block, represents the convolution layer weight of the second layer in the first residual block, represents the bias of the second layer in the first residual block, represents being processed by the one-dimensional convolution layer, represents being processed by the activation function, represents being processed by batch normalization; In the step of performing a residual skip connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block, the following relational expressions exist in the corresponding process: ; Among them, represents the output of the first residual block; In the step of iteratively repeating the steps of S101 and S102 for the output of the first residual block to obtain the output of the third residual block and performing global average pooling on the output features of the third residual block to obtain node embedding features, the following relational expressions exist in the corresponding process: ; Among them, represents the node embedding feature, represents the number of time steps, represents the output feature of the 3rd residual block at the th time step.
6. The method for detecting sleep apnea subtypes based on multi-modal information fusion according to claim 5, wherein In step 4, input the sleep heterogeneous graph into the graph convolutional layer for processing to obtain updated node feature vectors, which specifically includes the following sub-steps: Based on the graph attention network convolutional layer, use the LeakyReLU function to calculate the attention coefficients of the node embedding features of the homogeneous edges to obtain the attention coefficients of the homogeneous edge nodes; Use the softmax function to normalize the attention coefficients of the homogeneous edge nodes to obtain the normalized attention coefficients of the homogeneous edge nodes; Use the normalized attention coefficients of the homogeneous edge nodes to perform weighted aggregation on the features of the neighbor nodes to obtain updated homogeneous edge node features; Based on the graph sampling and aggregation network convolutional layer, sample the neighbor nodes of the heterogeneous edge nodes to obtain a set of sampled neighbor nodes. Based on the set of sampled neighbor nodes, use the mean aggregation function to perform an aggregation operation on the feature vectors of the sampled neighbor nodes to obtain the aggregated features of the neighbor nodes; Perform linear transformations on the feature vectors of heterogeneous edge nodes and the aggregated features of neighbor nodes respectively, and input them into the ReLU function for processing to obtain the updated heterogeneous edge node features; Perform a cross-edge type minimum aggregation operation on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain the updated node feature vector.
7. The method for detecting subtypes of sleep apnea based on multi-modal information fusion according to claim 6, characterized in that, Based on the graph attention network convolutional layer, use the LeakyReLU function to calculate the attention coefficients of the node embedding features of homogeneous edges, obtaining the attention coefficients of homogeneous edge nodes. The following relational expressions exist in the corresponding process: ; Among them, represents the node and its attention coefficient with its neighbor nodes ; represents the node 's feature vector, represents the node 's neighbor node 's feature vector, represents the attention mechanism parameter, represents the learnable weight matrix, represents the vector concatenation operation, represents the homogeneous edge node, represents the node 's homogeneous edge neighbor node; Use the softmax function to normalize the attention coefficients of homogeneous edge nodes, obtaining the normalized attention coefficients of homogeneous edge nodes. The following relational expressions exist in the corresponding process: ; Among them, represents a node and its neighbor nodes after normalizing the attention coefficient, indicating that it has been normalized; Use the normalized attention coefficients to perform weighted aggregation on the features of neighbor nodes to obtain the updated homogeneous edge node features. The following relational expressions exist in the corresponding process: ; Among them, represents the feature vector of the updated node , represents the set of neighbor nodes , represents the activation function; Based on the graph sampling and aggregation network convolutional layer, sample the neighbor nodes of heterogeneous edge nodes to obtain the set of sampled neighbor nodes. Based on the set of sampled neighbor nodes, use the mean aggregation function to perform an aggregation operation on the feature vectors of the sampled neighbor nodes to obtain the aggregated features of the neighbor nodes. The following relational expressions exist in the corresponding process: ; Among them, represents the aggregated features of neighbor nodes, represents the sampled neighbor nodes 's feature vectors, represents the sampled neighbor nodes, represents the set of neighbor nodes after sampling, represents being processed by the mean aggregation function; Perform linear transformations on the feature vectors of heterogeneous edge nodes and the aggregated features of neighbor nodes respectively, and input them into the ReLU function for processing to obtain the updated heterogeneous edge node features. The following relational expressions exist in the corresponding process: ; Among them, represents the updated heterogeneous edge node 's feature vector, represents being processed by the rectified linear unit activation function, represents the heterogeneous edge node 's feature vector, represents a learnable weight matrix for linearly transforming the feature vector of the heterogeneous edge node, represents a learnable weight matrix for linearly transforming the aggregated features of the neighbor nodes; Perform a cross-edge type minimum aggregation operation on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain the updated node feature vector. The following relational expressions exist in the corresponding process: ; Among them, represents the updated node feature vector, represents the minimum value operation.
8. The method for detecting sleep apnea subtype based on multi-modal information fusion according to claim 7, wherein In step 5, splice the updated node feature vectors to obtain a spliced feature matrix. Based on the multi-head self-attention mechanism and the multi-layer perceptron, construct a subtype classification module, and use the subtype classification module to process the spliced feature matrix to obtain the classification results of different sleep apnea subtypes. The specific steps are as follows: Splice the updated node feature vectors to obtain a spliced feature matrix; Based on the multi-head self-attention mechanism, perform linear transformations on the spliced feature matrix with each attention head respectively to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head; Use the Softmax function to process the query matrix of the attention head and the key matrix of the attention head to obtain the attention score matrix; Perform weighted fusion on the value matrix of the attention head and the attention score matrix to obtain the output of the attention head; Splice the outputs of the attention heads to obtain a spliced matrix, and perform a linear transformation on the spliced matrix to obtain the final graph feature; Perform full connection layer transformation processing and activation processing on the final graph feature in sequence to obtain the output of the full connection layer; Input the output of the last full connection layer into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes.
9. The method for detecting subtypes of sleep apnea based on multimodal information fusion according to claim 8, wherein Splice the updated node feature vectors to obtain a spliced feature matrix. The following relational expressions exist in the corresponding process: ; Among them, represents the spliced feature matrix, both represent the updated node feature vectors, represents the number of all nodes in the sleep heterogeneous graph; In the step of performing linear transformation on the concatenated feature matrix with each attention head respectively based on the multi-head self-attention mechanism to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head, the following relational expressions exist in the corresponding process: ; Among them, represents the query matrix of the th attention head, represents the key matrix of the th attention head, represents the value matrix of the th attention head, represents the weight matrix for generating the query matrix of the th attention head, represents the weight matrix for generating the key matrix of the th attention head, represents the weight matrix for generating the value matrix of the th attention head, represents the index of the attention head, represents the number of attention heads; In the step of using the Softmax function to process the query matrix of the attention head and the key matrix of the attention head to obtain the attention score matrix, the following relational expressions exist in the corresponding process: ; Among them, represents the th attention score matrix, represents the dimension of the attention head; In the step of performing weighted fusion on the value matrix of the attention head and the attention score matrix to obtain the output of the attention head, the following relational expressions exist in the corresponding process: ; Among them, represents the output of the th attention head; In the step of concatenating the outputs of the attention heads to obtain a concatenated matrix and performing linear transformation on the concatenated matrix to obtain the final graph feature, the following relational expressions exist in the corresponding process: ; Among them, represents the concatenated matrix, both represent the output of the attention head, represents the final graph feature, represents the learnable weight matrix for linearly transforming the concatenated matrix; In the step of sequentially performing fully connected layer transformation processing and activation processing on the final graph feature to obtain the output of the fully connected layer, the following relational expressions exist in the corresponding process: ; Among them, represents the output of the first fully connected layer, represents the weight matrix of the first layer of the fully connected layer, represents the bias vector of the first layer of the fully connected layer, represents the output of the th fully connected layer, output of the th fully connected layer, represents the weight matrix of the th layer of the fully connected layer, represents the bias vector of the th layer of the fully connected layer; represents the total number of layers of the fully connected layer in the multi-layer perceptron; In the step of inputting the output of the last fully connected layer into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes, the following relational expressions exist in the corresponding process: ; Among them, represents the classification results of different sleep apnea subtypes.
Citation Information
Patent Citations
Multi-mode multi-resolution sleep apnea automatic detection method and system
CN117838053A
Sleep apnea detection method and device based on feature fusion
CN118155826A