Sleep apnea subtype detection method based on multi-modal information fusion

By constructing sleep heterogeneous graphs and using graph convolutional layers and other technical means, the technical bottlenecks of existing models when dealing with sleep apnea multimodal signals are solved, and more accurate sleep apnea subtype detection is achieved.

CN120048499AActive Publication Date: 2025-05-27JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Patent Information

Application Number
CN202510522731.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

When the existing deep learning model deals with multimodal physiological signals of sleep apnea, it is difficult to fully explore the intrinsic connections and key features between signals, and there are technical bottlenecks when capturing complex spatiotemporal relationships, which affects the accurate detection and judgment of sleep apnea.

Method used

The sleep apnea subtype detection method based on multimodal information fusion is adopted. By obtaining the original multimodal physiological signals for preprocessing, sleep heterogeneous patterns are constructed, and features are extracted using residual convolutional neural network and graph convolutional layer. The subtype classification module is constructed in combination with the multi-head self-attention mechanism and multi-layer perceptron to realize the classification of different sleep apnea subtypes.

Benefits of technology

It significantly improves the complementary utilization efficiency of multimodal signals, can capture the temporal and spatial relationships of sleep breathing events more comprehensively and accurately, and improves the accuracy and clinical explanatory nature of sleep apnea subtype detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048499A_ABST
    Figure CN120048499A_ABST
Patent Text Reader

Abstract

The invention provides a sleep apnea subtype detection method based on multi-modal information fusion, and the method comprises the steps: obtaining an original multi-modal physiological signal, carrying out the preprocessing of the original multi-modal physiological signal, obtaining a preprocessed signal window, carrying out the node definition operation of the preprocessed signal window, obtaining a node set, and carrying out the detection of the sleep apnea subtype. Performing homogeneous edge construction and heterogeneous edge construction based on the node set to obtain a homogeneous edge set and a heterogeneous edge set respectively, and performing feature extraction on signal windows corresponding to nodes by using a residual convolutional neural network based on the node set to obtain node embedding features; the node embedding features, the homogeneous edge set and the heterogeneous edge set form a sleep heterogeneous graph. According to the method, the drawing edge rule is designed based on the physiological mechanism, so that the decision basis can be traced to specific physiological features, the clinical interpretability is higher, an intuitive diagnosis basis is provided for clinicians, and the pain point of poor interpretability of a traditional black box model is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical artificial intelligence and bio-signal processing, and particularly to a method for detecting subtypes of sleep apnea based on multi-modal information fusion. Background Art

[0002] Sleep Apnea is a quite common sleep disorder disease, which is widely affecting the health of numerous people around the world. According to the statistical data of relevant professional research, among the adult population, the incidence rate of sleep apnea ranges from 10% to 30%. Moreover, with the acceleration of the population aging process and the continuous increase in the number of obese people, its incidence rate shows a continuous upward trend. For clinical diagnosis and formulating precise treatment plans, the subtype classification of sleep apnea is of crucial significance. Specifically, different subtypes of sleep apnea, such as Obstructive Sleep Apnea (OSA), Central Sleep Apnea (CSA), and Mixed Sleep Apnea (MSA), have obvious differences in terms of pathogenesis, pathological characteristics, and the impact on the body, which also leads to significantly different corresponding treatment methods.

[0003] In recent years, deep learning models represented by Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM) have been extremely widely applied in the field of sleep apnea detection, and have effectively improved the detection performance in this field to a certain extent. Nevertheless, these deep learning models still have obvious limitations, prominently manifested as insufficient ability to handle signal heterogeneity and lack of modeling of spatio-temporal dependence. Specifically, physiological signals of different modalities, such as nasal airflow, chest movement, and blood oxygen saturation, each have unique sampling frequencies, amplitude ranges, and noise characteristics. When models such as CNN and LSTM process these heterogeneous signals, due to the large differences between the signals, it is difficult for the models to fully explore the potential internal connections between these signals and the key features they contain. At the same time, the spatio-temporal dependence presented by sleep breathing signals is extremely complex, which is reflected in the dynamic changes of the signals in the time dimension and the mutual correlation in the space dimension. However, the existing deep learning models currently have many technical bottlenecks in capturing these intricate spatio-temporal relationships, and it is difficult to comprehensively and accurately depict the entire dynamic change process of sleep breathing events from occurrence to development, thereby affecting the accurate detection and judgment of sleep apnea conditions. Summary of the Invention

[0004] In view of the above situation, the main object of the present invention is to propose a method for detecting subtypes of sleep apnea based on multi-modal information fusion to solve the above technical problems.

[0005] The present invention proposes a method for detecting subtypes of sleep apnea based on multi-modal information fusion, and the method includes the following steps: Step 1: Obtain the original multi-modal physiological signals, and preprocess the original multi-modal physiological signals to obtain a preprocessed signal window; Step 2: Perform node definition operations on the preprocessed signal window to obtain a node set. Based on the node set, construct homogeneous edges and heterogeneous edges respectively to obtain a homogeneous edge set and a heterogeneous edge set. The node set, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph; Step 3: Based on the node set, use a residual convolutional neural network (ResCNN) to extract features from the signal window corresponding to the nodes to obtain node embedding features; Step 4: Input the sleep heterogeneous graph into a graph convolutional layer for processing to obtain an updated node feature vector; Step 5: Concatenate the updated node feature vectors to obtain a concatenated feature matrix. Based on the multi-head self-attention mechanism and a multi-layer perceptron, construct a subtype classification module, and use the subtype classification module to process the concatenated feature matrix to obtain classification results of different sleep apnea subtypes.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention uses a heterogeneous graph structure to encode the spatio-temporal correlation of multi-modal physiological signals, constructs cross-modal connections through desaturation characteristic edges and mutual information edges, breaks through the limitation of traditional models that only rely on a single modality or simple splicing, and significantly improves the utilization efficiency of multi-modal signal complementarity; 2. Through the design of a hybrid convolutional layer that combines a graph attention network convolutional layer and a graph sampling and aggregation network convolutional layer, the present invention can process the time continuity of the same modality and the spatial correlation of different modalities respectively. Compared with traditional CNN / LSTM that can only capture time series or local spatial features, it realizes cross-modal and cross-level complex relationship modeling through multiple graph convolutional layers; 3. The present invention designs graph edge rules based on physiological mechanisms, making the decision basis traceable to specific physiological characteristics, with stronger clinical interpretability, providing an intuitive diagnostic basis for clinicians, and solving the pain point of poor interpretability of traditional black-box models.

[0007] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0008] Figure 1 The flowchart of the steps of the sleep apnea subtype detection method based on multi-modal information fusion proposed by the present invention; Figure 2 The overall structure diagram of the sleep apnea subtype detection method based on multi-modal information fusion proposed by the present invention. Specific embodiments

[0009] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0010] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention. However, it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0011] Please refer to Figure 1 , this embodiment provides a sleep apnea subtype detection method based on multi-modal information fusion. The method includes the following steps: Step 1: Obtain the original multi-modal physiological signals, and preprocess the original multi-modal physiological signals to obtain a preprocessed signal window.

[0012] In step 1, obtaining the original multi-modal physiological signals and preprocessing the original multi-modal physiological signals to obtain a preprocessed signal window specifically includes the following sub-steps: Obtain the original multi-modal physiological signals, and screen the original multi-modal physiological signals to obtain a blood oxygen saturation signal, a nasal airflow signal, and a chest movement signal; Use the interpolation technique to uniformly resample the blood oxygen saturation signal, the nasal airflow signal, and the chest movement signal to 20HZ to obtain a resampled blood oxygen saturation signal, a resampled nasal airflow signal, and a resampled chest movement signal; It should be noted that after resampling the signals, it can ensure the alignment in the time dimension.

[0013] Use a fourth-order Butterworth band-pass filter to filter the resampled nasal airflow signal and the resampled chest movement signal to obtain a filtered nasal airflow signal and a filtered chest movement signal. The following relationship exists in the corresponding process: ; Among them, represents the transfer function of the fourth-order Butterworth band-pass filter, represents the upper corner frequency of the fourth-order band-pass filter, represents the lower corner frequency of the fourth-order band-pass filter, represents the center corner frequency of the fourth-order band-pass filter and , represents the bandwidth and , represents the complex frequency variable; It should be noted that for the nasal airflow signal and the chest movement signal, a fourth-order Butterworth band-pass filter is used for filtering, and the cut-off frequency is set to 2 Hz. This filter can effectively suppress high-frequency noise and retain the mid-low frequency signal characteristics related to respiratory events. High-frequency noise may be generated by factors such as environmental interference and equipment errors. After filtering, the signal can be made smoother and more stable, highlighting the effective information.

[0014] The resampled blood oxygen saturation signal, the filtered nasal airflow signal, and the filtered chest movement signal are standardized to obtain the standardized blood oxygen saturation signal, the standardized nasal airflow signal, and the standardized chest movement signal. The following relational expressions exist in the corresponding process: ; represents the standardized signal, represents the input signal, represents the data mean, represents the data standard deviation; It should be noted that by standardizing the signal, the values of different signals are unified to the same scale range, eliminating the influence caused by differences in dimension and numerical distribution between signals.

[0015] The standardized blood oxygen saturation signal, the standardized nasal airflow signal, and the standardized chest movement signal are subjected to data screening and balancing to obtain the preprocessed signal window.

[0016] It should be noted that in actual physiological recordings, there may be some windows that contain a large number of null values. Such windows need to be discarded to ensure data quality. At the same time, the number of normal signal windows is usually much larger than that of respiratory event windows. This data imbalance will affect the detection effect of respiratory events. To solve this problem, the normal signal windows are undersampled, and the normal windows close to the respiratory event windows are preferentially selected. This can not only reduce the number of normal class data, balance the dataset, but also retain the context information related to normal signals and respiratory events, improving the recognition ability of respiratory events.

[0017] Step 2: Perform node definition operations on the preprocessed signal window to obtain a node set. Based on the node set, perform homogeneous edge construction and heterogeneous edge construction respectively to obtain a homogeneous edge set and a heterogeneous edge set. The node set, homogeneous edge set, and heterogeneous edge set constitute a sleep heterogeneous graph.

[0018] Please refer to Figure 2 , in Step 2, perform node definition operations on the preprocessed signal window to obtain a node set. Based on the node set, perform homogeneous edge construction and heterogeneous edge construction respectively to obtain a homogeneous edge set and a heterogeneous edge set. The node set, homogeneous edge set, and heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps: Based on the preprocessed signal window, define each 20-second signal window as a node and divide the node types according to the signal channels to obtain a node set; It should be noted that the node types are divided into blood oxygen saturation nodes, nasal airflow nodes, and chest movement nodes according to the signal channels.

[0019] Based on the node set, fully connect the nodes within the same signal channel to obtain a homogeneous edge set; It should be noted that in order to capture the local temporal continuity of the signals within the same channel, the nodes within the same channel are connected in a fully connected manner. The adjacent window nodes within the blood oxygen saturation channel are fully connected, the adjacent window nodes within the nasal airflow channel are fully connected, and the adjacent window nodes within the chest movement channel are fully connected.

[0020] Based on the node set, connect the respiratory event nodes with the nodes corresponding to the lowest blood oxygen saturation values to obtain a cross-channel edge set. The following relationship exists during the corresponding process: ; Among them, represents the cross-channel edge set, represents the node corresponding to the central event window characterizing the critical respiratory event moment in the nasal airflow or chest movement channel, represents the node corresponding to the nasal airflow channel, represents the node corresponding to the chest movement channel, represents the node corresponding to the lowest oxygen saturation value within the corresponding sample window in the blood oxygen saturation channel, represents the lowest blood sample saturation value.

[0021] Based on the node set and given a mutual information threshold, calculate the mutual information value between the nasal airflow nodes and the chest movement nodes, and construct mutual information edges for the nodes corresponding to the mutual information values exceeding the mutual information threshold to obtain a mutual information edge set. The following relationship exists during the corresponding process: ; Among them, Denote the mutual information edge set, represent the mutual information between the nasal airflow and chest movement nodes, represent the mutual information threshold; It should be noted that through the above connection method, it can ensure that only those connections with significant information associations are retained in the graph structure, avoiding the introduction of too many meaningless edges. Through the edges constructed based on mutual information, the co-variation relationship between the nasal airflow and chest movement signals can be captured, further enriching the multi-modal relationship expressed by the graph structure.

[0022] The cross-channel edge set and the mutual information edge set combine to form a heterogeneous edge set.

[0023] It should be noted that in the appendix Figure 2 in, denote the blood oxygen saturation, GATConv denote the convolutional operation based on the graph attention network for processing the information transfer of homogeneous edges; SAGEConv denote the convolutional operation based on the graph sampling and aggregation network for processing the information aggregation of heterogeneous edges.

[0024] Step 3: Based on the node set, use the residual convolutional neural network to extract features from the signal window corresponding to the node to obtain the node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph.

[0025] In Step 3, based on the node set, use the residual convolutional neural network to extract features from the signal window corresponding to the node to obtain the node embedding features. The node embedding features, the homogeneous edge set, and the heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps: S101: Input the signal window corresponding to the node into the first residual block and perform two-layer one-dimensional convolutional layer processing, batch normalization processing, and LeakyReLU function processing in sequence to obtain the output of the second layer of the first residual block. The following relational expressions exist in the corresponding process: ; Among them, denote the output of the first layer of the first residual block, denote the output of the second layer of the first residual block, denote the signal window corresponding to the node, denote the convolutional weight of the first layer in the first residual block, denote the bias of the first layer in the first residual block, denote the convolutional layer weight of the second layer in the first residual block, denote the bias of the second layer in the first residual block, denote after one-dimensional convolutional layer processing, denote after activation function processing, Indicates batch normalization processing.

[0026] S102. Perform a residual skip connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block. The following relational expression exists in the corresponding process: ; Among them, Indicates the output of the first residual block.

[0027] Repeat the steps of S101 and S102 for the output of the first residual block in an iterative manner to obtain the output of the third residual block. Perform a global average pooling operation on the output features of the third residual block to obtain node embedding features. The following relational expression exists in the corresponding process: ; Among them, Indicates the node embedding features, Indicates the number of time steps, Indicates the output features of the 3rd residual block at the th time step.

[0028] It should be noted that each node in the sleep heterogeneous graph corresponds to a node embedding feature respectively. The residual convolutional neural network adopts a hierarchical structure, including three residual blocks. Each residual block contains two layers, and each layer includes a one-dimensional convolutional layer, a batch normalization layer, and a LeakyReLU activation function. Among them, the kernel size of the one-dimensional convolutional layer is set to 3, and the stride is set to 1. The residual skip connection can alleviate the problem of gradient disappearance, enabling the network to learn multi-scale time patterns.

[0029] Step 4. Input the sleep heterogeneous graph into the graph convolutional layer for processing to obtain updated node feature vectors.

[0030] In Step 4, input the sleep heterogeneous graph into the graph convolutional layer for processing to obtain updated node feature vectors, which specifically includes the following sub-steps: Based on the graph attention network convolutional layer, use the LeakyReLU function to calculate the attention coefficients of the node embedding features of homogeneous edges to obtain the attention coefficients of homogeneous edge nodes. The following relational expression exists in the corresponding process: ; Among them, Indicates the node and its neighbor node 's attention coefficient, Indicates the feature vector of node , Indicates the neighbor node of node ​The eigenvector, represents the attention mechanism parameters, represents the learnable weight matrix, represents the vector concatenation operation, represents the homogeneous edge nodes, represents the node of the homogeneous edge neighbor nodes.

[0031] The attention coefficients of the homogeneous edge nodes are normalized using the softmax function to obtain the normalized attention coefficients of the homogeneous edge nodes. The following relational expressions exist in the corresponding process: ; Among them, represents the node and its neighbor node the normalized attention coefficient, represents being processed by normalization.

[0032] The features of the neighbor nodes are weighted and aggregated using the normalized attention coefficients of the homogeneous edge nodes to obtain the updated features of the homogeneous edge nodes. The following relational expressions exist in the corresponding process: ; Among them, represents the updated eigenvector of the node , represents the set of neighbor nodes , represents the activation function.

[0033] Based on the graph sampling and aggregation network convolutional layer, the neighbor nodes of the heterogeneous edge nodes are sampled to obtain the set of sampled neighbor nodes. Based on the set of sampled neighbor nodes, the eigenvectors of the sampled neighbor nodes are aggregated using the mean aggregation function to obtain the aggregated features of the neighbor nodes. The following relational expressions exist in the corresponding process: ; Among them, represents the aggregated features of the neighbor nodes, represents the eigenvector of the sampled neighbor node , represents the sampled neighbor node, represents the set of sampled neighbor nodes, represents being processed by the mean aggregation function.

[0034] It should be noted that for each node associated with a heterogeneous edge, a fixed number of neighbor nodes are sampled from the different types of neighbor node sets corresponding to this node.

[0035] The feature vectors of heterogeneous edge nodes and the aggregated features of neighbor nodes are linearly transformed respectively, and then input into the ReLU function for processing to obtain the updated heterogeneous edge node features. The following relational expressions exist in the corresponding process: ; Among them, represents the updated feature vector of the heterogeneous edge node ; represents the result after being processed by the rectified linear unit activation function, represents the feature vector of the heterogeneous edge node ; represents the learnable weight matrix for linearly transforming the feature vector of the heterogeneous edge node, represents the learnable weight matrix for linearly transforming the aggregated features of neighbor nodes.

[0036] The minimum cross-edge type aggregation operation is performed on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain the updated node feature vector. The following relational expressions exist in the corresponding process: ; Among them, represents the updated node feature vector, represents the minimum value operation.

[0037] It should be noted that the graph convolutional layer contains 4 layers. By stacking these layers, messages can be transmitted and fused multiple times in the graph, gradually mining more complex and in-depth relationships in the graph. Each layer of the graph convolutional layer will further update the embedding representation of the nodes on the basis of the previous layer, so as to capture richer information.

[0038] Step 5: Concatenate the updated node feature vectors to obtain the concatenated feature matrix. Based on the multi-head self-attention mechanism and the multi-layer perceptron, construct a subtype classification module, and use the subtype classification module to process the concatenated feature matrix to obtain the classification results of different sleep apnea subtypes.

[0039] In step 5, the updated node feature vectors are concatenated to obtain the concatenated feature matrix. Based on the multi-head self-attention mechanism and the multi-layer perceptron, construct a subtype classification module, and use the subtype classification module to process the concatenated feature matrix to obtain the classification results of different sleep apnea subtypes, which specifically include the following sub-steps: Concatenate the updated node feature vectors to obtain the concatenated feature matrix. The following relational expressions exist in the corresponding process: ; Among them, represents the concatenated feature matrix, Both represent the updated node feature vectors, Represents the number of all nodes in the sleeping heterogeneous graph.

[0040] Based on the multi-head self-attention mechanism, the concatenated feature matrix is ​​linearly transformed with each attention head to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head. The following relationship exists in the corresponding process: ; in, Indicates The query matrix of the attention heads, Indicates The key matrix of the attention heads, Indicates The value matrix of the attention head, Indicates The attention heads are used to generate the weight matrix of the query matrix, Indicates The attention heads are used to generate the weight matrix of the key matrix, Indicates The attention heads are used to generate the weight matrix of the value matrix, represents the index of the attention head, Represents the number of attention heads.

[0041] The Softmax function is used to process the query matrix of the attention head and the key matrix of the attention head to obtain the attention score matrix. The following relationship exists in the corresponding process: ; in, Indicates The attention score matrix, Represents the dimension of the attention head.

[0042] The value matrix of the attention head and the attention score matrix are weighted and fused to obtain the output of the attention head. The following relationship exists in the corresponding process: ; in, Indicates The output of an attention head; The output of the attention head is concatenated to obtain the concatenated matrix, and the concatenated matrix is ​​linearly transformed to obtain the final graph features. The following relationship exists in the corresponding process: ; in, represents the concatenated matrix, both represent the output of the attention head, represents the final graph feature, represents a learnable weight matrix for linearly transforming the concatenated matrix.

[0043] In the step of sequentially performing a fully connected layer transformation process and an activation process on the final graph feature to obtain the output of the fully connected layer, the following relational expressions exist in the corresponding process: ; where, represents the output of the first fully connected layer, represents the weight matrix of the first layer of the fully connected layer, represents the bias vector of the first layer of the fully connected layer, represents the output of the -th layer of the fully connected layer, represents the output of the -th layer of the fully connected layer, represents the weight matrix of the -th layer of the fully connected layer, represents the bias vector of the -th layer of the fully connected layer, represents the total number of layers of the fully connected layer in the multi-layer perceptron.

[0044] Input the output of the last fully connected layer into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes. The following relational expressions exist in the corresponding process: ; where, represents the classification results of different sleep apnea subtypes.

[0045] It should be noted that the classification results of different sleep apnea subtypes are a probability distribution matrix. Each row in the matrix corresponds to the predicted probability of each category (normal, hypopnea, obstructive sleep apnea, central sleep apnea, and mixed sleep apnea). The category with the highest probability value is selected as the final predicted sleep breathing event type, thereby completing the classification of sleep breathing events and providing an accurate judgment basis for sleep apnea subtype detection.

[0046] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0047] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0048] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A sleep apnea subtype detection method based on multimodal information fusion, characterized in that: The method comprises the following steps: Step 1: obtaining an original multimodal physiological signal, preprocessing the original multimodal physiological signal, and obtaining a preprocessed signal window; Step 2: Perform node definition operation on the preprocessed signal window to obtain a node set, and perform homogeneous edge construction and heterogeneous edge construction based on the node set to obtain a homogeneous edge set and a heterogeneous edge set respectively; Step 3: Based on the node set, the residual convolutional neural network is used to extract features of the signal window corresponding to the node to obtain the node embedding features. The node embedding features, homogeneous edge sets and heterogeneous edge sets constitute the sleep heterogeneous graph; Step 4: Input the sleep heterogeneous graph into the graph convolution layer for processing to obtain the updated node feature vector; Step 5: Concatenate the updated node feature vectors to obtain a concatenated feature matrix, construct a subtype classification module based on the multi-head self-attention mechanism and the multi-layer perceptron, and use the subtype classification module to process the concatenated feature matrix to obtain the classification results of different sleep apnea subtypes.

2. The sleep apnea subtype detection method based on multimodal information fusion according to claim 1, characterized in that: In the step 1, the original multimodal physiological signal is obtained, and the original multimodal physiological signal is preprocessed to obtain a preprocessed signal window, which specifically includes the following sub-steps: Acquire original multimodal physiological signals, filter the original multimodal physiological signals, and obtain blood oxygen saturation signals, nasal airflow signals, and chest movement signals; The blood oxygen saturation signal, the nasal airflow signal and the chest movement signal are uniformly resampled to 20 Hz by using the interpolation technology to obtain the resampled blood oxygen saturation signal, the resampled nasal airflow signal and the resampled chest movement signal; The resampled nasal airflow signal and the resampled chest motion signal are filtered by using a fourth-order Butterworth bandpass filter to obtain a filtered nasal airflow signal and a filtered chest motion signal; The resampled blood oxygen saturation signal, the filtered nasal airflow signal and the filtered chest motion signal are standardized to obtain a standardized blood oxygen saturation signal, a standardized nasal airflow signal and a standardized chest motion signal; Data screening and balancing are performed on the standardized blood oxygen saturation signal, the standardized nasal airflow signal and the standardized chest movement signal to obtain a preprocessed signal window.

3. The sleep apnea subtype detection method based on multimodal information fusion according to claim 2, characterized in that: In step 2, a node definition operation is performed on the preprocessed signal window to obtain a node set, and homogeneous edge construction and heterogeneous edge construction are performed based on the node set to obtain a homogeneous edge set and a heterogeneous edge set respectively. The node set, the homogeneous edge set and the heterogeneous edge set constitute a sleeping heterogeneous graph, which specifically includes the following sub-steps: Based on the preprocessed signal window, each 20-second signal window is defined as a node and the node types are divided according to the signal channel to obtain a node set; Based on the node set, the nodes in the same signal channel are fully connected to obtain a homogeneous edge set; Based on the node set, the respiratory event node is connected to the lowest blood oxygen saturation value node to obtain the cross-channel edge set. The following relationship exists in the corresponding process: ; in, represents the set of cross-channel edges, Represents the node corresponding to the central event window in the nasal airflow or chest movement channel that represents the key respiratory event moment, Indicates the nodes corresponding to the nasal airflow passages, Represents the node corresponding to the chest motion channel, The node representing the lowest oxygen saturation value in the corresponding sample window in the blood oxygen saturation channel. Indicates the lowest blood sample saturation value; Based on the node set and given mutual information threshold, the mutual information value between the nasal airflow node and the chest movement node is calculated, and the mutual information edge is constructed for the nodes corresponding to the mutual information value exceeding the mutual information threshold to obtain the mutual information edge set. The following relationship exists in the corresponding process: ; in, represents the mutual information edge set, represents the mutual information between nasal airflow and chest motion nodes, represents the mutual information threshold; The combination of cross-channel edge sets and mutual information edge sets constitutes a heterogeneous edge set.

4. The sleep apnea subtype detection method based on multimodal information fusion according to claim 3, characterized in that: In step 3, based on the node set, the residual convolutional neural network is used to extract features of the signal window corresponding to the node to obtain the node embedding features. The node embedding features, the homogeneous edge set and the heterogeneous edge set constitute a sleep heterogeneous graph, which specifically includes the following sub-steps: S101, input the signal window corresponding to the node into the first residual block, and perform two-layer one-dimensional convolution layer processing, batch normalization processing and LeakyReLU function processing in sequence to obtain the output of the second layer of the first residual block; S102, performing a residual jump connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block; Repeat steps S101 and S102 for the output of the first residual block in an iterative manner to obtain the output of the third residual block, perform a global average pooling operation on the output features of the third residual block, and obtain a node embedding feature; The node embedding features, homogeneous edge sets, and heterogeneous edge sets constitute the sleep heterogeneous graph.

5. The sleep apnea subtype detection method based on multimodal information fusion according to claim 4, characterized in that: The signal window corresponding to the node is input into the first residual block and processed by two layers of one-dimensional convolutional layers, batch normalization and LeakyReLU function in sequence to obtain the output of the second layer of the first residual block. The following relationship exists in the corresponding process: ; in, represents the output of the first layer of the first residual block, represents the output of the second layer of the first residual block, Indicates the signal window corresponding to the node, represents the convolution weights of the first layer in the first residual block, represents the bias of the first layer in the first residual block, represents the convolutional layer weights of the second layer in the first residual block, represents the bias of the second layer in the first residual block, It means that after one-dimensional convolution layer processing, It means that it has been processed by the activation function. Indicates that batch normalization has been performed; In the step of performing a residual jump connection between the input of the first residual block and the output of the second layer of the first residual block to obtain the output of the first residual block, the following relationship exists in the corresponding process: ; in, represents the output of the first residual block; In the step of iteratively repeating steps S101 and S102 on the output of the first residual block to obtain the output of the third residual block, performing a global average pooling operation on the output features of the third residual block to obtain the node embedding features, the following relationship exists in the corresponding process: ; in, represents the node embedding feature, represents the number of time steps, Indicates that the third residual block is in The output features of each time step.

6. The sleep apnea subtype detection method based on multimodal information fusion according to claim 5, characterized in that: In step 4, the sleep heterogeneous graph is input into the graph convolution layer for processing to obtain an updated node feature vector, which specifically includes the following sub-steps: Based on the convolutional layer of the graph attention network, the LeakyReLU function is used to calculate the attention coefficient of the node embedding feature of the homogeneous edge to obtain the attention coefficient of the homogeneous edge node; The softmax function is used to normalize the attention coefficients of homogeneous edge nodes to obtain the normalized attention coefficients of homogeneous edge nodes; The normalized attention coefficient of the homogeneous edge node is used to perform weighted aggregation on the features of the neighboring nodes to obtain the updated homogeneous edge node features; Based on the graph sampling and aggregation network convolutional layer, the neighbor nodes of the heterogeneous edge nodes are sampled to obtain the sampled neighbor node set. Based on the sampled neighbor node set, the feature vectors of the sampled neighbor nodes are aggregated using the mean aggregation function to obtain the aggregated features of the neighbor nodes. The feature vectors of heterogeneous edge nodes and the aggregated features of neighboring nodes are linearly transformed respectively and input into the ReLU function for processing to obtain the updated heterogeneous edge node features; A minimum aggregation operation across edge types is performed on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain an updated node feature vector.

7. The sleep apnea subtype detection method based on multimodal information fusion according to claim 6, characterized in that: Based on the convolutional layer of the graph attention network, the LeakyReLU function is used to calculate the attention coefficient of the node embedding feature of the homogeneous edge, and the attention coefficient of the homogeneous edge node is obtained. The following relationship exists in the corresponding process: ; in, Representation Node Its neighbor nodes The attention coefficient, Representation Node The characteristic vector of Representation Node Neighbor nodes The characteristic vector of represents the attention mechanism parameters, represents the learnable weight matrix, represents the vector concatenation operation, represents homogeneous edge nodes, Representation Node Homogeneous edge neighbor nodes; The softmax function is used to normalize the attention coefficients of homogeneous edge nodes to obtain the normalized attention coefficients of homogeneous edge nodes. The following relationship exists in the corresponding process: ; in, Representation Node Its neighbor nodes The normalized attention coefficient, It means that it has been normalized; The normalized attention coefficient is used to perform weighted aggregation on the features of neighbor nodes to obtain the updated homogeneous edge node features. The following relationship exists in the corresponding process: ; in, Represents the updated node The characteristic vector of Represents neighbor nodes A collection of represents the activation function; Based on the graph sampling and aggregation network convolutional layer, the neighbor nodes of the heterogeneous edge nodes are sampled to obtain the sampled neighbor node set. Based on the sampled neighbor node set, the feature vectors of the sampled neighbor nodes are aggregated using the mean aggregation function to obtain the aggregated features of the neighbor nodes. The following relationship exists in the corresponding process: ; in, represents the aggregation characteristics of neighbor nodes, Represents the sampled neighbor nodes The characteristic vector of represents the sampled neighbor nodes, represents the set of neighbor nodes after sampling, Indicates that it has been processed by the mean aggregation function; The feature vectors of heterogeneous edge nodes and the aggregated features of neighboring nodes are linearly transformed and input into the ReLU function for processing to obtain the updated heterogeneous edge node features. The following relationship exists in the corresponding process: ; in, Represents the updated heterogeneous edge node The characteristic vector of It means that it has been processed by the rectified linear unit activation function. Represents heterogeneous edge nodes The characteristic vector of represents the learnable weight matrix for linear transformation of the feature vectors of heterogeneous edge nodes, Represents a learnable weight matrix that linearly transforms the aggregated features of neighbor nodes; Perform the minimum aggregation operation across edge types on the updated homogeneous edge node features and the updated heterogeneous edge node features to obtain the updated node feature vector. The following relationship exists in the corresponding process: ; in, represents the updated node feature vector, Indicates the minimum value operation.

8. The sleep apnea subtype detection method based on multimodal information fusion according to claim 7, characterized in that: In step 5, the updated node feature vector features are spliced ​​to obtain a spliced ​​feature matrix, a subtype classification module is constructed based on a multi-head self-attention mechanism and a multi-layer perceptron, and the spliced ​​feature matrix is ​​processed using the subtype classification module to obtain classification results of different sleep apnea subtypes, which specifically includes the following sub-steps: The updated node feature vectors are concatenated to obtain a concatenated feature matrix; Based on the multi-head self-attention mechanism, the concatenated feature matrix is ​​linearly transformed with each attention head to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head; Use the Softmax function to process the query matrix of the attention head and the key matrix of the attention head to obtain the attention score matrix; Perform weighted fusion of the attention head value matrix and the attention score matrix to obtain the output of the attention head; The outputs of the attention heads are concatenated to obtain a concatenated matrix, and the concatenated matrix is ​​linearly transformed to obtain the final graph features; The final graph features are sequentially transformed and activated by the fully connected layer to obtain the output of the fully connected layer; The output of the last fully connected layer is input into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes.

9. The sleep apnea subtype detection method based on multimodal information fusion according to claim 8, characterized in that: The updated node feature vectors are concatenated to obtain the concatenated feature matrix. The following relationship exists in the corresponding process: ; in, represents the concatenated feature matrix, Both represent the updated node feature vectors, Represents the number of all nodes in the sleeping heterogeneous graph; Based on the multi-head self-attention mechanism, the concatenated feature matrix is ​​linearly transformed with each attention head to obtain the query matrix corresponding to the attention head, the key matrix corresponding to the attention head, and the value matrix corresponding to the attention head. In the corresponding process, the following relationship exists: ; in, Indicates The query matrix of the attention heads, Indicates The key matrix of the attention heads, Indicates The value matrix of the attention head, Indicates The attention heads are used to generate the weight matrix of the query matrix, Indicates The attention heads are used to generate the weight matrix of the key matrix, Indicates The attention heads are used to generate the weight matrix of the value matrix, represents the index of the attention head, Indicates the number of attention heads; In the step of using the Softmax function to process the query matrix of the attention head and the key matrix of the attention head to obtain the attention score matrix, the following relationship exists in the corresponding process: ; in, Indicates The attention score matrix, represents the dimension of the attention head; In the step of weighted fusion of the value matrix of the attention head and the attention score matrix to obtain the output of the attention head, the following relationship exists in the corresponding process: ; in, Indicates The output of an attention head; In the step of concatenating the outputs of the attention heads to obtain the concatenated matrix, performing a linear transformation on the concatenated matrix, and obtaining the final graph features, the following relationship exists in the corresponding process: ; in, represents the concatenated matrix, Both represent the output of the attention head, Represents the final graph features, represents the learnable weight matrix used to linearly transform the concatenated matrix; In the step of performing the fully connected layer transformation and activation processing on the final graph features in sequence to obtain the output of the fully connected layer, the following relationship exists in the corresponding process: ; in, represents the output of the first fully connected layer, represents the weight matrix of the first layer of the fully connected layer, represents the bias vector of the first layer of the fully connected layer, Indicates The output of the fully connected layer, Indicates The output of the fully connected layer, represents the fully connected layer The weight matrix of the layer, represents the fully connected layer The bias vector of the layer, Represents the total number of fully connected layers in a multilayer perceptron; In the step of inputting the output of the last fully connected layer into the Softmax function for processing to obtain the classification results of different sleep apnea subtypes, the following relationship exists in the corresponding process: ; in, Represents the classification results of different sleep apnea subtypes.

Citation Information

Patent Citations

  • Multi-mode multi-resolution sleep apnea automatic detection method and system

    CN117838053A

  • Sleep apnea detection method and device based on feature fusion

    CN118155826A

  • Intelligent pulmonary nodule grading method and system based on multi-modality feature fusion

    WO2025020719A1

Cited By

  • Intelligent sleep disorder diagnosis system based on multi-modal knowledge graph

    CN121040866A