Sleep Apnea Event Location Method and Device Based on Multi-Scale Convolutional Neural Network
A multi-scale CNN integrates nasal and pressure signals to enhance the precision of sleep apnea event detection, overcoming limitations of manual inspection and fixed time windows, achieving accurate localization and reducing fragmentation.
Patent Information
- Application Number
- CN202211467527.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-22
AI Technical Summary
In the prior art, the positioning of sleep breathing events is not accurate enough, it is difficult to select a suitable fixed time window and it is impossible to accurately locate the start and end time of sleep breathing events, resulting in fragment loss and fragmentation problems.
A multi-scale convolutional neural network is adopted to achieve accurate positioning of sleep breathing events through feature extraction and fusion of multi-modal data, combining multi-scale predictive feature layer generation, region generation and region of interest pooling.
It realizes accurate positioning of sleep breathing events, solves the problem of selection of fixed time windows and fragment loss problems during the cropping process, and can accurately inform the start and end time of the event.
Smart Images

Figure CN115758122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a method and device for locating sleep apnea events based on a multi-scale convolutional neural network. Background Art
[0002] Sleep is a type of spontaneous physiological activity affected by biological rhythms in the brain, which plays a crucial role in restoring the body functions and development of the human body. Sleep disordered breathing can reduce the quality of sleep and, in severe cases, lead to diseases such as memory decline and heart failure, seriously endangering health. Sleep apnea and hypopnea, as the main sleep apnea events, are important bases for evaluating the severity of sleep disordered breathing and have important clinical value. Currently, the diagnosis and recognition of sleep apnea events still mainly rely on the visual inspection of polysomnograms by experts, and this method of manual calibration is costly, inefficient and has poor consistency. Therefore, there is an urgent need for an accurate and efficient automated detection method to solve this problem.
[0003] The existing sleep apnea event detection methods mainly rely on classifiers for discrimination. After cropping the original signal according to a fixed time window, a feature set is extracted from the signal segment through feature engineering or feature learning methods, and finally a classifier is used for classification. However, the duration of sleep apnea events is uncertain, ranging from a few seconds to several minutes, which makes it difficult for the classifier-based method to not only select an appropriate fixed time window, but also prone to problems such as segment loss and fragmentation during the process of cropping samples; in addition, the above classifier-based method can only classify signal segments and cannot inform the start time and end time of sleep apnea events, making it difficult to achieve precise positioning of sleep apnea events. Summary of the Invention
[0004] The present invention provides a method and device for locating sleep apnea events based on a multi-scale convolutional neural network, so as to solve or at least partially solve the technical problem of inaccurate positioning of sleep apnea events in the prior art.
[0005] To solve the above technical problem, the first aspect of the present invention provides a method for locating sleep apnea events based on a multi-scale convolutional neural network, including:
[0006] S1: Extract the oral-nasal airflow signal, thoracic pressure signal and abdominal pressure signal from the original polysomnogram and splice the data of the three channels into a vector where X i represents the data extracted from the i-th polysomnogram, respectively represent the oral-nasal airflow signal, thoracic pressure signal and abdominal pressure signal in the i-th polysomnogram;
[0007] S2: preprocessing the data of the three channels extracted, specifically including: firstly resampling the data, filtering and standardizing each channel, and then labeling the sleep breathing events of the standardized data, including the event start time and event end time;
[0008] S3: Construct a multi-scale convolutional neural network, which includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module and a final classification module, wherein the feature extraction module is used to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features, the multi-scale prediction feature layer generation module is used to obtain multiple prediction feature layers based on the fused multi-modal features, the region generation module is used to generate candidate boxes on the prediction feature layer, the region of interest pooling module is used to obtain the feature map of the region corresponding to the candidate box, and unify the sizes of different feature maps, and the final classification module is used to classify the feature map output by the region of interest pooling module and output the offset;
[0009] S4: input the preprocessed data into the multi-scale convolutional neural network for training;
[0010] S5: Input the sleep data to be detected into the trained multi-scale convolutional neural network for positioning prediction, obtain classification results and offsets, and perform post-processing based on the classification results and offsets to obtain the final detection results.
[0011] In one embodiment, step S2 includes:
[0012] S2.1: resample the extracted data of the three channels, and use an 8th-order Butterworth filter to low-pass filter the data of the three channels respectively;
[0013] S2.2: For X i The data of channel k in Calculate the median separately and interquartile range Use the robust normalization method to normalize the data of channel k:
[0014]
[0015] in, Represents the standardized data;
[0016] S2.3: Perform sleep breathing event annotation on the standardized data, treat sleep apnea events and hypopnea events as sleep breathing events, and annotate the start and end time of each event.
[0017] In one implementation, during the training process of step S4, the processing process of the feature extraction module includes:
[0018] Extract the unimodal features of the input data through three parallel and independent unimodal feature extraction modules;
[0019] Concatenate the extracted unimodal features into multimodal features in the order of oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals, and then fuse the multimodal features through 2 residual convolution blocks and 2 downsampling residual convolution blocks to obtain the fused multimodal features.
[0020] In one implementation, during the training process of step S4, the processing process of the multi-scale prediction feature layer generation module includes:
[0021] Pass the fused multimodal features obtained by the feature extraction module through the Inception structure to obtain features Figure 1 , where the Inception structure consists of 3 branches: there is only 1 convolutional layer with a kernel size of 1*1 on the first branch; there are 2 convolutional layers on the second branch, with kernel sizes of 1*1 and 3*1 respectively; there are 2 convolutional layers on the third branch, with kernel sizes of 1*1 and 5*1 respectively; finally, concatenate the feature maps obtained from the 3 branches on the channel dimension to obtain features Figure 1 ;
[0022] Pass the features Figure 1 through 1 downsampling residual convolution block and then input the result into the Inception structure to obtain features Figure 2 ;
[0023] Pass the features Figure 2 through 1 downsampling residual convolution block and then input the result into the Inception structure to obtain features Figure 3 ;
[0024] Pass the features Figure 3 through 1 downsampling residual convolution block and then input the result into the Inception structure to obtain features Figure 4 ;
[0025] Pass the features Figure 1 ,2,3,4 into the feature pyramid network to obtain prediction feature layers 1,2,3,4.
[0026] In one implementation, the multiple prediction feature layers include prediction feature layers 1,2,3,4. During the training process of step S4, the processing process of the region generation module includes:
[0027] Generate preset boxes with different lengths on the prediction feature layers 1,2,3,4 respectively;
[0028] Slide a preset sliding window over the predicted feature layer, and use two convolutional layers with a convolutional kernel size of 1*1 to generate one probability value and two offsets for each sliding window. The probability value represents the probability that the preset box at this position is the target object, and the two offsets are (t x , t w ), where t x is the offset used to correct the center coordinates of the preset box, and t w is the offset used to correct the width of the preset box;
[0029] Use the generated offsets to correct the positions of the preset boxes to obtain candidate boxes;
[0030] Perform non-maximum suppression processing on the candidate boxes of each predicted feature layer to obtain the candidate boxes after processing for each predicted feature layer. Combine all the candidate boxes after processing for all predicted feature layers and select the N candidate boxes with the largest probability values as the output candidate boxes.
[0031] In one implementation, during the training process of step S4, the processing process of the region of interest pooling module includes:
[0032] Map the candidate boxes to different predicted feature layers according to the following formula:
[0033]
[0034] where w represents the width of the candidate box, scale is the scale factor, FS is the signal sampling rate, and k is the layer number of the predicted feature layer to which the candidate box is mapped;
[0035] Use the region of interest pooling method to perform scale unification processing on the feature maps of the regions corresponding to the candidate boxes.
[0036] In one implementation, the final classification module includes two fully connected layers. During the training process of step S4, the processing process of the final classification module includes:
[0037] Obtain two probability values through a fully connected layer according to the input feature map, which respectively represent the probability that the content in the candidate box corresponding to the feature map is the background and the probability that it is the foreground;
[0038] Obtain two offsets (t′ x , t′ w ) through another fully connected layer according to the input feature map. t′ x is the offset obtained by the final classification module for correcting the center coordinates of the candidate box, and t′ w is the offset obtained by the final classification module for correcting the width of the candidate box.
[0039] Based on the same inventive concept, the second aspect of the present invention provides a sleep apnea event localization device based on a multi-scale convolutional neural network, including:
[0040] An original data extraction module, configured to extract oral-nasal airflow signals, thoracic pressure signals, and abdominal pressure signals from an original polysomnogram and splice the data of the three channels into a vector where X i represents the data extracted from the i-th polysomnogram, respectively representing the oral-nasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the i-th polysomnogram;
[0041] A data preprocessing module, configured to preprocess the data of the three channels obtained by extraction, specifically including: first performing data resampling, and performing filtering processing and normalization processing channel by channel, and then performing sleep apnea event annotation on the normalized data, and the annotation content includes the event start time and the event end time;
[0042] A network construction module, configured to construct a multi-scale convolutional neural network, and the multi-scale convolutional neural network includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module, and a final classification module. Among them, the feature extraction module is configured to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features. The multi-scale prediction feature layer generation module is configured to obtain multiple prediction feature layers according to the fused multi-modal features. The region generation module is configured to generate candidate boxes on the prediction feature layer. The region of interest pooling module is configured to obtain the feature maps of the regions corresponding to the candidate boxes and unify the sizes of different feature maps. The final classification module is configured to classify the feature maps output by the region of interest pooling module and output the offset;
[0043] A training module, configured to input the preprocessed data into the multi-scale convolutional neural network for training;
[0044] A localization detection module, configured to input the sleep data to be detected into the trained multi-scale convolutional neural network for localization prediction, obtain the classification result and the offset, and perform post-processing according to the obtained classification result and the offset to obtain the final detection result.
[0045] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed, the method described in the first aspect is implemented.
[0046] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in the first aspect is implemented.
[0047] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:
[0048] 1) In order to maintain consistency with the data used by experts for discrimination, the present invention uses signals of oral and nasal airflow, thoracic pressure, and abdominal pressure as the source data for model input. For the above-mentioned multi-modal data, the feature extraction module proposed by the present invention fully considers single-modal feature extraction and multi-modal feature fusion, utilizes the complementarity between multi-modalities, eliminates redundancy between modalities, and enhances the feature representation ability of the model.
[0049] 2) The multi-scale prediction feature layer generation module designed by the present invention realizes multi-scale in both the network depth and network width of the prediction feature layer generation. The multi-scale in depth and width can enhance the richness of features and improve the detection ability of the model for sleep apnea events of various scales.
[0050] 3) The present invention uses the idea of an object detection framework to locate and classify sleep apnea events in signals, which can effectively solve the problem of time window selection and the problems of segment loss and event fragmentation caused during the cropping process when cropping samples based on a fixed time window. Compared with the aforementioned method based on classification discrimination, the present invention can not only classify the signals, but also inform the start time and end time of sleep apnea events, achieving precise positioning of sleep apnea events.
[0051] 4) The present invention provides a new idea for the detection of abnormal segments in time series data: using an object detection framework to automatically locate and classify abnormal segments in time series data. At the same time, in view of the differences in data formation and semantics between time series data and natural images, the present invention provides a multi-scale convolutional neural network for multi-scale and in-depth modeling of time series data, improving the positioning and classification effects of abnormal segments. Brief Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a flowchart of the method for locating sleep apnea events based on a multi-scale convolutional neural network provided in the embodiments of the present invention;
[0054] Figure 2 It is a structural diagram of the feature extraction module in the embodiments of the present invention;
[0055] Figure 3Structural diagram of the multi-scale prediction feature layer generation module in the embodiment of the present invention;
[0056] Figure 4 Schematic diagrams of two residual structures in the embodiment of the present invention;
[0057] Figure 5 Schematic diagram of the Inception structure in the embodiment of the present invention;
[0058] Figure 6 Block diagram of the structure of the sleep apnea event localization method device based on the multi-scale convolutional neural network in the embodiment of the present invention;
[0059] Figure 7 Schematic diagram of the structure of the computer-readable storage medium provided by the embodiment of the present invention;
[0060] Figure 8 Schematic diagram of the structure of the computer device provided by the embodiment of the present invention. Detailed implementation manners
[0061] A sleep apnea event localization method based on a multi-scale convolutional neural network proposed by the present invention selects the oral and nasal airflow, thoracic pressure, and abdominal pressure signals in the sleep monitoring data as the data basis, and designs an automatic sleep apnea event localization network model that integrates multi-modal data (i.e., the multi-scale convolutional neural network in S3). This model mainly includes the following four parts: extracting multi-modal features from the preprocessed time-series data based on the feature extraction module; generating a prediction feature layer with multi-scale in both network depth and network width based on the multi-scale prediction feature layer generation module; adaptively adjusting the size and position of the preset box designed according to the receptive field based on the region generation module; and classifying based on the scale-consistent features generated by the candidate segments. The model of the present invention adopts an end-to-end training mode and can simultaneously realize the classification and automatic localization of sleep apnea events.
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] Embodiment 1
[0064] The embodiment of the present invention provides a sleep apnea event localization method based on a multi-scale convolutional neural network, including:
[0065] S1: Extract the oral and nasal airflow signal, thoracic pressure signal, and abdominal pressure signal from the original polysomnogram, and splice the data of the three channels into a vector. Where $X$ i represents the data extracted from the $i$-th polysomnogram, which respectively represent the oral and nasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the $i$-th polysomnogram;
[0066] S2: Preprocess the data of the three extracted channels, specifically including: first, perform data resampling, and perform filtering and normalization processing channel by channel. Then, perform sleep apnea event annotation on the normalized data, and the annotation content includes the start time and end time of the event;
[0067] S3: Construct a multi-scale convolutional neural network. The multi-scale convolutional neural network includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module, and a final classification module. Among them, the feature extraction module is used to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features. The multi-scale prediction feature layer generation module is used to obtain multiple prediction feature layers based on the fused multi-modal features. The region generation module is used to generate candidate boxes on the prediction feature layer. The region of interest pooling module is used to obtain the feature maps corresponding to the candidate boxes and unify the sizes of different feature maps. The final classification module is used to classify the feature maps output by the region of interest pooling module and output the offset;
[0068] S4: Input the preprocessed data into the multi-scale convolutional neural network for training;
[0069] S5: Input the sleep data to be detected into the trained multi-scale convolutional neural network for localization prediction to obtain the classification result and the offset. Perform post-processing based on the obtained classification result and offset to obtain the final detection result.
[0070] Please refer to Figure 1 , which is the flowchart of the sleep apnea event localization method based on a multi-scale convolutional neural network provided in the embodiment of the present invention. S1 is the acquisition of the original data, S2 is the data preprocessing, S3 is the construction of the network, S4 is the training of the network, and S5 is the testing of the network.
[0071] In the specific implementation process, step S1 can be implemented through the following steps:
[0072] S1.1: Use a polysomnograph to synchronously detect multiple sleep physiological parameters of a person during sleep to obtain a polysomnogram. The polysomnogram contains multimodal data such as electroencephalogram signals, electrocardiogram signals, and oronasal airflow signals. Extract the oronasal airflow signal, thoracic pressure signal, and abdominal pressure signal, which are strongly correlated with sleep breathing events, from them as source data.
[0073] S1.2: The present invention uses the multimodal data as the source data for model input, and splices the three-channel data extracted in S1.1 into a vector where X i represents the data extracted from the i-th polysomnogram. To maintain data consistency, it is stipulated that respectively represent the oronasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the i-th polysomnogram.
[0074] In one implementation, step S2 includes:
[0075] S2.1: Resample the data of the three extracted channels, and use an 8th-order Butterworth filter to perform low-pass filtering on the data of the three channels respectively;
[0076] S2.2: For each channel k of the multi-channel data X i in the data calculate the median and the interquartile range respectively, and use the robust normalization method to perform normalization processing on the data of channel k:
[0077]
[0078] where, represents the normalized data;
[0079] S2.3: Perform sleep breathing event annotation on the normalized data, regard sleep apnea events and hypopnea events as sleep breathing events, and mark the start time and end time of each event.
[0080] In the specific implementation process, the resampling frequency can be set according to the situation, and the filtering frequency can also be set according to the actual situation. In this implementation, the three-channel source data is resampled to 32 HZ, and an 8th-order Butterworth filter is used to perform low-pass filtering on the data of the three channels respectively, and the cut-off frequency is set to 2.4 HZ.
[0081] In one implementation, due to the differences in the settings of sensors, filters, etc. used in the acquisition process, the oronasal airflow signal, thoracic pressure signal, and abdominal pressure signal extracted from the polysomnogram are three different modal signals. Therefore, in the feature extraction process, the features of each modality should be extracted separately first and then the multi-modal features are fused.
[0082] Therefore, during the training process of step S4, the processing process of the feature extraction module includes:
[0083] Extract the unimodal features of the input data through three parallel and independent unimodal feature extraction modules;
[0084] Concatenate the extracted unimodal features into multimodal features in the order of oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals, and then fuse the multimodal features through 2 residual convolutional blocks and 2 downsampling residual convolutional blocks to obtain the fused multimodal features.
[0085] Please refer to Figure 2 , for the structure diagram of the feature extraction module.
[0086] In the specific implementation process, it can be achieved through the following steps:
[0087] S4.1.1: Unimodal feature extraction for each channel. Input the preprocessed data channel by channel into three parallel and independent unimodal feature extraction modules. These data first pass through a convolutional layer with a kernel size of 7*1, and then pass through 3 residual convolutional blocks and 1 downsampling residual convolutional block to obtain the final unimodal features. The residual convolutional block is composed of 2 convolutional layers with a kernel size of 3*1. After each convolutional layer, the activation function Leaky ReLU is used to enhance the non-linear expression ability of the model, and there is also a residual connection between these two convolutional layers:
[0088]
[0089] Among them represents the input of the residual convolutional block, represents the output of the residual convolutional block, and Conv(x) represents the convolution and non-linear activation operation on x. The downsampling residual convolutional block adds the function of downsampling the signal by two times on the basis of the residual convolutional block.
[0090] S4.1.2: Multimodal feature fusion. Concatenate the unimodal features extracted in S4.1.1 into multimodal features in the order of oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals, and then fuse the multimodal features through 2 residual convolutional blocks and 2 downsampling residual convolutional blocks.
[0091] In one implementation, during the training process of step S4, the processing process of the multi-scale prediction feature layer generation module includes:
[0092] Pass the fused multimodal features obtained by the feature extraction module through the Inception structure to obtain features Figure 1, where the Inception structure consists of 3 branches: the first branch has only one convolutional layer with a kernel size of 1*1; the second branch has two convolutional layers with kernel sizes of 1*1 and 3*1 respectively; the third branch has two convolutional layers with kernel sizes of 1*1 and 5*1 respectively; finally, the feature maps obtained from the 3 branches are concatenated in the channel dimension to obtain the feature Figure 1 ;
[0093] The feature Figure 1 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the feature Figure 2 ;
[0094] The feature Figure 2 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the feature Figure 3 ;
[0095] The feature Figure 3 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the feature Figure 4 ;
[0096] The feature Figure 1 , 2, 3, 4 are input into the feature pyramid network to obtain the predicted feature layers 1, 2, 3, 4.
[0097] The duration of sleep apnea events is uncertain, ranging from a few seconds to several minutes. It is difficult to fit sleep apnea events with significantly different sizes using only a single-scale predicted feature layer. Therefore, predicting sleep apnea events of different scales through multi-scale predicted feature layers can effectively reduce the learning difficulty of the model and improve the model accuracy.
[0098] Please refer to Figures 3 to 5 , Figure 3 , which is the structure diagram of the multi-scale predicted feature layer generation module in the embodiment of the present invention, Figure 4 , which is the schematic diagram of two residual structures in the embodiment of the present invention; Figure 5 , which is the schematic diagram of the Inception structure in the embodiment of the present invention.
[0099] In the specific implementation process, it can be realized through the following steps:
[0100] S4.2.1: Generate the feature Figure 1 . The multi-modal feature obtained in S4.1.2 is passed through the Inception structure to obtain the feature Figure 1, where the Inception structure consists of three branches: the first branch has only one convolutional layer with a kernel size of 1*1; the second branch has two convolutional layers with kernel sizes of 1*1 and 3*1 respectively; the third branch has two convolutional layers with kernel sizes of 1*1 and 5*1 respectively; finally, the feature maps obtained from the three branches are concatenated in the channel dimension to obtain the feature Figure 1 .
[0101] S4.2.2: Generate features Figure 2 . The features obtained in S4.2.1 Figure 1 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the features Figure 2 .
[0102] S4.2.3: Generate features Figure 3 . The features obtained in S4.2.2 Figure 2 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the features Figure 3 .
[0103] S4.2.4: Generate features Figure 4 . The features obtained in S4.2.3 Figure 3 After passing through one downsampling residual convolutional block, the result is input into the Inception structure to obtain the features Figure 4 .
[0104] S4.2.5: Input the features Figure 1 , 2, 3, 4 into the feature pyramid network to obtain the prediction feature layers 1, 2, 3, 4.
[0105] In one implementation, the multiple prediction feature layers include the prediction feature layers 1, 2, 3, 4. During the training process of step S4, the processing process of the region generation module includes:
[0106] Generate preset boxes with different lengths on the prediction feature layers 1, 2, 3, 4 respectively;
[0107] Use a preset sliding window to slide on the prediction feature layer, and use two convolutional layers with a kernel size of 1*1 to generate one probability value and two offsets for each sliding window, where the probability value represents the probability that the preset box at this position is the target object, and the two offsets are (t x , t w ), t x is the offset used to correct the center coordinates of the preset box, and t w is the offset used to correct the width of the preset box;
[0108] Use the generated offsets to correct the positions of the preset boxes to obtain candidate boxes;
[0109] Perform non-maximum suppression (NMS) on the candidate boxes of each predicted feature layer to obtain the candidate boxes after processing for each predicted feature layer. Combine the candidate boxes after processing for all predicted feature layers, and then select the N candidate boxes with the largest probability values as the output candidate boxes.
[0110] The region generation module generates preset boxes of different scales for different predicted feature layers. The size and number of preset boxes are only related to the level of the predicted feature layer and have nothing to do with the sample data. Then, the region generation module generates prediction values and correction amounts for each preset box.
[0111] In the specific implementation process, it can be achieved through the following steps:
[0112] S4.3.1: The region generation module generates preset boxes with lengths of 224 (7s), 448 (14s), 672 (21s), and 1344 (42s) on the predicted feature layers 1, 2, 3, and 4 respectively. The generation interval is the downsampling multiple of the feature layer relative to the source data.
[0113] Among them, the generation interval refers to the interval between predicted feature layers (hereinafter referred to as feature layers). Taking feature layer 1 as an example, preset boxes with a length of 224 are generated on feature layer 1. The Figure 1 generation interval refers to the interval of the preset boxes generated on feature layer 1; the feature layer is Figure 3 the predicted feature layers 1 to 5 in, which are actually the feature maps output by different convolution modules; the source data refers to the Figure 1 input data in, that is, the initial length of the signal input into the model. The source data corresponding to different predicted feature layers is the same, which is the input signal or data.
[0114] S4.3.2: Use a sliding window with a length of 3 and a sliding distance of 1 to slide on the predicted feature layer. Use 2 convolutional layers with a convolutional kernel size of 1*1 to generate 1 probability value and 2 offset values for each time window. One probability value represents the probability that the preset box at this position is the target object, and the two offset values (t x , t w ) are the offset values t x for correcting the center coordinates of the preset box and the offset value t w for correcting the width of the preset box respectively. For a preset box with a coordinate position of (x1, x2), its center coordinate x a and width w a can be calculated as follows:
[0115]
[0116] wa = x2 - x1
[0117] S4.3.3: Use the generated offset to correct the position of the preset box to obtain the candidate box. For a preset box (x1, x2), calculate its center coordinate x using the above formula a and width w a , and then use the offset (t x , t w ) generated by the network to correct the coordinate position of the preset box:
[0118] x p = x a + t x * w a
[0119]
[0120] where x p represents the corrected center coordinate, and w p represents the corrected width, and then the candidate box coordinates can be calculated through the formula
[0121]
[0122]
[0123] S4.3.4: Perform non-maximum suppression on the candidate boxes of each predicted feature layer to remove the candidate boxes with too high overlap rate, and then mix the candidate boxes of all layers together and select the N candidate boxes with the largest probability value as the output candidate boxes of the region generation module.
[0124] In one implementation, during the training process of step S4, the processing process of the region of interest pooling module includes:
[0125] Map the candidate box to different predicted feature layers according to the following formula:
[0126]
[0127] where w represents the width of the candidate box, scale is the scale factor, FS is the signal sampling rate, and k is the layer number of the predicted feature layer to which the candidate box is mapped;
[0128] Adopt the region of interest pooling method to perform scale unification processing on the feature map of the region corresponding to the candidate box.
[0129] In one implementation, the final classification module includes two fully connected layers. During the training process of step S4, the processing process of the final classification module includes:
[0130] Two probability values are obtained from the input feature map through a fully connected layer, representing the probabilities that the content in the candidate box corresponding to the feature map is the background and the foreground, respectively;
[0131] Two offsets (t′ x , t′ w ) are obtained from the input feature map through another fully connected layer. t′ x is the offset for correcting the center coordinates of the candidate box obtained by the final classification module, and t′ w is the offset for correcting the width of the candidate box obtained by the final classification module.
[0132] The feature map obtained by the region of interest pooling module is input into two parallel fully connected layers. One fully connected layer outputs two probability values, representing the probabilities that the content in the candidate box is the background and the foreground, respectively; the other fully connected layer outputs two offsets (t′ x , t′ w ). Use these two offsets to further correct the position of the candidate box according to the methods in S4.3.2 and S4.3.3. Another position correction is performed here, so that the position of the candidate box can be more accurate. It should be noted that the meanings and functions of the offsets are the same, but the specific values are different.
[0133] In one implementation, in the above method for locating sleep breathing events based on a multi-scale convolutional neural network, post-processing the detection results, the specific steps include:
[0134] S5.1: Use the offset generated by the network to correct the position of the candidate box to obtain the final detection box, and perform softmax activation processing on the probability value generated by the network;
[0135] S5.2: Adjust the position of the detection box that exceeds the boundary to the signal boundary, and remove the detection box with too small size;
[0136] S5.3: Perform non-maximum suppression processing on all detection boxes of each data, and select the k detection boxes with the largest foreground probability as the final detection results.
[0137] Generally speaking, the method provided by the present invention mainly includes three parts:
[0138] The first part is multi-modal data extraction and preprocessing (steps S1, S2):
[0139] 1. Extract the oral-nasal airflow signal, thoracic pressure signal, and abdominal pressure signal strongly related to sleep breathing events from the original polysomnogram as source data. Concatenate the three-channel data into a vector where X idenotes the data extracted from the i-th polysomnogram. To maintain data consistency, it is stipulated that respectively denote the oro-nasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the i-th polysomnogram.
[0140] 2. To maintain the consistency of the input length, in this example, the three-channel source data is resampled to 32HZ, and the data of the three channels are respectively low-pass filtered using an 8th-order Butterworth filter, with the cut-off frequency set to 2.4HZ. The filtered multimodal data is respectively standardized using the robust scale method, and the standardized data is used as the dataset of the model.
[0141] 3. According to the annotation information of experts, sleep apnea events calibration is performed on the standardized data, specifically including: sleep apnea events and hypopnea events. In the present invention, these two events are uniformly regarded as sleep apnea events. When annotating, the start time and end time of each event should be indicated.
[0142] The second part is the network model design (step S3):
[0143] 1. In order to utilize the complementarity between multimodal data, eliminate the redundancy between modalities, and improve the feature representation ability of the model, this example designs a feature extraction module composed of residual convolution blocks. As Figure 2 shown, the feature extraction module consists of two parts: single-modal feature extraction and multimodal feature fusion. The multimodal data respectively obtains single-modal feature representations through the single-modal feature extraction part, and the single-modal features are concatenated in the order of oro-nasal airflow, thoracic pressure, and abdominal pressure to obtain unfused multimodal features. Finally, the feature vector is input into the multimodal feature fusion part to eliminate the redundancy between modalities and enhance the complementarity between modalities, obtaining multimodal fusion features. In addition, the single-modal feature extraction parts for different modality data are independent of each other and do not share weights.
[0144] 2. A multi-scale prediction feature layer generation module is designed in this example for the inherent uncertainty of the duration of sleep apnea events. The prediction feature layer generated by this module realizes multi-scale in both network depth and network width. The multi-scale in depth and width can enhance the richness of features. As Figure 3 shown, the multimodal features pass through several Inception structures and residual convolution blocks to obtain 4 feature maps of different scales, and then a feature pyramid network is used to fuse the feature maps of different scales to obtain a multi-scale prediction feature layer.
[0145] 3. To generate candidate boxes for sleep apnea events, a region generation module is designed in this example. This module needs to set the size of the preset boxes. Considering the receptive field sizes of different prediction feature layers and the distribution of the durations of real sleep apnea events comprehensively, in this example, the sizes of the preset boxes on prediction feature layers 1, 2, 3, and 4 are set to 224 (7s), 448 (14s), 672 (21s), and 1344 (42s) respectively. The region generation module generates foreground probability values and position correction offsets for each preset box through two fully connected layers, and uses the generated offsets to correct the positions of the preset boxes to obtain all candidate boxes. Non-maximum suppression processing is performed on the candidate boxes of each prediction feature layer to remove the candidate boxes with too high overlap rate. Finally, the N candidate boxes with the largest probability values are selected from all candidate boxes as the output candidate boxes of the region generation module.
[0146] 4. To enable the final classification module to process candidate boxes of different scales, this example adopts the Region of Interest (ROI) Pooling method to unify the scales of the candidate region feature maps. First, the candidate boxes output by the region generation module are mapped to the corresponding prediction feature layers to obtain candidate region feature maps, and then the candidate region feature maps with different scales are unified in scale by using the ROI Pooling method. In this example, the unified length is set to 24.
[0147] 5. To classify the candidate segments, a final classification module and a post-processing module are designed in this example. The final classification module generates class probability values and position correction offsets for each candidate feature map; the post-processing module uses the offsets to correct the positions of the candidate boxes, then adjusts the positions of the out-of-bounds candidate boxes to the signal boundaries and removes the candidate boxes with sizes less than 128 (4s). Finally, non-maximum suppression processing is performed on all candidate boxes of each data, and the k candidate boxes with the largest foreground probability are taken out as the final detection results.
[0148] The third part is the training and testing of the network model (Steps S4, S5)
[0149] 1. This example adopts an end-to-end training method to train the network model. All sample data are divided into a training set, a validation set, and a test set, and a 10-fold cross-validation method is used for training and validation. Specifically, the stochastic gradient descent method is used to train for 30 epochs, and during the training process, the learning rate starts from 0.01 and becomes the original
[0150] 2. Use the trained network model above to detect the test data.
[0151] Example 2
[0152] Based on the same inventive concept, this embodiment provides a sleep apnea event localization device based on a multi-scale convolutional neural network. Please refer to Figure 6 , the device includes:
[0153] An original data extraction module 201, configured to extract oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals from the original polysomnogram, and splice the data of the three channels into a vector where X i represents the data extracted from the i-th polysomnogram, respectively representing the oral and nasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the i-th polysomnogram;
[0154] A data preprocessing module 202, configured to preprocess the data of the three extracted channels, specifically including: first performing data resampling, and performing filtering processing and normalization processing channel by channel, and then performing sleep apnea event annotation on the normalized data, and the annotation content includes the event start time and the event end time;
[0155] A network construction module 203, configured to construct a multi-scale convolutional neural network. The multi-scale convolutional neural network includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module, and a final classification module. Among them, the feature extraction module is configured to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features. The multi-scale prediction feature layer generation module is configured to obtain multiple prediction feature layers according to the fused multi-modal features. The region generation module is configured to generate candidate boxes on the prediction feature layer. The region of interest pooling module is configured to obtain the feature maps corresponding to the candidate boxes and unify the sizes of different feature maps. The final classification module is configured to classify the feature maps output by the region of interest pooling module and output the offset;
[0156] A training module 204, configured to input the preprocessed data into the multi-scale convolutional neural network for training;
[0157] A localization detection module 205, configured to input the sleep data to be detected into the trained multi-scale convolutional neural network for localization prediction to obtain a classification result and an offset, and perform post-processing according to the obtained classification result and offset to obtain a final detection result.
[0158] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the sleep apnea event localization method based on a multi-scale convolutional neural network in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be repeated here. Any device adopted in the method of the first embodiment of the present invention belongs to the scope protected by the present invention.
[0159] Embodiment 3
[0160] Based on the same inventive concept, please refer to Figure 7 , the present invention also provides a computer-readable storage medium 300, on which a computer program 311 is stored, and when the program is executed, the method described in Embodiment 1 is implemented.
[0161] Since the computer-readable storage medium introduced in Embodiment 3 of the present invention is the computer-readable storage medium adopted for implementing the sleep apnea event localization method based on a multi-scale convolutional neural network in Embodiment 1 of the present invention, based on the method described in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium adopted by the method of Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0162] Embodiment 4
[0163] Based on the same inventive concept, the present application also provides a computer device, as Figure 8 shown, including a memory 401, a processor 402, and a computer program 403 stored on the memory and executable on the processor. When the processor executes the above program, the method in Embodiment 1 is implemented.
[0164] Since the computer device introduced in Embodiment 4 of the present invention is the computer device adopted for implementing the sleep apnea event localization method based on a multi-scale convolutional neural network in Embodiment 1 of the present invention, based on the method described in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer device, so it will not be elaborated here. Any computer device adopted by the method in Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0165] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0166] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.
[0167] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0168] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for sleep apnea event localization based on a multi-scale convolutional neural network, characterized in that Including: S1: Extract the oral-nasal airflow signal, thoracic pressure signal, and abdominal pressure signal from the original polysomnogram, and splice the data of the three channels into a vector. where X i represents the data extracted from the i-th polysomnogram, respectively representing the oral-nasal airflow signal, thoracic pressure signal, and abdominal pressure signal in the i-th polysomnogram; S2: Preprocess the data of the three channels obtained by extraction, specifically including: First, perform data resampling, and perform filtering processing and normalization processing channel by channel, and then perform sleep apnea event annotation on the normalized data, and the annotation content includes the event start time and the event end time; S3: Construct a multi-scale convolutional neural network. The multi-scale convolutional neural network includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module, and a final classification module. Among them, the feature extraction module is used to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features. The multi-scale prediction feature layer generation module is used to obtain multiple prediction feature layers according to the fused multi-modal features. The region generation module is used to generate candidate boxes on the prediction feature layer. The region of interest pooling module is used to obtain the feature maps of the regions corresponding to the candidate boxes and unify the sizes of different feature maps. The final classification module is used to classify the feature maps output by the region of interest pooling module and output the offset; S4: Input the preprocessed data into the multi-scale convolutional neural network for training; S5: Input the sleep data to be detected into the trained multi-scale convolutional neural network for localization prediction to obtain the classification result and the offset, and perform post-processing according to the obtained classification result and the offset to obtain the final detection result.
2. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, wherein Step S2 includes: S2.1: Resample the data of the three channels obtained by extraction, and use an 8th-order Butterworth filter to perform low-pass filtering on the data of the three channels respectively; S2.2: For X i For the data of channel k in Calculate the median and the interquartile range Use the robust normalization method to normalize the data of channel k: Among them, represents the standardized data; S2.3: Perform sleep apnea event annotation on the normalized data, regard sleep apnea events and hypopnea events as sleep apnea events, and annotate the start time and end time of each event.
3. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, characterized in that During the training process of step S4, the processing process of the feature extraction module includes: Extract the single-modal features of the input data through three parallel and independent single-modal feature extraction modules; Concatenate the extracted single-modal features into multi-modal features in the order of oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals, and then fuse the multi-modal features through 2 residual convolutional blocks and 2 downsampling residual convolutional blocks to obtain the fused multi-modal features.
4. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, characterized in that During the training process of step S4, the processing process of the multi-scale prediction feature layer generation module includes: Pass the fused multi-modal features obtained by the feature extraction module through the Inception structure to obtain feature map 1. The Inception structure consists of 3 branches: There is only 1 convolutional layer with a convolutional kernel size of 1*1 on the first branch; There are 2 convolutional layers on the second branch, and the convolutional kernel sizes are 1*1 and 3*1 respectively; There are 2 convolutional layers on the third branch, and the convolutional kernel sizes are 1*1 and 5*1 respectively; Finally, the feature maps obtained by the 3 branches are concatenated in the channel dimension to obtain feature map 1; Pass feature map 1 through 1 downsampling residual convolutional block and then input the result into the Inception structure to obtain feature map 2; The feature map 2 is input into the Inception structure after passing through 1 downsampling residual convolution block to obtain the feature map 3; The feature map 3 is input into the Inception structure after passing through 1 downsampling residual convolution block to obtain the feature map 4; The feature maps 1, 2, 3, and 4 are input into the feature pyramid network to obtain the prediction feature layers 1, 2, 3, and 4.
5. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, characterized in that, The multiple prediction feature layers include the prediction feature layers 1, 2, 3, and 4. During the training process of step S4, the processing process of the region generation module includes: Preset boxes with different lengths are generated respectively on the prediction feature layers 1, 2, 3, and 4; Sliding on the predictive feature layer using a preset sliding window, and using two convolutional layers with a convolutional kernel size of 1*1 to generate one probability value and two offsets for each sliding window, where the probability value represents the probability that the preset box is the target object, and the two offsets are (t x , t w ), t x is the offset used to correct the center coordinates of the preset box, and t w is the offset used to correct the width of the preset box; The positions of the preset boxes are corrected using the generated offsets to obtain candidate boxes; Non-maximum suppression processing is performed on the candidate boxes of each prediction feature layer to obtain the processed candidate boxes of each prediction feature layer. After merging all the processed candidate boxes of the prediction feature layers, the N candidate boxes with the largest probability values are selected as the output candidate boxes.
6. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, characterized in that During the training process of step S4, the processing process of the region of interest pooling module includes: Mapping the candidate boxes to different prediction feature layers according to the following formula: where w represents the width of the candidate box, scale is the scale factor, FS is the signal sampling rate, and k is the layer number of the prediction feature layer to which the candidate box is mapped; The region of interest pooling method is used to perform scale unification processing on the feature maps of the regions corresponding to the candidate boxes.
7. The sleep apnea event localization method based on a multi-scale convolutional neural network according to claim 1, characterized in that The final classification module includes two fully connected layers. During the training process of step S4, the processing process of the final classification module includes: Two probability values are obtained through a fully connected layer according to the input feature map, representing the probability that the content in the candidate box corresponding to the feature map is the background and the probability that it is the foreground respectively; Two offsets (t′ x , t′ w ) are obtained from the input feature map through another fully connected layer. t′ x is the offset for correcting the center coordinates of the candidate bounding box obtained by the final classification module, and t′ w is the offset for correcting the width of the candidate bounding box obtained by the final classification module.
8. A sleep apnea event localization device based on a multi-scale convolutional neural network, characterized in that, including: An original data extraction module, configured to extract oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals from an original polysomnogram, and splice the data of the three channels into a vector where X i represents the data extracted from the i-th polysomnogram, respectively representing the oral and nasal airflow signals, thoracic pressure signals, and abdominal pressure signals in the i-th polysomnogram; The data preprocessing module is used to preprocess the data of the three extracted channels. Specifically, it first performs data resampling, and then performs filtering processing and normalization processing channel by channel. Then, sleep apnea event annotation is performed on the normalized data, and the annotation content includes the event start time and the event end time; The network construction module is used to construct a multi-scale convolutional neural network. The multi-scale convolutional neural network includes a feature extraction module, a multi-scale prediction feature layer generation module, a region generation module, a region of interest pooling module, and a final classification module. Among them, the feature extraction module is used to extract the features of each modality from the input multi-modal data and then perform feature fusion to obtain the fused multi-modal features. The multi-scale prediction feature layer generation module is used to obtain multiple prediction feature layers according to the fused multi-modal features. The region generation module is used to generate candidate boxes on the prediction feature layers. The region of interest pooling module is used to obtain the feature maps of the regions corresponding to the candidate boxes and unify the sizes of different feature maps. The final classification module is used to classify the feature maps output by the region of interest pooling module and output the offsets; The training module is used to input the preprocessed data into the multi-scale convolutional neural network for training; The positioning detection module is used to input the sleep data to be detected into the trained multi-scale convolutional neural network for positioning prediction, obtain the classification result and the offset, and perform post-processing according to the obtained classification result and offset to obtain the final detection result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed, it implements the method described in any one of claims 1 to 7.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Breathing machine man-machine asynchronous classification method, terminal and storage medium
CN113539501A
Method and system for training neural network model based on multi-terminal data, and medium
CN114742100A