Obstructive sleep apnea detection method and system based on multi-scale convolutional neural network, storage medium and electronic equipment

By adopting multi-scale convolutional neural network and attention mechanism in OSA detection, the problem of limitations in the feature extraction function and insufficient ability of model to adapt to different scales in the prior art is solved, and efficient OSA detection performance is achieved.

CN120203520APending Publication Date: 2025-06-27HENAN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510391448.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as feature extraction function limitations and insufficient ability of model adaptation to different scales in obstructive sleep apnea (OSA) detection, resulting in poor detection performance.

Method used

Using a detection method based on multi-scale convolutional neural network (CNN) to build three different time scales of CNN models, dynamically pay attention to the importance of ECG signal fragments and neural network channels, and combine residual attention and channel attention modules to enhance feature extraction capabilities.

Benefits of technology

It significantly improves the feature characterization and detection performance of OSA detection, realizes effective fusion of data on different time scales and pays attention to key features, and improves the adaptability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120203520A_ABST
    Figure CN120203520A_ABST
Patent Text Reader

Abstract

The invention discloses an obstructive sleep apnea detection method and system based on a multi-scale convolutional neural network, a storage medium and electronic equipment. The obstructive sleep apnea detection method specifically comprises the following steps: loading single-lead electrocardiogram (ECG) data, and preprocessing the data; a P-R interval is determined; calculating by using Euclidean distance to obtain a two-dimensional distance array of the five-minute segment; further intercepting the 5-minute fragment into a 3-minute fragment and a 1-minute fragment; constructing three different CNN models, and respectively inputting the maximum distance and the minimum distance in the three scale fragments into the three CNN models; the convolution feature map of each model is optimized, and important information in the features is enhanced through a residual attention module; applying a channel attention module to optimize the attention capability of the model to the key features; and integrating the feature maps subjected to residual error and channel attention optimization, performing final classification output through a full connection layer, and predicting whether OSA (Obstructive Sleep Apnea) exists or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of convolutional neural networks, and in particular, to a method, system, storage medium, and electronic device for detecting obstructive sleep apnea based on a multi-scale convolutional neural network. Background Art

[0002] Currently, sleep apnea (SA) is one of the common sleep-related breathing disorders, including obstructive sleep apnea (OSA) and central sleep apnea (CSA). OSA is mainly caused by the repeated partial or complete collapse of the upper airway during sleep, resulting in reduced airflow and sleep interruption. In contrast, CSA is caused by insufficient ventilation and impaired gas exchange due to a lack of respiratory drive during sleep. OSA is also associated with hypoxemia and increased sympathetic nervous system activity. Research has shown that OSA can lead to problems such as hypertension, cardiovascular, and cerebrovascular diseases. Therefore, early detection of OSA is very important for reducing the onset of other related diseases.

[0003] In laboratory-based sleep studies, polysomnography (PSG) is considered the gold standard for diagnosing OSA. PSG uses electroencephalogram, electrocardiogram, electrooculogram, electromyogram, pulse oximetry, and airflow measurements to examine sleep and respiratory parameters as diagnostic criteria. PSG can provide detailed information about sleep structure, duration, and quality, with high diagnostic accuracy. However, in real life, PSG devices require a large number of connectors to record data and measure sleep in the laboratory, which may affect sleep and thus affect the detection of OSA. In addition, PSG is expensive, time-consuming, and impractical. Therefore, it is necessary to provide an alternative method for the early diagnosis and detection of OSA, while improving patient convenience and reducing costs.

[0004] In recent years, some researchers have proposed methods for detecting OSA using simple means, such as blood oxygen channels, electrocardiogram channels, electroencephalogram channels, etc. Although pulse oximeters can easily collect blood oxygen through simple data processing, they can only display the blood oxygen saturation (SpO2) level and cannot provide detailed information about sleep apnea or distinguish between CSA and OSA. Among them, electrocardiogram is one of the most commonly used signals for detecting OSA. Electrocardiogram signals can accurately detect subtle changes in the heart, such as heart rate variability (HRV) and R-R interval, which are particularly obvious during sleep apnea events. Therefore, using electrocardiogram signals to study obstructive sleep apnea syndrome has importance and practical significance. This can provide valuable insights into the condition and its impact on heart health and sleep quality.

[0005] The performance of the single-task learning methods described above is not satisfactory, so multi-task feature fusion is needed. Some methods use multi-task learning methods based on a combination of unsupervised feature learning and supervised feature learning to improve the feature extraction performance of the network. Some methods use multi-scale residual networks to extract features from the original signal and obtain sensitive features from various angles. Other methods use end-to-end spatiotemporal learning with multiple spatiotemporal blocks, each with multiple layers. However, these deep learning models still have some shortcomings. Their feature selection and feature extraction capabilities limit the OSA detection performance. In addition, these models usually use a unified network architecture for feature extraction, ignoring the need to design specialized deep learning models for segmentation at different scales, thereby limiting the model's ability to adapt to different scales. These defects may lead to limited comprehensiveness of feature extraction and further improve detection performance. Summary of the invention

[0006] The purpose of the present invention is to provide an obstructive sleep apnea detection method, system, storage medium and electronic device based on a multi-scale convolutional neural network, which can dynamically focus on the importance of different ECG signal segments and neural network channels, effectively enhance the feature extraction capability, and ultimately achieve excellent OSA detection performance.

[0007] The technical solution adopted by the present invention is:

[0008] A method for detecting obstructive sleep apnea based on a multi-scale convolutional neural network comprises the following steps:

[0009] Step 1: Loading ECG data, wherein each record of the ECG data includes a continuous digitized ECG signal, a set of apnea annotations, and a set of machine-generated QRS annotations;

[0010] Step 2: Preprocess the ECG data: first remove the unlabeled segments, then cut the segments and remove the noise to obtain segments with a time length of t1 minutes;

[0011] Step 3: Detect the positions of the R peak and P peak of the fragments obtained in the previous step, and determine the PR interval;

[0012] Step 4: Use the Euclidean distance to calculate the maximum distance X between each PR interval in each segment. max and the minimum distance X min , get the two-dimensional distance array corresponding to each fragment;

[0013] Step 5: Cut the middle part of each segment into a t2-minute segment by sampling, and then further cut out a t3-minute segment in the middle of the t2-minute segment, where t3 <t2<t1;

[0014] Step 6: Construct three different CNN models to meet the feature extraction requirements of different time scales, and input the X of the three ECG segments with different durations intercepted in Step 5 max and X min into the three CNN models respectively for adjustment and extraction to obtain the corresponding convolutional feature maps;

[0015] Step 7: Optimize the convolutional feature maps output by each CNN model, strengthen the important information in the features through the residual attention module, so that each CNN model can better capture local and global features;

[0016] Step 8: Integrate the three attention-weighted feature maps output by the residual attention, and apply the channel attention module to further weight the features of each channel to optimize the overall model's ability to focus on key features.

[0017] Step 9: Pass the weighted feature maps output by the channel attention through the fully connected layer for final classification output to predict whether obstructive sleep apnea (OSA) exists.

[0018] The specific steps of Step 2 are as follows: First, some unlabeled segments are deleted; then, a 1-minute segment corresponding to the label is selected as a reference in the ECG record, and 2-minute signal segments before and after the labeled segment are extracted to form a total length of 5 minutes, that is, t1 = 5 min; finally, the finite impulse response (FIR) filter is used to limit the signal frequency range to 3 - 45 Hz to remove noise.

[0019] The specific steps of Step 3 use the Hamilton algorithm to detect the R peaks and delete those signal segments near the start or end of the signal segment. Specifically, if the number of R peaks in the 5-minute signal segment is less than 40 or greater than 200, it means that the noise seriously interferes with this segment, and we directly discard this segment; according to the processed data, we determine the P wave position and the P-R interval; the P wave appears before the R peak, so the P wave is located in the window before each R peak; by locating the position of the P peak, the P-R interval of each segment can be located.

[0020] The specific steps of Step 4 are as follows:

[0021] By locating the P-R interval within each segment, a time series A with a length of m sampling points starting from the P peak is constructed m , as shown in Equation (1):

[0022]

[0023] In Equation (1), represents a subsequence of the time series, defined as a continuous segment with a length of m samples; then according to Am Define the distance profile matrix D of the ECG signal, which represents the distribution relationship of the distances between each pair of time series, as shown in Equation (2):

[0024]

[0025] where represents the Euclidean distance between two subsequences and The minimum value X min and the maximum value X max in each column of the matrix D are derived as the input segments of the model, denoted as [X min , X max ;

[0026] Normalize X min and X max in the range of [0, 1], and perform normalization according to Equation (3); X represents the set of all X max or X min ; in the 5-minute segment, min(X) represents the minimum value in the set X, and max(X) represents the maximum value in the set X;

[0027]

[0028] The specific steps of step 5 include the following steps: t1 = 5 min, t2 = 3 min, t3 = 1 min; then:

[0029] For each 5-minute segment, we intercept 540 sampling points from 181 to 720 in the middle of the segment to generate a 3-minute segment, and then further intercept 180 sampling points from 361 to 720 in the middle of the 3-minute segment to generate a 1-minute segment.

[0030] In step 6, three different convolutional neural networks CNN-1, CNN-2, and CNN-3 are specifically adopted to form a multi-scale CNN structure; the multi-scale CNN structure is improved based on the framework of AlexNet. AlexNet has 5 convolutional layers, 3 max-pooling layers, 2 normalization layers, 2 fully connected layers, and a SoftMax layer; each convolutional layer has a ReLU activation function, which does not set the gradient of negative inputs to zero, solves the problem of "neuron death" and helps to maintain the gradient flow; the non-linear Leaky ReLU activation function is expressed as Equations (6) and (7);

[0031]

[0032] where W1 and b1 represent the weights and biases of the shared fully connected layer, σ1 represents the ReLU activation function, Z avgDenotes the average value of all time step data on the d-th feature dimension, Z max Denotes the maximum value of each feature d within the time step range.

[0033] In step 6 mentioned above, corresponding adjustments were made to the configuration of the convolutional layer, the number of filters, and the pooling strategy among different CNNs: CNN-1 processes enhanced segments with a duration of 5 minutes and has 5 convolutional layers. Each layer consists of one-dimensional convolution, batch normalization, the Leaky ReLU activation function, and max pooling, except that the fourth layer does not include max pooling;

[0034] CNN-2 processes enhanced segments with a duration of 3 minutes and also has 5 convolutional layers. However, due to the different input lengths from CNN-1, the max pooling in the third layer was removed to ensure that the output feature maps have the same shape; in addition, to adapt to the 3-minute segments, we modified the number of one-dimensional convolutional filters and the size of the convolutional kernels to simplify the model;

[0035] CNN-3 processes enhanced segments with a duration of 1 minute: Since the data carried by the 1-minute segments is less than that of the 3-minute and 5-minute segments, we only use the first three layers of the CNN-2 model to further simplify the model to adapt to the 1-minute segments, ensuring the consistency of features at different time scales during fusion. Finally, the X max and X min in the 3 types of scale segments are respectively input into 3 types of CNN models.

[0036] An obstructive sleep apnea detection system based on a multi-scale convolutional neural network, comprising:

[0037] (1) An ECG data preprocessing module, used to filter and standardize the original ECG signal to remove noise and standardize the data format;

[0038] (2) A P-R interval determination module, used to detect the positions of the P wave and R wave in the ECG signal and calculate the P-R interval;

[0039] (3) A segment truncation module, used to divide the ECG signal into fixed-length segments according to a time window; each segment corresponds to a 5-minute signal and consists of 900 sampling points; during segment division, it is truncated according to the structure of the first 2 minutes, the marked area of 1 minute, and the last 2 minutes;

[0040] (3) A distance calculation module, used to calculate the Euclidean distance between time series segments in the ECG signal to quantify the similarity of different electrocardiogram patterns;

[0041] (4) Feature extraction module, which extracts multi-scale features based on three CNN models with different time scales; the models capture local and global features of time series signals at different scales to enhance the model's ability to express time correlation; Convolution optimization module, which extracts high-level features through multi-layer convolution operations, and combines Batch Normalization (BatchNorm), LeakyReLU activation, Max Pooling (MaxPooling) and Dropout layers to optimize the model's expression ability and prevent overfitting;

[0042] (5) Residual attention module, which is used to enhance the model's ability to focus on key features of ECG signals by combining residual connections and attention mechanisms, while retaining original feature information;

[0043] (6) Channel attention module, which is used to apply channel weights to multi-channel ECG feature maps, automatically assign attention weights to different channels, and enhance the attention to key channels; Feature fusion judgment module, which is used to fuse the multi-scale features extracted by the three time-scale CNN models and judge the final output.

[0044] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the device where the computer-readable storage medium is located executes the obstructive sleep apnea detection method based on a multi-scale convolutional neural network.

[0045] An electronic device, comprising: a memory and a processor. A program that can run on the processor is stored on the memory. When the processor executes the program, it implements the obstructive sleep apnea detection method based on a multi-scale convolutional neural network.

[0046] The present invention detects OSA by constructing a multi-scale convolutional neural network (CNN) model; the model adopts three different CNN structures to process ECG signals with different time lengths. By parallel processing ECG signal segments with different time scales, the model can capture local details and global trends simultaneously, significantly improving the feature representation ability. At the same time, the model also combines residual attention and channel attention to solve the problems of gradient disappearance and explosion, to identify and focus on key features of ECG signals. The present invention makes full use of the complementary advantages of long-term and short-term segments by using a multi-scale convolutional neural network to process data segments with different time scales. In addition, the model proposed by us introduces residual attention and channel attention to improve the detection performance of the model by integrating time information from adjacent ECG segments. Moreover, extensive experiments conducted by this application on public datasets including Apnea-ECG and UCD verify the performance of the model, providing an efficient and reliable solution for the automatic screening of OSA, and having important clinical practical value. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 is the flowchart of the present invention;

[0049] Figure 2 is the schematic diagram of the model structure described in the present invention;

[0050] Figure 3 The loss curve and accuracy curve diagram during the training of the present invention;

[0051] Figure 4 is the ROC curve diagram of the repeated training of the present invention. Specific embodiments

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0053] As Figure 1 、 2 and shown in 3, the present invention includes the following steps:

[0054] Step 1: Load ECG data. Each record of the ECG data includes a continuous digitized electrocardiogram signal, a set of apnea annotations, and a set of machine-generated QRS annotations.

[0055] Step 2: Data preprocessing;

[0056] During the preprocessing, we deleted some unlabeled segments. Then, a 1-minute segment corresponding to the label was selected as a reference in the electrocardiogram record, and 2-minute signal segments before and after the labeled segment were extracted to form a total length of 5 minutes. Then, a finite impulse response (FIR) filter was used to limit the signal frequency range to 3 - 45 Hz to remove noise.

[0057] Step 3: Detect the positions of the R peaks and P peaks, and determine the P-R interval; specifically including:

[0058] Use the Hamilton algorithm to detect R-peaks and remove those signals located near the start or end of the signal segment. If the number of R-peaks in a 5-minute signal segment is less than 40 or greater than 200, it indicates that noise seriously interferes with that segment, and we directly discard that segment. Based on the processed data, determine the P-wave position and the P-R interval. The P-wave appears before the R-peak, so the P-wave is located in the window before each R-peak. By locating the position of the P-peak, the P-R interval of each segment can be located.

[0059] Step 4: Calculate the maximum distance X between each P-R interval within each 5-minute segment using the Euclidean distance max and the minimum distance X min , to obtain the two-dimensional distance array of this 5-minute segment; specifically including:

[0060] By locating the P-R intervals within each segment, construct a time series A of length m sampling points starting from the P-peak m , as shown in Equation (1):

[0061]

[0062] In Equation (1), represents a subsequence of the time series, defined as a continuous segment of length m samples. Then, based on A m define the distance profile matrix D of the ECG signal, which represents the distribution relationship of the distances between each pair of time series, as shown in Equation (2):

[0063]

[0064] where represents the Euclidean distance between two subsequences and . The Euclidean distance is a measure of the distance between two points in a multi-dimensional space. Since the ECG signal is time series data, the Euclidean distance effectively quantifies the numerical difference between two time series on the same time slice, thus providing an intuitive measure of similarity. It can handle the continuous data in the ECG signal and reflect the influence of small changes in the ECG signal. This distance measurement helps to accurately detect and compare different electrocardiogram patterns.

[0065] Since the minimum and maximum values can effectively capture the extreme similarity and differences between time series, they provide one of the most representative features for each pair of time series. This method can significantly reduce the dimension of the input features while retaining the most important information, helping the model to better identify and classify signals. Therefore, the minimum and maximum values of the Euclidean distance for each pair of time series are considered. The minimum value X min and the maximum value X in each column of matrix Dmax The input segment derived as the model, denoted as [X min , X max .

[0066] Normalize X min and X max within the range of [0, 1], and perform normalization according to Equation (3). X represents the set of all X max or X min . In the 5-minute segment, min(X) represents the minimum value in the set X, and max(X) represents the maximum value in the set X.

[0067]

[0068] Step 5: Further segment the 5-minute segment into 3-minute and 1-minute segments; specifically including:[[]]

[0069] For each 5-minute segment, we intercept a total of 540 sampling points from the 181st to the 720th at the middle of the segment to generate a 3-minute segment, and then further intercept a total of 180 sampling points from the 361st to the 720th at the middle of the 3-minute segment to generate a 1-minute segment.

[0070] Step 6: Construct three different CNN models, and input X max and X min in these three scales of segments into the three CNN models respectively; specifically including:[[]]

[0071] Electrocardiogram signals are periodic, diverse, non-stationary, and vulnerable to noise. Using ECG signal segments of different lengths can provide richer information, thereby improving the accuracy of classification and prediction. Shorter segments can capture short-term detailed features, while longer segments help capture long-term trends and rhythm changes. In addition, by using segments of different lengths, we can increase the diversity of training data, prevent the model from overfitting, and improve the generalization ability of the model. We designed three different convolutional neural networks CNN-1, CNN-2, and CNN-3 because too many networks will lead to high computational costs, while too few networks may not be beneficial for feature extraction.

[0072] AlexNet is a deep convolutional neural network that achieved remarkable performance in the 2012 ImageNet Large Scale Visual Recognition Challenge. The AlexNet model has 5 convolutional layers, 3 max-pooling layers, 2 normalization layers, 2 fully-connected layers, and a SoftMax layer. Each convolutional layer has a ReLU activation function. AlexNet effectively extracts multi-level features and performs feature fusion through a deep learning structure, thereby improving the classification accuracy. AlexNet uses ReLU activation and dropout techniques between the convolutional layer and the fully-connected layer. Similarly, using ECG signals, we invented a convolutional neural network with one-dimensional convolutions. In addition, we introduced batch normalization into the network to stabilize the training process and accelerate convergence. We use the Leaky ReLU activation function instead of ReLU because it does not set the gradient of negative inputs to zero, solves the problem of "neuron death" and helps maintain the gradient flow, which is particularly suitable for processing complex data such as electrocardiogram signals.

[0073] ECG signals are essentially one-dimensional time series data, but can be extended to a two-dimensional-like data structure through different channel processing methods. To improve the classification performance, we adopted a multi-level feature extraction and fusion strategy. An improved version of AlexNet was used to better adapt to the classification of ECG signals. In addition, we used multiple convolutional neural networks for different scales of data to achieve more effective feature extraction. In this invention, we designed three different convolutional neural networks, CNN-1, CNN-2, and CNN-3, to process 5-minute, 3-minute, and 1-minute ECG segments respectively to meet the feature extraction requirements of different time scales. Among them, the CNN structure for longer time segments is more complex to capture long-term dependencies, while the CNN structure for shorter time segments is simplified to reduce computational complexity and adapt to less data volume. At the same time, the configuration of the convolutional layer, the number of filters, and the pooling strategy were adjusted accordingly between different CNNs to ensure the consistency of features at different time scales during fusion. The X max and X min in the 3-scale segments were respectively input into the 3 CNN models, and the model structures are shown in Figure 2 .

[0074] Step 7: Optimize the convolutional feature maps of each model, and strengthen the important information in the features through the Residual Attention Module to enable the model to better capture local and global features; specifically including:

[0075] The residual attention module introduces residual connections and attention mechanisms to enhance the model's perception of local features and attention calculation ability, which helps capture subtle feature changes and key information in the input data, thus better capturing long-term dependencies in the data. By introducing residual connections, the model can directly pass information from previous time steps to subsequent time steps, alleviating the problem of vanishing gradients and more effectively learning patterns and trends in time series data. The attention mechanism helps the model dynamically adjust the focus at different time points at each time step, thereby enhancing the model's performance. In the proposed model, the residual attention mechanism is used to enhance the model's attention to important features and improve the overall performance. By dynamically calculating the importance weights of input features, the residual attention mechanism enables the model to focus more on key feature regions relevant to the classification task. This dynamic weighting can effectively suppress background noise or irrelevant information, thus highlighting features meaningful to the task. In addition, with the help of residual connections, the model retains the original input information while integrating the enhanced features extracted by the attention module, achieving an organic combination of global and local information. The residual attention mechanism also provides the model with stronger generalization ability. Under different data distributions or sample features, it adaptively adjusts the importance distribution of features, thereby improving the model's robustness and classification accuracy for new samples.

[0076] Step 8: Apply the Channel Attention Module to further weight the features of each channel to optimize the model's ability to focus on key features; specifically including:

[0077] The channel attention module focuses on the importance of each channel. By dynamically adjusting the weights of each channel through pooling and fully connected layers, the model improves the efficiency of using information from different channels. This enables the model to more effectively integrate information from each channel, thereby improving the overall performance. The channel attention mechanism is used to extract the importance of each channel in the input tensor. By adopting the channel attention mechanism, we can effectively integrate the feature maps of three different scales processed by the residual attention mechanism. The feature map output by the residual attention serves as the input. Initially, the input tensor undergoes global max pooling and average pooling to capture two different feature descriptor vectors, as shown in Equations (4) and (5), representing the outputs of average pooling and max pooling respectively.

[0078]

[0079] Z max = max t∈{1,…,t} X td (5)

[0080] Average pooling obtains the pooling result by calculating the average value of all elements in the pooling window. In formula (4), T represents the total number of time steps in the current pooling window. X td represents the value of the t-th time step in the d-th element window. Z avg represents the average value of all time step data on the d-th feature dimension. Max pooling obtains the pooling result by selecting the maximum value of all elements in the pooling window. In equation (5), T represents the total number of time steps in the current pooling window. X td represents the value of the t-th time step in the d-th element window. Z avg represents the maximum value of each feature d within the time step range.

[0081] Next, these two feature description vectors generate two attention weights through a shared fully connected layer, where the ReLU activation function is used to introduce non-linearity. It can be expressed as equations (6) and (7).

[0082]

[0083] where W1 and b1 represent the weights and biases of the shared fully connected layer, and σ1 represents the ReLU activation function. Then, the two attention weight vectors A avg and A max are added together, as shown in equation (8), to consider the importance information within the channel. Then, the combined feature representation is mapped to an attention weight vector through another fully connected layer, where σ2 represents the sigmoid activation function. The sigmoid activation function is used to limit the weights between 0 and 1, thereby obtaining the final channel attention weights.

[0084] A = σ2(A avg + A max ) (8)

[0085] Finally, the input tensor X representing the residual weighted feature map is multiplied by the calculated channel attention weights to obtain the channel weighted feature map Y, as shown in equation (9).

[0086]

[0087] Step 9: Integrate the feature maps optimized by the residual and channel attention, and perform the final classification output through a fully connected layer to predict whether obstructive sleep apnea (OSA) exists; specifically including:

[0088] First, extract the optimized feature maps from the 1-minute, 3-minute, and 5-minute segments, and integrate the features of different time scales through feature weighted fusion; then, use global average pooling for dimensionality reduction to obtain the final feature vector; second, input the feature vector into the fully connected layer, and use the Leaky ReLU activation function for feature transformation, while adding a Dropout mechanism to prevent overfitting; finally, calculate the probability of OSA occurrence through the Softmax classifier, and output the final classification result based on a set threshold (such as 0.5), that is, OSA positive or OSA negative.

[0089] This invention was experimented on a computer with a 11th Gen Intel(R) Core(TM) i5-1155G7@2.50GHz processor, NVIDIA GeForce RTX 4090, and 128GB RAM memory configuration. According to the method of this invention, our model is very effective in improving the accuracy, sensitivity, specificity, and AUC of classification. Our invention can effectively distinguish apnea events and normal events in the data. We have verified the effectiveness of this model on the Apnea-ECG dataset and the UCD dataset. In segment-based classification, the average accuracy of the Apnea-ECG dataset is 92.82%, the sensitivity is 90.46%, and the specificity is 94.29%. While on the UCD dataset, the verification results show an accuracy of 90.73%, a sensitivity of 88.16%, and a specificity of 90.81%. The results indicate that on the two public datasets, the average accuracy of segment-based OSA detection reaches 92.82% and 90.73% respectively, which is 1.89% and 7.28% higher than other similar models. For the study of sleep apnea syndrome based on deep learning, it is crucial to compare with other advanced OSA detection methods. Table 1 shows the result comparison between this invention and other advanced SA detection methods in the segment-based classification scenario. Most of the models listed in Table 1 use CNN or CNN variants to calculate the input data, and the input data types are mostly R-peaks and R-R intervals. The classifier of our proposed model is a multi-scale CNN. After ECG processing, the input data type is [X min , X max . This model has achieved good results, with accuracy, specificity, and AUC superior to other models. This indicates that this invention is effective in data preprocessing strategies and multi-scale convolutional neural network architectures. The combination of X min and X max can effectively capture the dynamic range of the electrocardiogram signal, and the design of the multi-scale CNN enhances the model's ability to learn features of different time scales. Compared with traditional methods, this invention has achieved a good balance between simplifying input features and enhancing model performance.

[0090] The specific experimental results are as follows Figure 3 and Figure 4 shown below Figure 3 illustrates the loss curve and accuracy curve during the training process. In the figure, the horizontal axis is the epoch, and the vertical axis is the loss or accuracy. Different lines represent the loss or accuracy of training or validation. It can be observed from the figure that the losses of both the training set and the validation set decrease as the epoch increases, and there are slight differences between them. In contrast, the accuracies of both the training set and the validation set increase as the epoch increases. The validation loss drops rapidly and follows a trend similar to the training loss, indicating that the performance of the model on the validation set is also gradually improving without any obvious overfitting. The training accuracy increases rapidly in the first few epochs and then stabilizes near the highest value. The orange curve represents the accuracy of the validation set, which stabilizes after the 70th epoch and follows a trend similar to the training accuracy Figure 4 shows the ROC curve graph of the model proposed in the present invention when randomly repeating the training 10 times on the Apnea-ECG dataset. The ROC curve is an important tool for evaluating the performance of binary classifiers, and the AUC value quantitatively evaluates the overall performance under the curve. It can be clearly seen from the ROC curve shown in the figure that the model proposed in the present invention exhibits good stability and efficiency in the classification task. The model not only maintains consistent performance in different tests, but also further verifies its stability and reliability through the quantitative evaluation of the AUC value, indicating its high practical value in practical applications. The test comparison data of the present invention are specifically shown in the following table

[0091]

[0092]

[0093] The test results show that the proposed OSA detection model exhibits better performance compared with previous algorithms. On two public datasets, the average accuracies of segment-based OSA detection reach 92.82% and 90.73% respectively, which are 1.89% and 7.28% higher than the state-of-the-art models. The results indicate that the present invention achieves competitive OSA detection performance

[0094] An obstructive sleep apnea detection system based on a multi-scale convolutional neural network, comprising

[0095] an ECG data preprocessing module for filtering and normalizing the original ECG signal to remove noise and standardize the data format

[0096] a P-R interval determination module for detecting the positions of the P wave and R wave in the ECG signal and calculating the P-R interval

[0097] Segment extraction module, which is used to divide the ECG signal into segments of fixed length according to time windows. Each segment corresponds to a 5-minute signal and consists of 900 sampling points. When segmenting, it is intercepted according to the structure of the first 2 minutes, the annotation area of 1 minute, and the last 2 minutes;

[0098] Distance calculation module, which is used to calculate the Euclidean distance between time series segments in the ECG signal to quantify the similarity of different electrocardiogram patterns;

[0099] Feature extraction module, which extracts multi-scale features based on three CNN models with different time scales (5 minutes, 3 minutes, 1 minute); the models capture local and global features of time series signals at different scales to enhance the model's ability to express time correlation. Convolution optimization module, which extracts high-level features through multi-layer convolution operations, and combines batch normalization (BatchNorm), LeakyReLU activation, max pooling (MaxPooling), and Dropout layers to optimize the model's expression ability and prevent overfitting.

[0100] Residual attention module, which is used to enhance the model's attention ability to key features of the ECG signal by combining residual connections with the attention mechanism, while retaining the original feature information.

[0101] Channel attention module, which is used to apply channel weights to the multi-channel ECG feature map, automatically assign attention weights to different channels, and enhance the attention to key channels. Feature fusion judgment module, which is used to fuse the multi-scale features extracted by the three time-scale CNN models and judge the final output.

[0102] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it causes the device where the computer-readable storage medium is located to execute the obstructive sleep apnea detection method based on a multi-scale convolutional neural network as described above. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory, and other memories, etc.

[0103] An electronic device, including: a memory and a processor. A program that can run on the processor is stored on the memory. When the processor executes the program, it implements the obstructive sleep apnea detection method based on a multi-scale convolutional neural network as described above.

[0104] If the modules / units integrated in the electronic device described in this application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented.

[0105] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of blockchain nodes, etc.

[0106] Computer-readable instructions are stored in the computer-readable storage medium. The computer-readable instructions are executed by a processor in the electronic device to implement the obstructive sleep apnea detection method based on a multi-scale convolutional neural network described in any of the above embodiments.

[0107] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0108] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0109] The above-described embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

[0110] It should be noted that the terms "including" and "having" in the specification and claims of this application, as well as any of their variations, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0111] Note that the above is only a preferred embodiment of the present invention and the application of technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the specific embodiments described herein. Without departing from the concept of the present invention, more other effective embodiments can also be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An obstructive sleep apnea detection method based on a multi-scale convolutional neural network, characterized in that: The steps include: Step 1: Loading ECG data, wherein each record of the ECG data includes a continuous digitized ECG signal, a set of apnea annotations, and a set of machine-generated QRS annotations; Step 2: Preprocess the ECG data: first remove the unlabeled segments, then cut the segments and remove the noise to obtain segments with a time length of t1 minutes; Step 3: Detect the positions of the R peak and P peak of the fragments obtained in the previous step, and determine the PR interval; Step 4: Use the Euclidean distance to calculate the maximum distance X between each PR interval in each segment. max and the minimum distance X min , get the two-dimensional distance array corresponding to each fragment; Step 5: Cut the middle part of each segment into a t2-minute segment by sampling, and then further cut out a t3-minute segment in the middle of the t2-minute segment, where t3 <t2<t1; Step 6: Construct three different CNN models to meet the feature extraction requirements of different time scales, and transform the X-ray images of the three ECG segments of different lengths intercepted in step 5 into max and X min Input them into three CNN models for adjustment and extraction respectively to obtain the corresponding convolution feature maps; Step 7: Optimize the convolutional feature map output by each CNN model, and strengthen the important information in the features through the residual attention module, so that each CNN model can better capture local and global features; Step 8: Integrate the three attention-weighted feature maps output by the residual attention, and apply the channel attention module to further weight the features of each channel to optimize the overall model's ability to focus on key features; Step 9: The weighted feature map output by the channel attention is passed through the fully connected layer for final classification output to predict whether obstructive sleep apnea (OSA) exists.

2. The obstructive sleep apnea detection method based on a multi-scale convolutional neural network according to claim 1, characterized in that: The step 2 specifically includes the following steps: first, some unlabeled segments are deleted; then, a 1-minute segment corresponding to the label is selected as a reference in the ECG record, and a 2-minute signal segment before and after the label segment is extracted to form a total length of 5 minutes, that is, t1=5min; finally, a finite impulse response FIR filter is used to limit the signal frequency range to 3-45Hz to remove noise.

3. The obstructive sleep apnea detection method based on a multi-scale convolutional neural network according to claim 2, characterized in that: The step 3 specifically uses the Hamilton algorithm to detect the R peak and delete those signal segments located near the beginning or end of the signal segment. Specifically, if the number of R peaks in a 5-minute signal segment is less than 40 or greater than 200, it means that noise seriously interferes with the segment, and we directly discard the segment; based on the processed data, we determine the P wave position and PR interval; the P wave appears before the R peak, so the P wave is located in the window before each R peak; by locating the position of the P peak, the PR interval of each line segment can be located.

4. The obstructive sleep apnea detection method based on a multi-scale convolutional neural network according to claim 2, characterized in that: The step 4 specifically includes the following steps: By locating the PR interval within each line segment, a time series A with a length of m sampling points starting from the P peak is constructed. m , as shown in formula (1): In formula (1), T Pi,m Represents a subsequence of the time series, defined as a continuous segment of length m samples; then according to A m The distance profile matrix D of the ECG signal is defined, which represents the distribution relationship of the distance between each pair of time series, as shown in formula (2): in Represents two subsequences and The Euclidean distance between the minimum value X in each column of the matrix D min and the maximum value X max Derived as the input segment of the model, denoted as [X min , X max ]; X min and X max Normalization is performed in the range [0, 1] according to formula (3); X represents all X max or X min In the 5-minute segment, min(X) represents the minimum value in set X, and max(X) represents the maximum value in set X; 5. The obstructive sleep apnea detection method based on a multi-scale convolutional neural network according to claim 1, characterized in that: The step 5 specifically includes the following steps: t1 = 5 min, t2 = 3 min, t3 = 1 min; but: For each 5-minute segment, we cut 540 sampling points from 181 to 720 in the middle of the segment to generate a 3-minute segment, and then further cut 180 sampling points from 361 to 720 in the middle of the 3-minute segment to generate a 1-minute segment.

6. The obstructive sleep apnea detection method based on a multi-scale convolutional neural network according to claim 1, characterized in that: In the step 6, three different convolutional neural networks CNN-1, CNN-2 and CNN-3 are specifically used to form a multi-scale CNN structure; the multi-scale CNN structure is improved based on the framework of AlexNet, which has 5 convolutional layers, 3 maximum pooling layers, 2 normalization layers, 2 fully connected layers and a SoftMax layer; each convolutional layer has a ReLU activation function, which does not set the gradient of negative input to zero, solves the problem of "neuron death" and helps maintain the gradient flow; the nonlinear Leaky ReLU activation function is expressed as equations (6) and (7); Where W1 and b1 represent the weights and biases of the shared fully connected layer, σ1 represents the ReLU activation function, and Z avg Represents the average value of all time step data on the dth feature dimension, Z max Represents the maximum value of each feature d within the time step range.

7. The obstructive sleep apnea detection system based on a multi-scale convolutional neural network according to claim 1, characterized in that: In step 6, the configuration of the convolutional layers, the number of filters and the pooling strategy were adjusted accordingly between different CNNs: CNN-1 processed the enhanced clip with a duration of 5 minutes, with 5 convolutional layers, each of which consisted of one-dimensional convolution, batch normalization, leaky ReLU activation function and maximum pooling, but the fourth layer did not include maximum pooling. CNN-2 processes augmented clips of 3 minutes duration and also has 5 convolutional layers; however, since the input length is different from CNN-1, the maximum pooling in the third layer is removed to ensure that the output feature maps have the same shape; in addition, to adapt to the 3-minute clips, we modify the number of one-dimensional convolution filters and the size of the convolution kernel to simplify the model. CNN-3 processes enhanced segments with a duration of 1 minute: Since the 1-minute segment carries less data than the 3-minute and 5-minute segments, we only use the first three layers of the CNN-2 model and further simplify the model to adapt to the 1-minute segment to ensure that the features of different time scales can be consistent when fused. Finally, the X max and X min Input into three CNN models respectively.

8. The obstructive sleep apnea detection system based on a multi-scale convolutional neural network according to claim 1, characterized in that: include ECG data preprocessing module, used to filter and normalize the original ECG signal to remove noise and standardize the data format; A PR interval determination module, for detecting the positions of the P wave and the R wave in the ECG signal and calculating the PR interval; The segment capture module is used to divide the ECG signal into segments of fixed length according to the time window; each segment corresponds to a 5-minute signal and consists of 900 sampling points; when the segments are divided, they are captured according to the structure of the first 2 minutes, the marked area 1 minute, and the last 2 minutes; A distance calculation module is used to calculate the Euclidean distance between time series segments in the ECG signal, which is used to quantify the similarity of different ECG patterns; Feature extraction module, extracts multi-scale features based on CNN models of three different time scales; The model captures local and global features of time series signals at different scales to enhance the model's ability to express temporal correlations. The convolution optimization module extracts high-level features through multi-layer convolution operations, while combining batch normalization, LeakyReLU activation, maximum pooling, and Dropout layers to optimize the model's expressiveness and prevent overfitting. The residual attention module is used to enhance the model's ability to focus on the key features of ECG signals while retaining the original feature information through residual connections combined with the attention mechanism; Channel attention module, which is used to apply channel weights to multi-channel ECG feature maps, automatically assign attention weights to different channels, and enhance attention to key channels; The feature fusion judgment module is used to fuse the multi-scale features extracted by the three time-scale CNN models and judge the final output.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the device where the computer-readable storage medium is located executes the obstructive sleep apnea detection method based on a multi-scale convolutional neural network as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a program that can be run on the processor, and when the processor executes the program, the obstructive sleep apnea detection method based on a multi-scale convolutional neural network as described in any one of claims 1-7 is implemented.

Citation Information

Cited By

  • Sleep apnea detection method and system, electronic equipment and storage medium

    CN120477748A

  • Sleep quality monitoring method fusing multi-scale features

    CN120827347A