A two-stage detection method for sleep apnea

By combining the Mel spectral analysis of respiratory audio and heart sound signals, transfer learning and longitudinal box algorithms are used to build a sleep-wake detection model, which solves the problem of low sleep apnea detection accuracy in the prior art, and realizes more accurate sleep duration calculation and AHI evaluation, which improves the reliability of detection.

CN119279511BActive Publication Date: 2025-07-04NANJING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411490964.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-04
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The sleep apnea detection method based on sound signals in the prior art has low diagnostic accuracy and cannot accurately calculate the sleep duration, resulting in a low sleep apnea index and is difficult to meet clinical needs.

Method used

Combining respiratory audio signals and heart sound signals, a sleep-wake detection model is constructed using Mel spectrogram, transfer learning, longitudinal box algorithm and support vector machine. By calculating the RR interval and sleep duration, the sleep apnea hypoventilation index AHI is calculated.

Benefits of technology

It improves the accuracy and accuracy of sleep apnea detection, provides more comprehensive physiological information, reduces noise interference, and improves detection efficiency and reliability of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119279511B_ABST
    Figure CN119279511B_ABST
Patent Text Reader

Abstract

The present application discloses a two-stage detection method for sleep apnea, which relates to the field of respiratory detection and includes: collecting tracheal sound signals of a user's sleep respiration to generate a Mel spectrogram of the tracheal sound signals; wherein, the tracheal sound signals include respiratory audio signals and heart sound signals; constructing a sleep-wake detection model based on the VGG16 network to calculate the user's sleep duration; using the longitudinal box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram; extracting the S1 peak and S2 peak of potential sleep apnea segments, and calculating the RR interval according to the time interval between adjacent S1 peaks; obtaining the statistical features of the RR interval, and using a support vector machine SVM to detect apnea according to the statistical features; calculating the apnea-hypopnea index AHI of the user according to the number of sleep apnea times and the sleep duration statistically detected by the apnea detection result. Aiming at the low diagnostic accuracy in the prior art of sleep apnea detection methods based on sound signals, the present application improves the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of respiratory detection, and more specifically, to a two-stage detection method for sleep apnea. Background Art

[0002] Currently, sound perception has attracted extensive attention in the field of sleep apnea detection due to its excellent non-invasiveness and ease of use. However, most studies are affected by the limitations of the amount of information in respiratory sounds, resulting in insufficient diagnostic accuracy; and due to the lack of a sleep-wake detection step, the accurate sleep duration of patients cannot be obtained, resulting in a low calculated apnea-hypopnea index (AHI) of patients, which is difficult to meet clinical needs. Therefore, it is of great significance to develop an accurate detection method for sleep apnea based on sound signals.

[0003] Chinese Patent Application, Application No. 202310419815.2, discloses a method for identifying apnea snoring based on a two-stream multi-scale model, including: performing preprocessing on snoring data, including signal pre-emphasis and framing and windowing; then, extracting Mel-frequency cepstral coefficients (MFCCs) from the snoring data as snoring features; then, constructing a two-stream multi-scale model for classifying normal snoring and sleep apnea snoring, where the time-domain branch uses a one-dimensional convolutional neural network to process the MFCC features, and the frequency-domain branch uses a two-dimensional convolutional neural network to process the transpose of the MFCC features; finally, counting the number of sleep apnea snoring and calculating the apnea-hypopnea index of the patient.

[0004] Although the above method realizes the identification of sleep apnea events and AHI calculation to a certain extent, there are still some obvious disadvantages. First, the amount of information in a single respiratory sound is insufficient and is easily interfered by environmental noise, which may lead to missed or misjudged sleep apnea. In addition, the fixed convolutional kernel structure and single feature fusion method in the above method make it difficult for the model to learn the optimal feature representation when processing complex and variable all-night sound data, resulting in poor detection performance. More importantly, the above methods all ignore the wakefulness stage in all-night monitoring and directly use the monitoring duration instead of the actual sleep duration of the patient to calculate the AHI, resulting in a low AHI of the patient, i.e., the "dilution effect", which greatly limits its reliability and practicality in clinical applications. Summary of the Invention

[0005] 1. Technical Problems to be Solved

[0006] Aiming at the problem of low diagnostic accuracy in the existing sleep apnea detection method based on sound signals, this application provides a two-stage detection method for sleep apnea, which improves the detection accuracy by combining respiratory audio signals and heart sound signals and using Mel spectrogram, transfer learning, longitudinal box algorithm, heart sound feature extraction, and support vector machine, etc.

[0007] 2. Technical Solution

[0008] The object of the present application is achieved by the following technical solutions.

[0009] The present application provides a two-stage detection method for sleep apnea, including: collecting tracheal sound signals of a user's sleep breathing, preprocessing the tracheal sound signals to generate a Mel spectrogram of the tracheal sound signals; wherein, the tracheal sound signals include respiratory audio signals and heart sound signals; constructing a sleep-wake detection model based on the VGG16 network through transfer learning to calculate the user's sleep duration; using the longitudinal box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram to obtain potential sleep apnea segments; sampling the potential sleep apnea segments, extracting the S1 peak and S2 peak of the potential sleep apnea segments, and calculating the RR interval according to the time interval between adjacent S1 peaks; wherein, the S1 peak represents the first heart sound and the S2 peak represents the second heart sound; obtaining the statistical features of the RR interval, and using a support vector machine SVM to perform apnea detection according to the statistical features; calculating the apnea-hypopnea index AHI of the user according to the number of sleep apnea times statistically obtained from the apnea detection results and the sleep duration.

[0010] Among them, the tracheal sound signal is a sound signal of respiratory and cardiac activities collected by placing a sound sensor at the trachea of the human neck. It includes two main components: respiratory audio signals and heart sound signals. The respiratory audio signal reflects the sound generated when air flows through the respiratory tract, while the heart sound signal reflects the sound generated when the heart valves open and close and blood flows. The tracheal sound signal can provide important information about respiratory and cardiac functions for the detection and analysis of sleep apnea.

[0011] The VGG16 network is a classic convolutional neural network architecture and one of the excellent deep learning models in the ImageNet image classification competition. The VGG16 network consists of 16 convolutional layers and fully connected layers. Through transfer learning, the pre-trained VGG16 network can be applied to the sleep-wake detection task to utilize its powerful feature representation ability to distinguish between the sleep and wake states.

[0012] Vertical Box Algorithm AV-Box (Adaptive Vertical-Box Algorithm): It is an adaptive algorithm for time series signal segmentation. It slides a vertical rectangular window over the signal and determines the segmentation points of the signal according to the distribution of sample points within the window. The AV-Box algorithm can adaptively adjust the size and position of the window to adapt to the local characteristics of the signal. In sleep apnea detection, the AV-Box algorithm can be applied to the Mel spectrogram to locate potential sleep apnea segments by detecting sudden changes in energy distribution.

[0013] The S1 peak and S2 peak correspond to the peak points of the first heart sound (S1) and the second heart sound (S2) in the heart sound signal respectively. The first heart sound is caused by the closure of the atrioventricular valves (mitral valve and tricuspid valve), usually occurring at the beginning of ventricular contraction. The second heart sound is caused by the closure of the semilunar valves (aortic valve and pulmonary valve), usually occurring at the beginning of ventricular diastole. The localization and identification of the S1 peak and S2 peak are very important for the analysis of the heart sound signal and the calculation of the RR interval.

[0014] RR interval (RR Interval) refers to the time interval between the peaks of two consecutive R waves in an electrocardiogram. In heart sound signal analysis, the S1 peak can be regarded as an equivalent R wave peak, so the RR interval can also be understood as the time interval between two consecutive S1 peaks. The RR interval reflects the rhythm changes of the heart and can be used to evaluate heart rate variability and cardiac autonomic nerve function. In sleep apnea detection, the statistical characteristics of the RR interval can be used as an important basis for judging respiratory events.

[0015] Apnea-Hypopnea Index AHI (Apnea-Hypopnea Index): The Apnea-Hypopnea Index (AHI) is an important indicator for evaluating the severity of sleep apnea. It represents the number of apnea and hypopnea events that occur per hour on average during the patient's sleep. Apnea refers to the complete cessation of breathing for more than 10 seconds; hypopnea refers to a significant reduction in the amplitude of breathing for more than 10 seconds, accompanied by a decrease in blood oxygen saturation.

[0016] Furthermore, collect the tracheal sound signal of the user's sleep breathing, preprocess the tracheal sound signal to generate the Mel spectrogram of the tracheal sound signal, including: performing frame segmentation on the collected tracheal sound signal; performing short-time Fourier transform on the segmented tracheal sound signal to obtain the spectrogram of the tracheal sound signal; converting the spectrogram to a single-sided spectrum, and filtering the single-sided spectrum through a Mel filter bank to obtain the logarithmic Mel frequency spectrum of the tracheal sound signal; combining the logarithmic Mel frequency spectra into a Mel spectrogram.

[0017] Among them, the one-sided spectrum. The one-sided spectrum refers to the spectrum that only contains the positive frequency part in the spectrogram. In signal processing, by performing a Fourier transform on the time-domain signal, the frequency-domain representation of the signal, that is, the spectrogram, can be obtained. The spectrogram usually contains two parts: positive frequency and negative frequency. For a real-valued signal, its spectrum has conjugate symmetry, that is, the positive frequency part and the negative frequency part are mirror-symmetric. In the process of generating the Mel spectrogram of the tracheal sound signal, by performing a short-time Fourier transform on the framed tracheal sound signal, the spectrogram of the tracheal sound signal is obtained. Then, the spectrogram is converted into a one-sided spectrum, that is, only the positive frequency part is retained, and the negative frequency part is discarded. This can simplify the subsequent processing steps and reduce the computational amount.

[0018] Furthermore, collecting the tracheal sound signal of the user's sleep breathing, preprocessing the tracheal sound signal, and generating the Mel spectrogram of the tracheal sound signal further includes: adjusting the size of the Mel spectrogram using the bilinear interpolation method to obtain a Mel spectrogram with a unified size.

[0019] Furthermore, through transfer learning, constructing a sleep-wake detection model based on the VGG16 network and calculating the user's sleep duration includes: training the VGG16 network on the source domain database. After training, transfer all the network structures and the trained parameters before the last max-pooling layer of the VGG16 network to the sleep-wake detection model, and remove the last max-pooling layer and the fully connected layer of the VGG16 network; in the transferred network structure, sequentially add a new max-pooling layer, multiple convolutional layers, an LSTM layer, and a fully connected layer to form the sleep-wake detection model; input the Mel spectrogram with a unified size into the sleep-wake detection model, obtain the feature representation of the Mel spectrogram through forward propagation, and use the backpropagation algorithm to adjust the model parameters; add a softmax function after the last fully connected layer of the sleep-wake detection model to construct a sleep-wake state classifier, perform binary classification of the sleep-wake state on the input Mel spectrogram to obtain the sleep-wake detection result; according to the sleep-wake detection result, extract the start frame and end frame of the sleep state, calculate the time difference corresponding to the start frame and end frame to obtain the duration of the sleep state, and accumulate the durations of all sleep states to obtain the user's sleep duration.

[0020] Among them, the source domain database. In transfer learning, the source domain database refers to the original data set used to pre-train the model. It is usually a large-scale, well-annotated data set that is different from the domain (target domain) where the target task is located. The data samples and labels in the source domain database are used to train the base model so that it can learn general feature representations and patterns.

[0021] Further, the vertical box algorithm AV-Box is used to detect sleep apnea segments in the Mel spectrogram, obtaining potential sleep apnea segments, including: setting the right boundary center point coordinate of the rectangular window as (n, Yn), where n represents the sampling point and Yn represents the signal amplitude; setting the box length of the rectangular window as L and the box height as 2H, and initializing the right boundary center point coordinate as (n = L + 1, Yn); calculating the number of sampling points bLH(n) within the rectangular window; determining whether the number of sampling points bLH(n) is less than the threshold, if so, marking the current sampling point n as a change point; determining whether the current sampling point n is equal to the total number of sampling points of the Mel spectrogram, if so, outputting the position parameter of the current sampling point n, otherwise, moving the rectangular window forward by a unit distance, updating the right boundary center point coordinate as (n = ε + n, Yn), where ε is the unit distance, and repeating the determination of the number of sampling points until the current sampling point n is equal to the total number of sampling points of the Mel spectrogram.

[0022] Among them, the right boundary center coordinate (Right Boundary Center Coordinate), in the vertical box algorithm (AV-Box), the right boundary center coordinate refers to the coordinate of the center point of the right boundary of the rectangular window. It is used to determine the position of the rectangular window on the signal and its movement. The right boundary center coordinate consists of two parts: the sampling point n and the corresponding signal amplitude Yn. The sampling point n represents the discrete position of the signal on the time axis, that is, the nth sampling point. The signal amplitude Yn represents the amplitude value of the signal at the nth sampling point. The role of the right boundary center coordinate is to determine the position of the rectangular window on the signal and control the movement of the rectangular window. By continuously updating the right boundary center coordinate, the rectangular window can slide along the signal to capture the local features and changes of the signal. At each position, the algorithm calculates the number of sampling points within the rectangular window and determines whether the current position is a change point (the segmentation point of the signal) according to the threshold.

[0023] Further, using the vertical box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram and obtaining potential sleep apnea segments also includes: calculating the interval between adjacent change points output, determining whether the interval is less than the first threshold, if so, merging the corresponding two change points into one change point; calculating the segment duration between the merged adjacent change points, determining whether the segment duration is less than the second threshold or greater than the third threshold, if so, deleting the corresponding segment; where the second threshold is less than the third threshold; calculating the interval between adjacent segments after deleting the segment, determining whether the interval is greater than the fourth threshold, if so, marking the corresponding segment as a potential sleep apnea segment; where the fourth threshold is greater than the third threshold. The value range of the first threshold is 180 ms to 220 ms; the value range of the second threshold is the same as that of the first threshold, and the value range of the third threshold is 3 s to 7 s; the value range of the fourth threshold is 8 s to 12 s.

[0024] Further, sample the potential sleep apnea segments, extract the S1 peak and S2 peak of the potential sleep apnea segments, and calculate the RR interval according to the time interval between adjacent S1 peaks, including: performing Hilbert transform on the potential sleep apnea segments to calculate the energy envelope of the potential sleep apnea segments; comparing the energy envelope values with the threshold T, and retaining the energy envelope values greater than the threshold T; setting variables d12 and d21 to represent the time intervals between the S1-S2 peaks and the S2-S1 peaks respectively, initializing the two values with fixed values, and dynamically changing them using the weight factors k1 and k2, where k1 and k2 are set to the empirical values 0.95 and 0.05; within the time window corresponding to d12, obtain the peak with the largest energy envelope value and mark it as the S1 peak; within the time window corresponding to d21, obtain the peak with the largest energy envelope value and mark it as the S2 peak; calculate the time interval Δt between adjacent S1 peaks and S2 peaks, compare the time interval Δt with the variables d12 and d21, if the time interval Δt is less than d12, then mark the corresponding peak pair as the S1-S2 peak pair; if the time interval Δt is greater than d12 and less than d21, then mark the corresponding peak pair as the S2-S1 peak pair; calculate the time interval between adjacent S1 peaks as the RR interval.

[0025] Among them, the variable d12 represents the time interval between the S1 peak and the S2 peak, that is, the time difference between the first heart sound (S1) and the second heart sound (S2). It is a preset time window used to search for and locate the S1 peak and S2 peak in the energy envelope. By setting an appropriate value of d12, the S1 peak and S2 peak can be more accurately identified and extracted, avoiding misidentifying other non-heart sound peaks as heart sound peaks. The S1 peak refers to the peak point marked as the first heart sound (S1) in the energy envelope. The first heart sound is caused by the closure of the atrioventricular valves (mitral valve and tricuspid valve) and usually occurs at the beginning of ventricular contraction. Within the given time window (defined by the variable d12), find the peak point with the largest energy envelope value and mark it as the S1 peak. The S2 peak refers to the peak point marked as the second heart sound (S2) in the energy envelope. The second heart sound is caused by the closure of the semilunar valves (aortic valve and pulmonary valve) and usually occurs at the beginning of ventricular diastole. Similar to the S1 peak, within the given time window (defined by the variable d21), find the peak point with the largest energy envelope value and mark it as the S2 peak.

[0026] The S1-S2 peak pair refers to a pair of peaks consisting of an S1 peak and an S2 peak that are adjacent in time and meet specific conditions. By calculating the time interval Δt between adjacent S1 and S2 peaks and comparing it with the variable d12, if the time interval Δt is less than d12, the corresponding S1 and S2 peaks are marked as an S1-S2 peak pair. The S2-S1 peak pair refers to a pair of peaks consisting of an S2 peak and an S1 peak that are adjacent in time and meet specific conditions. Similar to the S1-S2 peak pair, by calculating the time interval Δt between adjacent S2 and S1 peaks and comparing it with the variables d12 and d21, if the time interval Δt is greater than d12 and less than d21, the corresponding S2 and S1 peaks are marked as an S2-S1 peak pair. The RR interval refers to the time interval between adjacent R-wave peaks. In heart sound signal analysis, the S1 peak can be regarded as an equivalent R-wave peak, so the RR interval can also be understood as the time interval between adjacent S1 peaks. By calculating the time difference between adjacent S1 peaks, the value of the RR interval can be obtained. The RR interval reflects the rhythm changes of the heart and is an important indicator for evaluating heart rate variability and cardiac autonomic nerve function. In the detection of sleep apnea, the change pattern of the RR interval can be used as one of the bases for judging respiratory events.

[0027] Furthermore, obtain the statistical characteristics of the RR interval. The statistical characteristics include: maximum value, minimum value, range, standard deviation, mean, skewness, and kurtosis.

[0028] Furthermore, calculate the apnea-hypopnea index AHI of the user through the following formula: AHI = N / T, where N represents the number of sleep apnea events of the user; T represents the sleep duration of the user.

[0029] 3. Beneficial effects

[0030] Compared with the prior art, the advantages of this application are as follows:

[0031] Using the Mel spectrogram to represent the time-frequency domain characteristics of tracheal sound signals and constructing a sleep-wake detection model through transfer learning can effectively identify the user's sleep state and accurately calculate the sleep duration; using the longitudinal box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram can quickly and accurately locate potential sleep apnea events, reduce the amount of data for subsequent processing, and improve the detection efficiency; by simultaneously collecting the respiratory audio signal and heart sound signal during the user's sleep, more comprehensive and accurate physiological information can be obtained; by extracting the first heart sound S1 peak and the second heart sound S2 peak in the potential sleep apnea segment and calculating the RR interval, the impact of sleep apnea events on the heart rhythm can be effectively evaluated, providing more physiological indicators; using the support vector machine SVM model to classify the statistical characteristics of the RR interval can accurately judge whether the potential sleep apnea segment is a true sleep apnea event, improving the detection accuracy. Description of the Drawings

[0032] Figure 1 It is a technical roadmap of a two-stage detection method for sleep apnea in this application;

[0033] Figure 2 It is a screening result diagram of potential sleep apnea segments in an embodiment of this application;

[0034] Figure 3 It is a waveform diagram of tracheal sound signals in an embodiment of this application;

[0035] Figure 4 It is a segmentation result diagram of S1 and S2 peaks in an embodiment of this application;

[0036] Figure 5 It is a traditional acoustic detection AHI analysis diagram;

[0037] Figure 6 It is a detection AHI analysis diagram in an embodiment of this application;

[0038] Figure 7 It is a sleep apnea estimation result diagram in an embodiment of this application. Detailed Implementation Modes

[0039] The following will describe this application in detail in combination with the drawings in the specification and specific embodiments.

[0040] Figure 1 It is a technical roadmap of a two-stage detection method for sleep apnea in this application. In the first stage, tracheal sound signals of a patient throughout the night are collected, and the signals are preprocessed, including operations such as denoising and normalization, to obtain preprocessed tracheal sound signals. The preprocessed tracheal sound signals are converted into Mel spectrograms. Specifically, the short-time Fourier transform (STFT) is used to perform time-frequency analysis on the signals, and then the spectrogram is mapped to the Mel frequency scale to obtain the Mel spectrogram. A sleep-wake detection model is constructed based on transfer learning. First, the VGG16 network is pre-trained on a large-scale sleep dataset, and then fine-tuned on the target domain data to obtain a feature extractor suitable for sleep-wake detection. The Mel spectrogram is input into this feature extractor to extract high-level features, and a classifier (such as a softmax layer) is used for sleep-wake state recognition. According to the recognition results, the actual sleep duration of the patient is calculated. The Adaptive Vertical Box (AV-Box) algorithm is used to screen potential sleep apnea segments in the Mel spectrogram. The AV-Box algorithm searches for regions with energy mutations on the Mel spectrogram by adaptively adjusting the size and position of the rectangular window and marks them as potential sleep apnea segments.

[0041] Second stage: Secondary confirmation of sleep apnea segments based on RR intervals. For each potential sleep apnea segment, synchronously analyze the corresponding heart sound signal segment. Use the peak detection algorithm to identify the S1 peak and S2 peak in the heart sound signal, and calculate the time interval between adjacent S1 peaks, that is, the RR interval. Extract the statistical features of the RR interval, including mean, standard deviation, maximum value, minimum value, etc. These statistical features can reflect the degree of influence of respiratory events on the cardiac rhythm. Input the statistical features of the RR interval into a pre-trained support vector machine (SVM) classifier for secondary confirmation of sleep apnea events. The SVM divides the potential sleep apnea segments into true sleep apnea events and non-sleep apnea events by finding the optimal classification hyperplane. According to the sleep apnea detection results in the second stage and the actual sleep duration obtained in the first stage, calculate the apnea-hypopnea index (AHI) of the patient. The AHI represents the average number of sleep apnea and hypopnea events occurring per hour of sleep and is an important indicator for evaluating the severity of OSA.

[0042] In the first stage, sleep-wake classification and respiratory sound event detection are performed based on tracheal sound signals to screen for potential sleep apnea segments. The specific steps are as follows: Set the frame length to 25 milliseconds and the frame shift to 10 milliseconds to ensure sufficient overlap between frames to capture the dynamic changes of the signal. Select the Hanning window as the analysis window and apply windowing to each frame of the signal to reduce the impact of spectral leakage. Perform a 512-point short-time Fourier transform (STFT) on each windowed frame of the signal to convert the time-domain signal into the frequency-domain signal and obtain the spectrogram of the sound signal. Since the spectrogram is symmetric, only half of the spectral information, that is, the single-sided spectrum, needs to be retained. Take the first 257 points (including the DC component) of the spectrogram to form the single-sided spectrum.

[0043] Design a set of 64-order Mel filters, and the center frequencies of the filters are distributed according to the Mel frequency scale. Filter the single-sided spectrum through the Mel filter bank to obtain the energy values in 64 frequency bands. Sum the amplitude values in each frequency band to obtain a 64-dimensional Mel frequency spectrum. Take the logarithm of the energy value in each frequency band to obtain the logarithmic Mel frequency spectrum. The purpose of taking the logarithm is to smooth the spectral line and make it closer to the auditory characteristics of the human ear.

[0044] Combine the log Mel spectrograms of every 96 frames into a Mel spectrogram, which serves as the time-frequency representation of the sound signal. The vertical axis of the Mel spectrogram represents the Mel frequency (a total of 64 frequency bands), and the horizontal axis represents the time frames (each unit consists of 96 frames). The number of Mel spectrograms depends on the length of the sound signal and can be adjusted according to the actual situation. To meet the input requirements of the subsequent deep learning model, it is necessary to adjust the size of the Mel spectrogram. Use bilinear interpolation to resize the Mel spectrogram to 224×224 while keeping the aspect ratio of the image unchanged. The resized Mel spectrogram can be directly input into the sleep-wake detection model based on transfer learning for feature extraction and classification.

[0045] In the second step, select the VGG16 network pre-trained on the large-scale image database ImageNet as the base model. Remove the top fully connected layer and the max pooling layer of the VGG16 network, and retain the remaining convolutional and pooling layer structures. Transfer the parameters of the pre-trained convolutional and pooling layers to the sleep-wake detection model as a feature extractor. On top of the transferred VGG16 network, add new convolutional layers, LSTM layers, Dropout layers, and fully connected layers to form a complete sleep-wake detection model. The convolutional kernels of the convolutional layers are all 3×3, and the ReLU activation function is selected to extract local features. After the convolutional layers, add 3 fully connected layers with 4096, 4096, and 256 neurons respectively to map the extracted features to a 256-dimensional embedding space. Insert Dropout layers between the fully connected layers to randomly discard some neurons to prevent overfitting of the network. Input the preprocessed Mel spectrogram into the transferred VGG16 network, and after a series of convolutional and pooling layer processes, extract the deep features of the tracheal sound signal. The convolutional layers extract the time-frequency features in the Mel spectrogram through local receptive fields and weight sharing. The pooling layers reduce the size of the feature map through downsampling operations to improve the robustness of the features. Map the extracted features to a 256-dimensional embedding space through the fully connected layers to obtain a compact representation of the tracheal sound signal.

[0046] The model is trained using labeled sleep-wake data, where the label represents the sleep or wake state corresponding to each frame. On top of the extracted 256-dimensional embedding features, convolutional layers and LSTM layers are added to capture temporal dependencies. The LSTM layer can effectively model long-term dependencies and capture the temporal evolution pattern of the sleep-wake state. After the LSTM layer, fully connected layers and a Softmax activation function are added to map the features to two categories: sleep and wake. The model is trained using the cross-entropy loss function and the Adam optimizer, continuously adjusting the network parameters to minimize the prediction error. The mel spectrogram of the test sample is input into the trained sleep-wake detection model to obtain the prediction results of the sleep or wake state for each frame. Based on the prediction results, the number of sleep frames and wake frames of the patient are counted, and the actual sleep duration of the patient is calculated.

[0047] In the third step, the tracheal sound signal is input, and the initial parameters of the AV-Box algorithm are set. The box length L and the box height 2H are set to determine the size of the rectangular window. The center point of the right boundary is initialized as (n = L + 1, Yn), where n represents the sampling point and Yn represents the signal amplitude. The number of sample points within the rectangular window is calculated. Taking the center point of the right boundary as a reference, the number of sample points within the rectangular window is counted and denoted as bLH(n). The number of sample points reflects the energy distribution of the signal within the current rectangular window.

[0048] The calculated number of sample points bLH(n) is compared with a preset threshold. If the number of sample points is below the threshold, the current point is determined as a change point, indicating the possible boundary of a respiratory event. It is judged whether the current change point is the end point of the signal.

[0049] If it is the end point, the position parameter of this point is output, indicating that a complete respiratory event segment has been detected. If it is not the end point, the rectangular window is moved forward by a unit distance, and the center point of the right boundary is updated to (n = n + ΔL, Yn), where ΔL represents the moving step. The rectangular window is continuously moved, the number of sample points is calculated, and the change point is judged until the end of the signal is reached. Through the iterative process, all potential respiratory event segments in the signal can be detected.

[0050] Check the time interval between adjacent change points. If the interval is less than 200 milliseconds, the respiratory event segments corresponding to these two change points are merged into one segment. The merging operation can eliminate over-segmentation caused by signal noise or transient respiratory changes.

[0051] Verify the duration of the merged respiratory event segments. If the duration of a segment is less than 200 milliseconds or greater than 5 seconds, it is determined that the segment does not belong to a valid respiratory event and is deleted. By screening based on duration, segments that are too short or too long can be removed, improving the accuracy of respiratory event detection. Analyze the time interval between adjacent respiratory event segments. If the interval is greater than 10 seconds, mark this interval as a sleep apnea segment. The sleep apnea segment represents an interruption of the respiratory event and is a key factor in determining the apnea-hypopnea index (AHI).

[0052] Figure 2 This is a diagram of the screening results of potential sleep apnea segments for an embodiment of the present application. The present application uses an adaptive longitudinal box algorithm to screen potential sleep apnea segments from tracheal sound signals. In the diagram, the black waveform is the tracheal sound signal, and the flat part between the red rectangular windows is the potential sleep apnea segment.

[0053] In the second stage, estimate the RR interval for potential sleep apnea segments, extract multi-dimensional statistical features, and use a support vector machine for classification to achieve precise detection of sleep apnea. Combine the sleep duration obtained in the first stage to calculate the patient's AHI. The specific steps are as follows: Downsample the sampling rate of the potential sleep apnea segments obtained in the first stage to 2 kHz to reduce the data volume and improve processing efficiency. Segment the potential sleep apnea segments with a window length of 2 seconds to obtain a series of heart sound signal segments. Perform Hilbert transform on each heart sound signal segment to obtain its analytic signal. Take the modulus of the analytic signal to obtain the energy envelope of the heart sound signal segment. The energy envelope reflects the instantaneous energy change of the heart sound signal and helps to locate the S1 and S2 peaks.

[0054] Define the initial threshold level T as 0.25 times the maximum energy value of the heart sound signal segment. Only retain the envelope signal with an energy envelope value greater than the threshold T to eliminate the energy peaks corresponding to the unwanted ripples in the data. Through threshold processing, low-energy noise and interference can be removed, highlighting the energy characteristics of the S1 and S2 peaks. Define variables d21 and d21 to represent the time intervals between the S1-S2 peaks and the S2-S1 peaks respectively. Initialize and d21 with fixed values, for example, it can be set to half of the average period of the heart sound signal. Introduce weight factors k1 and k2 to dynamically adjust the values of d12 and d21 to adapt to the rhythm changes of the heart sound signal.

[0055] Within each S1-S2 time distance, only retain the peak with the maximum energy, discard other peaks, and mark it as the S1 peak. Similarly, within each S2-S1 time distance, only retain the peak with the maximum energy, discard other peaks, and mark it as the S2 peak. Through peak screening, the S1 and S2 peaks in the heart sound signal can be accurately located, eliminating the interference of redundant peaks.

[0056] Calculate the time interval between adjacent S1 peaks and S2 peaks. Compare the time interval with variables d12 and d21. If the interval is close to d12, label the previous peak as the S1 peak and the subsequent peak as the S2 peak; if the interval is close to d21, label the previous peak as the S2 peak and the subsequent peak as the S1 peak. By comparing the time intervals, the S1 and S2 peaks can be accurately distinguished, ensuring the correct classification of the peaks. Calculate the time intervals between adjacent S1 peaks to obtain a series of RR interval values. The RR interval reflects the rhythm changes of the heart and is an important indicator for evaluating sleep apnea events.

[0057] Figure 3 This is the S1 peak segmentation result diagram of an embodiment of the present application; Figure 4 This is the S2 peak segmentation result diagram of an embodiment of the present application. As shown in Figure 3 and Figure 4 Figure 8 shows the S1 and S2 peak recognition result diagram in a 20-second-long jugular vein sound signal, where the red asterisks represent the S1 peaks and the red hollow circles represent the S2 peaks. It can be seen that the recognition results of the S1 and S2 peaks by this method are good. In this embodiment, the Hilbert envelope of the tracheal sound is calculated, the S1 and S2 peaks are extracted through the adaptive threshold of the envelope, and the peak timing characteristics are used for recognition to estimate the RR interval.

[0058] For each potential sleep apnea segment, calculate the following statistics of its corresponding RR interval sequence: Maximum value: The maximum value of the RR interval, which reflects the slowest heart rate point. Minimum value: The minimum value of the RR interval, which reflects the fastest heart rate point. Range: The difference between the maximum value and the minimum value, indicating the change range of the RR interval. Standard deviation: The standard deviation of the RR interval, which measures the dispersion degree of the RR interval. Mean value: The arithmetic mean of the RR interval, which reflects the average heart rate level. Skewness: The asymmetry of the RR interval distribution, which describes the offset direction and degree of the distribution. Kurtosis: The sharpness of the RR interval distribution, which describes the concentration degree of the distribution. Take the calculated statistics as the feature vector of the potential sleep apnea segment. Construct a support vector machine (SVM) classifier. Select a suitable SVM kernel function, such as a linear kernel, a Gaussian kernel, etc., and select the optimal kernel function according to the characteristics of the data. Set the parameters of the SVM, such as the penalty coefficient C, the parameters of the kernel function, etc., and optimize the parameter selection through methods such as cross-validation. Use the labeled training data to train the SVM classifier, and the label indicates whether each potential sleep apnea segment is a true sleep apnea event. Classify the potential sleep apnea segments.

[0059] The feature vectors of each potential sleep apnea segment are input into the trained SVM classifier. Based on the feature vectors, the SVM classifier classifies the potential sleep apnea segments into two categories: normal breathing and apnea. The classification result indicates whether each potential sleep apnea segment is a true sleep apnea event. The classification results are statistically analyzed to calculate the number of times sleep apnea events are determined during the entire sleep period. The number of sleep apnea events reflects the frequency of apnea occurrences in the patient during sleep. The actual sleep duration of the patient obtained in the second step is used as the denominator. The number of sleep apnea events obtained in the fourth step is used as the numerator. Calculate AHI = number of sleep apnea events / sleep duration (in hours). AHI represents the average number of sleep apnea events that occur per hour in the patient and is an important indicator for evaluating the severity of sleep apnea.

[0060] Figure 5 It is an analysis diagram of AHI detected by traditional acoustics; Figure 6 It is an analysis diagram of AHI detection in an embodiment of the present application. From Figure 5 and Figure 6 it can be seen by comparison that the AHI estimation result of the present application has higher consistency with the AHI result of PSG and is more adaptable to the requirements of clinical applications.

[0061] Figure 7 It is a diagram of the estimation result of the severity of sleep apnea in an embodiment of the present application. The method proposed in the present application is used to estimate the severity of sleep apnea in 108 local patients. It can be seen from the diagram that the present application is relatively accurate in distinguishing mild, moderate, and severe sleep apnea.

[0062] The present application's creation and its implementation manners are schematically described above. This description is not restrictive. Without departing from the spirit or basic characteristics of the present application, the present application can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present application's creation, and the actual structure is not limited thereto. Any reference signs in the claims should not limit the claimed claims. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of this creation, design similar structural manners and embodiments to this technical solution without creative efforts, they should all fall within the protection scope of this patent. In addition, the term "including" does not exclude other elements or steps, and the term "a" before an element does not exclude including "multiple" such elements. The multiple elements stated in the product claims can also be implemented by one element through software or hardware. The terms first, second, etc. are used to indicate names and do not indicate any specific order.

Claims

1. A two-stage detection system for sleep apnea, characterized in that, Including: A collection module that collects tracheal sound signals of the user's sleep breathing; A preprocessing module that preprocesses the tracheal sound signals to generate a Mel spectrogram of the tracheal sound signals; wherein, the tracheal sound signals include respiratory audio signals and heart sound signals; A sleep-wake detection module that constructs a sleep-wake detection model based on the VGG16 network through transfer learning and calculates the user's sleep duration; A sleep apnea detection module that uses the longitudinal box algorithm AV-BOX to detect sleep apnea segments in the Mel spectrogram to obtain potential sleep apnea segments; A heart sound feature extraction module that samples the potential sleep apnea segments, extracts the S1 peak and S2 peak of the potential sleep apnea segments, and calculates the RR interval according to the time interval between adjacent S1 peaks; wherein, the S1 peak represents the first heart sound and the S2 peak represents the second heart sound; An apnea detection module that obtains the statistical features of the RR interval, and uses a support vector machine (SVM) to perform apnea detection according to the statistical features; calculates the apnea-hypopnea index (AHI) of the user according to the number of sleep apneas and the sleep duration statistically obtained from the apnea detection results; Constructing a sleep-wake detection model based on the VGG16 network through transfer learning and calculating the user's sleep duration, including: Training the VGG16 network on the source domain database. After training, transfer all the network structures and the trained parameters before the last max pooling layer of the VGG16 network to the sleep-wake detection model, and remove the last max pooling layer and the fully connected layer of the VGG16 network; In the transferred network structure, sequentially add a new max pooling layer, multiple convolutional layers, an LSTM layer, and a fully connected layer to form the sleep-wake detection model; Input the Mel spectrograms with unified sizes into the sleep-wake detection model, obtain the feature representation of the Mel spectrograms through forward propagation, and adjust the model parameters using the backpropagation algorithm; Add a softmax function after the last fully connected layer of the sleep-wake detection model to construct a sleep-wake state classifier, perform binary classification of the sleep-wake state on the input Mel spectrograms, and obtain the sleep-wake detection results; According to the sleep-wake detection results, extract the start frame and end frame of the sleep state, calculate the time difference between the start frame and the end frame, obtain the duration of the sleep state, and accumulate the durations of all sleep states to obtain the user's sleep duration.

2. The two-stage detection system for sleep apnea according to claim 1, wherein: Collecting tracheal sound signals of the user's sleep breathing and preprocessing the tracheal sound signals to generate a Mel spectrogram of the tracheal sound signals, including: Performing frame segmentation on the collected tracheal sound signals; Performing short-time Fourier transform on the frame-segmented tracheal sound signals to obtain the spectrogram of the tracheal sound signals; Converting the spectrogram into a single-sided spectrum, and filtering the single-sided spectrum through a Mel filter bank to obtain the logarithmic Mel frequency spectrum of the tracheal sound signals; Combining the logarithmic Mel frequency spectra into a Mel spectrogram.

3. The two-stage detection system for sleep apnea according to claim 2, wherein: Collect the tracheal sound signal of the user's sleep breathing, preprocess the tracheal sound signal, and generate the Mel spectrogram of the tracheal sound signal. It further includes: Adjust the size of the Mel spectrogram using the bilinear interpolation method to obtain a Mel spectrogram with a unified size.

4. The two-stage detection system for sleep apnea according to claim 1, characterized in that: Use the longitudinal box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram to obtain potential sleep apnea segments, including: Set the coordinate of the center point of the right boundary of the rectangular window as (n, Yn), where n represents the sampling point and Yn represents the signal amplitude; Set the box length of the rectangular window as L and the box height as 2H, and initialize the coordinate of the center point of the boundary as (L + 1, Yn); Calculate the number of sampling points within the rectangular window; Judge whether the number of sampling points is less than the threshold. If so, mark the current sampling point as a change point; Judge whether the current sampling point is equal to the total number of sampling points of the Mel spectrogram. If so, output the position parameter of the current sampling point. Otherwise, move the rectangular window forward by a unit distance, update the coordinate of the center point of the boundary as (ɛ + n, Yn), where ɛ is the unit distance, and repeat the judgment of the number of sampling points until the current sampling point is equal to the total number of sampling points of the Mel spectrogram.

5. The two-stage detection system for sleep apnea according to claim 4, characterized in that: Use the longitudinal box algorithm AV-Box to detect sleep apnea segments in the Mel spectrogram to obtain potential sleep apnea segments. It further includes: Calculate the interval between adjacent change points output, and judge whether the interval is less than the first threshold. If so, merge the corresponding two change points into one change point; Calculate the segment duration between the merged adjacent change points, and judge whether the segment duration is less than the second threshold or greater than the third threshold. If so, delete the corresponding segment; where the second threshold is less than the third threshold; Calculate the interval between adjacent segments after deleting the segment, and judge whether the interval is greater than the fourth threshold. If so, mark the corresponding segment as a potential sleep apnea segment; where the fourth threshold is greater than the third threshold.

6. The two-stage detection system for sleep apnea according to claim 5, characterized in that: Sample the potential sleep apnea segment, extract the S1 peak and S2 peak of the potential sleep apnea segment, and calculate the RR interval according to the time interval between adjacent S1 peaks, including: Perform Hilbert transform on the potential sleep apnea segment and calculate the energy envelope of the potential sleep apnea segment; Compare the energy envelope value with the threshold T1, and retain the energy envelope value greater than the threshold T1; Set variables d12 and d21 to represent the time intervals of S1-S2 peaks and S2-S1 peaks respectively; Within the time window corresponding to d12, obtain the peak value with the largest energy envelope value and mark it as the S1 peak; within the time window corresponding to d21, obtain the peak value with the largest energy envelope value and mark it as the S2 peak; Calculate the time interval Δt between adjacent S1 peaks and S2 peaks, compare the time interval Δt with variables d12 and d21. If the time interval Δt is less than d12, mark the corresponding peak pair as the S1 - S2 peak pair; if the time interval Δt is greater than d12 and less than d21, mark the corresponding peak pair as the S2 - S1 peak pair. Calculate the time interval between adjacent S1 peaks as the RR interval.

7. The two - stage detection system for sleep apnea according to claim 6, wherein: Obtain the statistical features of the RR interval, and the statistical features include: maximum value, minimum value, range, standard deviation, mean, skewness, and kurtosis.

8. The two - stage detection system for sleep apnea according to claim 7, wherein: Calculate the apnea - hypopnea index AHI of the user through the following formula: AHI = N / T where N represents the number of sleep apnea events of the user; T represents the sleep duration of the user.

Citation Information

Patent Citations

  • Apnea snore recognition method based on double-flow multi-scale model

    CN116486839A

  • Sleep breathing monitoring method and device

    CN109431470A

  • Acoustic upper airway assessment system and method, and sleep apnea assessment system and method relying thereon

    US20170119303A1