A Sleep Apnea Detection Method Based on Multi-View Acoustic Spectrum Clustering

Through the combination of multi-view acoustic spectrum clustering and convolutional attention Utime network, the problems of low accuracy and insufficient real-time detection of sleep apnea in the prior art are solved, and high-precision apnea event detection and unsupervised learning are achieved, which improves the reliability and clinical application value of detection.

CN119601042BActive Publication Date: 2025-07-11NANJING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411454695.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-07-11
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

The prior art has low accuracy in sleep apnea detection, especially in complex noise environments, information is lost or misjudged, and it is difficult to achieve real-time detection and clinical application reliability.

Method used

The multi-view acoustic spectrum clustering method is used to construct three-dimensional tensors of time-frequency graphs, Mel spectrum graphs and wavelet graphs, combined with the pre-trained convolutional attention Utime network, and the apnea event is detected. The complementary and correlation information between the multi-view acoustic spectrum is used to learn robust acoustic feature representations.

Benefits of technology

It improves the accuracy and real-time nature of sleep apnea detection, can accurately identify apnea events in complex environments, and can adaptively discover the structure and category information of respiratory audio through unsupervised learning, which improves the reliability and clinical application value of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119601042B_ABST
    Figure CN119601042B_ABST
Patent Text Reader

Abstract

The present application discloses a sleep apnea detection method based on multi-view acoustic spectral clustering, belonging to the technical field of digital signal processing, including: collecting the sleep breathing audio data of a user and performing frame segmentation processing; for the audio data after frame segmentation processing, generating a time-frequency diagram of each frame through fast Fourier transform; using Mel filters to perform filtering processing on each frame of audio signal, and after superimposing the energies within the frequency bands of each filter, obtaining a Mel spectrogram; performing continuous wavelet transform on each frame of audio signal, and obtaining a wavelet diagram according to the wavelet coefficients; constructing a multi-view three-dimensional tensor according to the time-frequency diagram, Mel spectrogram and wavelet diagram; clustering the multi-view three-dimensional tensor to obtain the feature vectors of each time step; using the feature vectors as inputs and utilizing a pre-trained convolutional attention Utime network to perform sleep apnea detection on the sleep breathing audio data of the user; aiming at the low accuracy of sleep apnea detection in the prior art, the present application provides high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital signal processing technology. More specifically, it relates to a method for detecting sleep apnea based on multi-view acoustic spectral clustering. Background Art

[0002] Currently, audio signals have attracted extensive research in the field of sleep apnea detection due to their good non-invasiveness and ease of use. However, some studies directly estimate the AHI or OSA severity category without detecting specific OSA events, making it difficult to manually verify the output of the method and unable to meet clinical needs. At the same time, most studies have low time resolution for OSA detection and cannot achieve real-time detection. Therefore, it is of great significance to develop a real-time detection method for sleep apnea based on audio signals.

[0003] Application No. CN202310438705.0 proposed a method for identifying sleep apnea events by analyzing nocturnal sleep audio. It mainly includes an audio acquisition module, an audio signal feature extraction module, a feature annotation module, a training set construction module, and a statistical classification module. First, feature extraction is performed on the audio signal. The Mel Frequency Cepstral Coefficients (MFCC) are calculated. Through short-time Fourier transform (STFT) and Mel filter bank processing of the audio signal, 13-dimensional MFCC features are extracted. These include 1-dimensional signal frame energy and 12-dimensional DCT coefficients to capture information such as pitch, volume, and formants in the audio signal. Then, the feature annotation module is used to annotate the extracted MFCC features into two labels: normal sleep and apnea and hypopnea events. Finally, a training set for a time series classification algorithm model is constructed based on the annotated MFCC features, trained using the MiniRocket algorithm, and the classification task is completed through a linear classifier.

[0004] Although this method has achieved the recognition of sleep apnea events to a certain extent, there are still some obvious drawbacks. First, relying solely on MFCC features is not sufficient to comprehensively capture all the useful information in the audio signal. Especially in a complex noise environment, the selection of features with fewer dimensions may lead to information loss or misjudgment. In addition, although the MiniRocket algorithm performs well in time series classification, its fixed convolutional kernel structure and single feature aggregation method cannot effectively capture local spatio-temporal correlations and features with dynamic changes. This greatly limits the classification performance when dealing with complex and variable all-night audio data. Additionally, the system uses a fragment classification method, splitting the all-night audio into multiple fragments for processing. This approach may ignore the continuity information between fragments, thus affecting the real-time performance and accuracy of detection. Due to the neglect of the connection relationship between fragments, some short apnea events may be missed, or misjudgment may occur at the boundaries. Finally, the post-processing steps of the system are relatively simple, lacking further optimization and verification of the detection results, which also limits its reliability and practicality in clinical applications. Summary of the Invention

[0005] 1. Technical Problem to be Solved

[0006] Aiming at the problem of low detection accuracy of sleep apnea in the prior art, this application provides a sleep apnea detection method based on multi-view acoustic spectrum clustering, comprehensively utilizing time-frequency diagrams, Mel spectrograms, and wavelet Figure 3 types of acoustic spectrum features, constructing a multi-view Figure 3 dimensional tensor and performing clustering to obtain low-dimensional embedding vectors representing features at different time steps, and then using a pre-trained convolutional attention Utime network to detect apnea events, fully mining the complementary and associated information between multi-view acoustic spectra, learning more robust and discriminative acoustic feature representations, showing better apnea event detection performance on respiratory audio data, and having a high detection accuracy.

[0007] 2. Technical Solution

[0008] The objectives of this application are achieved through the following technical solutions.

[0009] One aspect of the present application provides a sleep apnea detection method based on multi-view acoustic spectral clustering, including: collecting the sleep breathing audio data of a user; performing frame division processing on the collected audio data; for the audio data after frame division processing, calculating the energy spectrum through fast Fourier transform, and generating a time-frequency diagram for each frame according to the energy spectra of all frames; filtering each frame of audio signal by using a Mel filter, and obtaining a Mel spectrogram after superimposing the energies within the frequency bands of each filter; performing continuous wavelet transform on each frame of audio signal, calculating the wavelet coefficients, and obtaining a wavelet diagram according to the wavelet coefficients; constructing a multi-view Figure 3 dimensional tensor according to the time-frequency diagram, Mel spectrogram and wavelet diagram; clustering the multi-view Figure 3 dimensional tensor to obtain the feature vectors at each time step; using the feature vectors as inputs and utilizing a pre-trained convolutional attention Utime network to perform sleep apnea detection on the sleep breathing audio data of the user.

[0010] Among them, the Mel filter is a filter bank designed based on the auditory characteristics of the human ear. The human ear has different perception abilities for sounds of different frequencies, and the Mel scale is a frequency scale proposed according to this characteristic of the human ear. The Mel filter bank consists of a series of triangular filters, each filter is equally spaced on the Mel scale, and has an unequal-width bandwidth on the frequency axis. The filter bandwidth in the low-frequency region is narrower, and the filter bandwidth in the high-frequency region is wider, which is consistent with the frequency selectivity of the human ear. Filtering the audio signal through the Mel filter bank can obtain a frequency-domain representation that is more matched with the human ear perception.

[0011] The time-frequency diagram is a commonly used method for representing the time-frequency domain of an audio signal. By performing short-time Fourier transform (STFT) on the audio signal, the energy distribution of the signal at different time points and frequency points can be obtained, forming a two-dimensional time-frequency matrix. The time-frequency diagram intuitively shows the variation law of the signal energy with time and frequency, and can capture the dynamic characteristics of the audio. In the solution, each frame of audio corresponds to a time-frequency diagram, which reflects the spectral structure within the frame.

[0012] The Mel spectrogram is a time-frequency representation obtained by transforming the frequency axis from a linear scale to a Mel scale based on the time-frequency diagram. Specifically, each column in the time-frequency diagram (corresponding to a frequency point) is filtered through the Mel filter bank, and the energy values output by the filter form a column of the Mel spectrogram. Due to the non-linear arrangement of the Mel filter bank on the frequency axis, the Mel spectrogram is more uniform in perception than the original time-frequency diagram and can better reflect the human ear's ability to distinguish different frequencies. The Mel spectrogram is commonly used in tasks such as speech recognition and audio classification.

[0013] The wavelet map is a time-frequency representation obtained based on wavelet transform. Different from the Fourier transform which uses fixed sine basis functions, the wavelet transform uses scalable and translatable wavelet basis functions to decompose signals. Through continuous wavelet transform, the wavelet coefficients of the signal at different scales (frequencies) and time points can be obtained, reflecting the local time-frequency characteristics of the signal. The wavelet transform has better time-frequency localization ability than the Fourier transform and can capture the instantaneous features and singularities of the signal. In the solution, wavelet transform is performed on each frame of audio to obtain the corresponding wavelet map, enriching the information of the multi-view representation.

[0014] The Utime network is a convolutional neural network architecture for sequence labeling tasks. It consists of two parts, an encoder and a decoder, forming a U-shaped network structure. The encoder extracts high-level features of the input sequence through convolution and pooling operations, and the decoder restores the resolution of the feature map through upsampling and skip connections, and finally generates prediction labels at each time step. The Utime network is widely used in fields such as speech recognition and named entity recognition, and can effectively model the long-range dependencies of sequence data. In the solution, the Utime network enhanced by convolutional attention is used to model the feature sequence of respiratory audio to achieve frame-level sleep apnea event detection.

[0015] Furthermore, according to the time-frequency map, Mel spectrogram, and wavelet map, a multi-view Figure 3 dimensional tensor is constructed, including: adjusting the sizes of the time-frequency map, Mel spectrogram, and wavelet map using bilinear interpolation; arranging the adjusted time-frequency map, Mel spectrogram, and wavelet map in the order of the audio data after frame processing to obtain three three-dimensional tensors, each tensor representing a view; for each view, flattening the image at each time step into a one-dimensional vector; calculating the similarity between each pair of time points in the one-dimensional vector in each view using a Gaussian kernel function to obtain the similarity matrix corresponding to the view; performing weighted averaging on the similarity matrices of the three views to obtain the fused similarity matrix S; according to the similarity matrix S, constructing a Laplacian matrix and solving for eigenvalues and eigenvectors to obtain the low-dimensional embedding vector U.

[0016] Among them, in multi-view learning, a view refers to different representations or description methods of the same data object. Each view depicts the characteristics of the data from different perspectives and provides complementary information. By fusing the information of multiple views, the internal structure and attributes of the data can be described more comprehensively and accurately. In the solution, the time-frequency diagram, Mel spectrogram, and wavelet diagram constitute three views of the sleep apnea audio data. They represent the audio signal from the perspectives of time-frequency energy distribution, perceived spectrum, time-frequency localization, etc., reflecting different characteristic aspects of the audio. The multi-view representation enhances the information volume and diversity of the data, which helps to improve the performance of subsequent clustering and detection tasks. The low-dimensional embedding vector U is a compact data representation obtained by dimensionality reduction of the similarity matrix through the spectral clustering algorithm. Specifically, first, the Laplacian matrix L is constructed based on the similarity matrix S, and then the eigenvalue decomposition of L is performed to obtain a set of eigenvalues and corresponding eigenvectors. The eigenvectors corresponding to the first k smallest non-zero eigenvalues are selected to form an \(N_t\times k\) matrix U, where \(N_t\) is the number of time steps and k is the dimension of the embedding space (usually much smaller than the dimension of the original data). Each row of U represents the low-dimensional embedding vector of a time step, reflecting the similarity relationship between this time step and other time steps.

[0017] Furthermore, based on the similarity matrix S, the low-dimensional embedding vector U is obtained by constructing the Laplacian matrix and solving the eigenvalues and eigenvectors, including: constructing the adjacency matrix W according to the similarity matrix S; constructing the degree matrix D by calculating the degree of each node according to the adjacency matrix W; obtaining the normalized Laplacian matrix L according to the degree matrix D and the adjacency matrix W; performing eigenvalue decomposition on the normalized Laplacian matrix L, selecting the eigenvectors corresponding to the first k smallest non-zero eigenvalues, and normalizing the eigenvectors to obtain the low-dimensional embedding vector U.

[0018] Among them, the adjacency matrix is a matrix representation used in graph theory to describe the graph structure. For a graph with N nodes, the adjacency matrix W is an N×N square matrix, where the element W(i, j) represents the connection weight between node i and node j. If there is an edge connecting node i and node j, then W(i, j) is the weight of the edge (for an unweighted graph, the weight is 1); if there is no edge connecting node i and node j, then W(i, j) is 0. The adjacency matrix is a symmetric matrix, that is, W(i, j) = W(j, i). The adjacency matrix completely characterizes the topological structure of the graph and the connection relationship between nodes. In the scheme, the fused similarity matrix S is assigned to the adjacency matrix W. Each node here corresponds to a time step, and the connection weight between nodes is the similarity between the corresponding time steps. Through the adjacency matrix W, the sleep apnea audio data is transformed into an undirected weighted graph, where nodes represent time steps and edges represent the similarity between time steps. The adjacency matrix W characterizes the structure and correlation of the audio data in the time dimension, providing a basis for subsequent spectral clustering.

[0019] The degree matrix is another matrix representation of a graph related to the adjacency matrix. For a graph with N nodes, the degree matrix D is an N×N diagonal matrix, where the diagonal element D(i, i) represents the degree of node i, that is, the number of edges connected to node i (for a weighted graph, it is the sum of the weights of the edges connected to node i); the non-diagonal elements D(i, j) are 0. The degree matrix characterizes the connection strength and importance of each node in the graph. In the scheme, the degree of each node (time step) is calculated according to the adjacency matrix W, that is, d_i = \sum_{j = 1}^{N_t}W(i, j), where N_t is the number of time steps. The degrees of the nodes are formed into a diagonal matrix to obtain the degree matrix D. The degree matrix D reflects the centrality and influence of each time step in the graph structure. Nodes with a large degree occupy important positions in the graph and have more connections with other nodes; nodes with a small degree are relatively isolated and have a weak association with other nodes. The degree matrix D and the adjacency matrix W are used together to construct a normalized Laplacian matrix, and then spectral clustering is realized.

[0020] Furthermore, cluster the multi-view Figure 3 dimensional tensor to obtain the feature vectors of each time step, including: using the K-means clustering algorithm to cluster the low-dimensional embedding vector U; for each cluster, using the cluster center as the feature vector of the corresponding time step.

[0021] Among them, the K-means clustering algorithm is a commonly used unsupervised machine learning algorithm for dividing a dataset into K clusters. The goal of the algorithm is to minimize the sum of the squared distances between each data point and its cluster center, thereby achieving compactness within the clusters and separation between the clusters. The basic steps of the K-means algorithm are as follows: Randomly select K data points as the initial cluster centers. For each data point in the dataset, calculate the distances between it and each cluster center, and assign it to the cluster represented by the nearest cluster center. For each cluster, recalculate the cluster center, that is, solve for the mean vector of all data points within the cluster. Repeat until the cluster centers no longer change significantly or the maximum number of iterations is reached.

[0022] In the solution, the K-means algorithm is used to cluster the low-dimensional embedding vector U. Each time step corresponds to a row vector of U, representing the coordinates of that time step in the spectral embedding space. The K-means algorithm divides these vectors into K clusters, and each cluster represents a group of semantically similar time steps. The clustering process can be regarded as finding the natural structure of the data distribution in the embedding space and grouping similar time steps into one category. The clustering results can be used to extract the feature vectors of each time step. Specifically, for each cluster, calculate the mean of all the vectors corresponding to the time steps within the cluster to obtain a cluster center vector. Use the center vector of the cluster to which each time step belongs as the feature vector of that time step.

[0023] Furthermore, the Gaussian kernel function is used to calculate the similarity between each pair of time points in the one-dimensional vectors of each view, obtaining the similarity matrix corresponding to the view, including: The expression of the similarity matrix is as follows:

[0024]

[0025]

[0026]

[0027] Among them, S X (i,j) represents the view of the time-frequency diagram; S Y (i,j) represents the view of the Mel spectrogram; S Y (i,j) represents the view of the wavelet diagram; σ represents the bandwidth parameter of the Gaussian kernel, vec(T X,i ) represents flattening the image T at the i-th time point into a one-dimensional vector; vec(T X,i ) represents flattening the image T at the i-th time point into a one-dimensional vector; vec(T Y,i ) represents flattening the image T at the i-th time point into a one-dimensional vector; vec(T Y,i ) represents flattening the image T at the i-th time point into a one-dimensional vector; vec(T Z,i ) represents representing flattening the image T at the i-th time point into a one-dimensional vector; vec(T Z,i ) represents flattening the image T at the i-th time point into a one-dimensional vector;

[0028] Furthermore, for the audio data after frame segmentation processing, through fast Fourier transform, calculate the energy spectrum, and generate a time-frequency graph for each frame based on the energy spectra of all frames, including:

[0029] Calculate the energy spectrum through the following formula:

[0030] X(i,k) = FFT[x i (m)]

[0031] E(i,k) = [X(i,k)] 2

[0032] Among them, i represents the i-th time point, k represents the k-th spectral line in the frequency domain; E represents the two-dimensional energy spectrum matrix, where the rows in the two-dimensional matrix represent frequency and the columns represent time; FFT represents fast Fourier transform; x represents the audio data at the i-th time point;

[0033] Furthermore, use the Mel filter to filter each frame of the audio signal. After superimposing the energies within each filter frequency band, obtain the Mel spectrogram, including:

[0034] Energy superposition is carried out through the following formula:

[0035]

[0036] Among them, S(i,m) represents the superimposed energy; H m (k) represents the frequency response of the m-th Mel filter; N represents the total number of spectral lines.

[0037] Furthermore, perform continuous wavelet transform on each frame of the audio signal, calculate the wavelet coefficients, and obtain the wavelet graph based on the wavelet coefficients, including:

[0038] Calculate the wavelet coefficients through the following formula:

[0039]

[0040] Among them, W(a,b) represents the wavelet coefficient; t represents time; a represents the scale factor; b represents the translation factor; w0 represents the center frequency of the mother wavelet;

[0041] Furthermore, based on the degree matrix D and the adjacency matrix W, obtain the normalized Laplacian matrix L, including:

[0042]

[0043] Among them, I represents the identity matrix.

[0044] Another aspect of the present application further provides a sleep apnea detection system based on multi-view acoustic spectrum clustering, which is used to execute a sleep apnea detection method based on multi-view acoustic spectrum clustering of the present application.

[0045] 3. Beneficial Effects

[0046] Compared with the prior art, the advantages of the present application are as follows:

[0047] By performing frame processing on sleep breathing audio data, calculating the energy spectrum to generate a time-frequency diagram, obtaining a Mel spectrogram using a Mel filter, and obtaining a wavelet diagram through continuous wavelet transform, the breathing audio signal can be comprehensively characterized from three different perspectives of the time domain, frequency domain, and time-frequency domain, and richer acoustic feature information can be extracted. By adjusting the sizes of the three spectrograms and arranging them in the order of audio frames to obtain a three-dimensional tensor, flattening each frame image into a one-dimensional vector and calculating the similarity using a Gaussian kernel function, and constructing a similarity matrix for each view, the correlation between different time steps can be measured while retaining the time structure information of the spectrogram, preparing a data basis for subsequent clustering.

[0048] By weighted averaging the similarity matrices of the three views to obtain a fusion matrix S, and constructing a Laplacian matrix based on S to solve the eigenvalues and eigenvectors to obtain a low-dimensional embedding vector U, the fusion and dimensionality reduction of multi-view acoustic spectrum features are realized. By fusing the similarity information of different views, the internal structure and data distribution of breathing audio in different feature domains can be comprehensively captured, and the low-dimensional embedding vector after dimensionality reduction can better highlight the discrimination between different categories of breathing segments.

[0049] The K-means algorithm is used to cluster the low-dimensional embedding vector U, and the cluster center is used as the feature vector corresponding to the time step. Combining the feature vectors of all time steps to obtain a two-dimensional matrix can adaptively discover the inherent structure and category information in breathing audio during the clustering process without pre-labeling the breathing segments, having the advantage of unsupervised learning. The feature matrix obtained by clustering contains the category membership information of each time step and can be used as a compact representation of the breathing segment for subsequent classification.

[0050] Inputting the feature matrix obtained by clustering into a pre-trained convolutional attention Utime network for sleep apnea event detection can make full use of the feature extraction and classification capabilities of the deep learning model to automatically learn highly discriminative high-level features. At the same time, the convolutional attention module can focus on the key regions in the input feature matrix that have a greater impact on the classification result, improving the representation ability of the features. The Utime network adopts a U-Net-like structure and introduces skip connections in the encoding and decoding processes, which can not only achieve hierarchical extraction of features but also integrate information of different scales, enhancing the feature representation and classification performance of the network. Brief Description of the Drawings

[0051] Figure 1 It is a technical flow chart of a sleep apnea detection method based on multi-view acoustic spectrum clustering in this application;

[0052] Figure 2 It is a schematic diagram of the convolutional attention Utime network in this application;

[0053] Figure 3 It is a sleep apnea detection result diagram in this application;

[0054] Figure 4 It is an AHI estimation result diagram in this application. Specific embodiments

[0055] The following combines the description of the accompanying drawings of the specification and specific embodiments to describe this application in detail.

[0056] Figure 1 It is a technical flow chart of a sleep apnea detection method based on multi-view acoustic spectrum clustering in this application. The real-time sleep apnea detection method based on multi-view acoustic spectrum clustering proposed in this application mainly includes three steps: The first step is multi-view acoustic spectrum clustering. First, collect the overnight sleep audio signal through a microphone and perform preprocessing and frame segmentation on it. Then, perform feature representation on each frame of audio from three perspectives: the time-frequency diagram, the Mel spectrogram, and the wavelet Figure 3 For each frame of audio, feature representations are obtained from three perspectives: the time-frequency diagram, the Mel spectrogram, and the wavelet. Specifically, the time-frequency diagram is obtained through the short-time Fourier transform (STFT), which shows the energy distribution of the signal in the time-frequency plane; the Mel spectrogram is obtained by mapping the spectrum to the Mel scale and filtering, which simulates the auditory characteristics of the human ear; the wavelet diagram is obtained through the continuous wavelet transform (CWT), which provides the time-frequency representation of the signal at different scales. Then, the feature representations of the three spectrograms are constructed into a three-dimensional tensor, and the similarity between different time steps in the tensor is calculated through a kernel function to obtain the similarity matrix of each view. Then, the similarity matrices of the three views are weighted and averaged for fusion to obtain a comprehensive similarity matrix. Finally, a Laplacian matrix is constructed based on the fused similarity matrix and the eigenvalues and eigenvectors are solved to obtain the low-dimensional embedding representation of the audio segments at different time steps. Through spectral clustering, the segments with similar embedding representations are clustered into one category, realizing the adaptive segmentation and representation learning of the overnight audio.

[0057] Second step: Sleep apnea detection. In this step, a convolutional attention Utime network is used to detect sleep apnea events in the audio segments belonging to each cluster in step S1. The network input is a feature matrix, which is formed by concatenating the embedding vectors of different segments obtained by clustering in step S1 in chronological order. The main body of the network adopts a U-Net-like structure, including multiple downsampling and upsampling modules, which can extract local and global context information at different scales. At the same time, skip connections are introduced between the encoder and the decoder to integrate features at different levels. In addition, convolutional attention modules are embedded in each downsampling and upsampling module. By generating attention weights with the same size as the input feature map, the features in the key regions are adaptively enhanced. The network outputs a breathing event prediction sequence of the same length as the input, indicating whether a breathing pause has occurred in the corresponding segment at each time step. This network adopts an end-to-end training method, uses the labeled breathing events as the supervision signal, optimizes the model parameters through the binary cross-entropy loss function, and finally realizes the second-level breathing apnea event detection based on audio signals.

[0058] Third step: Post-processing and AHI estimation. Since the diagnosis of sleep apnea requires counting the number of events occurring per hour, the output of step S2 needs to be post-processed. First, adjacent sleep apnea segments are merged, and segments with too short durations are filtered out. Then, the total duration of sleep apnea during the whole night is counted and divided by the total sleep time to obtain the average sleep apnea duration in hours, that is, the AHI index. Finally, according to the AHI index the SA severity of the subject is divided into four grades: normal, mild, moderate, and severe.

[0059] Specifically, the first step is to extract comprehensive acoustic features from the whole-night audio signal using multi-view acoustic spectral clustering from three perspectives: the time-frequency graph, the Mel spectrogram, and the wavelet graph. First, the whole-night audio signal is divided into multiple segments according to a length of 10 minutes, and each segment can be represented as x(n). Then, each segment is framed and windowed to obtain the frame signal x_i(n), where i represents the frame number. The frame length is set to 25 ms, and the corresponding number of sampling points is N = Fs×0.025 = 200 points; the frame shift is set to 10 ms, and the corresponding number of sampling points is M = Fs×0.01 = 80 points. Assuming that the total length of each 10-minute segment is L sampling points, then it can be divided into (L - N) / M + 1 frames, and each segment is about 60,000 frames.

[0060] Then, a fast Fourier transform (Fast Fourier Transform, FFT) is performed on each frame signal, and the energy spectrum is obtained: X(i,k) = FFT[x i (m)]

[0061] E(i,k) = [X(i,k)] 2

[0062] Among them, i represents the i-th time point, k represents the k-th spectral line in the frequency domain; E represents the two-dimensional energy spectrum matrix, where the rows in the two-dimensional matrix represent frequency and the columns represent time; FFT represents the fast Fourier transform; x represents the audio data at the i-th time point, so as to draw the time-frequency diagram of each frame.

[0063] Within the frequency range [0, 0.5Fs], K triangular filters are uniformly generated according to the Mel frequency scale, and the frequency response of the k-th filter is denoted as H k (f), k = 1, 2,..., K. For the FFT result X(i,k) of the i-th frame signal, multiply it by the frequency response of each filter and sum to obtain the energy S(i,m) of the k-th Mel frequency band of the i-th frame:

[0064]

[0065] Among them, H m (k) is the frequency response of the m-th Mel filter (there are a total of M = 13). Take the logarithm of S(i,m) and normalize it to obtain the Mel spectrogram.

[0066] Apply continuous wavelet transform (CWT) to each frame signal x i (m), and calculate the wavelet coefficients by adjusting the scale parameter a and the translation parameter b. The calculation formula for the wavelet coefficient W(a,b) is as follows:

[0067]

[0068] Among them, ψ represents the complex conjugate of the wavelet basis function, and the Morlet wavelet is selected as the wavelet basis function, and its expression is:

[0069]

[0070] The calculation formula for the wavelet coefficient W(a,b) can be further deduced as:

[0071]

[0072] After normalizing the wavelet coefficients, draw the wavelet diagram, and the color intensity of the image represents the magnitude of the wavelet coefficients, showing the change pattern of the signal in time and frequency.

[0073] Adjust the sizes of the time-frequency diagram, Mel spectrogram, and wavelet diagram generated for each frame, and arrange them in the original frame order to obtain three three-dimensional tensors T X 、T Y 、T Z ∈R 60000×32×32For each view, flatten the image at each time step into a one-dimensional vector, and use the Gaussian kernel function to calculate the similarity matrix s between each pair of time points X , s Y and s Z :

[0074]

[0075]

[0076]

[0077] where σ is the bandwidth parameter of the Gaussian kernel, and vec(T X,i ) represents flattening the image T at the i-th time point into a one-dimensional vector X,i .

[0078] Add the three similarity matrices and take the average to fuse them into a comprehensive similarity matrix S, and assign it to the adjacency matrix W. Calculate the degree of each node to construct the degree matrix D, and then calculate the normalized Laplacian matrix Solve the eigenvalues and eigenvectors of the Laplacian matrix L, select the eigenvectors corresponding to the first k = 2048 smallest non-zero eigenvalues, and form a new low-dimensional embedding U after standardization

[0079] Use the K-means algorithm for clustering, and set the number of clusters to 128. For each cluster, its center value is the one-dimensional feature vector at that time point, and combining them can obtain the 128-dimensional feature vector at each time step

[0080] Figure 2Schematic diagram of the convolutional attention Utime network for this application, sleep apnea detection based on the convolutional attention Utime network (CBAM-Utime). This network mainly consists of an encoding module, a decoding module, a convolutional attention module, and a skip connection unit. The specific implementation steps include: using the encoding module to extract the deep features of the input sequence layer by layer. The encoding module consists of six convolutional blocks, and each convolutional block contains a convolutional layer, a batch normalization layer, a max pooling layer, and an activation function. The convolutional kernel size of the convolutional block is 5; the number of filters is 32, 48, 64, 96, 128, and 256 respectively; the corresponding pooling sizes are 8, 6, 4, 2, 2, and 1 respectively. Using the decoding module to perform upsampling layer by layer. The decoding module consists of five upsampling blocks, and each upsampling block contains a convolutional layer, a batch normalization layer, an upsampling layer, and an activation function. The upsampling block shares the same convolutional kernel size and the number of filters with the corresponding convolutional block, and uses the pooling size of the convolutional block as the upsampling factor. The number of feature channels of the decoding module decreases layer by layer and finally has the same data dimension as the original input sequence. Using the convolutional attention module to calculate the feature weights from both the spatial and channel dimensions to enhance the network's representation ability for the input features. This module contains a channel attention sub-module and a spatial attention sub-module, and the input feature sequence passes through the two modules in sequence. Assume the input feature sequence is F ∈ R C×1×L , and the channel weight distribution after passing through the channel attention sub-module is W c ∈ R C×1×1 , and the spatial weight distribution after passing through the spatial attention sub-module is W s ∈ R C×1×L . The process of passing through these two sub-modules can be expressed as: where C is the number of channels of the feature sequence, L is the length of the feature sequence, is the element-wise multiplication operation, and F' is the final output of the convolutional attention module.

[0081] The skip connection is used to fuse the feature maps of the corresponding levels in the encoding module and the decoding module to comprehensively utilize the shallow and deep feature information. Specifically, the output of the i-th convolutional block in the encoding module is concatenated with the output of the (6 - i)-th upsampling block in the decoding module along the channel dimension. Taking the first convolutional block and the fifth upsampling block as an example, their output feature map sizes are: the first convolutional block: 32×(N t / 8), the fifth upsampling block: 32×(N t / 8). These two feature maps are concatenated along the channel dimension to obtain a 64×(N t / 8) feature map, which is used as the output of the skip connection. Similarly, the feature maps of other levels are concatenated in this way. The output of the skip connection is then adjusted in the number of channels through a convolutional layer to match the input channel number of the next layer of the decoding module.

[0082] The output of the last layer of the decoding module (with the same size as the original input sequence) is processed through two layers to obtain the final classification result: Linear layer: The feature map is transformed from 3K×N t to 2×N t , where 2 represents two states: normal breathing (0) and sleep apnea (1). Convolutional layer (Conv2d): The kernel size is 1x1, and the number of input and output channels is 2, which is equivalent to classifying and predicting frame by frame. Softmax activation function: Normalize the two logit values for each frame to obtain the probabilities of belonging to normal breathing and sleep apnea. Select the one with the larger probability as the prediction label for the current frame, and finally obtain a sequence of length N t , representing the breathing state per second, and realizing the detection of sleep apnea events at the second level.

[0083] Post-process the obtained second-level sleep apnea detection results to estimate the AHI of the subject. The specific implementation steps are as follows: The label for normal breathing is 0, and the label for sleep apnea is 1. Record a single 0 in a continuous 1 segment as an outlier and correct it to 1; Detect continuous 1 segments, and if the duration exceeds 10 seconds, record it as one sleep apnea event; Divide the number of sleep apnea times of the subject throughout the night by the number of sleep hours to obtain the AHI of the subject. Figure 3 This is the sleep apnea detection result graph of this application, Figure 4 This is the AHI estimation result graph of this application. From the graph, it can be seen that there is a high consistency between the method proposed in this application and PSG. Specifically, both the mean absolute error (MAE) and the root mean square error (RMSE) are low, indicating that the deviation between the estimated value and the actual value is small. The correlation coefficient is as high as 0.954, indicating a very strong linear relationship between the estimated value and the actual value. The intraclass correlation coefficient (ICC) and the coefficient of determination R 2 values are 0.929 and 0.85 respectively, further proving the stability and reliability of this method among different individuals.

[0084] The above has schematically described the present invention and its implementation manners. This description is not restrictive. Without departing from the spirit or basic characteristics of the present application, the present application can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Any reference signs in the claims should not limit the claims involved. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to the technical solution without creative efforts, they shall fall within the protection scope of the present application. In addition, the term "comprising" does not exclude other elements or steps, and the term "a" before an element does not exclude including "a plurality of" such elements. The plurality of elements stated in the product claims can also be implemented by one element through software or hardware. The terms such as "first" and "second" are used to denote names and do not indicate any particular order.

Claims

1. A sleep apnea detection method based on multi-view acoustic spectrum clustering, characterized in that Including: Collecting the sleep breathing audio data of the user, and performing frame processing on the collected audio data to obtain each frame of audio data; For all frames of audio data after frame processing, calculating the energy spectrum of all frames through fast Fourier transform, and generating the time-frequency diagram of each frame according to the energy spectrum of all frames; Using a Mel filter to perform filtering processing on each frame of audio data, and after superimposing the energies within each filter frequency band, obtaining a Mel spectrogram; Performing continuous wavelet transform on each frame of audio data, calculating wavelet coefficients, and obtaining a wavelet diagram according to the wavelet coefficients; Constructing a multi-view three-dimensional tensor according to the time-frequency diagram, Mel spectrogram and wavelet diagram; Clustering the multi-view three-dimensional tensor to obtain the feature vector of each time step; Taking the feature vector as the input, and using the pre-trained convolutional attention Utime network to perform apnea detection on the sleep breathing audio data of the user; Constructing a multi-view three-dimensional tensor according to the time-frequency diagram, Mel spectrogram and wavelet diagram, including: Using the bilinear interpolation method to adjust the sizes of the time-frequency diagram, Mel spectrogram and wavelet diagram; Arranging the adjusted time-frequency diagram, Mel spectrogram and wavelet diagram in the order of the audio data after frame processing to obtain 3 three-dimensional tensors, and each tensor represents a view; For each view, flattening the image of each time step into a one-dimensional vector; Using a Gaussian kernel function to calculate the similarity between each pair of time points in the one-dimensional vector in each view to obtain the similarity matrix of the corresponding view; Performing weighted average on the similarity matrices of the three views to obtain the fused similarity matrix S; According to the similarity matrix S, by constructing a Laplacian matrix and solving for eigenvalues and eigenvectors, obtaining a low-dimensional embedding vector U; Clustering the multi-view three-dimensional tensor to obtain the feature vector of each time step, including: Adopting the K-means clustering algorithm to cluster the low-dimensional embedding vector U; For each cluster, using the cluster center as the feature vector of the corresponding time step.

2. The apnea detection method based on multi-view acoustic spectrum clustering according to claim 1, characterized in that: According to the similarity matrix S, by constructing a Laplacian matrix and solving for eigenvalues and eigenvectors, obtaining a low-dimensional embedding vector U, including: Constructing an adjacency matrix W according to the similarity matrix S; According to the adjacency matrix W, constructing a degree matrix D by calculating the degree of each node; According to the degree matrix D and the adjacency matrix W, obtaining a normalized Laplacian matrix L; Performing eigenvalue decomposition on the normalized Laplacian matrix L, selecting the eigenvectors corresponding to the first k smallest non-zero eigenvalues, and normalizing the eigenvectors to obtain a low-dimensional embedding vector U.

3. The apnea detection method based on multi-view acoustic spectrum clustering according to claim 1, characterized in that: Using a Gaussian kernel function to calculate the similarity between each pair of time points in the one-dimensional vector in each view to obtain the similarity matrix of the corresponding view, including: The expression of the similarity matrix is as follows: Among them, represents the view of the time-frequency diagram; represents the view of the Mel spectrogram; represents the view of the wavelet diagram; represents the bandwidth parameter of the Gaussian kernel, represents the image at the th time point flattened into a one-dimensional vector; image at the th time point flattened into a one-dimensional vector; image at the th time point flattened into a one-dimensional vector.

4. The apnea detection method based on multi-view acoustic spectrum clustering according to claim 3, characterized in that: For the audio data after frame splitting processing, calculate the energy spectrum through fast Fourier transform, and generate the time-frequency diagram of each frame according to the energy spectra of all frames, including: Calculate the energy spectrum through the following formula: Among them, represents the th time point, represents the th spectral line in the frequency domain; represents the two-dimensional energy spectrum matrix, where the rows of the two-dimensional matrix represent frequency and the columns represent time; FFT represents the fast Fourier transform; x represents the audio data at the i-th time point.

5. The sleep apnea detection method based on multi-view acoustic spectrum clustering according to claim 4, wherein: Filter each frame of audio signal by using a Mel filter, and after superimposing the energies in each filter frequency band, obtain a Mel spectrogram, including: Energy superposition is carried out through the following formula: Among them, represents the superimposed energy; represents the frequency response of the m-th Mel filter; N represents the total number of spectral lines.

6. The sleep apnea detection method based on multi-view acoustic spectrum clustering according to claim 5, wherein: Perform continuous wavelet transform on each frame of audio signal, calculate the wavelet coefficients, and obtain a wavelet diagram according to the wavelet coefficients, including: Calculate the wavelet coefficients through the following formula: Among them, represents the wavelet coefficient; represents time; a represents the scale factor; b represents the translation factor; represents the central frequency of the mother wavelet.

7. The sleep apnea detection method based on multi-view acoustic spectrum clustering according to claim 6, wherein: Obtain the normalized Laplacian matrix L according to the degree matrix D and the adjacency matrix W, including: Among them, represents the identity matrix.

Citation Information

Patent Citations

  • Sleep apnea event identification system based on time sequence algorithm

    CN116343828A

  • Single-channel electroencephalogram sleep classification model construction method and model

    CN117828412A

  • Systems and methods for user interface comfort evaluation

    US20230364365A1