Sleep apnea detection method and system based on multi-mode audio signal

By combining multimodal analysis of tracheal tone and ambient tone, and using enhanced branched forward neural network to extract deep features, the problem of insufficient detection accuracy in the prior art is solved, and efficient and reliable sleep apnea detection is achieved.

CN120240965APending Publication Date: 2025-07-04HUNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510342624.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing sleep apnea detection methods lack comprehensive consideration of tracheal sounds and ambient sounds, resulting in insufficient detection accuracy and applicability, and traditional models are difficult to capture local details and global context, limiting the generalization ability of the model.

Method used

A learnable linear weight layer is used to combine multimodal analysis of tracheal tones and ambient tones, and deep features in audio signals are extracted using an enhanced branch forward (EBranchformer) neural network, and various types of apnea events are detected through local-global feature collaborative modeling.

Benefits of technology

It significantly improves the accuracy and sensitivity of sleep apnea detection, enhances the ability to recognize sleep apnea, and makes monitoring more reliable and economical and convenient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120240965A_ABST
    Figure CN120240965A_ABST
Patent Text Reader

Abstract

The invention discloses a sleep apnea detection method and system based on a multi-mode audio signal, and the method comprises the following steps: obtaining a trachea sound and an environment sound of a sleep record, and dividing the trachea sound and the environment sound into audio segments, respectively extracting trachea sound features and environment sound features from the trachea sound and the environment sound of each audio clip, calculating a corresponding trachea sound weight and a corresponding environment sound weight, and carrying out weighted fusion on the trachea sound features and the environment sound features of each audio clip to obtain corresponding multi-mode audio features; and detecting an abnormal breathing event for each multi-mode audio feature by using an enhanced branch pre-device, and evaluating the severity of sleep apnea according to the number of the detected abnormal breathing events. According to the invention, the accuracy and sensitivity of sleep apnea detection are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio detection, and particularly to a method and system for detecting sleep apnea based on multi-modal audio signals. Background Art

[0002] Sleep apnea syndrome (SAS) is a prevalent sleep disorder that affects the health of hundreds of millions of people worldwide. According to research, approximately 80% of SAS cases are not diagnosed in a timely manner, leading to serious health risks for patients, such as cardiovascular diseases and metabolic syndrome. Although traditional polysomnography (PSG) is considered the gold standard for diagnosing SAS, its high cost and complex operation process make it unacceptable to many patients. In recent years, monitoring methods based on audio signals have gradually attracted attention. By analyzing the sound signals generated during sleep (such as snoring sounds, breathing sounds, etc.), effective monitoring of SAS can be achieved.

[0003] However, most existing studies focus on the analysis of a single sound source, lacking comprehensive consideration of tracheal sounds and environmental sounds, which limits the accuracy and applicability of detection. Some methods directly splice or average multi-modal features, ignoring the dynamic changes in signal quality. In addition, traditional models (such as CNN, RNN) are difficult to simultaneously capture the local details of tracheal sounds (such as interrupted breathing airflow) and the global context of environmental sounds (such as the duration of snoring), resulting in a decrease in performance after multi-modal feature fusion. Although by constructing a complex neural network model, deep features in audio signals can be extracted, thereby improving the detection performance of SAS. However, current research still has insufficient application of large-scale public datasets, resulting in limited generalization ability of the model. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: aiming at the technical problems existing in the prior art, the present invention provides a method and system for detecting sleep apnea based on multi-modal audio signals, designs a learnable linear weight layer to combine multi-modal analysis of tracheal sounds and environmental sounds, and detects various types of apnea events through local-global feature collaborative modeling based on an enhanced branch former (EBranchformer), significantly improving the accuracy and sensitivity of detection.

[0005] To solve the above technical problems, the technical solution proposed by the present invention is:

[0006] A method for detecting sleep apnea based on multi-modal audio signals, comprising the following steps:

[0007] Obtain the tracheal sound and ambient sound of the sleep recording, segment both the tracheal sound and the ambient sound into audio segments, extract tracheal sound features and ambient sound features from the tracheal sound and ambient sound of each audio segment respectively, calculate the corresponding tracheal sound weight and ambient sound weight, and perform weighted fusion on the tracheal sound features and ambient sound features of each audio segment to obtain the corresponding multi-modal audio features;

[0008] Use an enhanced branch predictor to detect abnormal breathing events for each multi-modal audio feature respectively, and evaluate the severity of sleep apnea based on the number of detected abnormal breathing events.

[0009] Further, when extracting tracheal sound features and ambient sound features from the tracheal sound and ambient sound of each audio segment respectively, it specifically includes:

[0010] Use the trained speech feature extractor to extract frame-level tracheal sound speech features for the current audio segment of the tracheal sound, and perform average pooling on the frame-level tracheal sound speech features along the time axis of the current audio segment to obtain tracheal sound features;

[0011] Use the trained speech feature extractor to extract frame-level ambient sound speech features for the current audio segment of the ambient sound, and perform average pooling on the frame-level ambient sound speech features along the time axis of the current audio segment to obtain ambient sound features.

[0012] Further, when calculating the corresponding tracheal sound weight and ambient sound weight, it specifically includes:

[0013] Input the tracheal sound features and ambient sound features into the learnable linear layer respectively to obtain the tracheal sound weight and ambient sound weight. The expression of the learnable linear layer is as follows:

[0014] W = Softmax(MLP(F))

[0015] Where MLP is a multi-layer perceptron, and F is the tracheal sound feature or ambient sound feature.

[0016] Further, each layer of the enhanced branch predictor includes a convolutional branch and a self-attention branch. When using the enhanced branch predictor to detect abnormal breathing events for each multi-modal audio feature respectively, it includes:

[0017] Use the convolutional branch to capture the local time-frequency pattern of the current multi-modal audio feature and extract local features. The expression is as follows:

[0018] F conv = DWConv(ReLU(BN(Conv1d(F, kernel = 3, stride = 1)))

[0019] Among them, Conv1d represents one-dimensional convolution with a kernel size of 3 and a stride of 1; BN represents batch normalization; ReLU represents an activation function; DWConv represents depthwise separable convolution; F is the prediction result of the previous layer, initially being multi-modal audio features;

[0020] The global context dependence of the current multi-modal audio features is calculated using the attention branch, and the expression is as follows:

[0021] F attn = MultiHeadAttention(F, F, F)

[0022] Among them, MultiHeadAttention is the multi-head self-attention mechanism; MultiHeadAttention(F, F, F) means that the query matrix, key matrix, and value matrix under the multi-head attention mechanism are all replaced with the prediction result F of the previous layer, initially being multi-modal audio features;

[0023] The dynamic weights are calculated based on the output features of the convolution branch and the attention branch, and then the output features of the convolution branch and the attention branch are weighted and fused to obtain the prediction result, and the expression is as follows:

[0024] γ = σ(W g · [F conv ; F attn )

[0025] F out = γ · F conv + (1 - γ) · F attn

[0026] Among them, γ represents the dynamic weight, F out represents the prediction result corresponding to the current multi-modal audio features, W f is a learnable parameter matrix, and σ is the Sigmoid function;

[0027] The prediction result is subjected to multi-task classification to judge the probabilities of different respiratory events corresponding to the current multi-modal audio features. Different respiratory events include normal respiratory events and at least one abnormal respiratory event, and the expression is as follows:

[0028]

[0029] Among them, P represents the probability distribution of each type of respiratory event of the current multi-modal audio features, Softmax represents the activation function, FC represents a fully connected layer with the same number of dimensions as the number of categories of respiratory events, represents the prediction result of the last layer of the enhanced branch forecaster, and L represents the number of layers of the enhanced branch forecaster.

[0030] Further, after splitting both the tracheal sound and the ambient sound into audio segments, it further includes the step of data labeling in the model training phase, specifically including:

[0031] If the multi-task classification is a binary classification task, label the breathing events in each audio segment as one or more of normal or abnormal;

[0032] If the multi-task classification is a ternary classification task, label the breathing events in each audio segment as one or more of normal or low breathing or apnea;

[0033] If the multi-task classification is a quinary classification task, label the breathing events in each audio segment as one or more of normal or low breathing or obstructive apnea or central apnea or mixed apnea.

[0034] Further, when evaluating the severity of sleep apnea based on the number of detected abnormal breathing events, it includes:

[0035] Designate one type of abnormal breathing event from all the abnormal breathing events;

[0036] According to the designated sliding window size, calculate the average value of the predicted probabilities of the designated abnormal breathing event of the multi-modal audio features in all audio segments within the sliding window in sequence. The expression is as follows:

[0037]

[0038] where W is the sliding window size, is the smoothed predicted probability of the designated abnormal breathing event of the multi-modal audio feature of the i-th audio segment, p j is the predicted probability of the designated abnormal breathing event of the multi-modal audio feature of the j-th audio segment within the sliding window of [i - W / 2, i + W / 2];

[0039] Compare the smoothed predicted probability of the designated abnormal breathing event of the multi-modal audio feature of each audio segment with the designated threshold. If the smoothed predicted probability is greater than the designated threshold, the corresponding audio segment is an abnormal audio segment of the designated abnormal breathing event;

[0040] Merge the consecutive abnormal audio segments into new abnormal audio segments, and count the number of all abnormal audio segments longer than the designated duration to obtain the number of the designated abnormal breathing events;

[0041] Calculate the corresponding AHI index according to the number of the designated abnormal breathing events, and match the AHI index calculation result with the sleep apnea standard to obtain the severity of the designated abnormal breathing event.

[0042] Further, after obtaining the severity of the specified abnormal breathing event, it further includes:

[0043] Calculating the prediction confidence interval to determine the reliability of the severity evaluation result of the specified abnormal breathing event. The formula is as follows:

[0044] P severe =σ(2.5·AHI audio -40), where AHI audio is the AHI index, and σ is the Sigmoid function.

[0045] The present invention also provides a sleep apnea detection system based on multi-modal audio signals, including a microprocessor and a computer storage medium. The microprocessor executes a computer program in the computer storage medium to implement the sleep apnea detection method based on multi-modal audio signals.

[0046] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a microprocessor, the steps of the sleep apnea detection method based on multi-modal audio signals are implemented.

[0047] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a microprocessor, the steps of the sleep apnea detection method based on multi-modal audio signals are implemented.

[0048] Compared with the prior art, the advantages of the present invention are as follows:

[0049] The present invention combines multi-modal analysis of tracheal sound and ambient sound, and then uses an enhanced branchformer (EBranchformer) neural network to extract deep features in the audio signal to predict various types of apnea events, significantly improving the detection accuracy and sensitivity, enhancing the recognition ability of sleep apnea, and making the monitoring of sleep apnea more reliable.

[0050] By integrating the prediction sequences of the enhanced branchformer and performing statistical analysis of abnormal events, the present invention obtains the detection and analysis results of sleep apnea symptoms, overcomes the problem of strong dependence on medical devices in traditional methods, and provides an economical and convenient solution for monitoring sleep apnea symptoms. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic diagram of the principle of an embodiment of the present invention;

[0052] Figure 2 is a brief flowchart of an embodiment of the present invention;

[0053] Figure 3Schematic diagram of the post - processing process for the prediction results of multi - modal audio features in the embodiments of the present invention. Detailed implementation manners

[0054] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0055] Embodiment 1

[0056] Developing a multi - modal analysis based on audio signals, especially combining tracheal sounds and ambient sounds, has important theoretical significance and practical application value.

[0057] To meet this need, the present invention proposes a method for detecting sleep apnea based on multi - modal audio signals, designs a learnable linear weight layer to combine multi - modal analysis of tracheal sounds and ambient sounds, and uses an enhanced branch former (EBranchformer) neural network to extract deep features in audio signals, significantly improving the accuracy and sensitivity of detection and making the monitoring of sleep apnea more reliable.

[0058] As Figure 1 shown, the method of this embodiment includes the following steps:

[0059] S1) Audio Features Extraction and Multi - modal Feature Fusion stage, including:

[0060] Obtain the tracheal sounds and ambient sounds of the sleep recording, divide both the tracheal sounds and ambient sounds into audio segments, and mark each audio segment according to whether there is a certain abnormal breathing event. If there is no abnormal breathing event, it is marked as normal. Extract tracheal sound features and ambient sound features for the tracheal sounds and ambient sounds of each audio segment respectively, calculate the corresponding tracheal sound weights and ambient sound weights, and perform weighted fusion on the tracheal sound features and ambient sound features of each audio segment to obtain the corresponding multi - modal audio features;

[0061] S2) Sleep Apnea Syndrome (SAS) Audio Detection Stage Based on Enhanced Branch Former: Use the enhanced branch former to detect abnormal breathing events for each multi - modal audio feature respectively.

[0062] S3) Post - processing Stage: Evaluate the severity of sleep apnea according to the number of detected abnormal breathing events.

[0063] The following is an explanation of each step.

[0064] As Figure 2 shown, step S1 of this embodiment includes the following steps:

[0065] S11) Remove noise and silence segments from the input sleep recording to improve the accuracy of feature extraction, and then perform downsampling. Specifically, the original tracheal sound and ambient sound are recorded at a sampling rate of 48 kHz and the sampling rate is reduced to 8 kHz. The downsampled audio is divided into 40-second segments to obtain: audio segment A of tracheal sound Ti and audio segment A of ambient sound Ei ;

[0066] S12) Through a speech feature extractor (Acoustic Feature Extractor) such as common ones like Frank or Hubert, convert the audio from a time-domain signal to frequency-domain features, extract frame-level speech features, and perform average pooling on the frame-level features along the time axis to obtain segment-level features. The expression is as follows:

[0067]

[0068] where T represents the duration of the audio segment, and F a (t) represents the t-th frame-level feature of the audio segment;

[0069] Therefore, when extracting tracheal sound features and ambient sound features for the tracheal sound and ambient sound of each audio segment respectively, it includes:

[0070] Use the trained speech feature extractor to extract frame-level tracheal sound speech features for the current audio segment of tracheal sound, and perform average pooling on the frame-level tracheal sound speech features along the time axis of the current audio segment to obtain tracheal sound feature F tracheal ;

[0071] Use the trained speech feature extractor to extract frame-level ambient sound speech features for the current audio segment of ambient sound, and perform average pooling on the frame-level ambient sound speech features along the time axis of the current audio segment to obtain ambient sound feature F env ;

[0072] S13) Calculate the tracheal sound weight and ambient sound weight. Specifically, input the tracheal sound feature and ambient sound feature into a learnable linear layer respectively to obtain the tracheal sound weight and ambient sound weight. The expression of the learnable linear layer is as follows:

[0073] W = Softmax(MLP(F))

[0074] where MLP is a multi-layer perceptron, and F is the tracheal sound feature or ambient sound feature. In this embodiment, according to the input tracheal and ambient sound features, the output dimension of the learnable linear layer is set to 2 to obtain the weights of the tracheal sound and ambient sound;

[0075] S14) Multiply the tracheal sound features and environmental sound features by their corresponding weights and then sum them to obtain weighted audio features, which are used as the corresponding multimodal audio features to enhance the model's sensitivity to different modal audio features. The expression is as follows:

[0076]

[0077] where W t and W e represent the weights of tracheal sound and environmental sound respectively.

[0078] In step S2 of this embodiment, the enhanced branchformer is stacked by multiple identical modules. Each layer includes a convolutional branch (CNN Branch) and a self-attention branch (Attention Branch). The enhanced branchformer captures local time-frequency patterns through the convolutional branch, models global context dependencies through the attention branch, and then fuses the two features through a dynamic gating mechanism to establish a deep relationship between sleep apnea events and audio features, obtaining the prediction result of the audio segment. The prediction of each sequence of the weighted audio features for sleep apnea events is expressed as:

[0079] F out = EBranchformer(F fuse )

[0080] During this process, the flow of each component of the enhanced branchformer is as Figure 2 shown, including:

[0081] S21) Use the convolutional branch to capture the local time-frequency patterns of the current multimodal audio features and extract local features. The expression is as follows:

[0082] F conv = DWConv(ReLU(BN(Conv1d(F, kernel = 3, stride = 1))))

[0083] where Conv1d represents one-dimensional convolution, the kernel size is 3, and the stride is 1; BN represents batch normalization; ReLU represents the activation function; DWConv represents depthwise separable convolution; F is the prediction result of the previous layer, initially the multimodal audio feature F fused ; S22) Use the attention branch to calculate the global context dependencies of the current multimodal audio features. The expression is as follows:

[0084] F attn = MultiHeadAttention(F, F, F)

[0085] Among them, MultiHeadAttention is the multi-head self-attention mechanism; MultiHeadAttention(F,F,F) means replacing the query matrix, key matrix, and value matrix under the multi-head attention mechanism with the prediction result F of the previous layer, initially being the multi-modal audio feature F fused , the query matrix, key matrix, and value matrix of the multi-head attention mechanism have the following functional relationships:

[0086]

[0087] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively, and Softmax is the activation function;

[0088] S23) Use the Dynamic Gating Mechanism: fuse the F output by the convolutional branch and the attention branch conv and F attn These two features, specifically calculate the dynamic weight according to the output features of the convolutional branch and the attention branch, and then weighted fuse the output features of the convolutional branch and the attention branch to obtain the prediction result. The expression is as follows:

[0089] γ = σ(W g ·[F conv ; F attn )

[0090] F out = γ·F conv +(1 - γ)·F attn

[0091] Among them, γ represents the dynamic weight, F out represents the prediction result corresponding to the current multi-modal audio feature, W f is the learnable parameter matrix, and σ is the Sigmoid function;

[0092] S24) Perform multi-task classification on the prediction result to judge the probabilities of different breathing events corresponding to the current multi-modal audio feature. Different breathing events include normal breathing events and at least one abnormal breathing event. The expression is as follows:

[0093]

[0094] Among them, P represents the probability distribution of each type of breathing event of the current multi-modal audio feature, Softmax represents the activation function, FC represents the fully connected layer with the same number of dimensions as the number of categories of breathing events, represents the prediction result of the last layer of the enhanced branch pre-actor, and L represents the number of layers of the enhanced branch pre-actor.

[0095] The enhanced branch preprocessor supports multi-task classification. In this embodiment, a binary classification task can be performed first to distinguish abnormal breathing events from normal breathing events. At this time, the probability expressions for different breathing events are as follows:

[0096]

[0097] Among them, P binary represents the output probability distribution of the binary classification task, which is a two-dimensional vector. Each element represents the probability that the prediction result of the last layer of the enhanced branch preprocessor belongs to one of the two categories of normal / abnormal. The output dimension of FC1 is 2, corresponding to the scores of the two categories of normal / abnormal.

[0098] On this basis, more fine-grained tasks can also be performed, including distinguishing apnea and hypopnea events, or detecting various types of apnea events.

[0099] For example, a three-classification task is performed to distinguish normal / hypopnea / apnea events. At this time, the probability expressions for different breathing events are as follows:

[0100]

[0101] Among them, P tri represents the output probability distribution of the three-classification task, which is a three-dimensional vector. Each element represents the probability that the prediction result of the last layer of the enhanced branch preprocessor belongs to one of the three categories of normal / hypopnea / apnea. The output dimension of FC2 is 3, corresponding to the scores of the three categories of normal / hypopnea / apnea.

[0102] For example, a five-classification task is performed to distinguish normal / hypopnea / obstructive apnea / central apnea / mixed apnea events. At this time, the probability expressions for different breathing events are as follows:

[0103]

[0104] Among them, P fine represents the output probability distribution of the five-classification task, which is a five-dimensional vector. Each element represents the probability that the prediction result of the last layer of the enhanced branch preprocessor belongs to one of the five categories of normal / hypopnea / obstructive apnea / central apnea / mixed apnea. The output dimension of FC3 is 5, corresponding to the scores of the five categories of normal / hypopnea / obstructive apnea / central apnea / mixed apnea.

[0105] In step S3 of this embodiment, for the prediction probabilities of the foregoing multi-classification tasks, algorithms such as smoothing, binary reduction, and clustering reduction are used to reduce the volatility of the prediction results, ensure the stability of the monitoring results, and then perform the evaluation and statistical analysis of the SAS detection results, such asFigure 2 As shown in Figure 2 , when evaluating the severity of sleep apnea based on the number of detected abnormal breathing events, the following steps are included:

[0106] S31) Designate an abnormal breathing event from all abnormal breathing events. For the designated abnormal breathing event, there is a series of prediction probabilities {p i}, i = 1,..., N, where p o ∈[0,1] represents the prediction probability of the multimodal audio feature of the i-th segment under the designated abnormal breathing event;

[0107] S32) As Figure 3 shown, according to the designated sliding window size, calculate the average value of the prediction probabilities of the designated abnormal breathing event of the multimodal audio features of all audio segments within the sliding window in sequence. The expression is as follows:

[0108]

[0109] where W is the sliding window size, is the smoothed prediction probability of the designated abnormal breathing event of the multimodal audio feature of the i-th audio segment, and p j is the prediction probability of the designated abnormal breathing event of the multimodal audio feature of the j-th audio segment within the sliding window of [i - W / 2, i + W / 2];

[0110] S33) Compare the smoothed prediction probability of the designated abnormal breathing event of the multimodal audio feature of each audio segment with the designated threshold. If the smoothed prediction probability is greater than the designated threshold, the corresponding audio segment is an abnormal audio segment of the designated abnormal breathing event;

[0111] As Figure 3 shown, in this embodiment, the smoothed probability is converted into a binary label through binarization, 0 for normal and 1 for abnormal. The expression is as follows:

[0112]

[0113] where θ is the threshold, which is set to 0.5 in this embodiment. Therefore, when the smoothed prediction probability is greater than 0.5, the binary label is 1, indicating that there is a hypopnea or apnea or other designated abnormal breathing event in the corresponding audio segment, otherwise it is 0, indicating that there is no designated abnormal breathing event;

[0114] S34) Combine consecutive abnormal audio segments into new abnormal audio segments, and count the number of all abnormal audio segments longer than the designated duration to obtain the number of designated abnormal breathing events;

[0115] As Figure 3As shown in the figure, in this embodiment, after merging abnormal audio segments with a binary label of 1 and continuous in time, abnormal events with too short a duration are filtered out to avoid misjudgment caused by noise, and the expression is as follows:

[0116] Keep event if(t end -t start )≥T min

[0117] For the remaining abnormal events, continuous abnormal segments are merged into one event, and its center point is located. Finally, a list of abnormal centers is obtained, and the expression is as follows:

[0118]

[0119] Among them, represents the center point of the i-th abnormal event;

[0120] S35) Calculate the corresponding AHI index according to the number of specified abnormal breathing events, and match the AHI index calculation result with the sleep apnea standard to obtain the severity of the specified abnormal breathing events.

[0121] The calculation formula of the AHI index is as follows:

[0122]

[0123] Among them, N events represents the number of all specified abnormal breathing events. In this embodiment, it is the statistical value of the number of center points of all specified abnormal breathing events, denoted as T ttotal represents the number of all specified abnormal breathing events and all audio segments marked as normal breathing events (that is, the smoothed prediction probability is less than the threshold or less than the specified duration);

[0124] When predicting the severity, match the AHI index with the following formula:

[0125]

[0126] S36) Calculate the prediction confidence interval to determine the reliability of the severity evaluation result of the specified abnormal breathing event. The formula is as follows:

[0127] P severe =σ(2.5·AHI audio -40)

[0128] Among them, AHI audio is the AHI index, and σ is the Sigmoid function. If the calculated prediction confidence interval P SEVEREIf it is greater than the specified value, it is considered that the severity prediction result of the specified abnormal breathing event is reliable, and the prediction result and the corresponding information are recorded. Otherwise, the prediction result is discarded or modified to a normal event.

[0129] Through the above steps, combined with the audio data recorded by the user throughout the night, a personalized monitoring report can be generated and fed back to the user in real time. In this embodiment, the dataset of sleep recordings is used for model training, and new sleep recordings are used for model verification.

[0130] During the process of model training, after the audio signal features are extracted and the multi-modal features are fused, and both the tracheal sound and the ambient sound are segmented into audio segments, data marking is also performed, specifically including:

[0131] If the multi-task classification is a binary classification task, each breathing event in the audio segment is labeled as one or more of normal or abnormal.

[0132] If the multi-task classification is a three-classification task, each breathing event in the audio segment is labeled as one or more of normal or low breathing or apnea.

[0133] If the multi-task classification is a five-classification task, each breathing event in the audio segment is labeled as one or more of normal or low breathing or obstructive apnea or central apnea or mixed apnea.

[0134] Correspondingly, during the model training process, in the sleep apnea syndrome (SAS) audio detection stage based on the enhanced branch preprocessor, specifically, the enhanced branch preprocessor is trained using the labeled multi-modal audio signals. Through multiple iterations of training, the model parameters of the enhanced branch preprocessor are continuously adjusted and optimized, and finally, the prediction probabilities of the trained enhanced branch preprocessor for each breathing event of the multi-classification task meet the requirements.

[0135] In the process of model verification, the same steps as steps S1 to S3 are executed. The final results show that the method of this embodiment is superior to other latest methods in both qualitative and quantitative evaluations, and has better practicability and effectiveness.

[0136] Embodiment 2

[0137] This embodiment proposes a sleep apnea detection system based on multi-modal audio signals, including a microprocessor and a computer storage medium. The microprocessor executes the computer program in the computer storage medium to implement the sleep apnea detection method based on multi-modal audio signals described in Embodiment 1.

[0138] This embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a microprocessor, implements the steps of the method for detecting sleep apnea based on multimodal audio signals according to Embodiment 1.

[0139] This embodiment also provides a computer program product including a computer program, which, when executed by a microprocessor, implements the steps of the method for detecting sleep apnea based on multimodal audio signals according to Embodiment 1.

[0140] In summary, the present invention provides a method and system for detecting sleep apnea based on multimodal audio signals. First, multimodal analysis combining tracheal sound and ambient sound is adopted, and then a deep relationship between audio signals and sleep apnea events is established based on an enhanced branch predictor neural network, significantly improving the accuracy and sensitivity of detection and making the monitoring of sleep apnea more reliable. By post-processing to integrate segment-level prediction sequences and giving the SAS detection and analysis results, and finally by analyzing and sorting out SAS, the problem of strong dependence on medical devices in traditional methods is overcome, providing an economical and convenient home monitoring solution and promoting the health management of users.

[0141] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for detecting sleep apnea based on multimodal audio signals, characterized in that, It includes the following steps: Obtain the tracheal sound and ambient sound of the sleep recording, segment both the tracheal sound and the ambient sound into audio segments, extract tracheal sound features and ambient sound features for the tracheal sound and ambient sound of each audio segment respectively, calculate the corresponding tracheal sound weight and ambient sound weight, and perform weighted fusion on the tracheal sound features and ambient sound features of each audio segment to obtain the corresponding multimodal audio features; Use an enhanced branch pre-classifier to detect abnormal breathing events for each multimodal audio feature respectively, and evaluate the severity of sleep apnea according to the number of detected abnormal breathing events.

2. The sleep apnea detection method based on multi-modal audio signals according to claim 1, wherein When extracting tracheal sound features and ambient sound features for the tracheal sound and ambient sound of each audio segment respectively, it specifically includes: Use a trained speech feature extractor to extract frame-level tracheal sound speech features for the current audio segment of the tracheal sound, and perform average pooling on the frame-level tracheal sound speech features along the time axis of the current audio segment to obtain tracheal sound features; Use a trained speech feature extractor to extract frame-level ambient sound speech features for the current audio segment of the ambient sound, and perform average pooling on the frame-level ambient sound speech features along the time axis of the current audio segment to obtain ambient sound features.

3. The sleep apnea detection method based on multimodal audio signals according to claim 1, characterized in that When calculating the corresponding tracheal sound weight and ambient sound weight, it specifically includes: Input the tracheal sound features and ambient sound features into a learnable linear layer respectively to obtain the tracheal sound weight and ambient sound weight. The expression of the learnable linear layer is as follows: W = Softmax(MLP(F)) where MLP is a multi-layer perceptron and F is the tracheal sound feature or ambient sound feature.

4. The sleep apnea detection method based on multimodal audio signals according to claim 1, characterized in that, Each layer of the enhanced branch pre-classifier includes a convolutional branch and a self-attention branch. When using the enhanced branch pre-classifier to detect abnormal breathing events for each multimodal audio feature respectively, it includes: Use the convolutional branch to capture the local time-frequency pattern of the current multimodal audio feature and extract local features. The expression is as follows: F conv = DWConv(ReLU(BN(Conv1d(F, kernel=3, stride=1))) where Conv1d represents one-dimensional convolution, the kernel size is 3, and the stride is 1; BN represents batch normalization; ReLU represents an activation function; DWConv represents depthwise separable convolution; F is the prediction result of the previous layer, initially the multimodal audio feature; Use the attention branch to calculate the global context dependence of the current multimodal audio feature. The expression is as follows: F attn = MultiHeadAttention(F, F, F) where MultiHeadAttention is a multi-head self-attention mechanism; MultiHeadAttention(F, F, F) means replacing the query matrix, key matrix, and value matrix under the multi-head attention mechanism with the prediction result F of the previous layer, initially the multimodal audio feature; Calculate the dynamic weight according to the output features of the convolutional branch and the attention branch, and then perform weighted fusion on the output features of the convolutional branch and the attention branch to obtain the prediction result. The expression is as follows: γ = σ(W g ·[F conv ; F attn ) F out = γ·F conv +(1 - γ)·F attn Among them, γ represents the dynamic weight, and F out represents the prediction result corresponding to the current multi-modal audio feature, and W g is a learnable parameter matrix, and σ is the Sigmoid function; Perform multi-task classification on the prediction result to judge the probabilities of different breathing events corresponding to the current multimodal audio feature. Different breathing events include normal breathing events and at least one abnormal breathing event. The expression is as follows: Among them, P represents the probability distribution of each type of breathing event of the current multi-modal audio feature, Softmax represents the activation function, and FC represents the fully connected layer with the same dimension as the number of categories of breathing events. represents the prediction result of the last layer of the enhanced branch preprocessor, and L represents the number of layers of the enhanced branch preprocessor.

5. The sleep apnea detection method based on multimodal audio signals according to claim 4, wherein After segmenting both the tracheal sound and the ambient sound into audio segments, it also includes the step of data labeling in the model training stage, which specifically includes: If the multi-task classification is a binary classification task, label the breathing events in each audio segment as one or more of normal or abnormal; If the multi-task classification is a ternary classification task, label the breathing events in each audio segment as one or more of normal or low breathing or apnea; If the multi-task classification is a quinary classification task, label the breathing events in each audio segment as one or more of normal or low breathing or obstructive apnea or central apnea or mixed apnea.

6. The sleep apnea detection method based on multi-modal audio signals according to claim 4, wherein When evaluating the severity of sleep apnea according to the number of detected abnormal breathing events, it includes: Specify one abnormal breathing event from all abnormal breathing events; According to the specified sliding window size, calculate the average value of the predicted probabilities of the specified abnormal breathing events of the multi-modal audio features in all audio segments within the sliding window in sequence. The expression is as follows: where W is the size of the sliding window, is the smoothed predicted probability of the specified abnormal breathing event of the multimodal audio feature of the i-th audio segment, p j is the predicted probability of the specified abnormal breathing event of the multimodal audio feature of the j-th audio segment within the sliding window of [i - W / 2, i + W / 2]; Compare the smoothed predicted probability of the specified abnormal breathing event of the multi-modal audio features of each audio segment with the specified threshold. If the smoothed predicted probability is greater than the specified threshold, the corresponding audio segment is an abnormal audio segment of the specified abnormal breathing event; Merge consecutive abnormal audio segments into new abnormal audio segments, and count the number of all abnormal audio segments longer than the specified duration to obtain the number of the specified abnormal breathing events; Calculate the corresponding AHI index according to the number of the specified abnormal breathing events, and match the AHI index calculation result with the sleep apnea standard to obtain the severity of the specified abnormal breathing event.

7. The sleep apnea detection method based on multimodal audio signals according to claim 6, characterized in that After obtaining the severity of the specified abnormal breathing event, it further includes: Calculate the prediction confidence interval to determine the reliability of the severity evaluation result of the specified abnormal breathing event. The formula is as follows: P severe = σ(2.5·AHI audio - 40) where AHI audio is the AHI index and σ is the Sigmoid function.

8. A sleep apnea detection system based on multimodal audio signals, characterized in that, It includes a microprocessor and a computer storage medium, and the microprocessor executes the computer program in the computer storage medium to implement the sleep apnea detection method based on multi-modal audio signals according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the microprocessor, the steps of the sleep apnea detection method based on multi-modal audio signals according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the microprocessor, the steps of the sleep apnea detection method based on multi-modal audio signals according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Voice diagnosis calling system and method based on deep learning, and storage medium

    CN121122242A