Abnormal breathing event detection algorithm based on bidirectional long short-term memory network and attention mechanism

By using a bidirectional long and short-term memory network (BiLSTM) and attention mechanism detection algorithm in the detection of respiratory abnormal events, the problem of detecting apnea and hypoventilation abnormal events in the prior art is solved, and high-precision and low-cost detection effects are achieved.

CN120067789AInactive Publication Date: 2025-05-30SUZHOU RONGSHENG TECHNOLOGY INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104790.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently detect apnea and hypoventilation abnormal events, and the traditional time series prediction model performs poorly when processing nonlinear data and is not suitable for widespread applications in home environments.

Method used

The respiratory abnormal event detection algorithm based on bidirectional long and short-term memory network (BiLSTM) and attention mechanism is used to capture complex nonlinear relationships in the time series through the BiLSTM model, and the performance of the model is improved in combination with attention mechanism.

Benefits of technology

High-precision detection of respiratory abnormal events is achieved, which significantly improves the adaptability and prediction accuracy to nonlinear data, is suitable for applications in home environments, and reduces detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067789A_ABST
    Figure CN120067789A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal breathing event detection algorithm based on a bidirectional long-short-term memory network and an attention mechanism, and the algorithm comprises the steps: data description: in order to guarantee the actual applicability of data of a test set, extracting 672 pieces of data from the tail of each label for testing, in order to realize the balanced distribution of three event types, a traditional time sequence model generally assumes that the behavior of a time sequence is linear, for example, an autoregression model AR assumes that a current value can be represented by a linear combination of past values; the nonlinear relation of many complex real world data (such as heart rate and blood oxygen data of a subject) cannot be effectively captured and modeled, while the Bi LSTM can capture the complex nonlinear relation in time sequence data through a nonlinear activation function and a gating mechanism of an LSTM unit; and the adaptability and prediction accuracy of the model to real world data are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of detection algorithms, and more specifically, to a respiratory abnormality event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism. Background Art

[0002] Apnea and hypopnea are common sleep disorders, characterized by respiratory interruptions or attenuations lasting more than ten seconds during sleep, leading to intermittent hypoxemia and disruption of the sleep structure. Currently, diagnosing such syndromes requires patients to be monitored using polysomnography equipment in medical institutions, which causes significant physical and mental interference to patients and is not conducive to widespread use in a home environment. Therefore, developing an alternative apnea and hypopnea detection technology that is both comfortable and convenient has important practical significance.

[0003] Traditional time series prediction models (such as autoregressive model AR, moving average model MA, ARIMA, etc.) perform poorly in dealing with non-linear data, and due to their static parameters, dependence on short-term data, complex models, and high data requirements, these models are not suitable for detecting apnea and hypopnea abnormal events. Therefore, it is particularly important to explore and design a time series model that can accurately detect abnormal events based on existing data. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] In view of the problems existing in the prior art, the present invention provides a respiratory abnormality event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism to solve the technical problems mentioned in the background art.

[0006] (II) Technical Solutions

[0007] To achieve the above object, the present invention provides the following technical solution: A respiratory abnormality event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism, comprising the following steps:

[0008] Step 1: Data Description:

[0009] The data set contains data from 60 subjects, with a dimension of 60×2×180, where the first dimension represents the total number of subjects, the second dimension represents the number of channels (blood oxygen and heart rate), and the third dimension represents the temporal characteristics of 60 seconds. The label dimension of the training data is (n×1), and the label values represent no event (0), apnea (1), and hypopnea event (2) respectively. To ensure the practical applicability of the test set data, 672 data are extracted from the end of each label for testing to achieve an even distribution of the three event types;

[0010] Step 2: Modeling step: Modeling with a Bidirectional Long Short-Term Memory (BiLSTM) network. The modeling with the BiLSTM network includes the following steps: data preparation, constructing the BiLSTM layer, introducing the Attention mechanism, constructing the output layer, model compilation, adding a Dropout layer, automatically optimizing hyperparameters using Optuna, model training, and model evaluation and result output.

[0011] The present invention is further configured such that, in the data preparation step of Step 2

[0012] It includes loading, normalizing the data, uniquely encoding the labels, and splitting the dataset. Since the training data comes from 60 different subjects, to make the model training more stable, the team normalized 37,459 pieces of data. The normalized data needs to call the transpose method in the Numpy library to transform the data shape, swapping the positions of the number of channels and the number of features to adapt to the input format of the BiLSTM. The team uniquely encoded the labels, which not only converted the categories into independent and mutually exclusive vectors, eliminating possible numerical biases, but also transformed the categorical labels into a form that the BiLSTM model can understand and process. The proportion of the initial validation set was set to 0.1 in the experiment.

[0013] The present invention is further configured such that, in the constructing the BiLSTM layer step of Step 2

[0014] Call the Bidirectional wrapper in Keras.layers to make the entire incoming LSTM layer bidirectional, so as to capture the forward and backward dependencies of the training data. The key hyperparameter set is return_sequences = True, which makes the BiLSTM model output a result at each time step, providing continuity for subsequent stacking of multiple BiLSTM layers or the introduced Attention mechanism.

[0015] The present invention is further configured such that, in the introducing the Attention mechanism step of Step 2

[0016] To further improve the performance of the model, the team applied the attention mechanism after the BiLSTM layer to improve the model's performance. By applying the attention mechanism, the correlation matrix of the input sequence itself is calculated to identify the most important part of the sequence. The Attention layer serves as a bridge between the BiLSTM layer and the fully connected layer, using the Flatten method of Keras to flatten the output data of the BiLSTM layer into a one-dimensional vector for easy reception by the fully connected layer.

[0017] The present invention is further configured such that, in the constructing the output layer step of Step 2

[0018] The team adds a fully connected layer after the Attention layer as the output layer for classifying events. To ensure that the output dimension is consistent with the number of categories to be predicted, the first hyperparameter should be set to 3, and the softmax function is specified as the activation function to output the probability distribution of each category. The softmax activation function ensures that the sum of the three output values is 1, thus forming a valid probability distribution.

[0019] The present invention is further configured such that, in step two, the model is compiled

[0020] In the compilation stage, the optimizer, loss function, and evaluation metrics are defined. The model uses the Adam optimizer, which combines the advantages of the Momentum method and the RMSProp optimizer. It is a computationally efficient and memory - efficient optimization algorithm. At the same time, the model selects categorical cross - entropy as the loss function to match the multi - classification task of the model. The model selects the accuracy on the validation set as the evaluation metric for the effect.

[0021] The present invention is further configured such that, in step two, a Dropout layer is added

[0022] The regularization technique Dropout is introduced to prevent the BiLSTM model from being too deep and overfitting. To improve the performance and generalization ability of the model, the experiment sets to randomly discard 1 / 5 of the neurons for each bidirectional BiLSTM layer and DenseLayer.

[0023] The present invention is further configured such that, in step two, Optuna is used to automatically optimize the hyperparameters

[0024] Traditional hyperparameter tuning methods such as Grid Search or Random Search are inefficient in high - dimensional spaces. Optuna uses the Tree - structured Parzen Estimator (TPE) algorithm to find excellent hyperparameter combinations. The flexibility of Optuna allows the team to automatically search for the best hyperparameters by defining the function objective. Thus, the performance of the model is improved. In the experiment, the team takes the classification accuracy of the model in an unfamiliar data background as the custom function objective. Optuna will continuously optimize the hyperparameters of the model according to this function objective. In addition, the team introduces the adjustment mechanism of Optuna in aspects such as the number of BiLSTM and DenseLayer layers, the number of neurons in each LSTM and Dense layer, the dropout ratio of BiLSTM and Dense layer, the learning rate, and the batch size.

[0025] The present invention is further configured such that, in step two, model training

[0026] During the model training process, the team introduced strategies such as a learning rate scheduler, early stopping mechanism, and adding class weights to optimize the model performance and prevent overfitting. The initial number of training epochs was 100, and the batch size was 32. After each round of training, the validation set was used to evaluate the model performance. To address the problem of imbalanced label classes in the training data, the team called the compute_class_weight method in sklearn.utils to automatically adjust the weights of each class, making the weights of the minority classes higher and the weights of the majority classes lower to achieve balance.

[0027] The present invention is further configured such that, in step two, model evaluation and result output

[0028] After training is completed, the validation set is used to evaluate the model, and the loss and accuracy of the validation set are returned. The custom predict function is used to predict the test data to obtain the predicted labels, which are compared with the correct labels to obtain the accuracy.

[0029] (III) Advantageous effects

[0030] Compared with the prior art, the present invention provides a respiratory abnormality event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism, having the following advantageous effects:

[0031] I. High performance

[0032] 1. The model has the ability to handle non-linear relationships.

[0033] Traditional time series models usually assume that the behavior of time series is linear. For example, the autoregressive model AR assumes that the current value can be represented by a linear combination of past values. This cannot effectively capture and model the non-linear relationships of many complex real-world data (such as the heart rate and blood oxygen data of subjects, etc.). While BiLSTM can capture the complex non-linear relationships in time series data through the non-linear activation function and gating mechanism of the LSTM unit, significantly improving the adaptability and prediction accuracy of the model to real-world data.

[0034] The model has strong bidirectional information processing ability.

[0035] Compared with traditional unidirectional time series models (such as AR, MA, ARIMA), BiLSTM can utilize its bidirectional structure to process the information flow from the past to the future and from the future to the past simultaneously. This bidirectional information flow enables the model to integrate the context information before and after at each time point. Especially for medical events such as apnea and hypopnea, BiLSTM can more accurately capture the relevant features, thereby improving the recognition and classification accuracy of the events.

[0036] The model can capture both long-term and short-term dependencies simultaneously

[0037] BiLSTM is superior to traditional models and basic RNN structures because its design can effectively capture long-term and short-term dependencies in time series through forward and backward LSTM layers. This feature enables BiLSTM to maintain performance in the face of the vanishing gradient problem, thus better handling long time series data.

[0038] 4. The model can handle non-stationary data well

[0039] Different from traditional ARIMA models that require data stationarity, BiLSTM can adapt to the non-stationarity of time series. Its gating mechanism allows the model to perform effective modeling and prediction even when the statistical characteristics of the data change over time.

[0040] II. Low cost

[0041] 1. The model has a moderate complexity and low inference cost.

[0042] Compared with some complex deep learning models, the structure of the BiLSTM model is relatively simple and the number of parameters is moderate. On a relatively small-scale dataset, BiLSTM can complete training in a shorter time. Although the bidirectional structure requires additional calculations, due to its ability to better capture temporal information, it usually does not require too many layers or overly complex structures, resulting in a relatively low overall computational overhead.

[0043] 2. The model's ecological technology is mature, and the development and maintenance costs are low.

[0044] BiLSTM has rich implementation versions and tool support in the open source community. Existing deep learning frameworks (such as TensorFlow, PyTorch, Keras, etc.) all provide optimized implementations of LSTM and BiLSTM. This makes the development, debugging, and optimization processes of the model more efficient, saving development and maintenance costs.

[0045] III. Scalable

[0046] 1. The model has strong scalability and high flexibility.

[0047] BiLSTM can be easily combined with other advanced deep learning technologies (such as the attention mechanism) to further enhance its performance. Our team achieved automatic optimization of hyperparameters by introducing custom shapes for the model's input and output layers and integrating the Optuna framework, significantly improving the model's performance and adaptability.

[0048] IV. Easy to deploy

[0049] 1. The structural characteristics of BiLSTM enable it to run in an environment with low hardware requirements, such as embedded devices, mobile devices, or edge computing devices, which is very suitable for resource-constrained application scenarios. In addition, BiLSTM is optimized and supported on a variety of deep learning platforms and hardware accelerators, with fast inference speed, and can meet the application scenarios with high real-time requirements. Brief Description of the Drawings

[0050] Figure 1 It is the BiLSTM prediction flowchart in the present invention;

[0051] Figure 2 It is the TimesNet prediction flowchart in the present invention. Detailed Embodiment

[0052] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0053] It should be pointed out that unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0054] In the present invention, unless otherwise stated, the orientations such as "up, down" are usually in the directions shown in the drawings, or in the vertical, perpendicular or gravitational directions; similarly, for the convenience of understanding and description, "left, right" are usually in the left and right shown in the drawings; "inside, outside" refer to the inside and outside relative to the contour of each component itself, but the above orientation terms do not limit the present invention.

[0055] Please refer to Figure 1 - Figure 2 , the respiratory abnormality event detection algorithm based on the bidirectional long short-term memory network and the attention mechanism, includes the following steps:

[0056] Step 1: Data description:

[0057] The data set contains data from 60 subjects, with a dimension of 60×2×180, where the first dimension represents the total number of subjects, the second dimension represents the number of channels (blood oxygen and heart rate), the third dimension represents the temporal features of 60 seconds, the label dimension of the training data is (n×1), and the label values represent no event (0), apnea (1), and hypopnea event (2) respectively. To ensure the actual applicability of the test set data, 672 data are extracted from the end of each label for testing to achieve an even distribution of the three event types;

[0058] Step 2: Modeling Step: Modeling with the Bidirectional Long Short-Term Memory Network (BiLSTM) model. The BiLSTM model modeling includes the following steps: data preparation, constructing the BiLSTM layer, introducing the Attention mechanism, constructing the output layer, model compilation, adding the Dropout layer, automatically optimizing hyperparameters using Optuna, model training, and model evaluation and result output.

[0059] 1. Design a data standardization method.

[0060] In the data preparation stage, our team abandoned the traditional method of using StandardScaler for standardization and instead adopted a more detailed custom standardization strategy to achieve standard normal distribution processing of multi-dimensional data. This method not only speeds up the model's convergence rate but also significantly improves the overall performance of the model.

[0061] Combine the Attention mechanism with BiLSTM.

[0062] In the model construction stage, by customizing the structures of the input layer and output layer, our team innovatively combined the attention mechanism with the BiLSTM network. This combination enables the model to capture subtle patterns in the data more deeply and significantly improves the prediction accuracy.

[0063] Introduce a learning rate scheduler and regularization techniques.

[0064] In the model training stage, the Reduce on Plateau learning rate scheduler was introduced. This scheduler can automatically reduce the learning rate when the validation loss no longer decreases, optimizing the training process and accelerating the model's convergence. In addition, to prevent overfitting, the team adopted an early stopping strategy and dropout techniques, which effectively restricted the overfitting phenomenon and improved the model's stability and accuracy.

[0065] Call Optuna for hyperparameter optimization and customize the optimization objective

[0066] The team used the Optuna framework with the Tree-structured Parzen Estimator (TPE) algorithm for hyperparameter optimization. Compared with traditional grid search and random search methods, Optuna demonstrated superior performance. Through a large number of automated experiments, the team successfully found the optimal combination of hyperparameters. In addition, to better adapt to the actual task requirements, the team set the optimization objective as the accuracy of predicting abnormal event labels, rather than just focusing on the performance on the validation set, which further enhanced the model's adaptability and effectiveness for actual tasks.

[0067] The present invention mainly uses the BiLSTM model to detect abnormal events. However, the TimesNet model can also be used for similar technical purposes, although its effect is not as good as that of the BiLSTM model.

[0068] The modeling process of the Timesnet model is as follows

[0069] Step 1: Data standardization.

[0070] The data is processed separately according to two features: blood oxygen and heart rate. The StandardScaler library in Python is used to standardize these two features respectively, and then they are merged to prepare for subsequent processing steps.

[0071] Step 2: Data encoding.

[0072] Data encoding includes two parts: numerical encoding and positional encoding. Among them, numerical encoding uses a convolutional layer (Conv1d) to encode each time point of the data sequence, processing the data in the last dimension. Positional encoding uses sine and cosine functions to encode the positional information of the data, which not only solves the symmetry problem in the Transformer model but also gradually decays as the position gets farther away, increasing the dimension of the positional information.

[0073] Step 3: Time series modeling.

[0074] First, apply the Fourier transform (FFT) to convert the time series data to the frequency domain, observe the first k frequency points with the largest amplitudes, and use the periods of these frequency points (i.e., the reciprocals of the frequencies) to split the data of each period of the time series into n columns and arrange them side by side to form a two-dimensional image. On these two-dimensional data, use a convolutional layer with a 2D-kernel to extract features. For the two-dimensionalized data of each period, the features of different periods are merged through an adaptive aggregation method to obtain the output of the TimesBlock. Finally, layer normalization is performed on the output of the TimesBlock.

[0075] Step 4: Output classification results.

[0076] Use the GELU activation function to introduce non-linear processing to enhance the expression ability of the model, flatten the features, and finally map the flattened features to three categories (0: no event, 1: apnea, 2: hypopnea event) through a linear projection layer.

[0077] Although the TimesNet model provides a structured and innovative way for time series modeling and feature extraction, it still does not perform as well as the BiLSTM model in this application scenario. This is mainly due to the combined advantages of the bidirectional long short-term memory mechanism of BiLSTM and its ability to process non-linear time series data. In addition, the BiLSTM model shows higher efficiency and accuracy in dealing with long-term dependencies and complex time dynamics.

[0078]

[0079]

[0080] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An abnormal respiratory event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism, characterized by: The following steps are involved: Step 1: Data description: The dataset contains data from 60 subjects, with a dimension of 60×2×180, where the first dimension represents the total number of subjects, the second dimension represents the number of channels (blood oxygen and heart rate), and the third dimension represents the time series features of 60 seconds. The label dimension of the training data is (n×1), and the label values ​​represent no event (0), apnea (1), and hypopnea event (2). To ensure the practical applicability of the test set data, 672 data are extracted from the end of each label for testing to achieve a balanced distribution of the three event types. Step 2: Modeling step: BiLSTM modeling of bidirectional long short-term memory network. The modeling of the BiLSTM model of bidirectional long short-term memory network includes the following steps: data preparation, construction of BiLSTM layer, introduction of Attention mechanism, construction of output layer, model compilation, addition of Dropout layer, automatic optimization of hyperparameters using Optuna, model training, model evaluation and result output.

2. The abnormal respiratory event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 1 is characterized in that: Data preparation in step 2 Including loading, standardizing data, uniquely encoding labels and splitting data sets. Since the training data comes from 60 different subjects, in order to make the model training more stable, the team standardized 37,459 data. The standardized data needs to call the transpose method in the Numpy library to transform the data shape and exchange the positions of the number of channels and the number of features to adapt to the input format of BiLSTM. The team uniquely encodes the labels, converting the categories into independent and mutually exclusive vectors, eliminating possible numerical deviations, and converting the category labels into a form that the BiLSTM model can understand and process. The experiment sets the ratio of the initial validation set to 0.

1.

3. The abnormal respiratory event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 2 is characterized in that: In step 2, the BiLSTM layer is constructed Call the Bidirectional wrapper in Keras.layers to make the entire LSTM layer passed in bidirectional so that the forward and backward dependencies of the training data can be captured. The key hyperparameter set is return_sequences = True, which makes the BiLSTM model output a result at each time step, providing continuity for the subsequent stacking of multiple BiLSTM layers or the introduction of the Attention mechanism.

4. The abnormal breathing event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 3 is characterized in that: In step 2, the Attention mechanism is introduced To further improve the performance of the model, the team applied an attention mechanism after the BiLSTM layer to improve the performance of the model. The attention mechanism was applied to calculate the correlation matrix of the input sequence itself to identify the most important part of the sequence. The Attention layer served as a bridge between the BiLSTM layer and the fully connected layer. The Flatten method of Keras was used to flatten the output data of the BiLSTM layer into a one-dimensional vector for easy reception by the fully connected layer.

5. The abnormal respiratory event detection algorithm based on a bidirectional long short-term memory network and an attention mechanism according to any one of claims 2 to 4, characterized in that: In step 2, the output layer is constructed The team added a fully connected layer as the output layer after the Attention layer to classify events. To ensure that the output dimension is consistent with the number of categories to be predicted, the first hyperparameter should be set to 3, and softmax should be specified as the activation function to output the probability distribution of each category. The softmax activation function ensures that the sum of the three output values ​​is 1, thus forming a valid probability distribution.

6. The abnormal respiratory event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 5 is characterized in that: The model compilation in step 2 The optimizer, loss function, and evaluation index are defined in the compilation stage. The model uses the Adam optimizer, which combines the advantages of the momentum method (Momentum) and the RMSProp optimizer. It is an optimization algorithm with high computational efficiency and small memory usage. At the same time, the model selects categorical cross entropy (categorical_crossentropy) as the loss function to match the model's multi-classification tasks. The model selects the accuracy on the validation set as the evaluation index of the effect.

7. The abnormal respiratory event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 1 is characterized in that: In the step 2, Optuna is used to automatically optimize hyperparameters. Traditional hyperparameter tuning methods such as grid search or random search are inefficient in high-dimensional space. Optuna uses the tree structured Bayesian optimization (TPE) algorithm to find hyperparameter combinations with excellent performance. The flexibility of Optuna allows the team to automatically search for the best hyperparameters by defining function objectives, thereby improving the performance of the model. In the experiment, the team used the classification accuracy of the model in the context of unfamiliar data as a custom function objective. Optuna will continuously optimize the model's hyperparameters based on the secondary function objective. In addition, the team introduced Optuna's adjustment mechanism in terms of the number of BiLSTM and Dense layers, the number of neurons in each LSTM and Dense layer, the dropout ratio of BiLSTM and Dense layers, the learning rate, and the batch size.

8. The abnormal breathing event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 7 is characterized in that: Model training in step 2 During the model training process, the team introduced strategies such as learning rate scheduler, early stopping mechanism, and adding category weights to optimize model performance and prevent overfitting. The initial number of training rounds was 100 and the batch size was 32. After each round of training, the validation set was used to evaluate the model performance. In order to solve the problem of unbalanced label categories in the training data, the team called the compute_class_weight method in sklearn.utils to automatically adjust the weight of each category so that the weight of the minority class is higher and the weight of the majority class is lower to achieve balance.

9. The abnormal breathing event detection algorithm based on bidirectional long short-term memory network and attention mechanism according to claim 7 is characterized in that: Model evaluation and result output in step 2 After training is completed, the model is evaluated using the validation set, and the loss and accuracy of the validation set are returned. The custom predict function is used to predict the test data, obtain the predicted label, and compare it with the correct label to obtain the accuracy.