Respiratory signal estimation method based on ECG and PPG signal time-frequency fusion analysis
By using time-frequency fusion analysis in respiratory frequency prediction, combined with multi-scale S4-UNet and spectrum enhancement layer and other technical means, the problems of limited prediction accuracy and poor generalization in the existing technology are solved, and more efficient and robust respiratory frequency prediction is achieved.
Patent Information
- Application Number
- CN202510056968.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems with limited prediction accuracy and difficult to guarantee generalization in breathing frequency prediction, mainly because the model is difficult to capture the complex nonlinear relationship between the ECG and PPG signals and the breathing frequency, and the timing model cannot effectively deal with the long sequence problems caused by high sampling rates, resulting in information loss and data offset.
A breath signal estimation method based on time-frequency fusion analysis of ECG and PPG signals is proposed. By constructing a TF-RR network, the signal is fused from the perspectives of time and frequency domain. The time domain module adopts a multi-scale S4-UNet network, and the frequency domain module uses the spectrum enhancement layer and the cross attention layer to capture the frequency domain features, and combines the time domain and frequency domain features to predict.
It improves the accuracy and generalization ability of respiratory rate prediction, avoids information loss and data offset, and enhances the robustness and prediction performance of the model.
Smart Images

Figure CN119989264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological signal processing, and in particular to a respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals. Background Art
[0002] Respiratory rate (RR) is a key indicator for diagnosing abnormal health status of the human body, and is as important as heart rate, blood pressure and body temperature. RR is widely used in different fields. In hospital emergencies and clinical assessments, abnormal RR values are often used to identify patients' potential diseases and indicate dangerous situations, such as sepsis, cardiac arrest, allergic reactions, and respiratory diseases such as pneumonia. In identifying high-risk patient groups for cardiopulmonary arrest, the relative change of RR is more obvious than physiological parameters such as blood pressure and pulse rate. In addition, in special fields such as firefighting and sports and fitness, RR values can also be used as a heat stress indicator, and a significant increase in respiratory rate is a clear indicator of respiratory alkalosis.
[0003] Clinically, the main methods of respiratory measurement are carbon dioxide exhalation graph and spirometry based on gas exchange, but most of these devices are expensive and cumbersome to operate, and their invasive monitoring methods will cause certain damage to patients. In recent years, non-contact respiratory rate monitoring methods based on radar, infrared thermal imaging, etc. have become more popular, but their measured values are easily interfered by external environmental factors. Predicting respiratory rate using physiological signals modulated by respiration can effectively overcome the above problems and has become a research hotspot in this field. Among them, electrocardiogram (ECG) and photoplethysmography (PPG) signals are significantly affected by respiratory modulation, specifically in terms of baseline drift (BW), amplitude modulation (AM) and frequency modulation (FM). At the same time, these two signals can be measured non-invasively, are widely used and easy to collect, and have become strong candidates for measuring respiratory information in various environments.
[0004] In recent years, the rapid development of artificial intelligence technology and the continuous improvement of computing resources have made deep learning methods widely used in the medical field. At the same time, more and more respiratory signal-related data are publicly accessible on the Internet, making it possible to predict RR using deep learning methods. Some scholars proposed the RespNet model based on the U-Net architecture for RR prediction, using PPG signals to predict RR, with an MSE of 0.262 on the CapnoBase dataset. Other scholars have proposed a prediction scheme for estimating RR from downsampled PPG signals using convolutional neural networks (CNN). However, the above two methods have limited accuracy. Although using samples with longer durations as input can alleviate this problem, the longer time series leads to a larger prediction overhead for RR. In addition, the above studies only use a single physiological signal for respiratory rate estimation, which often makes it difficult to get rid of the influence of motion artifacts and individual differences inherent in the signal.
[0005] Although there have been many promising works on RR prediction based on deep learning, the current RR prediction still faces the problems of limited prediction accuracy and difficulty in ensuring generalization. The reasons are:
[0006] 1) Most current models estimate RR directly from ECG or PPG. However, this method requires the model to capture the complex nonlinear relationship between the input signal and RR, and the model training is difficult.
[0007] 2) The sampling frequency of ECG and PPG is usually over 100Hz, or even over 500Hz. Since traditional time series models cannot handle the long sequence problems caused by high sampling rates, most studies use downsampling strategies to process physiological signals. However, this strategy will cause some information loss to physiological signals, such as changes in the peak position of the R wave. This may cause the input data to be unable to capture respiratory modulation in a timely and accurate manner, limiting the prediction accuracy of the model;
[0008] 3) The above-mentioned deep learning-based RR prediction method only focuses on the temporal characteristics of physiological signals, while ignoring the important features in the frequency domain, which may be one of the reasons for the limited generalization ability of the model on unseen patients. On the other hand, the time domain model cannot get rid of the data offset caused by factors such as data acquisition equipment, environmental conditions or individual sample characteristics, resulting in poor performance of the time domain model on unseen data sets. Summary of the invention
[0009] In view of the shortcomings of the prior art, the present invention provides a respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals. The present invention proposes a TF-RR network to perform fusion analysis on ECG and PPG signals from both the time domain and frequency domain, and first use the fusion information to predict and estimate the respiratory signal.
[0010] The technical means adopted by the present invention are as follows:
[0011] A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals comprises the following steps:
[0012] Obtain training ECG signals, PPG signals, and corresponding breathing signals used as labels to construct a model training data set;
[0013] Constructing a respiratory signal prediction model for predicting a corresponding respiratory signal according to an ECG signal and a PPG signal, wherein the respiratory signal prediction model comprises a time domain module and a frequency domain module, wherein the time domain module and the frequency domain module respectively process the ECG signal and the PPG signal after coding and mapping, and the outputs of the time domain module and the frequency domain module are added to obtain a predicted respiratory signal;
[0014] Training a respiratory signal prediction model based on the model training data set;
[0015] The ECG signal and PPG signal to be estimated are obtained, and after encoding and mapping, they are input into the trained respiratory signal prediction model to obtain the predicted respiratory signal.
[0016] Furthermore, the time domain module adopts a multi-scale S4-UNet network, which is constructed by replacing the Encoder and Decoder based on the CNN convolution block in U-Net with S4-DilatedEncoder and S4-DilatedDecoder;
[0017] The S4-DilatedEncoder includes an S4layer and four parallel dilated convolutions; the S4-DilatedDecoder fuses the outputs of the S4-DilatedEncoders at the corresponding levels and uses a deconvolution layer to obtain a time-domain breathing signal.
[0018] Furthermore, the S4layer is implemented based on a linear state space model, which maps a one-dimensional input signal to a high-dimensional potential state and then projects it to a one-dimensional output signal.
[0019] Furthermore, the frequency domain module includes a spectrum enhancement layer and a cross-attention layer; the spectrum enhancement layer is used to convert the input features into frequency domain features by Fourier transform, and after enhancing the frequency domain features, output time domain features through inverse Fourier transform; the cross-attention layer is used to capture respiratory signal related features using the cross-attention mechanism based on the query generated by the rhythmic signal.
[0020] Furthermore, the ECG signal and the PPG signal are coded and mapped, including:
[0021] Input the ECG signal and the PPG signal into the Embedding layer for encoding, construct the first input feature of the model, and input the first input feature into the time series feature extraction module and the frequency domain feature extraction module;
[0022] The rhythm signal is input into the Embedding layer for encoding, the second input feature of the model is constructed, and the second input feature is input into the frequency domain feature extraction module.
[0023] Furthermore, obtaining the ECG signal, PPG signal and corresponding breathing signal used as a label for training includes:
[0024] Perform quality assessment on the original ECG signal and PPG signal, filter out low-quality segments based on the quality assessment results, and obtain high-quality signal data;
[0025] The high-quality signal data is preprocessed by a FIR bandpass filter, and the preprocessed data is sliced to generate ECG signals and PPG signals for training.
[0026] Further, training a respiratory signal prediction model based on the model training data set includes:
[0027] Obtain ECG and PPG signals in the training data set, generate a training sample time series after encoding and fusion, input the training sample time series into a respiratory signal prediction model, and obtain the respiratory signal prediction value output by the respiratory signal prediction model;
[0028] Calculate the loss function based on the mean square error and negative Pearson similarity loss of the actual value of the respiratory signal in the training data set and the predicted value of the respiratory signal;
[0029] The network parameters of the respiratory signal prediction model are optimized based on the loss function.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] The present invention provides a respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals. Based on the TF-RR network, ECG and PPG signals are fused and analyzed from the perspectives of time domain and frequency domain, and the fusion information is used to predict and estimate the respiratory signal. In the time domain, TF-RR adopts a Unet network based on the state space model S4. The network uses the characteristics of S4 that are suitable for processing extremely long data to avoid information loss caused by downsampling operations. The frequency domain module uses the frequency domain space after Fourier transformation to capture frequency domain features, and cross-attention is performed with learnable simulation signals to improve the model prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0033] Figure 1 1 is the network architecture of the respiratory signal prediction model (TF-RR) in an embodiment of the present invention.
[0034] Figure 2 a is the structure of the S4-Dilated encoder in an embodiment of the present invention.
[0035] Figure 2 b is the structure of the S4-Dilated decoder in an embodiment of the present invention.
[0036] Figure 3 This is the spectrum enhancement layer architecture in the embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0038] The present invention provides a respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals, which specifically includes the following steps:
[0039] S1. Obtain ECG signals, PPG signals and corresponding breathing signals used as labels for training, and construct a model training data set.
[0040] Two public data sets BIDMC and Capnobase are used in this application as training and test data for the model. The BIDMC data set is a subset of the MIMIC-II waveform database, which contains 53 adult patients aged 19-90 from ICU (including 32 women) with physiological signal records, and each record lasts for 8 minutes. Each record contains ECG, PPG, impedance respiratory signals collected at 125Hz and physiological data such as heart rate and respiratory rate collected at 1Hz, as well as respiratory annotations manually marked by impedance respiratory signals in an independent manner by two experts. Capnobase contains 8 minutes of records of 42 patients (13 adults, 29 children and newborns), and all patient signal collections are from elective surgery and routine anesthesia, where respiratory signals are measured by carbon dioxide. The data set records the ECG, PPG signals and respiratory signals of each subject at sampling frequencies of 300Hz, 100Hz and 25Hz respectively. In the embodiment of the present application, the ECG, PPG and respiratory signals of the Capnobase data set are all resampled to 125Hz, which is consistent with the BIDMC data.
[0041] Furthermore, the signal quality is tested, and the data is filtered and selected based on the test results.
[0042] Usually, there is a lot of noise in ECG and PPG signals, including power frequency interference caused by AC power supply, myoelectric interference caused by muscle contraction, motion artifacts, etc. In this regard, the present application first uses a template matching algorithm to evaluate the quality of the original ECG and PPG signals and filter out low-quality segments. The template matching algorithm (C. Orphanidou, T. Bonnici, P. Charlton, D. Clifton, D. Vallance, and L. Tarassenko, "Signal Quality Indices for the Electrocardiogram and Photoplethysmogram: Derivation and Applications to Wireless Monitoring," IEEE J. Biomed. Health Inform., pp. 1-1, 2014, doi: 10.1109 / JBHI.2014.2338351.) automatically searches for regular segments (irregular waveforms caused by artifacts) in ECG and PPG signal segments, and evaluates signal quality in two stages. The first stage identifies the heartbeat time in each signal to detect whether there is an extreme heart rate (low quality); the second stage constructs a template with the average value of different windows and calculates the correlation between each window and the template. Segments with a correlation coefficient less than the preset threshold (ECG: 0.66, PPG: 0.86) are considered low-quality segments and will be removed. In the end, 71.9% and 86.7% of the data remained in the BIDMC and Capnobase datasets, respectively.
[0043] Table 1. Preprocessing status of BIDMC dataset and Capnobas dataset.
[0044]
[0045] Furthermore, the ECG and PPG data are preprocessed.
[0046] There is a lot of information that is not related to breathing in the original ECG and PPG signals of good quality. For this reason, this application uses a FIR bandpass filter to filter to extract breathing-related information from ECG and PPG signals, and the bandpass frequency range is selected from 0.1-0.6Hz. The physiological signal of each patient is divided into sample slices of size 16s (i.e., 2000 time points) using a sliding window method, and the sliding step is 90% of the window length. The sample slices are used as training samples for the model, and the sample labels are the respiratory signals at the corresponding time of the sample slices. The main reasons for adopting the above strategy include two aspects: 1) Overlapping windows can increase the number of samples for model training; 2) The main respiratory frequency range in the data set is concentrated between 12-30 times / minute, and the sample slices require at least 2 signal cycles to obtain sufficient signals to extract the respiratory frequency. In addition, since the true respiratory frequency changes over time, sample slices that are too long are difficult to reflect instantaneous changes in respiratory conditions.
[0047] S2. Construct a respiratory signal prediction model for predicting a corresponding respiratory signal based on an ECG signal and a PPG signal. The respiratory signal prediction model includes a timing feature extraction module and a frequency domain feature extraction module. The timing feature extraction module and the frequency domain feature extraction module respectively process the ECG signal and the PPG signal after encoding and mapping, and add the outputs of the timing feature extraction module and the frequency domain feature extraction module to obtain a predicted respiratory signal.
[0048] This application proposes a deep learning model TF-RR model that integrates the time domain and frequency domain features of ECG and PPG to predict the corresponding respiratory signal. The TF-RR model architecture is as follows Figure 1 As shown in Figure 1, it consists of two main components: a time domain module based on multi-scale S4-UNet and a frequency domain module based on cross-attention. The two modules extract features in the time domain and frequency domain respectively. i (including ECG and PPG signals) are first encoded through the Embedding layer to project the original signal input into a high-dimensional feature E i =Embed(P i ), Subsequently, the eigenvectors are i Then they are processed separately in the time domain and frequency domain modules. The rhythm signal (Seasonalinit, parameters can be learned) is also the input of TF-RR, which will also pass through the encoder and output to the frequency domain module. Finally, the outputs of the time domain and frequency domain modules are added to get the final prediction result.
[0049] Furthermore, the time domain module based on the multi-scale S4-UNet is composed of a U-Net network based on S4, which mainly extracts features from different scales in the time domain. The multi-scale S4-UNet is formed by replacing the Encoder and Decoder based on the CNN convolution block in U-Net with those based on S4-DilatedEncoder and S4-DilatedDecoder. S4 is the core component of the time domain module and is suitable for capturing dependencies between long sequences. The multi-scale S4-UNet generally adopts a stacked structure of 2 layers of encoders / decoders, in which the input length of the upper encoder / decoder is half that of the lower layer, and the dimension is twice that of the lower layer, that is, the upper encoder / input sequence The output dimension of the lower layer encoder / decoder is (B, L / 2, 2D), where L is the sequence length and D is the feature dimension. Finally, the encoder downsamples the input twice to produce a compressed feature vector. The structure of the S4-Dilated encoder is as follows: Figure 2 As shown in a, it is implemented by S4layer and 4 parallel dilated convolutions. Among them, the convolution kernel size of the 4 dilated 1D convolutions is 3, the stride is 2, and the dilated convolution ratios are 1, 2, 4, and 8 respectively. Parallel dilated convolutions can capture the multi-scale feature information of ECG and PPG signals and perform batch normalization. Subsequently, the feature information is activated by Leak Relu (parameter is 0.3) and output after Concat. The S4-Dilated decoder is shown in Figure 2 As shown in b, the output of the corresponding layer encoder is fused and the deconvolution layer is used to obtain the time domain breathing signal.
[0050] In the four parallel one-dimensional convolutions (a) and one-dimensional deconvolutions (b), the dilatedrates are 1, 2, 4, and 8, the convolution kernels are all 3, and the stride is 2. The encoder and decoder use the same batch normalization and activation function (Leak Relu with a parameter of 0.3).
[0051] The S4 layer in this application is based on the linear state space model (SSM) (A.Gu, K.Goel, and C.Ré, "Efficiently Modeling Long Sequences with Structured State Spaces." arXiv, Aug.05, 2022. Accessed: Mar.20, 2024. [Online]. Available: http: / / arxiv.org / abs / 2111.00396). Define the one-dimensional input signal u(t) in the continuous state space, SSM first maps it to a high-dimensional potential state x(t), and then projects it to a one-dimensional output signal y(t), as shown in equation (1a). In S4, SSM is applied to each channel independently.
[0052] x′(t)=Ax(t)+Bu(t)#(1a)
[0053] y(t)=Cx(t)+Du(t)#(1b)
[0054] Among them, A, B, C, and D are the training parameters of the model. Using the bilinear discretization method for discretization, we can get:
[0055] Then we have:
[0056]
[0057] here At this time, the model has two calculation methods: linear recursion (2) and global convolution (3).
[0058] Furthermore, the frequency domain module based on cross attention consists of two parts: spectrum enhancement and cross attention, with both input and output dimensions being It mainly extracts features from the signal frequency domain. Similar to the time domain module, the features obtained after the Embedding layer enter the Spectral Enhancement Layer, and then form K and V of the Cross Attention Layer through Feed Forward. The periodic signal forms Q of the Cross Attention Layer after Embedding. The Spectral Enhancement Block and the Cross Attention Block will be introduced separately below. Figure 3As shown): The input time domain features are transformed into frequency domain features by Fourier transform (length is L), enhanced and then transformed back to time domain features by inverse Fourier transform, and the output dimension is the same as the input. Specifically, Fourier transform is first performed on each channel to obtain frequency domain information. In this application, only a small number of low-frequency components (the first M′ (M′<<L / 2)) are selected for enhancement to reduce noise interference (due to the conjugate symmetry of Fourier transform, only the first L / 2 frequency components need to be considered).
[0059]
[0060] In Fourier space, each complex component can be represented by amplitude and phase. Then the amplitude and phase information of the low-frequency component are extracted. The amplitude information is enhanced by the SpectrumConv layer.
[0061]
[0062] Where j∈{1,…,M′}. Thus, the enhanced real and imaginary information is obtained and can form a complex component
[0063]
[0064] Finally, before performing the inverse Fourier transform The M′ to L / 2 complex components are padded with zeros.
[0065]
[0066] Furthermore, although the respiratory signal is modulated by ECG and PPG signals, it is very difficult to directly extract the respiratory signal from the ECG and PPG waveforms. For this reason, this application simulates a respiratory rhythm signal as a query (Q) and uses a cross-attention mechanism to capture the relevant features of the respiratory signal. The respiratory signal has a limited frequency range and changes periodically. The signal can be constructed using a standard cosine signal. Among them, f∈[0,1], the sampling frequency and signal length of S remain the same as the sample, which are 125Hz and 2000 respectively, and the signal frequency range is the same as the breathing frequency, which can be achieved in the range (0.1~0.6Hz).
[0067] S=0.1+0.5*f#(8)
[0068] After the signal passes through the spectrum enhancement layer, it is transformed into k and v in the cross attention mechanism through different FCs. The simulated breathing signal S passes through the Embedding layer as q q, k, v are first converted into frequency domain representation Q, K, V by Fourier transform At this time, their frequency domain ranges are the same. Then Q, K, and V only retain M′ frequency components in the frequency domain for cross attention. The cross-attention process is expressed as:
[0069]
[0070] in Represents the output of cross attention.
[0071] Then, the output C is completed in the frequency domain through Padding, and then converted into a time domain signal through inverse Fourier transform. The process is as shown in formula 10:
[0072] O freq =FFT -1 (Padding(C))#(10)
[0073] It is sent to the feed-forward layer for further processing. Finally, it is combined with the time domain module output The sum of the results is the model output
[0074] S3. Training a respiratory signal prediction model based on the model training data set.
[0075] This paper uses a dataset containing N samples The TF-RR model proposed in this paper is trained and tested. are the ECG and PPG time series of the i-th sample, and the length of the time series is L. is the true value of the respiratory signal of sample i at L time points. The goal of the respiratory signal prediction model is to obtain the ECG and PPG signals P of length L. i In the above example, the corresponding breathing signal is predicted
[0076] Respiratory waveform reconstruction is a regression task. The commonly used loss function for regression tasks is MSE, which is calculated as follows (11), where B is the batch size, L is the sample length, and y ij and They represent the true value and predicted value of the j-th sample at sampling point i respectively.
[0077]
[0078] MSE focuses on the local difference between each sampling point. However, for waveform prediction, the similarity between the predicted waveform and the true waveform also needs to be considered. To this end, we introduce the negative Pearson similarity loss function L pearson (such as equation (12)). and Represents the sample mean of the true value and the predicted value.
[0079]
[0080] Finally, the loss function guiding the model training is as shown in Equation 13, where α is taken as 0.4.
[0081]
[0082] During the specific training, the model was built in the PyTorchLighting environment, and the model training used the Nvidia GTX3060GPU. During the training, the network used randomly initialized weights, and the network parameters were optimized using the stochastic gradient descent method. The learning rate was set to 0.001, and the iterations were 100 times.
[0083] S4. Obtain the ECG signal and PPG signal to be estimated, perform encoding mapping and then input them into the trained respiratory signal prediction model, so as to obtain the predicted respiratory signal.
[0084] In an application example of the present invention, the method of the present application was rigorously tested across subjects and across datasets on two datasets, BIDMC and Capnobase, and the TF-RR model was compared with the current optimal respiratory signal prediction and respiratory rate prediction models, see Table 2 and Table 3. The results show that the TF-RR model combining time domain and frequency domain features has good generalization and robustness, the respiratory rate prediction accuracy exceeds the current optimal model, and the COS waveform similarity is 10% higher than the optimal model. The TF-RR model is expected to be clinically applied in the future.
[0085] Table 2. Comparison of prediction accuracy of TF-RR, RespNet, and CapNet on respiratory signals
[0086]
[0087] Table 3 Comparison of prediction accuracy of TR-RR and other methods on respiratory rate
[0088]
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals, characterized in that: The following steps are involved: Obtain training ECG signals, PPG signals, and corresponding breathing signals used as labels to construct a model training data set; Constructing a respiratory signal prediction model for predicting a corresponding respiratory signal according to an ECG signal and a PPG signal, wherein the respiratory signal prediction model comprises a time domain module and a frequency domain module, wherein the time domain module and the frequency domain module respectively process the ECG signal and the PPG signal after coding and mapping, and the outputs of the time domain module and the frequency domain module are added to obtain a predicted respiratory signal; Training a respiratory signal prediction model based on the model training data set; The ECG signal and PPG signal to be estimated are obtained, and after encoding and mapping, they are input into the trained respiratory signal prediction model to obtain the predicted respiratory signal.
2. A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 1, characterized in that: The time domain module uses a multi-scale S4-UNet network, which is constructed by replacing the Encoder and Decoder based on the CNN convolution block in U-Net with S4-DilatedEncoder and S4-DilatedDecoder; The S4-DilatedEncoder includes an S4layer and four parallel dilated convolutions; the S4-DilatedDecoder fuses the outputs of the S4-DilatedEncoders at the corresponding levels and uses a deconvolution layer to obtain a time-domain breathing signal.
3. A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 2, characterized in that: The S4layer is implemented based on a linear state space model, which maps a one-dimensional input signal to a high-dimensional potential state and then projects it to a one-dimensional output signal.
4. The respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 1, characterized in that: The frequency domain module includes a spectrum enhancement layer and a cross-attention layer; the spectrum enhancement layer is used to convert the input features into frequency domain features by Fourier transform, and after enhancing the frequency domain features, the time domain features are output by inverse Fourier transform; the cross-attention layer is used to capture the relevant features of the respiratory signal using the cross-attention mechanism according to the query generated by the rhythmic signal.
5. The respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 1, characterized in that: Encode and map ECG signals and PPG signals, including: Input the ECG signal and the PPG signal into the Embedding layer for encoding, construct the first input feature of the model, and input the first input feature into the time series feature extraction module and the frequency domain feature extraction module; The rhythm signal is input into the Embedding layer for encoding, the second input feature of the model is constructed, and the second input feature is input into the frequency domain feature extraction module.
6. A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 1, characterized in that: Obtain ECG signals, PPG signals for training, and corresponding breathing signals used as labels, including: Perform quality assessment on the original ECG signal and PPG signal, filter out low-quality segments based on the quality assessment results, and obtain high-quality signal data; The high-quality signal data is preprocessed by a FIR bandpass filter, and the preprocessed data is sliced to generate ECG signals and PPG signals for training.
7. A respiratory signal estimation method based on time-frequency fusion analysis of ECG and PPG signals according to claim 1, characterized in that: Training a respiratory signal prediction model based on the model training data set includes: Obtain ECG and PPG signals in the training data set, generate a training sample time series after encoding and fusion, input the training sample time series into a respiratory signal prediction model, and obtain the respiratory signal prediction value output by the respiratory signal prediction model; Calculate the loss function based on the mean square error and negative Pearson similarity loss of the actual value of the respiratory signal in the training data set and the predicted value of the respiratory signal; The network parameters of the respiratory signal prediction model are optimized based on the loss function.
Citation Information
Cited By
Sleep apnea monitoring method and device, electronic monitoring equipment and storage medium
CN120284214A
Real-time pressure monitoring and intervention method and system based on deep learning
CN120581214A