Heart rate and respiration rate detection method and device
Through the fusion of multimodal feature extraction and deep learning models, the problem of unreliable detection results in the prior art relying on a single modal signal is solved, and high reliability and accuracy of heart rate and respiration rate detection is achieved.
Patent Information
- Application Number
- CN202510471097.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-27
AI Technical Summary
When performing heart rate and respiration rate detection, the prior art relies too much on a single modal signal, making it difficult to guarantee the reliability of the detection results.
The multimodal feature extraction module and deep learning sub-model are used to extract waveform features, time-frequency features and super-resolution frequency features of the electrocardiogram signal, blood pressure signal, blood oxygen saturation signal and respiratory signal, and the feature fusion is performed through the deep learning sub-model to determine the heart rate and respiratory rate.
Through the fusion of multimodal feature extraction and deep learning models, the dependence on a single signal form is reduced, and the reliability and accuracy of heart rate and respiration rate detection is improved.
Smart Images

Figure CN120036747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrocardiogram signal detection, and in particular to a method and device for detecting heart rate and respiratory rate. Background Art
[0002] In the field of medical monitoring, heart rate and respiratory rate are core physiological indicators for assessing individual health status, and play a key role in patient health management, disease prevention, and the formulation of clinical treatment strategies. By monitoring heart rate and respiratory rate in real time through ECG, important physiological signals such as arrhythmia, tachycardia or bradycardia, and abnormal changes in respiratory rate can be detected in a timely manner. It can not only help patients detect potential health problems early, but also support doctors to evaluate treatment effects, adjust treatment plans in a timely manner, and ensure the best treatment effect.
[0003] However, many current monitoring technologies rely mainly on a single ECG signal to assess heart and respiratory rates, which may encounter a variety of challenges in practice, including myoelectric interference caused by noise, baseline drift, power line interference, and problems with equipment and sensors such as electrode detachment, electrode positioning errors, and equipment failures. These challenges affect the quality of the signal and the accuracy of frequency estimation. In addition, the diversity of physiological and pathological conditions and individual differences in ECG signals can also lead to false alarms, which not only affect the accuracy of respiratory and heart rate estimation, but may also reduce the efficiency of the monitoring system and increase the burden on medical staff. In other words, it is difficult to ensure the reliability of the test results by detecting heart rate and respiratory rate using a single modal signal.
[0004] Therefore, in the process of detecting heart rate and respiratory rate, the prior art has the problem of difficulty in ensuring the reliability of the detection results due to over-reliance on a single modal signal. Summary of the invention
[0005] In view of this, it is necessary to provide a heart rate and respiratory rate detection method and device to solve the problem that in the process of heart rate and respiratory rate detection in the prior art, it is difficult to ensure the reliability of the detection results due to over-reliance on a single modal signal.
[0006] In order to solve the above problems, in a first aspect, the present invention provides a heart rate and respiratory rate detection method, comprising: Construct a signal monitoring model, which includes a multimodal feature extraction module and a deep learning sub-model; The waveform features, time-frequency features and super-resolution frequency features of the initial signal are extracted based on the multimodal feature extraction module; Based on the deep learning sub-model, waveform features, time-frequency features and super-resolution frequency features are fused, and the heart rate and respiratory rate of the initial signal are determined according to the fused features. The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0007] In a possible implementation, the deep learning sub-model includes a convolution layer, a backbone network layer, an average pooling layer, and a fully connected layer; based on the deep learning sub-model, waveform features, time-frequency features, and super-resolution frequency features are fused, and the heart rate and respiratory rate of the initial signal are determined respectively according to the fused features, including: The time domain sub-features and frequency domain sub-features of waveform features, time-frequency features and super-resolution frequency features are extracted based on the convolution layer respectively; Based on the backbone network layer, global feature extraction is performed on the time domain sub-features and frequency domain sub-features to obtain the feature map of the initial signal; The feature map is transformed based on the average pooling layer to obtain a one-dimensional feature vector of the initial signal; The heart rate and breathing rate of the initial signal are output based on the one-dimensional feature vector based on the fully connected layer.
[0008] In a possible implementation, the convolution layer includes a one-dimensional convolution layer and a two-dimensional convolution layer; based on the convolution layer, the time domain sub-features and the frequency domain sub-features of the waveform feature, the time-frequency feature, and the super-resolution frequency feature are respectively extracted, including: The time domain sub-features of waveform features and super-resolution frequency features are extracted based on one-dimensional convolutional layers; The frequency domain sub-features of the time-frequency features are extracted based on the two-dimensional convolutional layer.
[0009] In a possible implementation, the fully connected layer includes a first fully connected layer and a second fully connected layer; outputting the heart rate and respiratory rate of the initial signal according to the one-dimensional feature vector based on the fully connected layer includes: The first fully connected layer performs feature dimensionality reduction and nonlinear transformation on the one-dimensional feature vector, and the second fully connected layer outputs the heart rate and respiratory rate of the initial signal according to the one-dimensional feature vector after feature dimensionality reduction and nonlinear transformation.
[0010] In a possible implementation, the deep learning sub-model further includes a compression excitation layer; in the process of performing global feature extraction on the time domain sub-features and the frequency domain sub-features based on the backbone network layer to obtain a feature map of the initial signal, it also includes: The residual block of the backbone network layer is revised, and the feature channel weights of the one-dimensional convolutional layer and the two-dimensional convolutional layer are calibrated according to the adaptive feedback of the compressed excitation layer. The number of input channels of the residual block is adjusted to be equal to the number of feature output channels of the multimodal feature extraction module.
[0011] In a possible implementation, waveform features, time-frequency features, and super-resolution frequency features are fused based on the deep learning sub-model, and the heart rate and respiratory rate of the initial signal are determined respectively according to the fused features, and the following is also included before: Constructing a training sample set, the training sample set includes an initial signal sample and a heart rate label sample and a respiratory rate label sample corresponding to each initial signal sample; Assigning a weight to each initial signal sample based on the type of the initial signal sample; Based on the multimodal feature extraction module, the waveform features, time-frequency features and super-resolution frequency features of the initial signal samples are extracted respectively, and input into the initial deep learning sub-model, and the heart rate label samples and respiratory rate label samples of the initial signal samples are output. After iterative training, the weighted mean square error loss values of the heart rate label samples and the respiratory rate label samples are calculated based on the weighted mean square error loss function combined with the sample weight assignment results, and the initial deep learning sub-model corresponding to the minimum weighted mean square error loss value is determined as the target deep learning sub-model.
[0012] In a possible implementation, constructing a training sample set further includes: A time window is set, and the timestamps of the initial signal samples, the heart rate label samples, and the respiratory rate label samples are aligned based on the time window.
[0013] In a possible implementation, constructing a training sample set further includes: Perform enhancement processing on the initial training sample set to obtain the target training sample set; The enhancement processing means include at least one of adding noise, randomly erasing data, randomly deleting a channel and setting the channel data to zero, data scaling, and local signal clipping and random insertion.
[0014] In a possible implementation, a training sample set is constructed, and then the following steps are further included: When a group of initial signal samples includes an electrocardiogram signal / respiration signal, the heart rate label sample and the respiration rate label sample corresponding to the electrocardiogram signal are used as the heart rate label sample and the respiration rate label sample of the other initial signal samples in the group; When a group of initial signal samples only includes a blood oxygen signal and a blood pressure signal, the heart rate label samples and the respiratory rate label samples corresponding to the blood oxygen signal are used as the heart rate label samples and the respiratory rate label samples of the other initial signal samples in the group.
[0015] In a second aspect, the present invention further provides a heart rate and respiratory rate detection device, comprising: A signal monitoring model building module is used to build a signal monitoring model, which includes a multimodal feature extraction module and a deep learning sub-model; A feature extraction module, used to extract waveform features, time-frequency features and super-resolution frequency features of the initial signal based on the multimodal feature extraction module; The heart rate and respiratory rate detection module is used to fuse waveform features, time-frequency features and super-resolution frequency features based on the deep learning sub-model, and determine the heart rate and respiratory rate of the initial signal respectively according to the fused features; The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0016] The beneficial effect of adopting the above-mentioned embodiment is as follows: the present invention provides a heart rate and respiratory rate detection method, which extracts the waveform characteristics, time-frequency characteristics and super-resolution frequency characteristics of the initial signal through a multimodal feature extraction module. Since the initial signal can exist in various forms and various combinations, the dependence on a certain specific form of the initial signal is reduced; further, since the waveform characteristics, time-frequency characteristics and super-resolution frequency characteristics exist in various forms, they can reflect the characteristics of the initial signal from multiple dimensions, thereby ensuring the reliability of the detection results of the heart rate and respiratory rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of a flow chart of an embodiment of a heart rate and respiratory rate detection method provided by the present invention; Figure 2 A schematic diagram of a flow chart of an embodiment of determining the heart rate and respiratory rate of an initial signal based on a deep learning sub-model provided by the present invention; Figure 3 A schematic diagram of a flow chart of an embodiment of training an initial deep learning sub-model provided by the present invention; Figure 4 A schematic structural diagram of an embodiment of a heart rate and respiratory rate detection device provided by the present invention; Figure 5 This is a structural block diagram of an embodiment of an electrocardiogram and blood pressure monitor provided by the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0019] It should be understood that the schematic drawings are not drawn to scale. The flowchart used in the present invention shows the operations implemented according to some embodiments of the present invention. It should be understood that the operations of the flowchart can be implemented out of order, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art, under the guidance of the content of the present invention, can add one or more other operations to the flowchart, and can also remove one or more operations from the flowchart. Some of the block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor systems and / or microcontroller systems.
[0020] The descriptions of "first", "second", etc. involved in the embodiments of the present invention are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined as "first" or "second" may explicitly or implicitly include at least one of the features.
[0021] Reference to an "embodiment" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0022] In order to solve the problem in the prior art that it is difficult to ensure the reliability of detection results due to over-reliance on a single modal signal during heart rate and respiratory rate detection, the present invention provides a heart rate and respiratory rate detection method and device, which are described in detail below.
[0023] like Figure 1 As shown, Figure 1 A flow chart of an embodiment of a heart rate and respiratory rate detection method provided by the present invention includes: S101: constructing a signal monitoring model, where the signal monitoring model includes a multimodal feature extraction module and a deep learning sub-model; In some embodiments of the present invention, the executor of the signal monitoring model can be any one of a high-performance GPU (graphics processing unit), a dedicated deep learning accelerator (such as Tensor Processing Units (TPUs), Neural Processing Units (NPUs)), an FPGA (field programmable gate array), an ASIC (application-specific integrated circuit), and an embedded system, which can be selected and adjusted according to actual needs.
[0024] The executor of the deep learning sub-model can also be a high-performance GPU (graphics processing unit), a dedicated deep learning accelerator (such as Tensor Processing Units (TPUs), Neural Processing Units (NPUs)), FPGA (field programmable gate array), ASIC (application-specific integrated circuit) and any one of the embedded systems, which can be selected and adjusted according to actual needs.
[0025] S102: extracting waveform features, time-frequency features, and super-resolution frequency features of the initial signal based on a multimodal feature extraction module; S103: Fusing waveform features, time-frequency features, and super-resolution frequency features based on the deep learning sub-model, and determining the heart rate and respiratory rate of the initial signal respectively according to the fused features; The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0026] The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0027] In some embodiments of the present invention, the initial signal refers to physiological data that can be directly collected by sensors, wherein the multimodal feature extraction module can obtain three features based on an initial signal. That is to say, even if there is only one signal among the electrocardiogram signal, blood pressure signal, blood oxygen saturation signal and respiratory signal, after feature extraction by the multimodal feature extraction module, the three forms of features corresponding to the signal can be obtained.
[0028] Electrocardiogram (ECG or EKG) is a bioelectric signal that records the electrical activity of the heart. The electrocardiogram used in the present invention is the change of the electrical signal generated by the heart during contraction and relaxation, which is captured by the electrocardiogram electrodes on the surface of the skin.
[0029] The respiratory signal is a physiological signal that reflects the human respiratory activity. It is used to evaluate respiratory function and health status by detecting the dynamic data of chest movement, airflow changes or gas exchange during breathing. The respiratory signal used in the present invention is to apply a high-frequency current to the human chest through an ECG electrode and measure the voltage change at both ends to detect the impedance change signal caused by breathing. The respiratory signal and the ECG signal in the present invention are collected synchronously and exist at the same time.
[0030] The blood pressure signal is a warning signal sent by the human body through physiological reactions or symptoms when the blood pressure is abnormal (such as hypertension), prompting the need to monitor blood pressure in a timely manner and take intervention measures. The blood pressure signal used in the present invention is the oscillometric method (pressure oscillation method), that is, by measuring the pulse wave signal generated during the inflation and deflation of the cuff.
[0031] Blood oxygen saturation signal refers to the oxygenated hemoglobin (HbO 2 ) is data or waveform of the percentage of total hemoglobin that can be combined. The blood oxygen signal used in the present invention is a PPG (photoplethysmogram) signal.
[0032] Waveform characteristics refer to the graphical representation of respiratory-related parameters (such as pressure, flow rate, volume) recorded by ventilators, monitors and other equipment over time or in relationship to each other, which are used to evaluate respiratory status, mechanical ventilation effects and abnormal conditions.
[0033] Time-frequency features refer to the joint representation of a signal in the time domain and frequency domain, which can simultaneously reflect the changes of a signal at different times and frequencies. Time-frequency features are of great significance in signal processing, especially when processing non-stationary signals, and can provide richer information.
[0034] Super-resolution frequency features refers to the process of recovering or generating high-resolution data from low-resolution data through algorithms or techniques in image processing. This process is mainly used in image reconstruction, video enhancement, satellite image processing, medical image analysis and other fields.
[0035] Heart rate refers to the number of heartbeats per minute in a normal person's quiet state, also called resting heart rate. It is generally 60 to 100 beats per minute, and may vary from person to person due to age, gender or other physiological factors.
[0036] Respiratory rate, also known as respiratory frequency, refers to the number of breaths you take in one minute.
[0037] In this embodiment, firstly, the waveform features, time-frequency features and super-resolution frequency features of the initial signal are extracted by a multimodal feature extraction module. Since the initial signal can exist in various forms and combinations, the dependence on a specific form of the initial signal is reduced. Furthermore, since the waveform features, time-frequency features and super-resolution frequency features exist in various forms, they can reflect the characteristics of the initial signal from multiple dimensions, thereby ensuring the reliability of the detection results of the heart rate and respiratory rate.
[0038] In some embodiments of the present invention, in S102, it should be understood that the multimodal feature extraction module in the embodiments of the present invention includes three data extraction functions, namely, the waveform characteristics, time-frequency characteristics and super-resolution frequency characteristics of any initial signal can be extracted.
[0039] In some embodiments of the present invention, in order to extract high-quality waveform features, time-frequency features and super-resolution frequency features from the original initial signal. These features not only capture the basic characteristics of the signal, but also record the dynamic changes of the signal over time in detail, providing strong support for the accurate monitoring of heart rate and respiratory rate. In particular, the application of the minimum variance distortionless response (MVDR) algorithm brings significant improvement in the accuracy of frequency estimation, greatly enhancing the application potential in the field of heart rate and respiratory detection.
[0040] Specifically, the step of extracting waveform features includes: In order to effectively reduce the amount of calculation while retaining enough signal information for subsequent analysis, the construction of waveform features adopts a downsampling strategy for ECG and respiratory signals. By taking the average of every 4 data points in the preprocessed ECG and respiratory signals as a sample, this simplification step not only significantly reduces the amount of data, but also retains key waveform information, providing rich features for the deep learning model.
[0041] The steps of extracting short-time Fourier transform (STFT) time-frequency features include: The application of STFT is to reveal the frequency characteristics of the signal over time. By analyzing the signal in a short time window, STFT can show the dynamic changes of the signal frequency over time. In order to capture the local characteristics of the signal, the STFT window size is set to half the length of the input signal, and a non-overlapping area of 100 data points is set between adjacent windows. In addition, the frequency resolution is improved through zero padding operation, which enables the model to more accurately identify and analyze the frequency components in the signal.
[0042] The steps of extracting super-resolution frequency features include: In view of the limitation of the frequency resolution of the traditional Fourier transform and its sensitivity to noise, the present invention adopts the MVDR algorithm to achieve super-resolution fine-grained frequency estimation. The MVDR algorithm is based on the principle of adaptive beamforming and estimates the frequency components of the signal by calculating the power spectrum density of the signal. The core of the algorithm is to calculate the inverse matrix of the autocovariance matrix R and use the steering vector The weights are assigned and the power spectrum density at the target frequency is accurately estimated. The formula is as follows:
[0043] in, Indicates the frequency The power spectral density estimate at It is a steering vector calculated based on the wave number of the signal and the position of the sensor array, and is related to the direction of arrival (or frequency) of the signal; It is the inverse matrix of the signal autocovariance matrix R, which describes the correlation between the components of the signal; is the conjugate transpose of the steering vector. The MVDR algorithm can provide accurate estimation of the heartbeat and respiratory frequency components, effectively improving the accuracy of frequency estimation, thereby enhancing the practicality and accuracy of the feature.
[0044] In some embodiments of the present invention, in S103, the deep learning sub-model includes a convolutional layer, a backbone network layer, an average pooling layer and a fully connected layer; in order to fuse the waveform features, the time-frequency features and the super-resolution frequency features based on the deep learning sub-model, and determine the heart rate and the respiratory rate of the initial signal respectively according to the fused features; Figure 2 As shown, Figure 2 A flow chart of an embodiment of determining the heart rate and respiratory rate of an initial signal based on a deep learning sub-model provided by the present invention includes: S201: extracting waveform features and time domain sub-features of super-resolution frequency features, and frequency domain sub-features of time-frequency features based on the convolution layer; S202: Performing global feature extraction on the time domain sub-features and the frequency domain sub-features based on the backbone network layer to obtain a feature map of the initial signal; In some embodiments of the present invention, the backbone network layer is responsible for extracting multi-level feature representations from time domain sub-features and frequency domain sub-features, gradually abstracting the original data into high-dimensional features, and realizing the conversion from low-level features to high-level semantics.
[0045] A feature map refers to a multidimensional array obtained by processing time domain sub-features and frequency domain sub-features through the backbone network layer, usually a three-dimensional structure (height, width, number of channels).
[0046] S203: transforming the feature map based on the average pooling layer to obtain a one-dimensional feature vector of the initial signal; In some embodiments of the present invention, the average pooling layer is a downsampling operation in a convolutional neural network (CNN). Its core function is to perform mean calculation on the local area of the input feature map, thereby compressing the spatial dimensions (height and width), retaining the overall feature information, and reducing the amount of calculation and the number of parameters.
[0047] S204: Output the heart rate and breathing rate of the initial signal according to the one-dimensional feature vector based on the fully connected layer.
[0048] In some embodiments of the present invention, the fully connected layer (FC Layer) is a basic structure in a neural network, and its core feature is that each neuron is connected to all neurons in the previous layer. It is usually located at the end of the network and is responsible for mapping high-level features to target outputs (such as classification probabilities or regression values).
[0049] Furthermore, since two results, heart rate and respiratory rate, are output, two fully connected layers (fc1 and fc2) are used to complete the further feature extraction and final regression task: the fc1 layer is responsible for implementing feature dimensionality reduction and nonlinear transformation to prepare data for the final prediction results; the fc2 layer outputs the final heart rate and respiratory rate prediction results.
[0050] In this embodiment, the time domain sub-features and frequency domain sub-features of the waveform features, time-frequency features and super-resolution frequency features are respectively extracted through a deep learning sub-model, and the heart rate and respiratory rate of the initial signal are determined based on the time domain sub-features and frequency domain sub-features, thereby realizing a relationship mapping between the feature set of the initial signal - waveform features, time-frequency features and super-resolution frequency features, and the heart rate and respiratory rate, so that for any form of the feature set of the initial signal, the corresponding heart rate and respiratory rate can be obtained quickly and accurately.
[0051] In some embodiments of the present invention, in S201, the convolution layer includes a one-dimensional convolution layer and a two-dimensional convolution layer; in order to respectively extract the time domain sub-features and frequency domain sub-features of the waveform features, time-frequency features and super-resolution frequency features based on the convolution layer, specifically, the time domain sub-features of the waveform features and the super-resolution frequency features are respectively extracted based on the one-dimensional convolution layer; and the frequency domain sub-features of the time-frequency features are extracted based on the two-dimensional convolution layer.
[0052] In a specific embodiment, the convolution layer is defined to include three sub-convolution layers (conv1, conv2, and conv3), which process different types of input features respectively. Among them, conv1 and conv2 are 1D convolution layers, which are specifically used to process waveform features and super-resolution frequency features, while conv3 is a 2D convolution layer for processing time-frequency features. Then, the convolution layer can simultaneously learn the time domain and frequency domain sub-features of the signal, thereby providing more comprehensive data analysis capabilities.
[0053] In some embodiments of the present invention, in S202, in order to ensure the stability of the global feature extraction of time domain sub-features and frequency domain sub-features by the backbone network layer, a compression excitation layer is also set in the deep learning sub-model. Specifically, the residual block of the backbone network layer is revised, and the feature channel weights of the one-dimensional convolutional layer and the two-dimensional convolutional layer are calibrated according to the adaptive feedback of the compression excitation layer, and the number of input channels of the residual block is adjusted to be equal to the number of feature output channels of the multimodal feature extraction module.
[0054] In this embodiment, the number of input channels of the residual block of the backbone network layer is used as a benchmark to feedback-adjust the number of feature output channels of the multimodal feature extraction module, so that the backbone network layer can ensure the alignment of data when performing feature extraction, avoid data disorder problems, improve the reliability and stability of the feature map, and thus ensure the detection effect of heart rate and respiratory rate; by calibrating the feature channel weights of the one-dimensional convolutional layer and the two-dimensional convolutional layer, the representation capabilities of the time domain sub-features and the frequency domain sub-features can be adaptively controlled, highlighting important features and suppressing unimportant features, thereby improving the performance of the deep learning sub-model.
[0055] In some embodiments of the present invention, in S103, before the heart rate and respiratory rate detection are performed by the deep learning sub-model, the initial deep learning sub-model needs to be trained, such as Figure 3 As shown, Figure 3 A schematic diagram of a flow chart of an embodiment of training an initial deep learning sub-model provided by the present invention includes: S301: construct a training sample set, where the training sample set includes an initial signal sample and a heart rate label sample and a respiratory rate label sample corresponding to each initial signal sample; In some embodiments of the present invention, since an initial signal sample may include signals in multiple forms, the corresponding heart rate label samples and respiratory rate label samples may have different data. That is to say, for the ECG signal samples and blood pressure signal samples from the same user at a specific moment, there may be a problem of inconsistency between the heart rate label results obtained based on the ECG signal samples and the heart rate label results obtained based on the blood pressure signal samples. Therefore, for each heart rate label sample and respiratory rate label sample, it is also necessary to verify their availability and reliability.
[0056] In order to simplify the inspection steps for the heart rate label samples and the respiratory rate label samples, the reliability of the electrocardiogram signal, the blood pressure signal, the blood oxygen saturation signal and the respiratory signal is reduced in sequence, specifically: When a group of initial signal samples includes an electrocardiogram signal / respiration signal, the heart rate label samples and the respiration rate label samples corresponding to the electrocardiogram signal / respiration signal are used as the heart rate label samples and the respiration rate label samples of other initial signal samples in the group; When a group of initial signal samples only includes a blood oxygen signal and a blood pressure signal, the heart rate label samples and the respiratory rate label samples corresponding to the blood oxygen signal are used as the heart rate label samples and the respiratory rate label samples of the other initial signal samples in the group.
[0057] When a group of initial signal samples only includes blood pressure signals, in order to ensure the reliability of the samples, the group of initial signal samples and their corresponding heart rate label samples and respiratory rate label samples are directly eliminated. Alternatively, the heart rate label samples and respiratory rate label samples corresponding to the blood pressure signals are directly used as target labels.
[0058] Furthermore, in order to ensure the correctness of the heart rate label samples and the respiratory rate label samples, the labeling results of the training sample set can be manually checked, and the labeling results can be secondarily checked by other methods, which are not limited here.
[0059] In this embodiment, by re-verifying the heart rate label samples and the respiratory rate label samples, the reliability of the training sample set is effectively guaranteed, thereby ensuring the training effect of the subsequent deep learning sub-model.
[0060] S302: assigning a weight to each initial signal sample based on the type of the initial signal sample; In some embodiments of the present invention, first, the signal quality of the electrocardiogram signal samples, blood pressure signal samples, blood oxygen saturation signal samples and respiratory signal samples is determined, and a confidence level is defined for the corresponding heart rate label samples and respiratory rate label samples on the signal samples with good signal quality.
[0061] Generally, the confidence level of the heart rate of the ECG signal is the highest, and the weight of the heart rate label sample of the ECG signal sample is assigned to 1; the confidence level of the heart rate of the blood oxygen saturation signal sample is second best, and the weight of the heart rate label sample of the blood oxygen saturation signal sample is assigned to 0.8; the confidence level of the heart rate of the blood pressure signal is relatively low, and the weight of the heart rate label sample of the blood pressure signal sample is assigned to 0.3. In other embodiments, the specific weight assignment can be adaptively adjusted according to actual needs, and is not limited here. In particular, for the initial signal sample with poor signal quality, its weight is directly assigned to 0 to avoid its influence on the model training effect.
[0062] S303: Based on the multimodal feature extraction module, the waveform features, time-frequency features and super-resolution frequency features of the initial signal samples are respectively extracted, and input into the initial deep learning sub-model, the heart rate label samples and the respiratory rate label samples of the initial signal samples are output, the training is iteratively performed, and the weighted mean square error loss values of the heart rate label samples and the respiratory rate label samples are calculated based on the weighted mean square error loss function combined with the sample weight assignment results, and the initial deep learning sub-model corresponding to the minimum weighted mean square error loss value is determined as the target deep learning sub-model.
[0063] In this embodiment, by assigning weights to the samples and calculating the weighted mean square error loss value of each sample based on the weighted mean square error loss function, the weight distribution of samples of different modalities to the weighted mean square error loss function is optimized, and the contribution of sample data of different modalities to the final estimation result is adaptively adjusted, thereby improving the optimization effect of the target deep learning sub-model.
[0064] In some embodiments of the present invention, in S301, in the process of constructing a training sample set, for initial signal samples that can be directly obtained, such as: electrocardiogram signal samples, blood pressure signal samples, blood oxygen saturation signal samples and respiratory signal samples, feature extraction is required to obtain corresponding label samples, and the electrocardiogram signal samples are used as an example for explanation: The input data for feature extraction of ECG signals is one or more lead ECG signals and respiratory signals. The specific steps of feature extraction are as follows: 1. Signal sampling rate setting: The signal sampling rate is set to 1000Hz, which means that the ECG signal contains 1000 data points per second. At the same time, the window size is set to include 10 seconds of data length, so that each data window can generate a set of features, which helps to capture sufficient dynamic information while maintaining time resolution.
[0065] 2. Time interval setting for label value reliability: In order to ensure the reliability of heartbeat and respiration label values, a 60-second time interval standard is set. When the difference between the average sampling time of the data window and the sampling time of the heartbeat and respiration rate labels is less than or equal to 60 seconds, these label values can be considered to be associated with the generated features, thereby ensuring the accuracy of the analysis.
[0066] 3. ECG signal segmentation: In order to improve processing efficiency and maintain the integrity of signal details, the continuous ECG signal is segmented into multiple sub-windows, each containing 5000 data points. The data points between adjacent sub-windows do not overlap, and each interval is 50 data points, thereby constructing a (101, 5000) two-dimensional matrix. Through the multiple sub-window strategy, a large amount of data can be processed efficiently while retaining the necessary signal details for feature extraction, preparing for the subsequent use of super-resolution algorithm processing.
[0067] 4. Steering vector generation: Set the heart rate range to 0-5 Hz, the breathing rate range to 0-2.5 Hz, and generate the steering vector The steering vector is the key to constructing super-resolution frequency features, which can accurately reflect the characteristics of the signal within a specific frequency range. This step is the basis for achieving accurate frequency estimation, especially when applying advanced algorithms such as minimum variance distortionless response (MVDR) for signal analysis.
[0068] In order to improve the quality of ECG signals, ECG signals need to be preprocessed. First, the third-order partial derivative smoothing filter is used to eliminate the respiratory interference and other non-heartbeat related interference components in the ECG signals. The third-order partial derivative smoothing filter can enhance the characteristics of the heartbeat signal and facilitate subsequent feature extraction. The working principle of the filter is defined by the following formula:
[0069] in, represents six consecutive signal sample points, represents the time interval between samples, represents the denoised signal.
[0070] Then, a Bass low-pass filter is used to further remove high-frequency noise, retaining only the components within the heartbeat signal frequency range (0-5 Hz). This step specifically removes unnecessary high-frequency components from the signal and focuses on signal changes related to the heartbeat, thereby improving the quality and usability of the signal.
[0071] Finally, the filtered signal is normalized by subtracting the mean of the signal and dividing it by its standard deviation. This processing step ensures the uniformity of the scale of ECG signals from different leads, eliminates data differences for subsequent deep learning training, and improves model robustness and accuracy.
[0072] In order to improve the quality of the respiratory signal, it is also necessary to preprocess the respiratory signal. Specifically, the differential operation of adjacent data points is used to suppress baseline drift and other low-frequency interference, and highlight the signal changes caused by respiratory activity. This method effectively strengthens the signal components of respiratory periodicity. Similar to the heartbeat signal, the Bass low-pass filter is also used to remove high-frequency noise in the respiratory signal, but only retains the signal components within the respiratory frequency range (0-2.5 Hz).
[0073] Through the above preprocessing steps, the quality of the signal can be significantly improved, providing a reliable and clear data basis for subsequent multimodal analysis. These steps ensure the consistency and accuracy of the data, which is the prerequisite for efficient feature extraction.
[0074] Furthermore, after the waveform features, time-frequency features, and super-resolution frequency features of the initial signal are extracted by the multimodal feature extraction module, in order to correctly label its heart rate label samples and respiratory rate label samples, it is also necessary to set a time window and perform timestamp alignment processing on the initial signal samples, heart rate label samples, and respiratory rate label samples based on the time window. The specific steps are as follows: 1. Align the timestamps of heartbeat and respiratory features: Accurately match the timestamps of heartbeat and respiratory features with the reference label data to ensure the synchronization of the features and their corresponding labels. Specifically, for each independent data window, calculate the time difference between the center moment of the window and the reference label moment. If this difference is less than the preset threshold, the heartbeat or respiratory rate value in the reference label is directly used as the label of the window feature. This step enhances the accuracy of subsequent analysis by ensuring a high degree of consistency between the data and the label.
[0075] 2. Application of MVDR algorithm in frequency label estimation: When the original label data is missing or unreliable, the MVDR algorithm is used to estimate the heart rate or respiratory rate. By screening the frequency components in the MVDR features that meet the normal heart rate and respiratory range, and calculating their average power spectrum density, the frequency corresponding to the maximum power is selected as the new heart rate or respiratory rate label. This algorithm significantly improves the accuracy of label estimation, especially under non-ideal conditions.
[0076] 3. Synchronization of blood pressure and blood oxygen saturation feature extraction: In order to maintain consistency with the electrocardiogram (ECG) signal features and facilitate subsequent deep learning processing, zero padding is performed on the blood pressure signal (500Hz sampling rate) and the blood oxygen saturation signal (208Hz sampling rate) to match the feature size of the ECG signal. Considering the close correlation between the fluctuations of blood pressure and blood oxygen saturation and cardiac activity, these signal change cycles can be regarded as heartbeat frequencies.
[0077] 4. Rough estimation and smoothing of respiratory rate: The MVDR algorithm is used to estimate the respiratory rate from the blood pressure signal and the blood oxygen saturation signal, and the exponential filter is combined to smooth the label values of the heart rate and respiratory rate. This not only reduces the estimation error, but also increases the continuity and reliability of the prediction, thereby optimizing the performance of the overall monitoring system.
[0078] In some embodiments of the present invention, in order to increase the number of training sample sets and better adapt to actual conditions, it is also necessary to enhance the initial training sample set to obtain a target training sample set; The enhancement processing means include at least one of adding noise, randomly erasing data, randomly deleting a channel and setting the channel data to zero, data scaling, and local signal clipping and random insertion.
[0079] Specifically, adding noise includes adding random noise to ECG, blood pressure and blood oxygen saturation signals to simulate the interference that the signals may encounter during acquisition and transmission. This method can significantly enhance the robustness of the model to noise in practical applications and ensure that the model can maintain efficient performance in a noisy environment.
[0080] Random data erasure involves deleting a small piece of data at a random position in the signal to simulate signal interruption or data loss. This method forces the model to learn how to extract useful information from incomplete data and is an effective way to improve the model's ability to handle abnormal data.
[0081] Randomly deleting channels and setting the data of the changed channels to zero includes: randomly selecting data of one or more channels on the signal and setting the data of this channel to 0 to simulate the situation of lead detachment, channel loss, etc. This method forces the model to learn how to extract useful information from incomplete data, which is an effective way to improve the model's ability to handle abnormal data.
[0082] Data scaling involves randomly scaling the amplitude of the signal to simulate the variation in signal strength from different patients or different devices. This method helps the model adapt to the variation in signal strength from different sources and enhances the model's adaptability when processing multi-source data.
[0083] Local signal cropping and random insertion involves randomly cropping a small piece from the signal and placing it in the original size data to simulate the local changes of the signal. This method encourages the model to focus on the local features of the signal, which helps improve the accuracy of the model in parsing details.
[0084] Through the above data enhancement technology, not only the generalization ability of the model is improved, but also the accuracy and robustness of the model in complex real-world scenarios are ensured.
[0085] In summary, the present application extracts the waveform features, time-frequency features and super-resolution frequency features of the initial signal through a multimodal feature extraction module. Since the initial signal can exist in multiple forms and multiple combinations, the dependence on a certain specific form of the initial signal is reduced; further, since the waveform features, time-frequency features and super-resolution frequency features exist in multiple forms, they can reflect the characteristics of the initial signal from multiple dimensions, thereby ensuring the reliability of the detection results of the heart rate and respiratory rate. By assigning weights to the samples and calculating the weighted mean square error loss value of each sample based on the weighted mean square error loss function, the weight distribution of the weighted mean square error loss function for samples of different modes is optimized, and the contribution of sample data of different modes to the final estimation result is adaptively adjusted, thereby improving the optimization effect of the target deep learning sub-model. Through data enhancement technology, not only the generalization ability of the model is improved, but also the accuracy and robustness of the model in complex real-world scenarios are ensured.
[0086] In order to better implement the heart rate and respiratory rate detection method in the embodiment of the present invention, the embodiment of the present invention also provides a heart rate and respiratory rate detection device, such as Figure 4 As shown, Figure 4 This is a schematic structural diagram of an embodiment of a heart rate and respiratory rate detection device provided by the present invention. The heart rate and respiratory rate detection device 400 includes: A signal monitoring model construction module 401 is used to construct a signal monitoring model, the signal monitoring model includes a multimodal feature extraction module and a deep learning sub-model; A feature extraction module 402 is used to extract waveform features, time-frequency features and super-resolution frequency features of the initial signal based on the multimodal feature extraction module; A heart rate and respiratory rate detection module 403 is used to fuse waveform features, time-frequency features and super-resolution frequency features based on a deep learning sub-model, and determine the heart rate and respiratory rate of the initial signal according to the fused features; The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0087] like Figure 5 As shown, Figure 5 This is a structural block diagram of an embodiment of an electrocardiovascular and blood pressure monitor provided by the present invention. The electrocardiovascular and blood pressure monitor 500 includes a processor 501 , a memory 502 and a display 503 . Figure 5 Only some components of the ECG and blood pressure monitor 500 are shown, but it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead.
[0088] In some embodiments, the processor 501 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 502, such as the heart rate and respiratory rate detection method of the present invention.
[0089] In some embodiments, the processor 501 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 501 may be local or remote. In some embodiments, the processor 501 may be implemented in a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination thereof.
[0090] In some embodiments, the memory 502 may be an internal storage unit of the ECG and blood pressure monitor 500, such as a hard disk or memory of the ECG and blood pressure monitor 500. In other embodiments, the memory 502 may also be an external storage device of the ECG and blood pressure monitor 500, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the ECG and blood pressure monitor 500.
[0091] Furthermore, the memory 502 may include both an internal storage unit of the electrocardiovascular and blood pressure monitor 500 and an external storage device. The memory 502 is used to store application software installed in the electrocardiovascular and blood pressure monitor 500 and various data.
[0092] In some embodiments, the display 503 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 503 is used to display information on the ECG and blood pressure monitor 500 and to display a visual user interface. The components 501-503 of the ECG and blood pressure monitor 500 communicate with each other via a system bus.
[0093] In one embodiment, when the processor 501 executes the heart rate and respiratory rate detection program in the memory 502, the following steps may be implemented: Construct a signal monitoring model, which includes a multimodal feature extraction module and a deep learning sub-model; The waveform features, time-frequency features and super-resolution frequency features of the initial signal are extracted based on the multimodal feature extraction module; Based on the deep learning sub-model, waveform features, time-frequency features and super-resolution frequency features are fused, and the heart rate and respiratory rate of the initial signal are determined according to the fused features. The initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
[0094] It should be understood that: when the processor 501 executes the heart rate and respiratory rate detection program in the memory 502, in addition to the above functions, other functions can also be implemented. For details, please refer to the description of the corresponding method embodiment above.
[0095] Furthermore, the embodiment of the present invention does not specifically limit the type of the ECG and blood pressure monitor 500 mentioned. The ECG and blood pressure monitor 500 may be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of portable electronic devices include but are not limited to portable electronic devices equipped with IOS, Android, Microsoft or other operating systems. The above-mentioned portable electronic devices may also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the ECG and blood pressure monitor 500 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0096] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0097] The heart rate and respiratory rate detection method, device, electronic device and storage medium provided by the present invention are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A heart rate and respiratory rate detection method, characterized in that: include: Constructing a signal monitoring model, wherein the signal monitoring model includes a multimodal feature extraction module and a deep learning sub-model; Based on the multimodal feature extraction module, the waveform features, time-frequency features and super-resolution frequency features of the initial signal are respectively extracted; Based on the deep learning sub-model, the waveform features, the time-frequency features and the super-resolution frequency features are subjected to feature fusion, and the heart rate and the respiratory rate of the initial signal are respectively determined according to the fused features; Wherein, the initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.
2. The heart rate and respiratory rate detection method according to claim 1, characterized in that: The deep learning sub-model includes a convolution layer, a backbone network layer, an average pooling layer and a fully connected layer; the waveform feature, the time-frequency feature and the super-resolution frequency feature are fused based on the deep learning sub-model, and the heart rate and the respiratory rate of the initial signal are determined according to the fused features, including: Extracting the waveform feature and the time domain sub-features of the super-resolution frequency feature, and the frequency domain sub-features of the time-frequency feature based on the convolution layer; Performing global feature extraction on the time domain sub-features and the frequency domain sub-features based on the backbone network layer to obtain a feature map of the initial signal; Transforming the feature map based on the average pooling layer to obtain a one-dimensional feature vector of the initial signal; The heart rate and breathing rate of the initial signal are output according to the one-dimensional feature vector based on the fully connected layer.
3. The heart rate and respiratory rate detection method according to claim 2, characterized in that: The convolution layer includes a one-dimensional convolution layer and a two-dimensional convolution layer; the time domain sub-features and frequency domain sub-features of the waveform feature, the time-frequency feature and the super-resolution frequency feature are extracted based on the convolution layer, respectively, including: Extracting the time domain sub-features of the waveform feature and the super-resolution frequency feature respectively based on the one-dimensional convolution layer; The frequency domain sub-features of the time-frequency features are extracted based on the two-dimensional convolutional layer.
4. The heart rate and respiratory rate detection method according to claim 2, characterized in that: The fully connected layer includes a first fully connected layer and a second fully connected layer; the step of outputting the heart rate and the respiratory rate of the initial signal according to the one-dimensional feature vector based on the fully connected layer includes: The one-dimensional feature vector is subjected to feature dimensionality reduction and nonlinear transformation according to the first fully connected layer, and the second fully connected layer outputs the heart rate and respiratory rate of the initial signal according to the one-dimensional feature vector after the feature dimensionality reduction and nonlinear transformation.
5. The heart rate and respiratory rate detection method according to claim 3, characterized in that: The deep learning sub-model also includes a compression excitation layer; in the process of performing global feature extraction on the time domain sub-feature and the frequency domain sub-feature based on the backbone network layer to obtain the feature map of the initial signal, it also includes: The residual block of the backbone network layer is revised, and the feature channel weights of the one-dimensional convolution layer and the two-dimensional convolution layer are calibrated according to the adaptive feedback of the compression excitation layer, and the number of input channels of the residual block is adjusted to be equal to the number of feature output channels of the multimodal feature extraction module.
6. The heart rate and respiratory rate detection method according to claim 1, characterized in that: Before fusing the waveform feature, the time-frequency feature, and the super-resolution frequency feature based on the deep learning sub-model, and determining the heart rate and the respiratory rate of the initial signal respectively according to the fused features, the method further includes: Constructing a training sample set, wherein the training sample set includes an initial signal sample and a heart rate label sample and a respiratory rate label sample corresponding to each initial signal sample; Assigning a weight to each of the initial signal samples based on the type of the initial signal sample; Based on the multimodal feature extraction module, the waveform features, the time-frequency features and the super-resolution frequency features of the initial signal samples are respectively extracted and input into the initial deep learning sub-model, and the heart rate label samples and the respiratory rate label samples of the initial signal samples are output. The training is iteratively performed, and the weighted mean square error loss values of the heart rate label samples and the respiratory rate label samples are calculated based on the weighted mean square error loss function combined with the sample weight assignment results, and the initial deep learning sub-model corresponding to the minimum weighted mean square error loss value is determined as the target deep learning sub-model.
7. The heart rate and respiratory rate detection method according to claim 6, characterized in that: The constructing of the training sample set further includes: A time window is set, and timestamp alignment processing is performed on the initial signal sample, the heart rate label sample, and the respiratory rate label sample based on the time window.
8. The heart rate and respiratory rate detection method according to claim 6, characterized in that: The constructing of the training sample set further includes: Perform enhancement processing on the initial training sample set to obtain the target training sample set; The enhancement processing means include at least one of adding noise, randomly erasing data, randomly deleting a channel and setting the channel data to zero, data scaling, and local signal clipping and random insertion.
9. The heart rate and respiratory rate detection method according to claim 5, characterized in that: The step of constructing a training sample set further includes: When a group of initial signal samples includes the electrocardiogram signal / the respiratory signal, the heart rate label sample and the respiratory rate label sample corresponding to the electrocardiogram signal are used as the heart rate label sample and the respiratory rate label sample of the other initial signal samples in the group; When a group of initial signal samples only includes the blood oxygen signal and the blood pressure signal, the heart rate label sample and the respiratory rate label sample corresponding to the blood oxygen signal are used as the heart rate label sample and the respiratory rate label sample of the other initial signal samples in the group.
10. A heart rate and respiratory rate detection device, characterized in that: include: A signal monitoring model construction module, used to construct a signal monitoring model, wherein the signal monitoring model includes a multimodal feature extraction module and a deep learning sub-model; A feature extraction module, used to extract waveform features, time-frequency features and super-resolution frequency features of the initial signal based on the multimodal feature extraction module; A heart rate and respiratory rate detection module, configured to fuse the waveform features, the time-frequency features and the super-resolution frequency features based on the deep learning sub-model, and determine the heart rate and respiratory rate of the initial signal respectively according to the fused features; Wherein, the initial signal includes at least one of an electrocardiogram signal, a blood pressure signal, a blood oxygen saturation signal and a respiratory signal.