A motion interference resistant ear-wearable heart rate respiration real-time monitoring method and system
By collecting audio and IMU data through ear-worn devices and using signal processing and deep learning technologies, the problem of accurately monitoring heart rate and respiratory rate during exercise is solved, and accurate estimation of heart rate and respiratory rate during exercise is achieved. The device is comfortable to wear and easy to clean.
Patent Information
- Application Number
- CN202411705770.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Existing motion monitoring equipment is difficult to accurately monitor heart rate and respiratory rate simultaneously during exercise, and existing ear-worn devices are easily disturbed by motion and cannot effectively eliminate noise interference.
An ear-worn wearable device is used to collect audio data and IMU data from the ear canal. Through multi-sensor collaboration, signal processing and deep learning techniques are used to extract heartbeat and breathing features, reconstruct the electrocardiogram spectrum and respiratory waveform, and combine peak-to-peak detection to obtain estimated results of heart rate and respiratory rate.
It realizes simultaneous monitoring of heart rate and respiratory rate during exercise, effectively eliminating motion interference and providing more accurate heartbeat and respiratory information. The device is easy to wear, less invasive and easy to clean.
Smart Images

Figure CN119791616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and more particularly to an ear-worn heart rate and respiration real-time monitoring method and system that is resistant to motion interference. Background Art
[0002] With improved quality of life and advances in sensor technology, people are increasingly interested in using activity monitoring devices to record their exercise. During exercise, heart rate and respiratory rate are key indicators of exercise effectiveness. Guided by these indicators, people can better choose the right exercise for them, determine the duration and intensity of their exercise, and ultimately achieve better fitness results with scientific data.
[0003] In the prior art, a common heart rate monitoring sensor is a PPG (photoplethysmography) sensor, which generally needs to fit closely to the skin and has weak resistance to motion interference. A common respiratory monitoring sensor is an IMU (inertial measurement unit), which monitors breathing by monitoring the rise and fall of the chest. For example, patent application CN202110267923.3 provides a human respiratory rate monitoring device integrated in a mask, which integrates a relatively independently arranged breathing valve and a detection module with the mask. The breathing valve is connected to the breathing valve hole of the mask by snapping, and the detection module is connected to the mask by gluing. The mask is replaceable, and the breathing valve and detection module can be easily disassembled and reused. However, this solution cannot achieve simultaneous monitoring of heart rate and respiratory rate, and during exercise, the mask will burden the user's breathing.
[0004] Patent application CN201410630915.0 provides an ear-worn heart rate monitoring device and method. The monitoring device comprises: a sensing unit, including a light source and a brightness sensor, disposed within the human auricle for detecting the human body's pulse wave photoplethysmography signal and transmitting it to an acquisition and processing unit; an acquisition and processing unit for acquiring and calculating the pulse wave photoplethysmography signal to obtain heart rate detection information; a data transmission and reception unit for transmitting heart rate detection information and receiving configuration information or control instructions; and a battery and power supply unit for providing power to the ear-worn heart rate monitoring device. This solution utilizes an ear-worn wearable design, shielding ambient light through the structure of the human ear, without compressive body contact, and suitable for long-term wear. However, this solution cannot simultaneously monitor heart rate and respiratory rate during exercise, and uses the peak value of a digital signal sequence to obtain the heartbeat interval, and thus the heart rate, which is not immune to motion interference.
[0005] In summary, the existing technology mainly has the following defects:
[0006] 1) Among existing exercise monitoring solutions, wristband smartwatches based on PPG are the preferred choice. However, physical activity can cause the PPG sensor to shift relative to its original position, changing the light path through the sensor and generating motion artifacts. Furthermore, wristband sports watches cannot monitor respiration during exercise.
[0007] 2) In monitoring solutions based on breathing belts and sports chest straps, RR (heart rate) is primarily monitored using a breathing belt or sports chest strap worn around the abdomen. Breathing belts cannot effectively counteract the effects of exercise, while sports chest straps need to be worn close to the skin, making them cumbersome to wear, easily stained by sweat, and difficult to clean, making them unsuitable for average people.
[0008] 3) Among ear-worn device monitoring methods, IMU-based respiratory monitoring infers respiratory rate by capturing the subtle head movements produced by a person breathing while at rest. However, this monitoring method is easily disrupted by large movements and cannot effectively monitor during exercise. Furthermore, the audio signal processing methods used are not robust enough to accurately detect respiratory rate during exercise. Summary of the Invention
[0009] The purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide an ear-worn heart rate and respiration real-time monitoring method and system that is resistant to motion interference.
[0010] According to a first aspect of the present invention, a method for real-time heart rate and respiration monitoring using an ear-worn device that is resistant to motion interference is provided. The method comprises:
[0011] For target users, use ear-worn wearable devices to collect audio data and IMU data from the ear canal;
[0012] extracting heartbeat features and breathing features from the audio data;
[0013] Inputting the respiratory characteristics into a respiratory waveform reconstruction model to obtain a reconstructed respiratory waveform;
[0014] Inputting the heartbeat features into a spectrum reconstruction model to obtain a reconstructed electrocardiogram spectrogram, and converting the reconstructed electrocardiogram spectrogram from a frequency domain representation to a time domain representation to obtain a time domain electrocardiogram waveform;
[0015] Preprocessing the IMU data to obtain an IMU chest cavity fluctuation waveform;
[0016] Based on the time-domain electrocardiogram waveform, the reconstructed respiratory waveform, and the IMU chest rise and fall waveform, estimation results of the respiratory frequency and heart rate are obtained through peak-to-peak value detection.
[0017] According to a second aspect of the present invention, there is provided an ear-worn heart rate and respiration real-time monitoring system that is resistant to motion interference, comprising an ear-worn wearable device and a terminal device, wherein the ear-worn wearable device is used to collect audio data and IMU data in the ear canal of a target user and transmit the data to the terminal device; the terminal device executes: extracting heartbeat features and breathing features from the audio data; inputting the breathing features into a breathing waveform reconstruction model to obtain a reconstructed breathing waveform; inputting the heartbeat features into a spectrum reconstruction model to obtain a reconstructed electrocardiogram spectrum, and converting the reconstructed electrocardiogram spectrum from a frequency domain representation to a time domain representation to obtain a time domain electrocardiogram waveform; preprocessing the IMU data to obtain an IMU chest cavity fluctuation waveform; and obtaining estimation results of the breathing frequency and heart rate through peak-to-peak value detection based on the time domain electrocardiogram waveform, the reconstructed breathing waveform and the IMU chest cavity fluctuation waveform.
[0018] Compared with existing technologies, the present invention offers the advantage of providing a motion-interference-resistant, ear-worn heart rate and respiration real-time monitoring solution. This solution utilizes an ear-worn device and a multi-sensor collaborative approach to simultaneously monitor heart rate and respiration rate, extracting effective heartbeat and respiration information while effectively eliminating the effects of motion interference. The ear-worn wearable device is less invasive, easier to wear, and easier to clean than wearable clothing or chest straps.
[0019] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0021] Figure 1 This is a flow chart of a method for real-time heart rate and respiration monitoring with earwear that is resistant to motion interference according to an embodiment of the present invention;
[0022] Figure 2 1 is a schematic diagram of the overall framework of a method for real-time heart rate and respiration monitoring with earwear that is resistant to motion interference according to an embodiment of the present invention;
[0023] Figure 3 is an overall framework diagram of an earhook headset according to an embodiment of the present invention;
[0024] Figure 4 is a structural diagram of an earhook device according to one embodiment of the present invention;
[0025] Figure 5 is an overall framework diagram of a neckband headset according to an embodiment of the present invention;
[0026] Figure 6 is an appearance diagram of a neckband headset according to an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of a main control board of a neckband headset according to one embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of the result after empirical wavelet decomposition according to one embodiment of the present invention;
[0029] Figure 9 is a schematic diagram of an ECG spectrum mapping model according to one embodiment of the present invention;
[0030] Figure 10 is a schematic diagram of a respiratory waveform mapping model according to an embodiment of the present invention;
[0031] Figure 11 Schematic diagram of a user personal interface, a personal information drop-down interface, and a sports record interface according to an embodiment of the present invention;
[0032] Figure 12 Schematic diagram of the application home page interface, exercise progress interface, exercise pause interface, and exercise end recording interface according to one embodiment of the present invention;
[0033] Figure 13 Schematic diagram of a fast heart rate warning interface, a rapid breathing warning interface, and an exercise intensity prompt interface according to one embodiment of the present invention;
[0034] Figure 14 is a schematic diagram of overall user performance according to one embodiment of the present invention;
[0035] Figure 15 is a schematic diagram of an empirical cumulative distribution function according to an embodiment of the present invention;
[0036] Figure 16 is a schematic diagram of the results of long-term tracking of heart rate and respiration according to one embodiment of the present invention;
[0037] In the figure, Max Pool-maximum pooling; Conv-convolution; Global Average Pooling-global average pooling; Fully Connected Layer-fully connected layer; Transposed Conv-transposed convolution; Conv Block-convolution block. DETAILED DESCRIPTION
[0038] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0039] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0040] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0041] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0042] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0043] In summary, the present invention first designs an ear-worn wearable device to collect user audio data and IMU data. After obtaining the user data, the signal is preprocessed to obtain the basic input of a deep learning model. Finally, the model reconstructs the ECG (electrocardiogram) curve and respiratory waveform curve, and then applies heart rate extraction algorithms and respiratory rate extraction algorithms to obtain heart rate and respiratory rate estimates.
[0044] Specifically, combined Figure 1 and Figure 2 As shown, the provided ear-worn heart rate and respiration real-time monitoring method with anti-motion interference includes the following steps:
[0045] Step S110: Design an ear-worn wearable device to collect the user's audio data and IMU data.
[0046] To accommodate different user preferences, ear-worn wearable devices can be designed with different configurations. For example, two types of wearable devices are available: ear-hook and neck-hook. Ear-hook devices contain an in-ear microphone and an ear IMU, while neck-hook devices contain an in-ear microphone and a neck IMU. When worn, the neck IMU is positioned close to the chest, while the ear IMU is located on the user's head. Both devices can obtain audio data and IMU data from the ear canal.
[0047] (1) Ear-hook wearable devices
[0048] In one embodiment, the ear-hanging wearable device is an ear-hanging earphone. The device contains an ESP32 single-chip microcomputer, several sensors, and a 3D support shell designed by the inventor. See Figure 3 The overall framework of the ear-hanging earphone is shown in FIG. 1, in which the connection and communication of the ESP32 single-chip microcomputer with the multiple sensors and the mobile terminal are as follows.
[0049] The ESP32 single-chip microcomputer and the MEMS (Micro Electro Mechanical System) microphone sensor are connected using an I2S digital audio bus. The ESP32 acquires the audio data collected by the microphone through the SCK (Serial Clock), WS (Word Select), and SD (Serial Data) pins.
[0050] The ESP32 single-chip microcomputer and the IMU sensor are connected using an I2C bus. The ESP32 communicates bidirectionally with the IMU sensor through the SCL (Serial Clock) and SDA (Serial Data) pins to read acceleration, gyroscope (and magnetometer) data.
[0051] The ESP32 single-chip microcomputer and the mobile terminal are wirelessly connected. For example, a Wi-Fi-based wireless connection is used, and the ESP32 single-chip microcomputer continuously sends audio data to the mobile terminal through the HTTP protocol.
[0052] The structure of the ear-hanging earphone is shown in FIG. 2, in which a miniature microphone sensor is provided at the ear entry of the device to continuously collect audio signals in the user's ear canal. At the same time, the IMU sensor in the device can also continuously collect the user's body posture information, which is then transmitted to the mobile terminal device through Wi-Fi for subsequent heart rate and respiration rate detection. Figure 4
[0053] (2) Neck-hanging wearable device
[0054] In one embodiment, the neck-hanging (or hanging) wearable device is a neck-hanging earphone. The neck-hanging earphone includes a neck-hanging ring, earbuds connected to the neck-hanging ring, a main control board circuit, a battery unit, an IMU sensor, and a microphone sensor. The microphone sensor is arranged in the earbuds, and the main control circuit board is arranged at one end of the neck-hanging ring. The IMU sensor and the battery unit are arranged at the other end of the neck-hanging ring. See Figure 5 The device contains an ESP32 single-chip microcomputer and several sensors. The connection and communication of the ESP32 single-chip microcomputer with the multiple sensors and the mobile terminal are as follows.
[0055] The ESP32 microcontroller and the MEMS microphone sensor are connected using the I2S digital audio bus. The ESP32 obtains the audio data collected by the microphone through the SCK (Serial Clock), WS (Word Select), and SD (Serial Data) pins.
[0056] The ESP32 microcontroller and the IMU sensor are connected using the I2C bus. The ESP32 communicates bidirectionally with the IMU sensor via the SCL (Serial Clock) and SDA (Serial Data) pins to read data from the accelerometer, gyroscope, and magnetometer.
[0057] A wireless connection is used between the ESP32 microcontroller and the mobile terminal. For example, a wireless connection based on Wi-Fi is used, and the ESP32 microcontroller continuously sends audio data to the mobile terminal via the HTTP protocol.
[0058] The overall appearance of the neckband headphones is as follows Figure 6 As shown, the main control board structure is as follows Figure 7 As shown in the figure, there is a miniature microphone sensor at the ear of the device, which can continuously collect audio signals from the user's ear canal. At the same time, the IMU sensor in the device can also continuously collect the user's body posture information, and then transmit various data to the mobile terminal device via Wi-Fi for subsequent heart rate and respiratory rate detection.
[0059] It should be understood that the ear-worn device designed in this invention primarily uses an in-ear microphone and an IMU to collect user data and transmit it to a mobile device for calculation. In addition to ear-hook and neck-hook designs, other forms can also be used, for example, using multiple devices to obtain a microphone and an IMU respectively. Furthermore, the earbuds and neck loop can be connected wirelessly or via a data cable. The wireless ear-worn design ensures stable wearing and stable signal transmission, thereby reducing interference during exercise.
[0060] Step S120: Process the collected audio data to extract heartbeat features and breathing features.
[0061] Audio signals contain multiple sources of aliasing. Signal decomposition methods such as empirical mode decomposition (EMD) and discrete wavelet transform (DWT) can be used. EMD cannot effectively separate information from different modes, making modal aliasing more likely. DWT uses fixed-ratio sub-bands to divide frequencies and lacks adaptive filtering capabilities. Therefore, in one embodiment, the empirical wavelet transform (EWT) is used to adaptively select frequency bands to overcome modal aliasing caused by discontinuous time-frequency scales in the signal.
[0062] Specifically, after obtaining the audio data, we first use empirical wavelet decomposition to obtain data in different frequency bands. For the signal f(t), we obtain the normalized FFT spectrum of the audio, search for M maximum points of the spectrum, and sort them in descending order. Then, we use the minimum point between two adjacent maximum points as the boundary ω. n To divide the spectrum, where ω0=0, ω n =π, each segment can be expressed as Λ n =[ω n-1 ,ω n ]. n As the center, define a converter T n , with a width of 2τ n Using EWT, the signal is decomposed into approximate coefficients and detail coefficients. The approximate coefficients usually only contain some trends of the signal, which is not useful for HR (heart rate) and RR (respiratory rate) estimation. Therefore, for Given the following empirical wavelet function:
[0063]
[0064] in:
[0065] β(x)=x 4 (35-84x+70x 2 -20x 3 )(2)
[0066] EWT can be viewed as a filter bank, detail coefficient It is obtained by the inner product with the empirical wavelet:
[0067]
[0068] Therefore, the empirical coefficient:
[0069]
[0070] like Figure 8As shown, after EWT processing, the corresponding component coefficients are selected according to the specific frequency band. The component of the 1-3 Hz frequency band contains the heart beat period (Heartbeart period coefficients), which is a sinusoidal curve, and can be used to determine the rough position of the heart beat. The component of the 3-20 Hz frequency band contains the harmonic information of the heart beat sound (Heartbeartharmonic coefficients), and the curve has more details, which is convenient for supplementing details when reconstructing ECG. The sampling rate is set to 10 KHz, and the last detail coefficient is the breathing sound coefficient (breathing sound coefficients), which contains information of 20-5000 Hz, covering the frequency distribution of the breathing sound, and can be used as the input of the subsequent breathing frequency extraction.
[0071] Although EWT can extract the harmonic information of the heart beat sound, it is still disturbed by some noise. Therefore, the extraction of the heart beat is further realized by a deep learning method. For example, the original audio data (Original audio data), the heart beat period, and the heart beat harmonic coefficient are down-sampled to 1 kHz, and then they and the breathing sound coefficient are converted to Mel Spectrogram (Mel Spectrogram) and logarithm. For example, the window length of the fast Fourier transform of the heart beat and breathing signal is 128 and 2048 respectively, the window length of the short-time Fourier transform is 128 and 2048 respectively, the window overlap length is 51 and 256 respectively, and the number of Mel filters is 64 and 128 respectively. Then the first three are stacked as the input of the heart rate estimation model, and the remaining one is used as the input of the breathing estimation model, which refers to the breathing wave reconstruction model containing the spatial feature extraction module and the temporal feature extraction module in the following.
[0072] In step S130, the collected IMU data is processed to obtain a chest fluctuation waveform reflecting the user's motion information.
[0073] The IMU signal is used for breathing frequency estimation. Since the breathing sound of a person in a resting state is weak, the breathing sound collected by the in-ear microphone is very weak, which can easily lead to poor breathing frequency estimation effect. Therefore, the IMU can be used to capture the subtle movements of a person when breathing, such as chest fluctuation and head movement caused by breathing, as an additional breathing frequency estimation signal.
[0074] The breathing frequency of a person is mainly in the range of 10-40 BPM, so the IMU signal can be band-pass filtered or wavelet decomposed, etc. For example, the band-pass filtering range is 0.2-0.8 Hz. After filtering, the IMU signal filters out most of the motion noise and random noise, and retains the low-frequency breathing information.
[0075] It should be understood that the preprocessing of audio data and IMU data can adopt various signal processing methods such as wavelet transform, empirical mode decomposition, bandpass filtering, etc., as long as the information containing the frequency of heartbeat and breathing information can be extracted.
[0076] Step S140: input the heartbeat features into a spectrum reconstruction model to obtain a reconstructed electrocardiogram spectrum.
[0077] In signal preprocessing, a three-channel stacked heart rate estimation input spectrum is obtained. In one embodiment, an ECG spectrum reconstruction model is designed to map the multi-channel MEL spectrogram to the ECG STFT spectrogram and eliminate noise in the audio spectrogram. Figure 9 This is an ECG spectrum mapping model that uses UNet as its base model and incorporates a channel-attention mechanism to effectively extract features from the spectrum. In the encoder, the feature map undergoes repeated 3x3 convolutions, LeakReLU activation functions, and normalization layers before undergoing channel-attention extraction. For example, SENet uses global average pooling to compress the spatial features of each channel into a scalar. After passing through two fully connected layers, the channel-attention weight is obtained using a sigmoid function. Using the channel-attention mechanism, the model adaptively selects valid feature channels from the original spectrum and EWT additional features, improving model performance. In the decoder, the data passes through successive upconvolution blocks, halving the number of feature maps. After each upconvolution, the feature map is merged with the corresponding feature map from the encoder, followed by a convolution block and batch normalization. In the final layer, a 1x1 convolution is used to map the feature map to the spectrum of a single output channel.
[0078] Step S150: input the respiratory characteristics into the respiratory waveform reconstruction model to obtain a reconstructed respiratory waveform.
[0079] The distribution of breathing sounds varies depending on individual characteristics and physical activities, and is often masked by background noise. Therefore, a breathing curve reconstruction model (or breathing curve mapping model) is designed to effectively reconstruct the breathing waveform from the audio spectrogram, such as Figure 10 As shown in FIG, the respiratory waveform reconstruction model includes a spatial feature extraction module (SpatialFeatureExtraction) and a temporal feature extraction module (TemporalFeatureExtraction).
[0080] (1) Spatial feature extraction module
[0081] Log-Mel Spectrogram first passes through a layer of convolution module to obtain shallow feature information. The structure of the convolution module is similar to Figure 9The same as in . Then it goes through two layers of deformable convolution (DCN) modules to extract the deep feature representation of breathing sound. Given a convolution kernel with K sampling positions, w k and p k Denote the weight and predefined offset of the kth position respectively. 3x3 convolution means K=9. x(p) and y(p) denote the input feature map and output feature map at position p respectively. The formula of deformable convolution is as follows:
[0082]
[0083] Where Δp k is the learnable offset of the kth position, Δp k Is a real number with an unlimited range of values. k This is achieved by applying a single convolutional layer to the same input feature map, with a 2K-channel output and kernel weights initialized to 0. The learning rate of this single convolutional layer is set to 0.1 times that of other conventional layers. After passing through two DCN layers, the output is connected to an upconvolution block. The structure of the upconvolution block is similar to that of the heart rate estimation model. However, this upconvolution block only expands the time dimension while continuously halving the frequency bins. The final intermediate feature is a bin dimension of 1, 512 channels, and a time length of 250. This feature is then fed into the temporal information capture module.
[0084] (2) Temporal feature extraction module
[0085] A complete breath consists of two phases: inhalation and exhalation. These phases are represented by the rising and falling edges of a real breathing curve, respectively, and by two continuous regions of concentrated energy in the audio time-frequency domain. Because the sounds of inhalation and exhalation are very similar, convolution cannot effectively reconstruct the phase of the breathing curve. In other words, the model incorrectly reconstructs inhalation and exhalation as the falling and rising edges of the breathing curve. A bidirectional LSTM can be used to capture the long-term context of breathing, allowing the model to determine whether the current phase is inhalation or exhalation based on contextual information. The Bi-LSTM has 8 layers, and the LSTM hidden layer dimension is 256. The output is connected to a fully connected layer to reduce the LSTM output channels to 1D to obtain a reconstructed breathing curve.
[0086] In summary, the present invention considers the recovery of respiratory waveform signals. Respiration rate is typically derived from the fluctuations of the respiratory waveform, which reflects the fluctuations in the chest cavity during breathing and, in turn, the user's breathing. The present invention can recover a true and valid respiratory waveform signal from audio data and IMU data.
[0087] Step S160 , estimating the user's respiratory frequency and heart rate using the reconstructed electrocardiogram spectrogram and respiratory waveform.
[0088] Figure 2 The model training or reasoning can use the respiratory waveform and ECG collected by the existing sports chest belt to train the spectrum reconstruction model and the respiratory waveform reconstruction model.
[0089] After model inference, the reconstructed ECG spectrogram and respiratory waveform are obtained. Since the reconstructed ECG spectrogram lacks phase information, the Griffin-Lim algorithm is used to convert the frequency domain representation back to the time domain. A sliding window with a length of 10 seconds and a step size of 5 seconds is applied to segment the reconstructed ECG and respiratory curves. Peak-to-peak detection is used for each window to calculate the time interval between peaks. The z-scores of these intervals are then calculated, and outliers are filtered out. The HR (heart rate) and RR of each window are determined by averaging the time intervals after filtering out outliers. Finally, a moving average filter is applied to smooth the HR and RR curves.
[0090] The preprocessed IMU signal is first monitored for amplitude. If it exceeds a certain value, the user is considered to be in motion. At this point, the IMU signal is too noisy and the IMU signal estimation is discarded. If it does not exceed a certain value, the user is considered to be at rest. The RR of the filtered IMU signal is then calculated and averaged with the RR value calculated from the audio to obtain the final RR value.
[0091] Furthermore, in order to display the estimated results of HR and RR on the user's device, such as a mobile phone, tablet computer, etc. In terms of software, a convenient and easy-to-use mobile phone application software has been developed. By connecting with the smart ear-worn device, the user can view the physiological index data during exercise in real time. Figure 11-13 This app features an intuitive and user-friendly interface, ensuring users can easily access the information they need. Users can receive personalized health and fitness advice and guidance based on their physical condition and exercise history, enabling them to conduct exercise training more scientifically.
[0092] See also Figure 11 As shown in the figure, the user's personal homepage interface displays their physical health status and exercise records. Each exercise record includes the type, duration, and date of the exercise, as well as physiological data related to the exercise process, such as heart rate changes and respiratory rate changes. This record helps users understand their exercise performance and health status, providing a reference for formulating future exercise plans.
[0093] See also Figure 12As shown, on the home page of the application, the user can select the type and duration of the exercise needed, and obtain real-time heart rate, respiratory rate and other data through the smart Bluetooth earphone. On the page during the exercise, the user can real-time understand the current exercise duration, heart rate, respiratory rate and heart rate variability and other information. In addition, the user can pause the exercise record at any time, and save the record at the end.
[0094] Referring to Figure 13 As shown, through real-time monitoring of the user's heart rate and respiratory rate, the user's exercise stage, such as warm-up or fat-burning stage, can be determined, and warnings can be issued according to the situation, such as warning the user if the heart rate is too high or the breathing is too fast, and suggesting the user to adjust or end the exercise in time. Such a function helps the user better understand his own physical condition during exercise, and avoids excessive exercise or health risks. Overall, the present application can provide users with comprehensive exercise monitoring and guidance services, so that they can exercise more safely and scientifically, and achieve better fitness results.
[0095] To further verify the effect of the present application, experiments were conducted.
[0096] Figure 14 The MAE of HR and RR estimates for different users and activities is given, where Figure 14 (a) is the estimate of HR, Figure 14(b) is the RR estimation. Due to device placement issues, data from users 4 and 9 while cycling and walking were excluded. The results reveal the following findings. First, there is no correlation between HR and RR estimates. For HR estimation, users U2-U5, U11, and U12 achieved an average MAE exceeding 6 BPM. However, their RR estimates were quite accurate, with an average MAE less than 2.26 RPM. This discrepancy is primarily due to the different physical properties of heartbeat and respiration. Masking significantly amplifies low-frequency heartbeat sounds while minimally affecting high-frequency respiration sounds. Therefore, the reduced occlusion caused by improper earbud closure has a more severe impact on heartbeat sounds, resulting in higher MAE for heart rate estimation, while RR estimation performance remains stable. Second, even the same activity has different effects on HR and RR estimation. The present invention provides excellent HR estimation at rest, with an average MAE of 0.93 BPM, but poor RR estimation. This is due to the reduced breathing intensity at rest, which makes respiration difficult to distinguish from background noise. Conversely, the present invention provides excellent RR estimation, with an average MAE of 1.44 RPM, but poor HR estimation during rowing. This is because the synchronization of breathing with rowing motion stabilizes HR, while the body noise generated by rowing masks low-frequency heartbeats, impairing HR measurement as described in Section 2. During running, the present invention exhibits the highest MAE for both heart rate and RR due to the strong sound of footsteps masking the heartbeat. Furthermore, body motion during running causes hardware vibrations, which creates additional interference and masks breathing sounds.
[0097] Next, the empirical cumulative distribution functions (ECDFs) of the HR and RR estimates were also calculated, see Figure 15 As shown, Figure 15 (a) is the ECDF of HR, Figure 15 (b) is the ECDF of RR. It can be seen that in the worst case, the absolute errors of HR and RR estimation for about 60% of the samples are less than 12.5BPM and 2.5RPM respectively.
[0098] Furthermore, among the activities considered, running and rowing had a significant effect on HR estimates, while resting and cycling had the smallest effects. For RR estimates, rowing and cycling had the smallest effects compared to the other three activities.
[0099] The long-term tracking performance was evaluated while users performed various aerobic exercises under real-world conditions. Users were rowing, cycling, running, and walking in the gym, with rest periods between activities depending on their activity level. Figure 16 The long-term tracking results shown in Figure 16 (a) is the long-term tracking of HR estimation, Figure 16(b) shows the long-term tracking of RR. The average MAE for HR and RR estimates is 3.46 BPM and 2.40 RPM, respectively, with MAPE values of 2.78 and 12.89, respectively. The MAE for HR estimates during each activity phase is 4.00, 3.47, 3.93, and 3.44 BPM, respectively, while the MAE for RR estimates is 1.35, 0.88, 4.24, and 2.50 RPM. HR performs well. Although some fluctuations are observed in RR estimates, the overall trend is acceptable.
[0100] It should be noted that those skilled in the art may make appropriate changes or modifications to the above embodiments without departing from the spirit and scope of the present invention. For example, the number of input channels of the U-Net spectrum mapping model can be set according to actual needs and is not limited to three-channel input. The reconstruction model can be a spectrum-to-spectrum mapping model or a spectrum-to-curve mapping model. The specific structure of the model, such as the number of convolution blocks, activation function, and convolution kernel size, can be flexibly set.
[0101] In summary, compared with the prior art, the present invention has the following advantages:
[0102] 1) The ear-worn motion-noise-resistant heart rate and respiratory rate monitoring solution proposed in this invention uses a microphone sensor and an IMU sensor to collect heart and respiratory sounds in the user's ear canal, and body movement information caused by breathing in the head or neck. It also combines signal processing and deep learning technology to suppress various noises introduced by movement, reconstruct ECG curves and respiratory curves, and ultimately obtain heart rate and respiratory rate estimation results, thereby improving the accuracy of respiratory rate estimation.
[0103] 2) In signal processing, the present invention utilizes a variety of different signal processing methods, such as empirical mode decomposition, discrete wavelet transform, and bandpass filtering, to obtain information on heartbeat and breathing in different frequency bands in the audio, and converts them into spectrograms respectively. Then, through a deep learning model, they are mapped into ECG spectrum and respiratory waveform respectively. Then, using inverse short-time Fourier transform, the ECG spectrum is inverted into an ECG waveform. Finally, the motion peak monitoring algorithm is used to calculate the heart rate and respiratory rate.
[0104] 3) This invention incorporates two different hardware configurations: ear-hook and neck-hook. These configurations are tailored to suit different user preferences, allowing users to choose their preferred style. Compared to existing wear methods, these configurations are more convenient and less intrusive. The ear-hook design is compact, small, and easy to wear. Furthermore, ear-worn devices offer a variety of configurations, meeting the needs of multiple users. The neck-hook design provides greater comfort and yields higher-quality data.
[0105] 4) Compared to sports watches, this device can monitor more physiological indicators, including not only heart rate but also respiratory rate, and can combine multiple sensors for more accurate respiratory monitoring. Compared to sports chest straps, this device does not come into contact with the skin, is less invasive, easier to clean and maintain, and more convenient to wear.
[0106] 5) During exercise, body movement, clothing friction, footsteps, external environmental noise, etc. will bring a lot of interference signals. The present invention can effectively eliminate the influence of motion interference and extract effective heartbeat information and breathing information.
[0107] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0108] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.
[0109] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0110] Computer readable program instructions for carrying out operations of the present application can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.
[0111] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0112] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0113] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0114] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.
[0115] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A method for real-time heart rate and respiration monitoring using earwear that is resistant to motion interference, comprising: For target users, use ear-worn wearable devices to collect audio data and IMU data from the ear canal; extracting heartbeat features and breathing features from the audio data; Inputting the respiratory characteristics into a respiratory waveform reconstruction model to obtain a reconstructed respiratory waveform; Inputting the heartbeat features into a spectrum reconstruction model to obtain a reconstructed electrocardiogram spectrogram, and converting the reconstructed electrocardiogram spectrogram from a frequency domain representation into a time domain representation using a Griffin-Lim algorithm to obtain a time domain electrocardiogram waveform, wherein a sliding window is applied to segment the reconstructed electrocardiogram spectrogram and the respiratory waveform, peak-to-peak detection is applied to each window to calculate the time interval between peaks and filter out outliers, the heart rate HR and respiratory rate RR of each window are determined by averaging the time intervals after filtering out outliers, and a moving average filter is applied to smooth the heart rate HR and respiratory rate RR curves; Preprocessing the IMU data and performing amplitude monitoring. If the amplitude exceeds a set value, the user is judged to be in motion. If the amplitude does not exceed the set value, the user is judged to be in a resting state. The heart rate RR of the filtered IMU data is calculated and averaged with the heart rate value calculated based on the audio data to obtain a final heart rate value, thereby obtaining an IMU chest rise and fall waveform. Based on the time-domain electrocardiogram waveform, the reconstructed respiratory waveform, and the IMU chest rise and fall waveform, obtaining estimation results of the respiratory rate and heart rate through peak-to-peak detection; The breathing characteristics and the heartbeat characteristics are obtained according to the following steps: Taking a logarithmic Mel-spectrogram of the audio data to obtain a heartbeat feature of the first channel; Decomposing the audio data by wavelet decomposition to obtain heartbeat cycle coefficients, heartbeat harmonic coefficients and respiratory sound coefficients; For the heartbeat period coefficient and the heartbeat harmonic coefficient, respectively, logarithmic Mel-spectrograms are obtained to obtain a second channel heartbeat feature and a third channel heartbeat feature, and the heartbeat features are combined with the first channel heartbeat feature to form the heartbeat feature; For the breathing sound coefficient, using a pre-emphasized logarithmic Mel-spectrogram to obtain the breathing feature; Decomposing the audio data by wavelet to obtain the heartbeat cycle coefficient, the heartbeat harmonic coefficient, and the breathing sound coefficient comprises the following steps: For the audio data, using empirical wavelet decomposition to obtain data of different frequency bands; The corresponding component coefficient is selected according to the set frequency band. The components of the 1-3 Hz frequency band are used to determine the heartbeat cycle coefficient, the components of the 3-20 Hz frequency band are used to determine the heartbeat harmonic coefficient, and the information of 20-5000 Hz is used to determine the breathing sound coefficient.
2. The method according to claim 1, characterized in that The spectrum reconstruction model is built based on the UNet model. In the encoder, the feature map undergoes repeated convolution, LeakReLU activation function, and normalization layer, and then a channel attention mechanism is extracted. In the decoder, the data passes through continuous upconvolution blocks, halving the number of feature maps. After each upconvolution, the feature map is merged with the feature map corresponding to the encoder, and then passes through the convolution block and batch normalization. In the last layer, a 1x1 convolution is used to map the feature map to the spectrum of a single output channel.
3. The method according to claim 1, characterized in that The respiratory waveform reconstruction model includes a temporal feature extraction module and a spatial feature extraction module. The temporal feature extraction module includes a bidirectional long short-term memory network and a fully connected layer. The bidirectional long short-term memory network is used to capture the long-term relationship between breathing. In the fully connected layer, the output channel of the bidirectional long short-term memory network is reduced to 1 dimension to obtain the reconstructed respiratory waveform. The spatial feature extraction module includes a convolution layer and a deformable convolution module. The convolution layer is used to extract shallow feature information of breathing sounds. The deformable convolution module is used to extract deep features of breathing sounds. For a given convolution kernel with K sampling positions, the deformable convolution formula is: in, It is The learnable offset of the position, and Respectively represent weights and predefined offsets for each position, and Respectively indicate the position The input feature map and output feature map at .
4. The method according to claim 1, wherein The IMU data preprocessing includes: performing wavelet decomposition or bandpass filtering on the IMU data to filter out motion noise and random noise, and obtain the IMU chest cavity fluctuation waveform.
5. The method according to claim 1, wherein The ear-worn wearable device is an earhook headset, which includes an earhook, an earplug, a main control board, an IMU sensor and a microphone sensor, and the microphone sensor is arranged at the ear inlet of the earplug.
6. The method according to claim 1, characterized in that The ear-worn wearable device is a neck-hanging headset, which includes a neck hanging loop, an earplug connected to the neck hanging loop, a main control board circuit, a battery unit, an IMU sensor and a microphone sensor. The microphone sensor is arranged in the earplug, the main control board circuit is arranged at one end of the neck hanging loop, and the IMU sensor and the battery unit are arranged at the other end of the neck hanging loop.
7. A motion-interference-resistant ear-worn heart rate and respiration real-time monitoring system, comprising an ear-worn wearable device and a terminal device, wherein: The ear-worn wearable device is used to collect audio data and IMU data in the ear canal of a target user and transmit the data to the terminal device; the terminal device performs: extracting heartbeat features and breathing features from the audio data; Inputting the respiratory characteristics into a respiratory waveform reconstruction model to obtain a reconstructed respiratory waveform; The heartbeat features are input into a spectrum reconstruction model to obtain a reconstructed electrocardiogram spectrogram, and the reconstructed electrocardiogram spectrogram is converted from a frequency domain representation to a time domain representation using the Griffin-Lim algorithm to obtain a time domain electrocardiogram waveform, wherein a sliding window is applied to segment the reconstructed electrocardiogram spectrogram and the respiratory waveform, and peak-to-peak detection is used for each window to calculate the time interval between peaks and filter out abnormal values. The heart rate HR and respiratory rate RR of each window are determined by the average value of the time interval after filtering out abnormal values, and then a moving average filter is applied to smooth the heart rate HR and respiratory rate RR curves; the IMU data is preprocessed and amplitude monitoring is performed. If the amplitude exceeds the set value, it is determined that the user is in motion. If the amplitude does not exceed the set value, it is determined that the user is in a resting state. The heart rate RR of the filtered IMU data is calculated and averaged with the heart rate value calculated based on the audio data as the final heart rhythm value to obtain an IMU chest fluctuation waveform; Based on the time-domain electrocardiogram waveform, the reconstructed respiratory waveform, and the IMU chest rise and fall waveform, obtaining estimation results of the respiratory rate and heart rate through peak-to-peak detection; The breathing characteristics and the heartbeat characteristics are obtained according to the following steps: Taking a logarithmic Mel-spectrogram of the audio data to obtain a heartbeat feature of the first channel; Decomposing the audio data by wavelet decomposition to obtain heartbeat cycle coefficients, heartbeat harmonic coefficients and respiratory sound coefficients; For the heartbeat period coefficient and the heartbeat harmonic coefficient, respectively, logarithmic Mel-spectrograms are obtained to obtain a second channel heartbeat feature and a third channel heartbeat feature, and the heartbeat features are combined with the first channel heartbeat feature to form the heartbeat feature; For the breathing sound coefficient, using a pre-emphasized logarithmic Mel-spectrogram to obtain the breathing feature; Decomposing the audio data by wavelet to obtain the heartbeat cycle coefficient, the heartbeat harmonic coefficient, and the breathing sound coefficient comprises the following steps: For the audio data, using empirical wavelet decomposition to obtain data of different frequency bands; The corresponding component coefficient is selected according to the set frequency band. The components of the 1-3 Hz frequency band are used to determine the heartbeat cycle coefficient, the components of the 3-20 Hz frequency band are used to determine the heartbeat harmonic coefficient, and the information of 20-5000 Hz is used to determine the breathing sound coefficient.
8. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Ear-wearing type heart rate monitoring device and method
CN105640532A
Human respiratory rate monitoring device integrated on mask
CN113080932A