Multi-view fusion millimeter wave radar breath and heartbeat monitoring method
Through a millimeter-wave radar monitoring method with multi-view angle fusion, combined with channel attention convolution and cross-domain fusion convolution Transformer module, the problems of signal compression, loss and noise introduction in the prior art are solved, and more accurate heart rate and respiration rate prediction are achieved.
Patent Information
- Application Number
- CN202510274227.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has problems of signal compression, loss or noise introduction when monitoring heart rate and respiration rates using time and frequency domain or IQ components, resulting in feature extraction errors and nonlinear distortion.
The millimeter wave radar monitoring method with multi-view angle fusion is adopted to generate the final fusion characteristics for predicting breathing rate and heart rate through data preprocessing, channel attention convolution module, cross-domain fusion convolution Transformer module and dynamic weights.
It effectively solves the problems of signal compression, loss and noise introduction, improves the prediction accuracy of heart rate and breathing rate, and achieves better vital sign monitoring effects.
Smart Images

Figure CN120078397A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of intelligent perception, non-contact monitoring of millimeter-wave radar, new medical equipment, and artificial intelligence deep learning, and specifically relates to a method for monitoring respiration and heartbeat using millimeter-wave radar with multi-view fusion. Background Art
[0002] Vital signs such as heart rate (HR) and respiratory rate (RR) are crucial for home healthcare. By collecting continuous vital sign data in a personal home environment and monitoring the respiratory rate and heart rate reliably and accurately, cardiovascular diseases can be evaluated and diagnosed. In addition, potential health problems can be detected by continuously detecting changes in HR / RR, and early detection of health deterioration can be facilitated. Moreover, new applications such as fatigue monitoring and emotion recognition rely on vital sign perception. As one of the non-contact monitoring devices, millimeter-wave radar can relieve discomfort, increase the acceptable monitoring time, and eliminate the psychological stress caused by contact measurement devices and measurement errors caused by physical limitations. Moreover, compared with vision-based non-contact monitoring methods, millimeter-wave radar is not affected by lighting conditions, does not generate images, and there is no privacy problem. Therefore, millimeter-wave radar is very suitable for home monitoring.
[0003] In some existing works, for the monitoring of respiratory rate and heart rate, previous studies focused on the amplitude or phase in time-domain and frequency-domain data to predict the respiratory rate and heart rate. However, time-domain and frequency-domain information are low-dimensional space projections of IQ components, and different degrees of distortion will occur, including intensity changes and missing cycles. Therefore, some recent works have started to study from IQ components to extract potential respiration and heart rate characteristics. However, IQ data also has drawbacks. IQ data are two components of the signal obtained by decomposing the original signal, and there may be a mismatch between the I and Q components, resulting in errors in the feature extraction process, and there will also be problems such as more noise causing nonlinear distortion and artifacts.
[0004] Therefore, solving the above defects of using time-domain and frequency-domain or IQ components to better predict vital sign information is an urgent need for the present invention. Summary of the Invention
[0005] Object of the Invention: When traditional methods use time-domain and frequency-domain data to monitor HR and RR, since these are simplified representations of IQ components (in-phase component and quadrature component, which are the real part and imaginary part of complex data respectively), they can be regarded as low-dimensional space projections of the signal. This will cause the phase and some high-frequency information of the signal to be compressed or lost. While directly using IQ components retains more original signal information, it also introduces more noise, and there may be a mismatch between the I and Q components, resulting in errors in the feature extraction process. In order to solve the above defects of using time-domain and frequency-domain or IQ components and better predict the heart rate and respiratory rate.
[0006] Technical solution: To achieve the above object, the technical solution adopted by the present invention is as follows: A method for monitoring respiration and heartbeat by a millimeter-wave radar with multi-view fusion, comprising the following steps: Step 1: Perform data preprocessing on the collected original millimeter-wave radar data to obtain time-domain data, frequency-domain data, real-number data, and imaginary-number data. The processed data is stored in the same file correspondingly for convenient unified reading of the data.
[0007] Step 2: Use the backbone network of the channel attention convolutional module to parallel process the time-domain data, frequency-domain data, real-number data, and imaginary-number data obtained from the preprocessing of the original radar data, extract the effective features of each part, and obtain the rich correlation between each part of the data and respiration and heartbeat.
[0008] Step 3: Pass through the cross-domain fusion convolutional Transformer module, and add a one-dimensional convolutional layer in front of the Transformer encoder as a pre-layer. The pre-layer preprocesses the low-dimensional space features (time-domain features and frequency-domain features) or high-dimensional space features (real-number features and imaginary-number features), and perform pre-fusion and local encoding on the input features of different domains, and then input them into the Transformer encoder.
[0009] Step 4: Since the low-dimensional space features (time-domain features and frequency-domain features) and high-dimensional space features (real-number features and imaginary-number features) are different perspective expressions of the latent space features of respiration and heartbeat, use weight fusion to dynamically adjust the importance of each part to generate the final fusion feature.
[0010] Step 5: Because the features of each modality cannot be effectively aligned in the latent space during fusion, which affects the performance of the model. Use the proposed fusion distance alignment harmonic index loss with dynamic weights, mean square error, and negative Pearson correlation loss for multi-loss joint learning as the optimization function for predicting respiration rate and heart rate.
[0011] Step 6: The final fusion feature is input into the cross-domain fusion convolutional Transformer module again. After average pooling, the pooled result is input into two regression models. By sharing the underlying features, the model simultaneously performs two regression tasks to predict the respiration rate and heart rate.
[0012] Step 1 specifically includes: First, for the original millimeter-wave radar data r(m, n, k), where m is the number of frames, each frame has 4 chirps, and the average value of the chirps is taken in each frame. n is the fast-time index, and k is the antenna index. Taking 1 in the k dimension means only taking the data of one of the antennas. Then, perform Range-FFT on the second dimension of the fast-time index dimension of r(m, n) to obtain the Range matrix R(m,n) and the value of the maximum range bin. In this way, the position of the human chest is obtained. Then, use the value of the obtained maximum range bin to sum this column on n and slide the time window on m, and organize the data into a matrix of (15,3750). The matrix uses the phase demodulation algorithm to extract the phase sequence in the distance unit where the chest is located, and the time-domain data X of the respiratory and heartbeat signals can be observed. T , and then use the fast Fourier transform on the basis of the time-domain data to obtain the frequency data of the phase data change. X F . At the same time, apply np.real and np.imag to the reorganized matrix to separate the real part X R and the imaginary part X I . Finally, store the processed time-domain data, frequency-domain data, real part data, and imaginary part data in the same file for convenient subsequent reading.
[0013] Step 2 specifically includes: Parallelly process the time-domain data X T , the frequency-domain data X F , the real part data X R , and the imaginary part data X I through the channel attention convolution module. Each CNN layer in the module is followed by an SE-Block. The Squeeze-and-Excitation (SE) block compresses the spatial dimension through the adaptive average pooling layer, then uses two fully connected networks for feature recalibration and dimension recovery, and finally outputs the channel weights for adjusting the original features through the Sigmoid function.
[0014] Step 3 specifically includes: It is characterized in that it combines the advantages of convolutional neural networks and Transformer technology through the cross-domain fusion convolutional Transformer module, and performs in-depth feature encoding on the input data through the cross-domain fusion convolution and Transformer encoding strategy. The low-dimensional space features or high-dimensional space features after the cat process are input into the channel fusion layer. The Transformer encoder receives the features after the channel fusion layer, uses the multi-head attention mechanism (n_head is 8) to capture the long-range dependencies in the data, and outputs the attention-weighted features , 。
[0015] Step 4 specifically includes: , , which is the expression of different perspectives of the potential space features of RR and HR. Therefore, we can use the weight fusion layer to dynamically adjust , importance. The weight fusion layer is normalized by F.softmax to obtain , , the fusion weights of the two groups of features, and the two groups of features are weighted and summed using the weights to generate the final fused features 。
[0016] Step 5 specifically includes: After passing through the cross-domain fusion convolutional Transformer module, low-dimensional space features and high-dimensional space features are obtained. There is a problem. Although the self-attention mechanism can capture the complex relationships between different inputs within low-dimensional or high-dimensional features, there are different feature distributions and ranges between the two features. The obtained and , directly input into the weight fusion module may cause the features of each modality to fail to be effectively aligned in the potential space during fusion, thus affecting the performance of the model. Therefore, the 2-Wasserstein distance between them is selected to adjust their respective potential spaces. There are two reasons for using the Wasserstein distance. On the one hand, minimizing the distance will make the two distributions close, which can reduce the impact of distance misalignment. On the other hand, the Wasserstein distance can provide useful gradients. Calculate the mean and variance of the features of each channel in each batch, and then obtain the fusion alignment distance according to formula (1).
[0017] (1) where ∥·∥Frob is the Frobenius norm, defined as the square root of the sum of the squares of the absolute values of the matrix elements.
[0018] Then, the original alignment distance is also calculated using the above formula with the original time-domain and frequency-domain data for auxiliary enhancement , to help the potential space alignment of the fused features. These two distances are input into the following formula (2) (2) where and is the output calculated according to (1), α and β, with initial weights of 2, 2, which can be updated during backpropagation. The two distances are transformed through an exponential function, where the negative sign ensures that as the distance increases, its influence decreases. The harmonic form numerically favors smaller distances, helping to reduce the impact of larger outliers on the overall metric.
[0019] Overall Loss: By combining three loss functions, the overall loss function (3) is obtained.
[0020] (3) where L M is the mean squared error, and L NP is the negative Pearson correlation loss, and the weight γ is set to 5.
[0021] Step 6 specifically includes: The final fused feature is input into the cross-domain fusion convolutional Transformer module again, and the input data is downsampled through one-dimensional average pooling to reduce its dimension. Then, the pooled results are respectively passed to two regression models for different prediction tasks, and finally the respiratory rate and heart rate are obtained as the prediction results.
[0022] Beneficial effects: 1) The present invention proposes a novel end-to-end cross-space multi-view fusion model, which combines multi-channel convolutional attention and cross-space multi-layer fusion modules to fuse low-dimensional space projections (time domain and frequency domain) and high-dimensional space vectors (real part and imaginary part). This is the first work to combine the real part, imaginary part, time domain, and frequency domain.
[0023] 2) For multi-view fusion, the present invention proposes a fusion distance alignment harmonic exponential loss with dynamic weights, and the performance is improved through multi-loss joint learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is the system framework diagram of the method of the present invention.
[0025] Figure 2 is the system flow chart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the above objects, features, and advantages of the present invention more obvious and understandable. The technical solutions in the present invention will be further described in detail below with reference to the accompanying drawings: The specific implementation system framework of the method of the present invention is as Figure 1 shown, and the process is as follows: Step 1) First, for the original millimeter-wave radar data r(m, n, k), where m is the number of frames, each frame has 4 chirps, and the average value of the chirps is taken in each frame. n is the fast-time index, and k is the antenna index. Taking 1 in the k dimension means only taking the data of one of the antennas. Then, perform Range-FFT on the second dimension of the fast-time index dimension of r(m, n) to obtain the Range matrix R(m,n) and the value of the maximum range bin. In this way, the position of the human chest is obtained. Then, use the obtained value of the maximum range bin to take the sum of this column in n and slide the time window in m, and organize the data into a matrix of (15,3750). The matrix uses the phase demodulation algorithm to extract the phase sequence in the distance cell where the chest is located, and the time-domain data X of the respiratory and heartbeat signals can be observed. T , and then use the fast Fourier transform on the basis of the time-domain data to obtain the frequency data of the phase data change. X F . At the same time, apply np.real and np.imag to the reorganized matrix to separate the real part X R and the imaginary part X I . Finally, store the processed time-domain data, frequency-domain data, real part data, and imaginary part data in the same file for convenient subsequent reading.
[0027] Step 2) The backbone network integrated with the channel attention convolution module predicts the respiratory rate and heart rate in parallel. Each backbone network integrated with the channel attention convolution module processes the preprocessed time-domain data X T , frequency-domain data X F , real number data X R , and imaginary number data X I in parallel. Among them, X T ∈ R W*M , X F ∈ R W*N , X R ∈ R W*E , X I ∈ R W*J , W is the window length of 15, M, N, E, J are 3750, 1875, 3750, 3750 respectively, representing the number of samples in the window length. Since the frequency-domain data is obtained by the fast Fourier transform of the time-domain data, and because the time-domain data is a real input signal and the negative frequency component is the complex conjugate of the positive frequency component, the length N of the frequency-domain data takes half of the length of the input data. In this module, an SE-Block follows each CNN layer. The Squeeze-and-Excitation (SE) block compresses the spatial dimension through the adaptive average pooling layer, then uses two fully connected networks for feature recalibration and dimension restoration, and finally outputs the channel weights for adjusting the original features through the Sigmoid function.
[0028] Step 3) The cross-domain fusion convolutional Transformer module combines the advantages of convolutional neural networks and Transformer technology. Through cross-domain convolution and Transformer encoding strategies, in-depth feature encoding is performed on the input data. The low-dimensional space features after the cat processing or high-dimensional space features , are input into the channel fusion layer, which is a one-dimensional convolutional layer (kernel size is 1, stride is 1). This layer does not change the spatial dimension of the data and only performs a linear transformation at the channel level to optimize the feature representation between channels. The Transformer encoder receives the features after the channel fusion layer and uses the multi-head attention mechanism (n_head is 8) to capture the long-range dependencies in the data and outputs the attention-weighted features , .
[0029] Step 4) , , which are the expressions of different perspectives of the RR and HR latent space features. Therefore, we can use the weight fusion layer to dynamically adjust , the importance. The weight fusion layer is normalized using F.softmax to obtain , , the fusion weights of the two groups of features. The weights are used to perform weighted summation on the two groups of features to generate the final fusion features .
[0030] Step 5) Loss function design Reconstruction loss: To accurately predict the respiratory rate and heart rate, we use the common regression loss function L2-norm to calculate the mean square error between the predicted value and the label value as the reconstruction loss L M .
[0031] (4) where y pred represents the predicted value, y gt represents the label value, and i represents the number of samples Negative Pearson correlation coefficient related loss: In addition, the negative Pearson correlation loss L NP is used to constrain the prediction and reduce the error between the predicted value and the label value.
[0032] (5) where μ represents taking the average and σ represents the standard deviation SD Fusion distance alignment harmonic index loss of dynamic weights: After passing through the cross-domain fusion convolutional Transformer module, low-dimensional spatial features are obtained. and high-dimensional spatial features After that, there is a problem. Although the self-attention mechanism can capture the complex relationships between different inputs within low-dimensional or high-dimensional features, there are different feature distributions and ranges between the two features. The obtained and , when directly input into the weight fusion module, it may cause the features of each modality to fail to be effectively aligned in the latent space during fusion, thus affecting the performance of the model. To solve this problem, we choose to minimize the 2-Wasserstein distance between them to adjust their respective latent spaces. There are two reasons for using the Wasserstein distance. On the one hand, minimizing the distance will make the two distributions close, which can reduce the impact of distance misalignment. On the other hand, the Wasserstein distance can provide useful gradients. We calculate the mean and variance of the features for each channel in each batch, and then obtain the fusion alignment distance according to formula (6) with the obtained mean and variance .
[0033] (6) where ∥·∥Frob is the Frobenius norm, defined as the square root of the sum of the squares of the absolute values of the matrix elements.
[0034] Then we also calculate the original alignment distance using the above formula with the original time-domain and frequency-domain data Auxiliary enhancement , to help align the latent spaces of the fused features. These two distances are input into the following formula (7) (7) where and are the outputs calculated according to (3), and α and β, with the initial weights being 2, 2, which can be updated during backpropagation. The two distances are transformed through the exponential function, where the negative sign ensures that when the distance increases, its impact decreases. The harmonic form makes the numerical value tend to the smaller distance, which helps to reduce the impact of larger outliers on the overall metric.
[0035] Overall Loss: By combining three loss functions, the overall loss function (8) is obtained.
[0036] (8) where, L M is the mean squared error, L NP is the negative Pearson correlation loss, and the weight γ is set to 5.
[0037] Step 6) The final fused features are input into the cross-domain fusion convolutional Transformer module again. One-dimensional average pooling is used to downsample the input data to reduce its dimension. Then, the pooled results are respectively passed to two regression models for different prediction tasks, and finally the respiratory rate and heart rate are obtained as the prediction results.
[0038] A novel end-to-end cross-space multi-view fusion model is proposed, which combines channel attention convolution and cross-space feature multi-layer fusion module to fuse low-dimensional space projections (time domain and frequency domain) and high-dimensional space vectors (real part and imaginary part). This is the first work to combine the real part, imaginary part, time domain and frequency domain. In addition, aiming at the misalignment of multi-view fusion latent space features, a new fusion distance alignment harmonic index loss with dynamic weights is proposed. The multi-view fusion model can better complete the vital sign monitoring task because it complements each other's information.
Claims
1. A multi-view fusion millimeter wave radar method for monitoring breathing and heartbeat, characterized by: The following steps are involved: Step 1) Preprocess the collected millimeter-wave radar raw data to obtain time domain data, frequency domain data, real data, and imaginary data. The processed data are stored in the same file to facilitate unified data reading; Step 2) Use the backbone network of the channel attention convolution module to parallelly process the time domain data, frequency domain data, real data, and imaginary data obtained from the radar raw data preprocessing, extract the effective features of each part, and obtain the rich correlation between each part of the data and breathing and heartbeat; Step 3) By cross-domain fusion convolutional Transformer modules, a one-dimensional convolutional layer is added before the Transformer encoder as a pre-layer; The front layer preprocesses the low-dimensional spatial features after concat processing Or high-dimensional space features , and pre-fuse and locally encode the input features of different domains, and then input them into the Transformer encoder; Step 4) Since the low-dimensional space features and high-dimensional space features are expressions of the breathing and heartbeat latent space features from different perspectives, the importance of each part is dynamically adjusted by weight fusion to generate the final fusion feature; Step 5) Since the features of each modality are not effectively aligned in the latent space during fusion, which affects the performance of the model, the proposed dynamic weighted fusion distance alignment harmonic index loss, mean square error and negative Pearson correlation loss are used to perform multi-loss joint learning as the optimization function for predicting respiratory rate and heart rate; Step 6) The final fused features are input into the cross-domain fusion convolutional Transformer module again, and then after average pooling, the pooled results are input into two regression models. By sharing the underlying features, the model performs two regression tasks at the same time to predict the respiratory rate and heart rate.
2. The method for monitoring breathing and heartbeat using a multi-view fusion millimeter-wave radar according to claim 1, characterized in that: First, the millimeter-wave radar raw data r(m, n, k), where m is the number of frames, each frame has 4 chirps, the average number of chirps is taken in each frame, n is the fast time index, k is the antenna index, and 1 is taken in the k dimension, indicating that only the data of one antenna is taken. Then, the second dimension of the fast time index dimension of r(m, n) is Range-FFT to obtain the Range matrix R(m, n) and the value of the maximum range bin; thereby obtaining the position of the human chest cavity; then, using the obtained maximum range bin value, take this column on n and slide the time window on m, organize the data into a (15,3750) matrix, and use the phase demodulation algorithm to extract the phase sequence in the distance unit where the chest is located, so that the time domain data X of the respiratory heartbeat signal can be observed. T , and then use fast Fourier transform to obtain the frequency data X of phase data change based on the time domain data F ; Apply np.real and np.imag to the reorganized matrix to separate the real part X R and the imaginary part X I ; Finally, the processed time domain data, frequency domain data, real data, and imaginary data are processed for subsequent reading.
3. The method for monitoring breathing and heartbeat by multi-view fusion millimeter wave radar according to claim 2 is characterized in that: Parallel processing of time domain data X through channel attention convolution modules T , frequency domain data X F , real data X R , imaginary data X I ; By setting independent input paths for the time domain, frequency domain, and real and imaginary parts of complex signals, deep joint learning of different dimensions of signals is achieved. The multi-input design can effectively extract the independent features of each signal pattern, and the signal of each input path passes through multiple layers of convolutional layers, followed by a Squeeze-and-Excitation (SE) block, which realizes the alternating stacking of convolutional layers and SE blocks; the SE block obtains the global information of each channel through global average pooling, and then generates the adaptive weight of the channel through the fully connected layer to dynamically adjust the attention of each channel.
4. The method for monitoring breathing and heartbeat using a multi-view fusion millimeter-wave radar according to claim 3, characterized in that: The cross-domain fusion convolutional Transformer module combines the advantages of convolutional neural networks and Transformer technology. Through the cross-domain fusion convolution and Transformer encoding strategy, the input data is deeply encoded with features; the low-dimensional spatial features after cat processing Or high-dimensional space features , input channel fusion layer; Transformer encoder receives the features after the channel fusion layer, uses the multi-head attention mechanism (n_head is 8), captures the long-distance dependencies in the data, and outputs the features weighted by attention. , .
5. The method for monitoring breathing and heartbeat using a multi-view fusion millimeter-wave radar according to claim 4, characterized in that: , , is the expression of different perspectives of the RR and HR latent space features, so we can use the weight fusion layer to dynamically adjust , importance; The weight fusion layer is normalized with F.softmax to obtain , ,The fusion weight of the two sets of features, use the weight to perform weighted summation on the two sets of features to generate the final fusion feature .
6. The method for monitoring breathing and heartbeat using a multi-view fusion millimeter-wave radar according to claim 5, characterized in that: After cross-domain fusion convolution modules, low-dimensional spatial features are obtained and high-dimensional space features After that, there is a problem. Although the self-attention mechanism can capture the complex relationship between different inputs within low-dimensional or high-dimensional features, the two features have different feature distributions and ranges. and , directly input into the weight fusion module, may lead to the failure of effective alignment of the features of each modality in the latent space during fusion, thus affecting the performance of the model; therefore, we choose to minimize the 2- Wasserstein distance between them to adjust their respective latent spaces; there are two reasons for using Wasserstein distance. On the one hand, minimizing the distance will make the two distributions close, which can reduce the impact of distance misalignment. On the other hand, Wasserstein distance can provide useful gradients; calculation The feature mean and variance of each channel in each batch, and then the obtained mean and variance are used to obtain the fusion alignment distance according to formula (1): ; (1) where ∥·∥Frob is the Frobenius norm, defined as the square root of the sum of the squares of the absolute values of the matrix elements; Then we use the original time domain and frequency domain data to calculate the original alignment distance using the above formula Auxiliary Enhancement , helps to align the fused feature potential space, and these two distances are input into the following formula (2) (2) in and are the outputs calculated according to (1), α and β, with initial weights of 2,2, which can be updated in the back propagation; the two distances are transformed by an exponential function, where the negative sign ensures that their influence decreases as the distance increases; the harmonic form makes the value tend to be smaller, which helps to reduce the impact of larger outliers on the overall metric; Overall Loss: By combining the three loss functions, we get the overall loss function (3); (3) Among them, L M is the mean square error, L NP is the negative Pearson correlation loss and the weight γ is set to 5.
7. The method for monitoring breathing and heartbeat using a multi-view fusion millimeter-wave radar according to claim 6, characterized in that: The final fused features are input into the cross-domain fusion convolutional Transformer module again, and the input data is downsampled through one-dimensional average pooling to reduce its dimension. The pooled results are then passed to two regression models for different prediction tasks, and finally the respiratory rate and heart rate are obtained as the prediction results.
Citation Information
Cited By
Attention-based thoracico-abdominal movement signal segmentation method and system and storage medium
CN120899237A