Non-contact real-time pressure detection method and system based on millimeter wave radar
By adopting one-dimensional and two-dimensional neural network fusion in millimeter-wave radar-based pressure detection technology, reconstructing breathing and heartbeat signals and extracting multi-time scale features, the problem of susceptibility to interference and single feature extraction in the prior art is solved, achieving higher detection accuracy and reliability.
Patent Information
- Application Number
- CN202510348478.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing pressure detection technology based on millimeter wave radar is susceptible to environmental noise and subject's limb movement, and the feature extraction is single, resulting in limited detection accuracy.
A non-contact real-time pressure detection method based on millimeter wave radar is adopted to obtain millimeter wave radar echo signals in real time, reconstruct breathing and heartbeat signals, and use one-dimensional and two-dimensional neural network fusion to extract multi-time scale features for real-time pressure detection.
It improves the accuracy of heart rate estimation, enhances the accuracy and reliability of pressure detection, and can detect pressure states more conveniently, objectively, reliably and timely.
Smart Images

Figure CN120189119A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of information technology, and more specifically, relates to a non-contact real-time pressure detection method and system based on millimeter-wave radar. Background Art
[0002] Pressure is one of the fundamental problems in today's society. Factors such as high workload and the variability of modern life severely affect people's mental health, causing various negative impacts. At the same time, pressure is also one of the main factors leading to various diseases, such as: depression, stroke, sleep disorders, heart attacks, and cardiac arrests. Therefore, it is particularly important to accurately detect pressure at its initial stage.
[0003] Due to the negative impact of pressure on health and the widespread demand for pressure detection, inventions related to pressure detection have attracted extensive attention from engineers and psychologists. Most of the existing pressure detection methods are based on audio-video processing technology, that is: by extracting speech features, face image features (such as: facial expressions and pupil diameters, etc.), gesture features, etc. in audio-video to judge the pressure state of an individual. Although this technology has reached a relatively high recognition accuracy, there are still many problems. On the one hand, it is easily affected by objective factors such as dim light, uneven illumination, facial deflection and tilt, resulting in detection errors; on the other hand, subjective factors such as the disguise of the test subject (such as: covering up their true state by adjusting intonation, speech rate or concealing facial expressions) will also cause great interference to the test results. In addition, this technology also faces privacy protection issues.
[0004] At the same time, due to the advantages of physiological signals such as uniqueness and stability, pressure state detection based on physiological signals has been continuously developed and has gradually become an invention hotspot in the academic and industrial circles in recent years. Generally speaking, there are two general methods for pressure detection based on physiological signals: (1) Contact detection based on wearable devices (such as: electrocardiogram ECG, electromyogram EMG, electroencephalogram EEG, etc.): This method has been widely used due to the convenience of the device, but there are also certain problems. For example: long-term wearing may cause skin discomfort and is not suitable for damaged skin (such as: painful rashes, burns, urticaria, etc.). In addition, it is impossible to distinguish whether the source of pressure is device discomfort or psychological factors; (2) Non-contact detection based on millimeter-wave radar: This method monitors the millimeter-wave radar echo signal reflected by the human body, extracts vital sign signals such as breathing and heartbeat, and thus conducts pressure detection in a non-contact and non-intrusive manner. This method has the advantages of low power consumption, low cost, and easy use, but there are also some problems: in terms of signal acquisition, millimeter-wave radar is easily interfered by factors such as environmental noise and the limb movements of the test subject, resulting in poor signal quality; in terms of pressure detection, existing methods are mostly limited to traditional feature extraction and do not pay attention to the application of deep learning features, resulting in limitations in the real-time performance, robustness, and transferability of the algorithm. Summary of the Invention
[0005] Aiming at the defects of the existing technology, the purpose of this application is to provide a non-contact real-time pressure detection method and system based on millimeter-wave radar, aiming to solve the problems faced by the existing pressure detection technology based on millimeter-wave radar, such as being vulnerable to factors such as environmental noise and the movement of the subject's limbs, and the single feature extraction, resulting in limited detection accuracy.
[0006] To achieve the above purpose, in the first aspect, this application provides a non-contact real-time pressure detection method based on millimeter-wave radar, including: Real-time obtain the echo signal of the millimeter-wave radar reflected from the user to be measured, and reconstruct the respiration signal and heartbeat signal of the user to be measured based on the echo signal, and obtain the estimated value of the respiration signal; Input the reconstructed heartbeat signal into a one-dimensional neural network to obtain the first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrogram, input the CWT spectrogram into a YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram; then brighten the spectrum of the marked frequency band region, and input the brightened CWT spectrogram into a two-dimensional convolutional neural network to obtain the second estimated value of the heartbeat signal; determine the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; Extract the multi-time scale features of the respiration and heartbeat of the user to be measured based on the reconstructed respiration signal and heartbeat signal; Perform real-time pressure detection on the user to be measured based on the multi-time scale features of respiration and heartbeat, the estimated value of the respiration signal, and the final estimated value of the heartbeat signal.
[0007] It can be understood that this application combines the powerful positioning ability of YOLO to accurately locate the frequency band most relevant to the heart rate. The spectrum brightening enhances the spectrum located by YOLO (both retains the effective information and appropriately weakens the out-of-band noise). Therefore, after the enhanced spectrum is input into the two-dimensional convolutional neural network, the accuracy of heart rate estimation can be effectively improved, thus ensuring the accuracy of the above pressure detection scheme.
[0008] In a possible implementation manner, the one-dimensional neural network includes: a plurality of convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block, and a Dense connection block; A plurality of convolutional neural network backbone blocks are used to capture the local features of the reconstructed heartbeat signal at different scales through different convolutional kernels; Pyramid input module, which is used to perform feature learning on signals input to each convolutional neural network backbone block at different scales, and then combine the signals obtained from the feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block can capture local features at corresponding scales; Transformer encoding block, which is used to combine the multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by multiple convolutional neural networks at different scales, and capture the features of the most representative heartbeat signals in the local features at different scales; Dense connection block, which is used to combine the output signals of the Transformer encoding block to obtain the first estimated value of the heartbeat signal.
[0009] It should be noted that the one-dimensional neural network adopted in this application uses multiple convolutional neural network backbone blocks to extract local features of the heartbeat signal, uses the pyramid structure to extract multi-scale features, and uses Transformer to extract global features, effectively making up for the deficiencies of traditional convolutional neural networks in multi-scale and long-range dependence modeling, and can greatly improve the accuracy of the first estimated value of the heartbeat signal to improve the accuracy and reliability of pressure detection.
[0010] In a possible implementation, reconstruct the respiratory signal and heartbeat signal of the user to be measured based on the echo signal, and obtain the estimated value of the respiratory signal, including: Extract the corresponding beat signal from the echo signal; the beat signal carries the respiratory signal and heartbeat signal of the user to be measured; Based on the normalized least mean squares (NLMS) algorithm of polynomial fitting, preliminarily suppress the motion artifacts of the beat signal; Perform multi-level filtering on the beat signal after preliminarily suppressing motion artifacts using continuous wavelet transform (CWT); the multi-level filtering includes: performing first-level band-pass filtering on the beat signal in the wavelet domain to obtain the first-level filtered respiratory signal and heartbeat signal, and then combining the peak-seeking and valley-seeking algorithms and spectrum evaluation to obtain the first-level respiratory signal estimation and heartbeat signal estimation; performing second-level band-pass filtering on the first-level filtered respiratory signal and heartbeat signal to obtain the second-level filtered respiratory signal and heartbeat signal, and then combining the peak-seeking and valley-seeking algorithms and spectrum evaluation to obtain the second-level respiratory signal estimation and heartbeat signal estimation; performing third-level band-pass filtering on the second-level filtered heartbeat signal to obtain the third-level filtered heartbeat signal, and then combining the peak-seeking and valley-seeking algorithms and spectrum evaluation to obtain the third-level heartbeat signal estimation; among them, the center frequency of the next-level band-pass filtering is set with reference to the respiratory signal estimation and heartbeat signal estimation of the previous level; Take the respiration signal obtained after secondary filtering and the heartbeat signal obtained after tertiary filtering as the respiration signal and heartbeat signal of the user to be measured for reconstruction.
[0011] Take the estimated respiration signal of the second stage obtained during the process of reconstructing the respiration signal of the user to be measured as the estimated value of the respiration signal.
[0012] In a possible implementation, when the three-axis acceleration data of the user to be measured can be obtained in real time, the multi-stage filtering performed for reconstructing the heartbeat signal of the user to be measured based on the echo signal is as follows: perform first-stage band-pass filtering on the beat signal in the wavelet domain to obtain the heartbeat signal after the first-stage filtering, and then combine the peak and valley finding algorithm and spectrum evaluation to obtain the first-stage estimated heartbeat signal; perform second-stage band-pass filtering on the heartbeat signal after the first-stage filtering to obtain the heartbeat signal after the second-stage filtering, and then combine the three-axis acceleration data of the user to be measured and use the photoplethysmogram Q-PPG deep learning algorithm to estimate the heart rate of the heartbeat signal after the second-stage filtering to obtain the second-stage estimated heartbeat signal; perform third-stage band-pass filtering on the heartbeat signal after the second-stage filtering to obtain the heartbeat signal after the third-stage filtering, and then combine the peak and valley finding algorithm and spectrum evaluation to obtain the third-stage estimated heartbeat signal; take the heartbeat signal obtained after the third-stage filtering as the heartbeat signal of the user to be measured for reconstruction.
[0013] In a possible implementation, the method further includes: Convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrogram, and input the CWT spectrogram into the YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram, so as to obtain the third estimated value of the heartbeat signal; Determine the final estimated value of the heartbeat signal based on the multi-stage estimated heartbeat signals, the first estimated value of the heartbeat signal, the second estimated value of the heartbeat signal, and the third estimated value of the heartbeat signal obtained during the multi-stage filtering process.
[0014] In a possible implementation, extract the multi-time scale features of the respiration and heartbeat of the user to be measured, including: Based on the reconstructed respiration signal and heartbeat signal, use the first deep learning network to extract the respiration features and heartbeat features of the user to be measured under the first time span based on the detection window of the first time span; Based on the reconstructed respiration signal and heartbeat signal, use the second deep learning network to extract the respiration features and heartbeat features of the user to be measured under the second time span based on the detection window of the second time span; Take the respiration features and heartbeat features of the user to be measured under different time spans as the multi-time scale features.
[0015] In a possible implementation, the first deep learning network includes: one-dimensional convolution, a causal analyzer, dilated convolution, and residual connection; the one-dimensional convolution is used to capture the time-domain change characteristics of the time series data corresponding to the breathing signal and the heartbeat signal within the first time span; the causal analyzer is used to limit the convolution process to the past and present data to ensure that the prediction result is consistent with the real scenario; the dilated convolution is used to assist in downsampling and provide a larger receptive field; the residual connection speeds up the convergence rate by allowing gradient propagation without passing through an activation function to extract the breathing characteristics and heart rate characteristics within the corresponding time span; The second deep learning network is a self-supervised learning network, which is used to learn and extract the spatio-temporal characteristics of the breathing signal and the heartbeat signal within the corresponding time span.
[0016] In a second aspect, the present application provides a non-contact real-time pressure detection system based on a millimeter-wave radar, including: A signal reconstruction module, which is used to obtain in real time the echo signal of the millimeter-wave radar reflected from the user to be measured, and reconstruct the breathing signal and the heartbeat signal of the user to be measured based on the echo signal, and obtain an estimated value of the breathing signal; A heart rate final estimation module, which is used to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrogram, input the CWT spectrogram into a YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram; then brighten the marked frequency band region, and input the brightened CWT spectrogram into a two-dimensional convolutional neural network to obtain a second estimated value of the heartbeat signal; determine the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; A pressure detection module, which is used to extract the multi-time scale characteristics of the breathing and heartbeat of the user to be measured based on the reconstructed breathing signal and heartbeat signal; and perform real-time pressure detection on the user to be measured based on the multi-time scale characteristics of the breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.
[0017] In a possible implementation, the one-dimensional neural network used by the heartbeat final estimation module includes: a plurality of convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block, and a Dense connection block; the plurality of convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales through different convolutional kernels; the pyramid input module is used to perform feature learning on signals input to each convolutional neural network backbone block at different scales, and then combine the signals obtained by feature learning with each convolutional neural network backbone block, so that each convolutional neural network backbone block can capture local features at corresponding scales; the Transformer encoding block is used to combine the multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the plurality of convolutional neural networks at different scales, and capture the features of the most representative heartbeat signals in the local features at different scales; the Dense connection block is used to combine the output signals of the Transformer encoding block to obtain the first estimated value of the heartbeat signal.
[0018] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0020] In a fifth aspect, the present application provides a computer program product, and when the computer program product runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0021] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0022] Generally speaking, compared with the prior art, the above technical solution conceived by the present application has the following beneficial effects: The present application provides a non-contact real-time pressure detection method and system based on millimeter-wave radar, which combines the advantages of non-intrusive detection through audio and video analysis and physiological signal detection by wearable devices, and can detect the pressure state more conveniently, objectively, reliably and timely. In terms of the detection method, first, clutter suppression and echo selection are performed on the millimeter-wave echo signal, and then the original beat signal of the user to be detected is extracted. Secondly, motion artifact suppression is performed on the beat signal, respiration and heart rate estimation are carried out, and one-dimensional and two-dimensional neural network fusion is used to further improve the accuracy of heart rate estimation. Finally, respiration and heart rate signals are fused, and at the same time, a multi-scale time detection window is used. Based on a deep learning network, pressure detection is performed through a classifier. As a non-contact real-time pressure detection method, it has the advantages of flexible and diverse detection devices, objective and accurate detection results, and strong real-time pressure detection ability. It has the advantages of strong flexibility and good real-time performance, makes up for the defects of existing physiological signal-based technologies, improves the robustness and accuracy of detection, and provides technical support for the wide application of pressure detection in real life.
[0023] The present application provides a non-contact real-time pressure detection method and system based on millimeter-wave radar. When using one-dimensional and two-dimensional neural network fusion, first, in terms of the one-dimensional neural network, multiple convolutional neural network backbone blocks are used to extract local features of the heart rate signal, a pyramid structure is used to extract multi-scale features, and a Transformer is used to extract global features, effectively making up for the deficiencies of traditional convolutional neural networks in multi-scale and long-range dependence modeling, and can greatly improve the accuracy of the first estimated value of the heart rate signal. Secondly, combined with the powerful positioning ability of YOLO, the frequency band most relevant to the heart rate can be accurately located, and the spectrum brightening enhances the spectrum located by YOLO (both retaining effective information and appropriately weakening out-of-band noise). Therefore, after the enhanced spectrum is input into the two-dimensional convolutional neural network, the accuracy of heart rate estimation can be effectively improved. Therefore, the present application combines one-dimensional and two-dimensional neural network fusion to improve the accuracy of heart rate estimation, thereby ensuring the accuracy of pressure detection and the objective accuracy of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the non-contact real-time pressure detection method based on millimeter-wave radar provided by the embodiment of the present application; Figure 2 It is a block diagram of the non-contact real-time pressure detection based on millimeter-wave radar provided by the embodiment of the present application; Figure 3 It is a flowchart of the non-contact real-time pressure detection based on millimeter-wave radar provided by the embodiment of the present application; Figure 4 It is a multi-process real-time detection framework diagram based on ZeroMQ provided by the embodiment of the present application; Figure 5Flowchart of beat signal extraction for millimeter-wave radar provided by an embodiment of this application; Figure 6 Flowchart of the NLMS algorithm based on polynomial fitting provided by an embodiment of this application; Figure 7 Flowchart of the Q-PPG algorithm provided by an embodiment of this application; Figure 8 Overall architecture diagram of the neural network fusion estimation algorithm for one-dimensional (1D) waveform estimation and two-dimensional (2D) spectrum estimation provided by an embodiment of this application; Figure 9 Architecture diagram of the YOLOv8 pre-trained model provided by an embodiment of this application; Figure 10 Architecture diagram of the DenseNet-201 network provided by an embodiment of this application; Figure 11 Flowchart of the residual-temporal convolution network (Res-TCN) provided by an embodiment of this application; Figure 12 Flowchart of the self-supervised learning (SSL) network provided by an embodiment of this application; Figure 13 Visualization interface diagram of real-time pressure detection provided by an embodiment of this application; Figure 14 Architecture diagram of the non-contact real-time pressure detection system based on millimeter-wave radar provided by an embodiment of this application; Figure 15 Architecture diagram of the electronic device provided by an embodiment of this application. Detailed implementation manners
[0025] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application, and are not used to limit this application.
[0026] The terms "first" and "second" in the specification and claims of this document are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first estimated value and the second estimated value are used to distinguish different estimated values, rather than to describe the specific order of the estimated values.
[0027] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0028] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0029] Aiming at the deficiencies of the prior art, the purpose of the present application is to provide a non-contact real-time pressure detection method and system based on a millimeter-wave radar. It aims to solve the interference problems such as lighting factors and human camouflage faced by the existing audio-visual-based pressure detection technology, as well as the problems such as being vulnerable to environmental noise, the movement of the subject's limbs, and single feature extraction faced by the millimeter-wave radar-based pressure detection technology, so as to improve the pressure detection accuracy.
[0030] Exemplarily, the present application provides a non-contact real-time pressure detection method based on a millimeter-wave radar, which has the advantages of high reliability, strong robustness, low power, and strong real-time performance. As Figure 1 shown, it includes the following steps: Step S101, continuously obtain the echo signal of the millimeter-wave radar reflected from the user to be measured, and reconstruct the breathing signal and heartbeat signal of the user to be measured based on the echo signal, and obtain an estimated value of the breathing signal; Step S102, input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, input the CWT spectrogram into a YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram; then brighten the spectrum of the marked frequency band region, and input the brightened CWT spectrogram into a two-dimensional convolutional neural network to obtain a second estimated value of the heartbeat signal; determine the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; In one example, the one-dimensional neural network includes: a plurality of convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block, and a Dense connection block; The plurality of convolutional neural network backbone blocks are used to capture local features of the preliminarily estimated heartbeat signal at different scales through different convolutional kernels; The pyramid input module is used to perform feature learning at different scales on the signals input to each convolutional neural network backbone block, and then combine the signals obtained by feature learning with each convolutional neural network backbone block, so that each convolutional neural network backbone block can capture local features at corresponding scales; The Transformer encoding block is used to combine the multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the multiple convolutional neural networks at different scales, and capture the features of the most representative heartbeat signals in the local features at different scales; The Dense connection block is used to obtain the first estimated value of the heartbeat signal by combining the output signal of the Transformer encoding block.
[0031] It can be understood that the one-dimensional neural network at this time is slightly different from the traditional one-dimensional convolutional neural network. The one-dimensional neural network adopted in this application uses multiple convolutional neural network backbone blocks to extract local features, uses a pyramid structure to extract multi-scale features, and uses a Transformer to extract global features, effectively making up for the deficiencies of the traditional convolutional neural network in multi-scale and long-range dependence modeling, and can greatly improve the accuracy of the first estimated value of the heartbeat signal to improve the accuracy and reliability of pressure detection.
[0032] In one embodiment, the above YOLO model can select the YOLOv8 model. The above two-dimensional convolutional neural network can select the DenseNet-201 network.
[0033] In another example, the method further includes the following steps: converting the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, and inputting the CWT spectrogram into the YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram, so as to obtain the third estimated value of the heartbeat signal.
[0034] Further, determining the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal includes: Determining the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal, the second estimated value of the heartbeat signal, and the third estimated value of the heartbeat signal.
[0035] It should be noted that due to the powerful localization ability of YOLO, it can accurately locate the frequency band most relevant to the heart rate. The spectrum brightening enhances the spectrum located by YOLO (both preserving the effective information and appropriately weakening the out-of-band noise). Therefore, after the enhanced spectrum is input into the two-dimensional convolutional neural network, the accuracy of heart rate estimation is significantly improved. Thus, the influence of the above second estimated value on the final estimated accuracy of the entire heartbeat signal is more critical than the influence of the third estimated value on the final estimated accuracy of the entire heartbeat signal. Through subsequent experimental comparisons in this application, it is also found that the accuracy of the final heartbeat estimated value determined based on the above first estimated value and second estimated value is better than that of the final estimated value determined based on the first estimated value and third estimated value. Further, the scheme based on the first estimated value, second estimated value, and third estimated value will be slightly improved compared to the scheme combining the first estimated value and second estimated value. Therefore, this application can ultimately use this scheme as a further optimization scheme.
[0036] Step S103: Extract multi-time scale features of the respiration and heartbeat of the user to be measured based on the reconstructed respiration signal and heartbeat signal; Step S104: Perform real-time stress detection on the user to be measured based on the multi-time scale features of respiration and heartbeat, the estimated value of the respiration signal, and the final estimated value of the heartbeat signal.
[0037] In addition, in the embodiments of this application, since the YOLO model can be reused when determining the second estimated value and the third estimated value, in subsequent embodiments of this application, the third estimated value may be described first and the second estimated value may be described later in combination with the order of appearance of the YOLO model. The above order adjustment is easy to understand for those skilled in the art and will not be specifically described later.
[0038] Specifically, as Figure 2 shown, the method provided by this application is based on the ZeroMQ framework and relies on four processes to carry out real-time millimeter-wave data acquisition, millimeter-wave data processing, vital sign monitoring, and stress state detection: In one example, the above processes can be implemented through the following modules: The first process, the millimeter-wave data acquisition module, acquires millimeter-wave echo data through a UDP (User Datagram Protocol) port; The second process, the millimeter-wave data processing module, extracts the original beat signal from the millimeter-wave echo signal; The third process, the vital sign monitoring module, monitors the respiration and heartbeat states; The fourth process, the stress state detection module, identifies the stress state and non-stress state.
[0039] In a possible embodiment, the system framework provided by the embodiments of the present application may be as follows: First, the detection system emits low-power millimeter waves to the user to be detected through the millimeter wave module, and detects the echo signal generated by the reflection of the signal from the human body (such as the chest, etc.); Second, after clutter suppression and echo selection of the echo signal, the original beat signal is extracted; Then, based on motion artifact suppression, the reconstructed respiration and heartbeat signals are obtained, and the respiration and heartbeat frequency estimation is carried out. Combining one-dimensional and two-dimensional neural networks can further improve the accuracy of heart rate estimation; Finally, using multi-modal multi-time scale fusion and based on a deep learning network, high-precision pressure detection is realized. Specifically, it includes four basic modules: Figure 3 As shown in the figure: First, the detection system emits low-power millimeter waves to the user to be detected through the millimeter wave module, and detects the echo signal generated by the reflection of the signal from the human body (such as the chest, etc.); Second, after clutter suppression and echo selection of the echo signal, the original beat signal is extracted; Then, based on motion artifact suppression, the reconstructed respiration and heartbeat signals are obtained, and the respiration and heartbeat frequency estimation is carried out. Combining one-dimensional and two-dimensional neural networks can further improve the accuracy of heart rate estimation; Finally, using multi-modal multi-time scale fusion and based on a deep learning network, high-precision pressure detection is realized. Specifically, it includes four basic modules: (1) Millimeter wave data acquisition module. This module captures and receives millimeter wave data through a UDP port and transmits the captured data to the host computer. First, key parameters such as the detection range and accuracy of the millimeter wave radar are initialized. Second, the UDP port is monitored through the Socket interface to capture UDP packets in real time; at the same time, the binary data is converted into int16 format. Finally, the converted data enters a 30-second queue for subsequent data processing.
[0040] During system initialization, when the data is less than 30 seconds, all data (regardless of length) is directly transmitted to the subsequent processing module. When the data reaches 30 seconds, data processing and transmission will use the 30-second queue as a window to ensure the continuous operation of the real-time system and the efficient processing of data, as shown in Figure 4 the figure. After the subsequent processing is completed, the system will immediately extract the latest data from the queue and perform analysis. This loop is seamlessly connected, effectively balancing the real-time data response speed and processing efficiency, and ensuring the stable operation of the system under different loads.
[0041] (2) Millimeter wave data processing module. This module processes a single chirp, suppresses clutter, and finally extracts the original beat signal. This module includes three units: (2-1) Single chirp processing unit. First, each frame of the received signal contains a chirp, and the signals of multiple frames are received in sequence to form an echo matrix. Then, FFT (fast fourier transform) is used to generate a range-time graph, which is represented by a matrix as shown in Figure 5 the figure. In this range-time graph, the vertical bright bars represent possible periodic vital sign signals, and the positions of the bright bars on the horizontal axis from left to right represent the range units from near to far.
[0042] (2-2) Clutter suppression unit. There are mainly two types of clutter in radar applications: static clutter generated by stationary objects and non-static clutter generated by moving objects. This system uses adaptive background subtraction to eliminate static clutter and singular value decomposition to eliminate non-static clutter. The range-time diagram after clutter suppression is marked as matrix .
[0043] (2-3) Beat signal extraction unit. After clutter suppression, the distance between the detected target and the radar is confirmed, and the original vital sign signal is extracted from the distance unit where the detected target is located. Specifically, first, the detected target is identified by locating the distance unit with the maximum energy, and the corresponding distance unit index is marked as . Subsequently, the original signal is extracted at the distance unit . The phase of this signal is extracted to represent the original beat signal. The beat signal is a mixture of respiration, heartbeat, and noise signals, and its extraction method is as follows: . Here, the arctan operation is used to determine the phase, and the unwrap operation is used to avoid phase wrapping and ensure the continuity of the phase.
[0044] (3) Vital sign monitoring module. This module is responsible for monitoring the respiration and heartbeat states. Given that the chest vibration caused by respiration is stronger than that caused by heartbeat, resulting in more challenges in heartbeat monitoring than in respiration monitoring, this application uses motion artifact suppression and time-frequency fusion techniques for heart rate estimation; while respiration monitoring uses traditional motion artifact suppression and partial multistage filtering, as shown in Figure 3 (underlined and highlighted). This module includes two units: (3-1) Motion artifact suppression unit. As shown in Figure 3 , this application suppresses motion noise through a combination of traditional motion artifact suppression algorithms, multistage filtering algorithms, and improved motion artifact suppression algorithms. Since the improved motion artifact suppression algorithm relies on the real-time acquisition of triaxial acceleration data, considering that the lack of this data may occur due to equipment limitations and other reasons in real application scenarios, the improved motion artifact suppression algorithm is regarded as an optional component rather than an essential component in this application. The impact of using and not using the improved motion artifact suppression algorithm on the heart rate estimation accuracy will be further discussed later in this application.
[0045] It should be noted that reconstructing the respiration signal and heartbeat signal of the user to be measured based on the echo signal includes: extracting the corresponding beat signal from the echo signal; the beat signal carries the respiration signal and heartbeat signal of the user to be measured; based on the normalized least mean squares (NLMS) algorithm of polynomial fitting, preliminarily suppressing the motion artifacts of the beat signal; using continuous wavelet transform (CWT) to perform multi-level filtering on the beat signal after preliminarily suppressing the motion artifacts; the multi-level filtering includes: performing first-level band-pass filtering on the beat signal in the wavelet domain to obtain the respiration signal and heartbeat signal after the first-level filtering, and then combining the peak-finding and valley-finding algorithm and spectrum evaluation to obtain the first-level respiration signal estimation and heartbeat signal estimation; performing second-level band-pass filtering on the respiration signal and heartbeat signal after the first-level filtering to obtain the respiration signal and heartbeat signal after the second-level filtering, and then combining the peak-finding and valley-finding algorithm and spectrum evaluation to obtain the second-level respiration signal estimation and heartbeat signal estimation; performing third-level band-pass filtering on the heartbeat signal after the second-level filtering to obtain the heartbeat signal after the third-level filtering, and then combining the peak-finding and valley-finding algorithm and spectrum evaluation to obtain the third-level heartbeat signal estimation; wherein, the center frequency of the next-level band-pass filtering is set with reference to the respiration signal estimation and heartbeat signal estimation of the previous level. Taking the respiration signal obtained after the second-level filtering and the heartbeat signal obtained after the third-level filtering as the reconstructed respiration signal and heartbeat signal of the user to be measured.
[0046] Taking the second-level respiration signal estimation obtained during the process of reconstructing the respiration signal of the user to be measured as the estimated value of the respiration signal. (3-1-1) Traditional motion artifact suppression algorithm. The present application first introduces the normalized least mean squares NLMS algorithm based on polynomial fitting, as Figure 6 shown, to preliminarily suppress the motion artifacts. This algorithm does not require prior knowledge of the statistical characteristics of the input signal and noise, fits the noise with a polynomial, and dynamically adjusts the parameters with NLMS. Specifically, the key to this algorithm lies in the selection of the reference signal, and the present application uses polynomial fitting of noise :
[0047] wherein, is the power of the variable , are constant coefficients. In this embodiment, the waveform data of the 30s window is split into 3 10s windows for polynomial fitting respectively; within the waveform data of each 10s window, is 50. Then NLMS is used, with the filter coefficient and the reference signal Computation-related output signal :
[0048] Estimation error Can be expressed as:
[0049] According to the least mean square criterion, this error gradually converges. The filter weights The update formula is:
[0050] Wherein, Is the step size of the adaptive filtering algorithm, while Is a constant. Finally, Is the combination of the respiration and heartbeat signals after noise suppression.
[0051] (3-1-2) Multistage filtering algorithm. Continuous wavelet transform CWT is used for three-stage filtering, as Figure 3 Shown. First, a filter with a relatively wide bandwidth is used for filtering. Based on the preliminary frequency estimation, with this frequency as the center, narrowband filtering is used to further remove noise. For the sake of illustration, the heartbeat signal is taken as an example.
[0052] First, perform the first-stage filtering. Use CWT to reconstruct the heartbeat signal, , That is: perform band-pass filtering in the wavelet domain, and then perform the inverse CWT operation. The frequency range of this filter is [0.7, 2.5] Hz, the center frequency is 1.6 Hz, and the half bandwidth is 0.9 Hz. Then, combine the traditional peak and valley seeking algorithm and the spectrum evaluation method to carry out heart rate estimation, denoted as . The corresponding heartbeat frequency calculation formula is .
[0053] Next, perform the second-stage filtering. Further use CWT to reconstruct the heartbeat signal, . Compared with the initial filter, this filter has a narrower bandwidth, and its half bandwidth HBW2 ≤ 0.9 Hz, and the frequency range is , Hz:
[0054] In this application, the default value of HBW2 is 0.4 Hz. Then, combine the traditional peak and valley seeking algorithm and the spectrum evaluation method to carry out heart rate estimation, denoted as . The corresponding heartbeat frequency calculation formula is .
[0055] Finally, perform third-level filtering. Use CWT to reconstruct the heartbeat signal again. . Similar to the intermediate filtering, the filter has a narrow bandwidth. The default value of HBW3 is 0.4 Hz, and the cut-off frequencies of the band-pass filter are:
[0056] Next, combine the traditional peak and valley finding algorithm and the spectrum evaluation method to carry out heart rate estimation, denoted as ; The corresponding calculation formula for the heartbeat frequency is .
[0057] It should be noted that as Figure 3 shown, if there is no real-time triaxial acceleration data, then it is directly obtained from the second-level filtering; however, if there is real-time triaxial acceleration data, then it is obtained by the improved motion artifact suppression algorithm. In the improved motion artifact suppression algorithm, and are the input waveform data and the output heartbeat frequency value respectively.
[0058] Furthermore, when the triaxial acceleration data of the user to be measured can be obtained in real time, the multi-level filtering performed on the heartbeat signal of the user to be measured reconstructed based on the echo signal is as follows: perform first-level band-pass filtering on the beat signal in the wavelet domain to obtain the heartbeat signal after the first-level filtering, and then combine the peak and valley finding algorithm and the spectrum evaluation to obtain the first-level heartbeat signal estimation; perform second-level band-pass filtering on the heartbeat signal after the first-level filtering to obtain the heartbeat signal after the second-level filtering, and then combine the triaxial acceleration data of the user to be measured and use the photoplethysmogram Q-PPG deep learning algorithm to estimate the heart rate of the heartbeat signal after the second-level filtering to obtain the second-level heartbeat signal estimation; perform third-level band-pass filtering on the heartbeat signal after the second-level filtering to obtain the heartbeat signal after the third-level filtering, and then combine the peak and valley finding algorithm and the spectrum evaluation to obtain the third-level heartbeat signal estimation; use the heartbeat signal obtained after the third-level filtering as the reconstructed heartbeat signal of the user to be measured.
[0059] (3-1-3) Improved motion artifact suppression algorithm. When the triaxial acceleration data is available in real time, the present application uses a simplified version of the Q-PPG (Quantized-PPG) deep learning algorithm to further suppress motion artifacts, as Figure 7 shown. This algorithm includes two stages: First, in the initialization stage, use the seed network to generate a concise framework. Specifically, the seed network is a slightly modified version of TEMPONet, consisting of three convolutional blocks and three fully connected layers, as Figure 7As shown. Each convolutional block contains three convolutional layers. To expand the search space during subsequent architecture exploration, the seed network was modified based on the original TEMPONet: the dilation parameter of all convolutional layers was set to 1, while the filter size was increased to maintain the same receptive field. As Figure 7 shown, the initial loss function is defined as , where and represent the task loss term and the regularization loss term, respectively. Among them, the task loss is used to measure the prediction error of a specific task (such as heart rate estimation) to ensure the accuracy of the model. The regularization loss imposes a complexity constraint by controlling the mask value; the coefficient λ represents the regularization strength, which is used to adjust the relative importance of the two loss terms to maintain a balance between the heart rate tracking error and the model complexity. The default value of λ in this application is .
[0060] Secondly, in the optimization stage, two neural architecture search (NAS) tools, namely MorphNet and Pruning-In-Time (PIT), were used to optimize the network architecture. Specifically, on the one hand, MorphNet introduced mask parameters (i.e., α masks) in the convolutional layers and fully connected layers, gradually removing the weights with smaller mask values and their associated output channels, thereby optimizing the number of output channels. On the other hand, PIT introduced mask parameters (i.e., β masks) at the time steps of the convolutional layers, adjusting the dilation rate of the convolutional filters to affect the receptive field, thereby reducing the computational load while maintaining accuracy.
[0061] (3-2) Time-frequency fusion estimation unit. Compared with the respiratory signal, the heartbeat signal is relatively weak. Therefore, this application uses a neural network hybrid model that combines one-dimensional waveform estimation and two-dimensional spectrum estimation. As Figure 8 shown, based on the fusion of the time domain and the frequency domain, the accuracy of heart rate estimation is further improved.
[0062] (3-2-1) In the one-dimensional neural network model, the reconstructed heartbeat signal is first normalized to the range of [-1, 1], and then processed by the one-dimensional neural network to estimate the heart rate (i.e., the first estimated value of the heartbeat signal).
[0063] Specifically, the one-dimensional neural network model based on 1D waveform estimation gives full play to the advantages of convolutional operations and Transformer. It extracts local features through convolutional layers to effectively capture detailed information, and extracts global features through Transformer to effectively capture context relationships. This not only makes up for the deficiency of convolution in remote dependence modeling but also reduces the computational burden of Transformer. The model consists of four parts: a Convolutional Neural Networks (CNN) backbone block, a pyramid input block, a Transformer encoding block, and a Dense connection block, as Figure 8 shown.
[0064] First, the CNN backbone block. At the bottom layer of the network, in order to achieve a good balance between the complexity and performance of the model, four consecutive CNN backbone blocks are proposed to compress and extract local features. These backbone blocks are designed to effectively extract local features while significantly reducing the subsequent computational burden. Each backbone block consists of 1D convolution, batch normalization (BN), a rectified linear unit (ReLu) activation function, average pooling, and dropout. Among them, for 1D convolution, four different-sized convolutional kernels are proposed, namely 1×15, 1×11, 1×5, and 1×3. As the input changes, the size of the convolutional kernel gradually decreases to capture features at different scales. Similarly, average pooling downsamples the features with kernel sizes of 5, 5, 2, and 2 respectively, retaining important feature information while effectively reducing the size; and dropout aims to further improve the generalization ability, with a default value of 0.2 to effectively prevent overfitting.
[0065] Second, the pyramid input block. It aims to fuse multi-scale heartbeat signals by introducing feature learning at different scales, obtain a multi-level receptive field, and obtain more original signal information to further improve the model performance. Specifically, this module uses nearest neighbor interpolation technology to downsample the input signal to different degrees, namely 1 / 15, 1 / 5, and 1 / 2. These downsampled signals are combined with the corresponding CNN backbone blocks to participate in subsequent convolutional operations.
[0066] Third, the Transformer encoding block. Since the heartbeat signal is still affected by small-amplitude limb movements or noises such as breathing in the stationary state, if the local features learned by CNN are directly used for heart rate estimation, the estimation results are often not ideal. In addition, the heartbeat signal is essentially a time series signal, and it is not enough to only extract local features. To solve this problem, the Transformer multi-head attention mechanism is proposed to effectively learn global features and capture the dependencies between heartbeats, thus forming a powerful complement to CNN. Specifically, the processing flow of this module is as follows: compress the input heartbeat signal and exchange the positions of each dimension; after 6 layers of Transformer encoding, capture the long-range dependencies in the signal through the self-attention mechanism; calculate the average value of the multi-channel size and output it to extract the most representative features. The hyperparameters of this module are set as follows: d_model (model dimension) is 12, nhead (number of heads of multi-head attention) is 6, dim_feedforward (dimension of the feed-forward network) is 24, and dropout is 0.2.
[0067] Fourth, the Dense connection block. Two Dense layers are used to predict the final heart rate estimation result.
[0068] The hyperparameters of the 1D neural network are set as follows: epochs (number of iterations) is 1000, learning rate is 0.001, batch size is 32; optimizer is Adam, and its hyperparameters are set as beta_1 (exponential decay rate of the first moment estimate of the gradient) is 0.9, beta_2 (exponential decay rate of the second moment estimate of the gradient) is 0.999, and epsilon (numerical stability) is 1e-08.
[0069] (3-2-2) In the two-dimensional neural network model, the reconstructed heartbeat signal is converted into a two-dimensional CWT spectrogram, and a cascaded network composed of YoLov8 and DenseNet-201 is used for heart rate estimation Specifically, On the one hand, input the CWT spectrogram into the YOLOv8 pre-trained model. The architecture of the YOLOv8 pre-trained model is shown in Figure 9 as shown, frame the core frequency band area related to the heart rate, and output the heart rate estimation value (i.e., the third estimation value of the heartbeat signal). The hyperparameters of this network are set as follows: batch size is 10, epochs is 200, and learning rate is 0.01.
[0070] On the other hand, linear transformation is adopted to adjust the pixel values within the frequency band framed by YOLOv8 to enhance the contrast of the CWT spectrogram. :
[0071] Among them, is used to control the contrast (>1 indicates enhanced contrast, <1 indicates reduced contrast), is used to control the brightness (>0 indicates enhanced brightness, <0 indicates reduced brightness), and clip is used to limit the result within the range of [0, 255] to avoid overflow. In this embodiment, Subsequently, the brightened CWT spectrum is input into the fine-tuned DenseNet-201 network. The architecture of the DenseNet-201 network is shown in Figure 10 as shown. During the fine-tuning process, the underlying layer of the DenseNet-201 pre-trained model is retained, while the top layer is re-trained using the two-dimensional CWT spectrogram in this application to output the heart rate estimation value (i.e., the second estimation value of the heartbeat signal). The hyperparameters of this network are set as: batch size is 16, epochs is 800, and learning rate is 0.0025. It should be noted that the main reason for using CWT to generate the two-dimensional spectrogram in this application is that although CWT, short-time Fourier transform (STFT), and smoothed pseudo-Wigner–Ville distribution (SPWVD) are all applicable to the analysis of time-varying signals, CWT has higher accuracy and more efficient processing ability compared with STFT and SPWVD.
[0072] (3-2-3) Preferably, the final estimation value of the heartbeat signal is determined based on the multi-level heartbeat signal estimations, the first estimation value of the heartbeat signal, the second estimation value of the heartbeat signal, and the third estimation value of the heartbeat signal obtained during the multi-level filtering process.
[0073] Exemplarily, the heart rate estimation values from the one-dimensional and two-dimensional networks ( ) and the obtained by multi-level filtering during the motion artifact suppression process are input into the ridge regression ensemble model to generate the final heart rate estimation value (i.e., the final estimation value of the heartbeat signal).
[0074] (4)Pressure state detection module. This module aims to identify the pressure state based on multi-modal multi-scale fusion. In this application, two deep learning networks are adopted: the Residual-Temporal Convolution Network (Res-TCN) and the Self-Supervised Learning network (SSL). The reconstructed respiratory signal and heartbeat signal are processed with a 30-second long detection window and a 10-second short detection window respectively. After being analyzed by the two networks, the multi-time scale features of 30 seconds and 10 seconds are fused, and the respiratory and heart rate features are fused to improve the accuracy of pressure recognition. This module includes three units: (4-1)Res-TCN unit. Res-TCN consists of four residual blocks: one-dimensional convolution, causal analyzer, dilated convolution, and residual connection, as Figure 11 shown. Among them, the one-dimensional convolution captures the time-domain change features of time series data; the causal analyzer ensures that the prediction results are consistent with the real-world scenario by restricting the convolution process to past and present data; the dilated convolution helps with downsampling and provides a larger receptive field; while the residual connection speeds up the convergence rate by allowing gradient propagation without passing through the activation function.
[0075] Specifically, each residual block in the network consists of three residual sub-blocks, and each sub-block contains three layers, representing the batch normalization layer, the activation function layer, and the convolution layer respectively. Among them, the parameters of the convolution layer are represented by Res-U (F, K, S, D), where F represents the number of filters, K represents the size of the filters, S represents the stride, and D represents the dilation rate.
[0076] (4-2)SSL unit. SSL includes two parts: the self-supervised learning part and the pressure detection part, as Figure 12 shown.
[0077] First, in the self-supervised learning part, a multi-task CNN model is trained using automatically generated labels to extract high-level abstract features of the respiratory or heartbeat time series. Specifically, this module takes the recognition of the reconstructed respiratory signal and heartbeat signal and their six signal transformations as pre-training tasks, enabling the network to learn spatio-temporal features and thus improving its generalization ability. The weights obtained from the training are transferred to the pressure detection part.
[0078] Next, in the pressure detection part, the convolution layer remains frozen, and the weights are the same as those obtained in the self-supervised learning part, as Figure 12 shown; while the fully connected layer is fine-tuned to perform pressure detection.
[0079] (4-3)Classifier unit. As Figure 3As shown, based on multi-modal and multi-scale fusion, this application uses a classifier for pressure recognition.
[0080] (4-3-1) First, aggregate the features of multiple time scales of the respiration and heartbeat signals, including: (i) Res-TCN features extracted from a 30-second detection window and (ii) SSL features extracted from a 10-second detection window. Specifically, after aggregation, a total of 8 features are obtained in each 30-second detection window, including 4 respiration-related features and 4 heartbeat-related features. Taking the respiration features as an example, 1 feature comes from Res-TCN and 3 features come from SSL.
[0081] (4-3-2) Then, input these 8 features, the estimated value of the respiration signal, and the final estimated value of the heartbeat signal into a random forest classifier to classify the state of the user to be tested into two types: non-stress and stress, and perform visualization, as Figure 13 shown. The recognition result is updated once per second.
[0082] (5) To verify the reliability of the non-contact real-time pressure detection method and system proposed in this application, 15 subjects were recruited in this embodiment to participate in two tests with a duration of 5 minutes each (one test was under a non-stress state and the other was a stress state induced by a mental arithmetic experiment; the interval between the two tests was 2 minutes). After each test, the subjects filled out a simplified version of the State-Trait Anxiety Inventory (STAI). This questionnaire contains 6 questions for evaluating whether the subjects are in a stress state: a questionnaire score greater than 15 points is regarded as a stress state, and a score less than or equal to 15 points is regarded as a non-stress state. During the experiment, an IWR1642BOOST+DCA1000EVM millimeter-wave radar device was used synchronously to collect millimeter-wave data for vital sign detection and pressure recognition. At the same time, a wearable electrocardiogram device MAX-ECG-MONITOR was used to collect electrocardiogram data as a benchmark for millimeter-wave heart rate estimation; a Shimmer3 IMU device was used to collect triaxial acceleration data to suppress motion artifacts. The visualization effect of this embodiment is as Figure 13 shown, and the experimental results are shown in Table 1 and Table 2.
[0083] (5-1) The heart rate estimation results of this embodiment are shown in Table 1. Three metrics are used in this embodiment for performance evaluation: MAE (mean absolute error), RMSE (root mean square error), and Corr (Correlation). As can be seen from Table 1, in the motion artifact suppression stage and the heart rate estimation stage, the method proposed in this embodiment shows superior performance in both scenarios of without real-time three-axis accelerometer signals (wo ACCs) and with real-time three-axis accelerometer signals (w ACCs).
[0084] Table 1 Experimental results of motion artifact suppression and heart rate estimation
[0085] Note: w ACCs and wo ACCs represent scenarios with and without real-time three-axis accelerometer signals respectively; CorNET represents heart rate estimation using the CorNET one-dimensional network; 1D DCNN represents heart rate estimation using 1D DCNN; 1D (CNN + pyramid + Transformer) represents heart rate estimation using the one-dimensional neural network model proposed in this embodiment; VGG16 + STFT represents generating a spectrogram using STFT and inputting it into the VGG16 network for heart rate estimation; DenseNet + STFT represents generating a spectrogram using STFT and inputting it into the DenseNet network for heart rate estimation; ResNet + CWT represents generating a spectrogram using CWT and inputting it into the ResNet network for heart rate estimation; DenseNet + CWT represents generating a spectrogram using CWT and inputting it into the DenseNet network for heart rate estimation; 2D YoLo represents generating a spectrogram using CWT and inputting it into the YOLO model for heart rate estimation; 2D YoLo+频谱提亮+DenseNet represents generating a spectrogram using CWT, inputting it into the YOLO model, then enhancing the spectrum, and then inputting it into the DenseNet network for heart rate estimation; 2D (2D YoLo +2D YoLo+频谱提亮+DenseNet ) represents generating a spectrogram using CWT, inputting it into the YOLO model for heart rate estimation, then enhancing the spectrum, and then inputting it into the DenseNet network for heart rate estimation again, and integrating the two heart rate estimation results to obtain the final heart rate estimation; 1D+2D YoLoIndicates that heart rate estimation is performed using 1D and then 2D YoLo for heart rate estimation, and the results of the two heart rate estimations are integrated to obtain the final heart rate estimation; 1D+2D YoLo+频谱提亮+DenseNet Indicates that heart rate estimation is performed using 1D and then 2D YoLo+频谱提亮+DenseNet for heart rate estimation, and the results of the two heart rate estimations are integrated to obtain the final heart rate estimation; 1D+2D+hr 1~3 Indicates that heart rate estimation is performed using 1D and then 2D for heart rate estimation, and then the three heart rate estimation results obtained in the multi-stage filtering process are integrated to obtain the final heart rate estimation.
[0086] (5-1-1) In the case of wo ACCs, first, in the motion artifact suppression stage, the experimental results show that compared with six state-of-the-art traditional algorithms, the method proposed in this embodiment can effectively suppress motion noise. Specifically, the MAE of the method proposed in this embodiment is 4.501 bpm (beats per minute), which is better than the best-performing LMS based on polynomial fitting among the six state-of-the-art comparison algorithms, whose MAE is 4.545 bpm. Subsequently, in the heart rate estimation stage, the 1D+2D+hr 1~3 method proposed in this embodiment is significantly better than the six state-of-the-art comparison algorithms. The lowest MAE of the six state-of-the-art comparison algorithms is 4.320 bpm (1D DCNN), while the MAE of the method proposed in this embodiment is 3.780 bpm. In addition, Table 1 also shows the following points: 1) The 1D (CNN+pyramid+Transformer) method proposed in this embodiment (MAE is 4.059 bpm) is better than two traditional 1D DCNN methods: CorNET (MAE is 5.642 bpm) and 1D DCNN (MAE is 4.320 bpm). The reason why the 1D method proposed in this embodiment is significantly better than the traditional method is that in this embodiment, CNN is used to extract local features, the pyramid structure is used to extract multi-scale features, and Transformer is used to extract global features, effectively making up for the deficiencies of traditional DCNN networks in multi-scale and long-range dependence modeling.
[0087] 2) YOLO is better than the traditional two-dimensional DCNN network. The MAE of YOLO in this embodiment is 4.366 bpm, while the lowest MAE of the traditional two-dimensional method is 4.411 bpm (DenseNet+CWT). The reason why the YOLO method used in this embodiment is significantly better than the traditional 2D DCNN network is that for YOLO, heart rate estimation is regarded as a spectral localization problem, and this perspective can make full use of the high-precision localization ability of YOLO to improve the accuracy of heart rate estimation.
[0088] 3) The combination of YOLO + spectral brightening + DenseNet (MAE is 4.105 bpm) is superior to YOLO alone (MAE is 4.366 bpm) or DenseNet alone (MAE is 4.411 bpm). The reason is that: the powerful positioning ability of YOLO can accurately locate the frequency band most relevant to the heart rate, and spectral brightening enhances the spectrum located by YOLO (both retaining the effective information and appropriately weakening the out-of-band noise). Therefore, after the enhanced spectrum is input into DenseNet, the accuracy of heart rate estimation is significantly improved.
[0089] 4) The method of 1D + 2D (MAE is 3.815 bpm) is superior to 1D alone (MAE is 4.059 bpm) or 2D alone (MAE is 4.070 bpm). The reason is that the fusion of 1D time-domain features and 2D frequency-domain features can improve the accuracy of heart rate estimation.
[0090] 5) 1D + 2D + hr 1~3 (MAE is 3.780 bpm) is superior to all other methods. The reason is that this method not only fuses the time-domain and frequency-domain features extracted by the deep learning network, but also makes full use of the time-domain features of the pre-filtered signals in the previous three times, thus obtaining the highest heart rate recognition accuracy.
[0091] (5-1-2) In terms of w ACCs, first, in the motion artifact suppression stage, the experimental results show that compared with three leading deep learning (DL) algorithms, the DL method proposed in this embodiment has significant advantages. Specifically, the best MAE of the three leading methods is 3.779 bpm (Q-PPG), while the MAE of the method proposed in this embodiment is 3.603 bpm. On the one hand, the performance of this embodiment is superior to the three leading DL comparison methods; on the other hand, compared with the traditional motion artifact suppression method, this embodiment is superior to the best-performing LMS based on polynomial fitting among the six leading comparison algorithms, whose MAE is 4.545 bpm. This result shows that combining the traditional method (NLMS based on polynomial fitting) with the DL framework (Q-PPG) can further enhance the motion artifact suppression performance. Second, in the heart rate estimation stage, the experimental results once again confirm the effectiveness of the 1D + 2D + hr 1~3 method proposed in this embodiment. The best MAE of the six leading comparison algorithms is 3.281 bpm (DenseNet + CWT), while the MAE of the method proposed in this embodiment is 2.824 bpm. At the same time, the relevant results once again verify the five conclusions drawn from the wo ACCs situation in Table 1.
[0092] In summary, the method proposed in this embodiment fully considers the actual application scenarios, including wo and w ACCs, and provides a practical solution. The proposed method shows excellent performance in both the motion artifact suppression and heart rate estimation stages, demonstrating its potential in the vital sign monitoring system.
[0093] (5-2) The pressure detection results of this embodiment are shown in Table 2. To evaluate the pressure detection performance, this embodiment uses the five-fold and leave-one-out cross-validation methods to test on the self-built dataset: on the one hand, the five-fold cross-validation ignores the differences of individual subjects and aims to determine the generalization ability of the algorithm on the entire dataset; on the other hand, the leave-one-out cross-validation method focuses on learning from unrelated subjects and aims to evaluate the generalization ability of the algorithm among different subjects.
[0094] Table 2 Experimental results of five-fold and leave-one-out cross-validation
[0095] Note: B and H represent respiratory and heartbeat data respectively.
[0096] (5-2-1) As shown in Table 2, this embodiment uses four metrics for performance evaluation: accuracy, precision, F1-score, and area under curve (AUC). The experimental results show that for the five-fold cross-validation, the values of the four relevant metrics in the last row of Table 2 all exceed 0.940, indicating that this embodiment can accurately identify the pressure; similarly, for the leave-one-out cross-validation, the four relevant metrics in the last row of Table 2 are all greater than 0.953, indicating that this embodiment can accurately identify the pressure state of unknown subjects based on the data of known individuals.
[0097] (5-2-2) Table 2 also shows the comparison between the method proposed in this embodiment and four state-of-the-art methods. Generally speaking, these four state-of-the-art methods can be divided into two categories: traditional feature extraction methods (Handcrafted) such as EQ-Radio, and deep learning methods (DL), such as Stressalyzer, Res-TCN, and SSL. The experimental results in Table 2 show that the method proposed in this embodiment is significantly better than these four state-of-the-art methods. For example, for the five-fold cross-validation, the pressure recognition accuracy of the method proposed in this embodiment is 0.940, which is significantly higher than that of the traditional feature extraction method EQ-Radio (see the third row of Table 2) and the best-performing Res-TCN in the deep learning algorithms (see the seventh row of Table 2), and their accuracies are 0.713 and 0.820 respectively.
[0098] (5-2-3) In addition, the data in the last three rows of Table 2 show that the performance of the multi-modal multi-scale fusion method proposed in this embodiment is significantly better than that of the two single-data methods. For example, for five-fold cross-validation, the accuracy of the method proposed in this embodiment is 0.940, which is significantly higher than the accuracies of the other two single-data methods, which are 0.793 and 0.873 respectively.
[0099] Figure 14 This is the architecture diagram of the non-contact real-time pressure detection system based on millimeter-wave radar provided by the embodiment of the present application; as Figure 14 shown, it includes: A signal reconstruction module 1410, configured to obtain in real time the echo signal of the millimeter-wave radar reflected from the user to be measured, and reconstruct the respiration signal and heartbeat signal of the user to be measured based on the echo signal, and obtain an estimated value of the respiration signal; A heart rate final estimation module 1420, configured to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, and input the CWT spectrogram into a YOLO model to mark the frequency band region related to the heart rate of the user to be measured in the CWT spectrogram; then brighten the spectrum of the marked frequency band region, and input the CWT spectrogram with the brightened spectrum into a two-dimensional convolutional neural network to obtain a second estimated value of the heartbeat signal; determine the final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; A pressure detection module 1430, configured to extract multi-time scale features of the respiration and heartbeat of the user to be measured based on the reconstructed respiration signal and heartbeat signal; and perform real-time pressure detection on the user to be measured based on the multi-time scale features of the respiration and heartbeat, the estimated value of the respiration signal, and the final estimated value of the heartbeat signal.
[0100] It should be understood that the above system is used to execute the method in the above embodiment. For the corresponding program modules in the system, their implementation principles and technical effects are similar to those described in the above method. The working process of the system can refer to the corresponding process in the above method, which will not be elaborated here.
[0101] Based on the method in the above embodiment, the embodiment of the present application provides an electronic device, as Figure 15 shown, the electronic device may include: a processor 1510, a communication interface 1520, a memory 1530, and a communication bus 1540, wherein the processor 1510, the communication interface 1520, and the memory 1530 communicate with each other through the communication bus 1540. The processor 1510 can call the logical instructions in the memory 1530 to execute the method in the above embodiment.
[0102] In addition, when the logical instructions in the above-mentioned memory 1530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application.
[0103] Based on the method in the above embodiment, an embodiment of this application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, it causes the processor to execute the method in the above embodiment.
[0104] Based on the method in the above embodiment, an embodiment of this application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the method in the above embodiment.
[0105] It can be understood that the processor in the embodiments of this application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0106] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this application shall be included in the protection scope of this application.
Claims
1. A non-contact real-time pressure detection method based on millimeter wave radar, characterized in that: include: Acquire in real time the echo signal of the millimeter-wave radar reflected from the user to be measured, reconstruct the breathing signal and heartbeat signal of the user to be measured based on the echo signal, and obtain an estimated value of the breathing signal; Input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrum graph, and input the CWT spectrum graph into a YOLO model to mark the frequency band area related to the heart rate of the user to be measured in the CWT spectrum graph; Then, the marked frequency band area is spectrally brightened, and the CWT spectrum after spectrally brightening is input into the two-dimensional convolutional neural network to obtain the second estimated value of the heartbeat signal; Determining a final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; Extract multi-time scale features of breathing and heartbeat of the user to be tested based on the reconstructed breathing signal and heartbeat signal; Based on the multi-time scale characteristics of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal, real-time stress detection is performed on the user to be tested.
2. The method according to claim 1, characterized in that The one-dimensional neural network includes: a plurality of convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block and a Dense connection block; The multiple convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales through different convolution kernels; The pyramid input module is used to perform feature learning of different scales on the signals input to each convolutional neural network backbone block, and then combine the signals obtained by feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block captures local features of corresponding scales; The Transformer encoding block is used to combine the multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the multiple convolutional neural networks at different scales, and capture the most representative features of the heartbeat signals among the local features at different scales; The Dense connection block is used to obtain a first estimated value of the heartbeat signal in combination with the output signal of the Transformer encoding block.
3. The method according to claim 1, characterized in that Reconstructing the breathing signal and the heartbeat signal of the user to be tested based on the echo signal and obtaining an estimated value of the breathing signal, including: Extracting a corresponding beat signal from the echo signal; the beat signal carries a breathing signal and a heartbeat signal of the user to be tested; A normalized minimum mean square error (NLMS) algorithm based on polynomial fitting is used to preliminarily suppress motion artifacts of the beat signal; The continuous wavelet transform CWT is used to perform multi-stage filtering on the beat signal after preliminary suppression of motion artifacts; the multi-stage filtering includes: performing a first-stage bandpass filtering on the beat signal in the wavelet domain to obtain a breathing signal and a heartbeat signal after the first-stage filtering, and then combining a peak-finding algorithm and a spectral evaluation to obtain a first-stage breathing signal estimation and a heartbeat signal estimation; performing a second-stage bandpass filtering on the breathing signal and the heartbeat signal after the first-stage filtering to obtain a breathing signal and a heartbeat signal after the second-stage filtering, and then combining a peak-finding algorithm and a spectral evaluation to obtain a second-stage breathing signal estimation and a heartbeat signal estimation; performing a third-stage bandpass filtering on the heartbeat signal after the second-stage filtering to obtain a heartbeat signal after the third-stage filtering, and then combining a peak-finding algorithm and a spectral evaluation to obtain a third-stage heartbeat signal estimation; wherein the center frequency of the next-stage bandpass filtering is set with reference to the breathing signal estimation and the heartbeat signal estimation corresponding to the previous stage; The breathing signal obtained after the secondary filtering and the heartbeat signal obtained after the tertiary filtering are used as the reconstructed breathing signal and heartbeat signal of the user to be tested; The second-level breathing signal estimation obtained in the process of reconstructing the breathing signal of the user to be tested is used as the estimated value of the breathing signal.
4. The method according to claim 3, characterized in that When the three-axis acceleration data of the user to be measured can be obtained in real time, the multi-stage filtering performed on the heartbeat signal of the user to be measured based on the echo signal is as follows: performing a first-stage bandpass filtering on the difference signal in the wavelet domain to obtain a heartbeat signal after the first-stage filtering, and then combining the peak-finding algorithm and the spectrum evaluation to obtain a first-stage heartbeat signal estimation; performing a second-stage bandpass filtering on the heartbeat signal after the first-stage filtering to obtain a heartbeat signal after the second-stage filtering, and then combining the three-axis acceleration data of the user to be measured, using the photoelectric plethysmography Q-PPG deep learning algorithm to perform heart rate estimation on the heartbeat signal after the second-stage filtering to obtain a second-stage heartbeat signal estimation; performing a third-stage bandpass filtering on the heartbeat signal after the second-stage filtering to obtain a heartbeat signal after the third-stage filtering, and then combining the peak-finding algorithm and the spectrum evaluation to obtain a third-stage heartbeat signal estimation; and using the heartbeat signal obtained after the third-stage filtering as the reconstructed heartbeat signal of the user to be measured.
5. The method according to claim 3 or 4, characterized in that: Also includes: Convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrogram, and input the CWT spectrogram into a YOLO model so that the model marks a frequency band region in the CWT spectrogram that is related to the heart rate of the user to be measured, so as to obtain a third estimated value of the heartbeat signal; A final estimated value of the heartbeat signal is determined based on the multi-stage heartbeat signal estimation obtained in the multi-stage filtering process, the first estimated value of the heartbeat signal, the second estimated value of the heartbeat signal and the third estimated value of the heartbeat signal.
6. The method according to claim 1, characterized in that The step of extracting multi-time scale features of breathing and heartbeat of the user to be tested includes: Based on the reconstructed breathing signal and heartbeat signal, a first deep learning network is used to extract breathing features and heartbeat features of the user to be tested in the first time span based on a detection window of the first time span; Based on the reconstructed breathing signal and heartbeat signal, a second deep learning network is used to extract the breathing characteristics and heartbeat characteristics of the user to be tested in the second time span based on the detection window of the second time span; The breathing characteristics and heartbeat characteristics of the user under test in different time spans are used as multi-time scale features.
7. The method according to claim 6, characterized in that The first deep learning network includes: a one-dimensional convolution, a causal analyzer, a dilated convolution and a residual connection; the one-dimensional convolution is used to capture the time domain variation characteristics of the time series data corresponding to the respiratory signal and the heartbeat signal within the first time span; the causal analyzer is used to limit the convolution process to the past and present data to ensure that the prediction result is consistent with the real scene; the dilated convolution is used to help down sampling and provide a larger receptive field; the residual connection speeds up the convergence speed by allowing the gradient to propagate without passing through the activation function, so as to extract the respiratory features and heart rate features within the corresponding time span; The second deep learning network is a self-supervised learning network, which is used to learn and extract the spatiotemporal features of the respiratory signal and the heartbeat signal within the corresponding time span.
8. A non-contact real-time pressure detection system based on millimeter wave radar, characterized in that: include: A signal reconstruction module, used to obtain in real time the echo signal of the millimeter-wave radar reflected from the user to be measured, and reconstruct the breathing signal and heartbeat signal of the user to be measured based on the echo signal, and obtain an estimated value of the breathing signal; The heart rate final estimation module is used to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrum map, and input the CWT spectrum map into a YOLO model to mark the frequency band area related to the heart rate of the user to be measured in the CWT spectrum map; Then, the marked frequency band area is spectrally brightened, and the CWT spectrum after spectrally brightening is input into the two-dimensional convolutional neural network to obtain the second estimated value of the heartbeat signal; Determining a final estimated value of the heartbeat signal based on the first estimated value of the heartbeat signal and the second estimated value of the heartbeat signal; A pressure detection module, used to extract multi-time scale features of breathing and heartbeat of the user to be tested based on the reconstructed breathing signal and heartbeat signal; And based on the multi-time scale characteristics of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal, real-time stress detection is performed on the user to be tested.
9. The system according to claim 8, characterized in that The one-dimensional neural network used in the final heartbeat estimation module includes: multiple convolutional neural network backbone blocks, pyramid input blocks, Transformer encoding blocks and Dense connection blocks; the multiple convolutional neural network backbone blocks are used to capture the local features of the reconstructed heartbeat signal at different scales through different convolution kernels; the pyramid input module is used to perform feature learning of different scales on the signals input to each convolutional neural network backbone block, and then combine the signals obtained by feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block captures the local features of the corresponding scale; the Transformer encoding block is used to combine the multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the multiple convolutional neural networks at different scales, and capture the most representative features of the heartbeat signal among the local features at different scales; the Dense connection block is used to obtain the first estimated value of the heartbeat signal in combination with the output signal of the Transformer encoding block.
10. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Motion noise interference eliminating method suitable for wearable heart rate monitoring device
CN104161505A
Heart rate monitoring method based on multi-channel parallel filtering and spectral peak weighted selection algorithm
CN109864713A
Non-contact real-time vital sign monitoring system and method based on millimeter wave radar
CN112401863A
Non-contact identity recognition method and system based on millimeter-wave radar
CN112686094A
Non-contact fatigue detection method and system
CN113420624A
Cited By
Sea surface weak target detection method based on multi-modal time-frequency diagram fusion
CN120722304A