Millimeter wave radar-based non-contact real-time pressure detection method and system

WO2026199722A1PCT designated stage Publication Date: 2026-10-01HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/099553
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2025-06-06
Publication Date
2026-10-01

Smart Images

  • Figure CN2025099553_01102026_PF_FP_ABST
    Figure CN2025099553_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a millimeter wave radar-based non-contact real-time pressure detection method and system. The non-contact real-time pressure detection method comprises: acquiring in real time a millimeter wave echo signal reflected from a user to be detected, reconstructing a respiratory signal and a heartbeat signal of said user, and obtaining an estimated value of the respiratory signal; inputting the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimated value of the heartbeat signal; converting the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, inputting the two-dimensional CWT spectrogram into a YOLO model, and marking a frequency band region related to a heart rate therein; subsequently, performing spectrum brightening on the marked frequency band region, inputting same into a two-dimensional neural network to obtain a second estimated value of the heartbeat signal, and determining a final estimated value of the heartbeat signal; extracting, on the basis of the reconstructed respiratory signal and heartbeat signal, multi-time-scale features; and performing, on the basis of the multi-time-scale features, the estimated value of the respiratory signal, and the final estimated value of the heartbeat signal, real-time pressure detection on said user. By means of the present application, an objective and accurate pressure detection method is provided.
Need to check novelty before this filing date? Find Prior Art

Description

A non-contact real-time pressure detection method and system based on millimeter-wave radar [Technical Field]

[0001] This application belongs to the field of information technology, and more specifically, relates to a non-contact real-time pressure detection method and system based on millimeter-wave radar. [Background Technology]

[0002] Stress is one of the fundamental problems in today's society. High workloads and the variability of modern life severely impact people's mental health, causing various negative effects. Furthermore, stress is a major contributing factor to various diseases, such as depression, stroke, sleep disorders, heart attacks, and cardiac arrest. Therefore, accurate detection of stress in its initial stages is particularly important.

[0003] Due to the negative health effects of stress and the widespread demand for stress detection, stress detection-related inventions have attracted considerable attention from engineers and psychologists. Most existing stress detection methods are based on audio and video processing technology, that is, judging an individual's stress state by extracting speech features, facial image features (such as facial expressions and pupil diameter), and gesture features from audio and video. Although this technology has achieved a high recognition accuracy rate, many problems remain. On the one hand, it is easily affected by objective factors such as dim lighting, uneven lighting, facial turning, and tilting, leading to detection errors. On the other hand, subjective factors such as the subject's deception (e.g., adjusting tone of voice, speaking speed, or concealing facial expressions to mask their true state) can also significantly interfere with the test results. Furthermore, this technology also faces privacy protection issues.

[0004] Meanwhile, due to the uniqueness and stability of physiological signals, stress detection based on physiological signals has been continuously developing and has gradually become a hot topic of invention in academia and industry in recent years. In general, stress detection methods based on physiological signals generally include two types: (1) Contact detection based on wearable devices (such as ECG, EMG, EEG, etc.): This method has been widely used due to the convenience of the devices, but it also has certain problems. For example, prolonged wear may cause skin discomfort and is not suitable for damaged skin (such as painful rashes, burns, urticaria, etc.). In addition, it is impossible to distinguish whether the source of stress is device discomfort or psychological factors; (2) Non-contact detection based on millimeter-wave radar: This method extracts vital signs such as breathing and heartbeat by monitoring the millimeter-wave radar echo signal reflected by the human body, thereby conducting stress detection in a non-contact and non-disruptive manner. This method has advantages such as low power consumption, low cost and easy use, but it also has some problems: In terms of signal acquisition, millimeter-wave radar is easily affected by environmental noise and the subject's limb movements, resulting in poor signal quality; In terms of pressure detection, existing methods are mostly limited to traditional feature extraction and do not focus on the application of deep learning features, which limits the algorithm's real-time performance, robustness and transferability. [Summary of the Invention]

[0005] To address the shortcomings of existing technologies, this application aims to provide a non-contact real-time pressure detection method and system based on millimeter-wave radar. This method addresses the problems of existing millimeter-wave radar-based pressure detection technologies being susceptible to interference from environmental noise and subject limb movements, as well as the limitations in detection accuracy due to single feature extraction methods.

[0006] To achieve the above objectives, in a first aspect, this application provides a non-contact real-time pressure detection method based on millimeter-wave radar, comprising:

[0007] The system acquires the echo signal of the millimeter-wave radar reflected from the user under test in real time, reconstructs the respiratory and heartbeat signals of the user under test based on the echo signal, and obtains the estimated value of the respiratory signal.

[0008] The reconstructed heartbeat signal is input into a one-dimensional neural network to obtain a first estimate of the heartbeat signal. The reconstructed heartbeat signal is then converted into a two-dimensional continuous wavelet transform (CWT) spectrogram, which is input into a YOLO model to mark the frequency band regions in the CWT spectrogram that are related to the heart rate of the user being tested. The marked frequency band regions are then highlighted, and the highlighted CWT spectrogram is input into a two-dimensional convolutional neural network to obtain a second estimate of the heartbeat signal. The final estimate of the heartbeat signal is determined based on the first and second estimates.

[0009] Multi-timescale features of the user's breathing and heartbeat were extracted based on the reconstructed breathing and heartbeat signals.

[0010] Real-time stress detection of the user is performed based on the multi-timescale characteristics of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.

[0011] Understandably, this application combines YOLO's powerful positioning capabilities to accurately locate the frequency band most relevant to heart rate. The spectrum brightening enhances the spectrum located by YOLO (preserving effective information while appropriately reducing out-of-band noise). Therefore, after the enhanced spectrum is input into the two-dimensional convolutional neural network, it can effectively improve the accuracy of heart rate estimation, thereby ensuring the accuracy of the above-mentioned stress detection scheme.

[0012] In one possible implementation, the one-dimensional neural network includes: multiple convolutional neural network backbone blocks, pyramid input blocks, Transformer coding blocks, and Dense connection blocks;

[0013] Multiple convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales through different convolutional kernels;

[0014] The pyramid input module is used to learn features at different scales from the signals input to each convolutional neural network backbone. Then, the signals obtained from feature learning are combined with each convolutional neural network backbone so that each convolutional neural network backbone can capture local features at the corresponding scale.

[0015] Transformer encoding blocks are used to combine multi-head attention mechanisms to learn global features corresponding to local features of heartbeat signals output by multiple convolutional neural networks at different scales, and to capture the most representative heartbeat signal features among the local features at different scales.

[0016] The Dense connection block is used to combine the output signal of the Transformer coded block to obtain the first estimate of the heartbeat signal.

[0017] It should be noted that the one-dimensional neural network used in this application extracts local features of the heartbeat signal using multiple convolutional neural network backbone blocks, extracts multi-scale features using a pyramid structure, and extracts global features using a Transformer. This effectively makes up for the shortcomings of traditional convolutional neural networks in multi-scale and long-range dependency modeling, and can greatly improve the accuracy of the first estimate of the heartbeat signal, thereby improving the accuracy and reliability of stress detection.

[0018] In one possible implementation, the respiratory and heartbeat signals of the user under test are reconstructed based on the echo signals, and an estimated value of the respiratory signal is obtained, including:

[0019] Extract the corresponding beat signal from the echo signal; the beat signal carries the respiratory and heartbeat signals of the user being tested.

[0020] The normalized least mean squares (NLMS) algorithm based on polynomial fitting is used to initially suppress motion artifacts in beat signals.

[0021] A continuous wavelet transform (CWT) is used to perform multi-stage filtering on the beat signal after initial suppression of motion artifacts. The multi-stage filtering includes: performing a first-stage bandpass filter on the beat signal in the wavelet domain to obtain the first-stage filtered respiratory and heartbeat signals, which are then combined with peak-and-valley finding algorithms and spectral evaluation to obtain first-stage respiratory and heartbeat signal estimates; performing a second-stage bandpass filter on the first-stage filtered respiratory and heartbeat signals to obtain the second-stage filtered respiratory and heartbeat signals, which are then combined with peak-and-valley finding algorithms and spectral evaluation to obtain second-stage respiratory and heartbeat signal estimates; and performing a third-stage bandpass filter on the second-stage filtered heartbeat signal to obtain the third-stage filtered heartbeat signal, which is then combined with peak-and-valley finding algorithms and spectral evaluation to obtain the third-stage heartbeat signal estimate. The center frequency of the next stage bandpass filter is set with reference to the respiratory and heartbeat signal estimates of the previous stage.

[0022] The respiratory signal obtained after secondary filtering and the heartbeat signal obtained after tertiary filtering are used as the reconstructed respiratory and heartbeat signals of the user under test.

[0023] The second-level respiratory signal estimate obtained during the reconstruction of the user's respiratory signal is used as the estimated value of the respiratory signal.

[0024] In one possible implementation, when the triaxial acceleration data of the user under test can be acquired in real time, the multi-stage filtering performed to reconstruct the user's heartbeat signal based on the echo signal is as follows: A first-stage bandpass filter is performed on the beat signal in the wavelet domain to obtain the first-stage filtered heartbeat signal. This is then combined with a peak-and-valley finding algorithm and spectral evaluation to obtain a first-stage heartbeat signal estimate. A second-stage bandpass filter is performed on the first-stage filtered heartbeat signal to obtain a second-stage filtered heartbeat signal. This is then combined with the user's triaxial acceleration data and a photoelectric volumetric mapping (Q-PPG) deep learning algorithm to estimate the heart rate of the second-stage filtered heartbeat signal, obtaining a second-stage heartbeat signal estimate. A third-stage bandpass filter is performed on the second-stage filtered heartbeat signal to obtain a third-stage filtered heartbeat signal. This is then combined with a peak-and-valley finding algorithm and spectral evaluation to obtain a third-stage heartbeat signal estimate. The heartbeat signal obtained after the third-stage filtering is used as the reconstructed heartbeat signal of the user under test.

[0025] In one possible implementation, the method further includes:

[0026] The reconstructed heartbeat signal is converted into a two-dimensional continuous wavelet transform (CWT) spectrum. The CWT spectrum is then input into the YOLO model to mark the frequency band regions in the CWT spectrum that are related to the heart rate of the user being tested, so as to obtain the third estimate of the heartbeat signal.

[0027] The final estimate of the heartbeat signal is determined based on the multi-level heartbeat signal estimates, the first estimate of the heartbeat signal, the second estimate of the heartbeat signal, and the third estimate of the heartbeat signal obtained during the multi-level filtering process.

[0028] In one possible implementation, multi-timescale features of the user's breathing and heartbeat are extracted, including:

[0029] Based on the reconstructed respiratory and heartbeat signals, a first deep learning network is used to extract the respiratory and heartbeat features of the user under test in the first time span using a detection window based on the first time span.

[0030] Based on the reconstructed respiratory and heartbeat signals, a second deep learning network is used to extract the respiratory and heartbeat features of the user under test in the second time span using a detection window based on the second time span.

[0031] The respiratory and heart rate characteristics of the test users at different time spans are used as multi-time-scale features.

[0032] In one possible implementation, the first deep learning network includes: one-dimensional convolution, a causal analyzer, dilated convolution, and residual connections; the one-dimensional convolution is used to capture the temporal variation features of the time series data corresponding to respiratory and heart rate signals within a first time span; the causal analyzer is used to restrict the convolution process to past and present data, ensuring that the prediction results are consistent with the real situation; the dilated convolution is used to help downsampling and provide a larger receptive field; the residual connections accelerate convergence by allowing gradient propagation without the need for activation functions, in order to extract respiratory and heart rate features within the corresponding time span;

[0033] The second deep learning network is a self-supervised learning network, used to learn and extract the spatiotemporal features of respiratory and heartbeat signals within the corresponding time span.

[0034] Secondly, this application provides a non-contact real-time pressure detection system based on millimeter-wave radar, comprising:

[0035] The signal reconstruction module is used to acquire the echo signal of the millimeter-wave radar reflected from the user under test in real time, and reconstruct the respiratory signal and heartbeat signal of the user under test based on the echo signal, and obtain the estimated value of the respiratory signal.

[0036] The final heart rate estimation module is used to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimate of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrogram, input the CWT spectrogram into a YOLO model to mark the frequency band regions in the CWT spectrogram that are related to the heart rate of the user being tested; then highlight the marked frequency band regions, and input the spectrum-highlighted CWT spectrogram into a two-dimensional convolutional neural network to obtain a second estimate of the heartbeat signal; and determine the final estimate of the heartbeat signal based on the first and second estimates of the heartbeat signal.

[0037] The stress detection module is used to extract multi-timescale features of the user's breathing and heartbeat based on the reconstructed breathing and heartbeat signals; and to perform real-time stress detection on the user based on the multi-timescale features of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.

[0038] In one possible implementation, the one-dimensional neural network used in the final heartbeat estimation module includes: multiple convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block, and a Dense connection block; the multiple convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales through different convolutional kernels; the pyramid input block is used to learn features at different scales from the signals input to each convolutional neural network backbone block, and then combine the signals obtained from feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block captures local features at the corresponding scale; the Transformer encoding block is used to learn global features corresponding to the local features of the heartbeat signal output by multiple convolutional neural networks at different scales by combining a multi-head attention mechanism, and capture the most representative heartbeat signal features among the local features at different scales; the Dense connection block is used to obtain the first estimate of the heartbeat signal by combining the output signal of the Transformer encoding block.

[0039] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof.

[0040] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.

[0041] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof.

[0042] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0043] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art:

[0044] This application provides a non-contact real-time pressure detection method and system based on millimeter-wave radar, which combines the advantages of non-intrusive detection via audio-visual analysis and physiological signal detection via wearable devices. It enables more convenient, objective, reliable, and timely detection of pressure status. Regarding the detection method, firstly, clutter suppression and echo selection are performed on the millimeter-wave echo signal to extract the original beat signal of the user to be detected. Secondly, motion artifact suppression is applied to the beat signal, and respiratory and heart rate estimation is performed. One-dimensional and two-dimensional neural networks are fused to further improve the accuracy of heart rate estimation. Finally, respiratory and heart rate signals are fused, and a multi-scale time detection window is used, based on a deep learning network and a classifier, to perform pressure detection. As a non-contact real-time pressure detection method, it has advantages such as flexible and diverse detection equipment, objective and accurate detection results, and strong real-time pressure detection capabilities. It possesses high flexibility and good real-time performance, overcoming the shortcomings of existing physiological signal-based technologies, improving the robustness and accuracy of detection, and providing technical support for the widespread application of pressure detection in real life.

[0045] This application provides a non-contact real-time pressure detection method and system based on millimeter-wave radar. When fusing one-dimensional and two-dimensional neural networks, firstly, in the one-dimensional neural network aspect, multiple convolutional neural network backbone blocks are used to extract local features of the heartbeat signal, a pyramid structure is used to extract multi-scale features, and a Transformer is used to extract global features. This effectively compensates for the shortcomings of traditional convolutional neural networks in multi-scale and long-range dependency modeling, significantly improving the accuracy of the first estimate of the heartbeat signal. Secondly, combined with YOLO's powerful positioning capabilities, the most relevant frequency band to heart rate can be accurately located. Spectrum brightening enhances the spectrum located by YOLO (preserving effective information while appropriately reducing out-of-band noise). Therefore, the enhanced spectrum, when input into the two-dimensional convolutional neural network, effectively improves the accuracy of heart rate estimation. Thus, this application combines one-dimensional and two-dimensional neural networks to improve the accuracy of heart rate estimation, thereby ensuring the precision of pressure detection and the objectivity and accuracy of the detection results. [Attached Image Description]

[0046] Figure 1 is a flowchart of a non-contact real-time pressure detection method based on millimeter-wave radar provided in an embodiment of this application;

[0047] Figure 2 is a block diagram of non-contact real-time pressure detection based on millimeter-wave radar provided in an embodiment of this application;

[0048] Figure 3 is a flowchart of a non-contact real-time pressure detection based on millimeter-wave radar provided in an embodiment of this application;

[0049] Figure 4 is a diagram of a multi-process real-time detection framework based on ZeroMQ provided in an embodiment of this application.

[0050] Figure 5 is a flowchart of millimeter-wave radar beat signal extraction provided in an embodiment of this application;

[0051] Figure 6 is a flowchart of the NLMS algorithm based on polynomial fitting provided in an embodiment of this application;

[0052] Figure 7 is a flowchart of the Q-PPG algorithm provided in an embodiment of this application;

[0053] Figure 8 is an overall architecture diagram of the neural network fusion estimation algorithm for one-dimensional (1D) waveform estimation and two-dimensional (2D) spectrum estimation provided in the embodiments of this application.

[0054] Figure 9 is an architecture diagram of the YOLOv8 pre-trained model provided in an embodiment of this application;

[0055] Figure 10 is an architecture diagram of the DenseNet-201 network provided in an embodiment of this application;

[0056] Figure 11 is a flowchart of the residual-temporal convolution network (Res-TCN) provided in the embodiments of this application;

[0057] Figure 12 is a flowchart of the self-supervised learning (SSL) network provided in the embodiments of this application;

[0058] Figure 13 is a diagram of the real-time pressure detection visualization interface provided in an embodiment of this application;

[0059] Figure 14 is a diagram of the architecture of a non-contact real-time pressure detection system based on millimeter-wave radar provided in an embodiment of this application.

[0060] Figure 15 is an architecture diagram of the electronic device provided in an embodiment of this application.

Detailed Implementation Methods

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0062] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first estimate" and "second estimate," etc., are used to distinguish different estimates, not to describe a specific order of estimates.

[0063] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0064] The embodiments of this application are described below with reference to the accompanying drawings.

[0065] To address the shortcomings of existing technologies, this application aims to provide a non-contact real-time pressure detection method and system based on millimeter-wave radar. It seeks to solve the interference problems faced by existing audio-visual-based pressure detection technologies, such as interference from lighting conditions and human camouflage, as well as the problems faced by millimeter-wave radar-based pressure detection technologies, such as susceptibility to environmental noise and subject limb movements, and limited feature extraction, thereby improving pressure detection accuracy.

[0066] For example, this application provides a non-contact real-time pressure detection method based on millimeter-wave radar, which has advantages such as high reliability, strong robustness, low power consumption, and strong real-time performance. As shown in Figure 1, it includes the following steps:

[0067] Step S101: Real-time acquisition of the echo signal of the millimeter-wave radar reflected from the user under test, and reconstruction of the respiratory signal and heartbeat signal of the user under test based on the echo signal, and obtaining the estimated value of the respiratory signal.

[0068] Step S102: The reconstructed heartbeat signal is input into a one-dimensional neural network to obtain a first estimate of the heartbeat signal; the reconstructed heartbeat signal is converted into a two-dimensional CWT spectrogram, and the CWT spectrogram is input into a YOLO model to mark the frequency band regions in the CWT spectrogram that are related to the heart rate of the user being tested; then the marked frequency band regions are highlighted, and the spectrum-highlighted CWT spectrogram is input into a two-dimensional convolutional neural network to obtain a second estimate of the heartbeat signal; the final estimate of the heartbeat signal is determined based on the first estimate and the second estimate of the heartbeat signal.

[0069] In one example, a one-dimensional neural network includes: multiple convolutional neural network backbone blocks, pyramid input blocks, Transformer encoding blocks, and Dense connection blocks;

[0070] The multiple convolutional neural network backbone blocks are used to capture local features of the initially estimated heartbeat signal at different scales through different convolutional kernels;

[0071] The pyramid input module is used to perform feature learning at different scales on the signals input to each convolutional neural network backbone block, and then combine the signals obtained from feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block can capture local features at the corresponding scale.

[0072] The Transformer encoding block is used to combine a multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the multiple convolutional neural networks at different scales, and to capture the most representative features of the heartbeat signals among the local features at different scales.

[0073] The Dense connection block is used to combine the output signal of the Transformer coding block to obtain a first estimate of the heartbeat signal.

[0074] Understandably, the one-dimensional neural network used here is slightly different from the traditional one-dimensional convolutional neural network. The one-dimensional neural network used in this application extracts local features using multiple convolutional neural network backbone blocks, extracts multi-scale features using a pyramid structure, and extracts global features using a Transformer. This effectively makes up for the shortcomings of traditional convolutional neural networks in multi-scale and long-range dependency modeling, and can greatly improve the accuracy of the first estimate of the heartbeat signal, thereby improving the accuracy and reliability of stress detection.

[0075] In one embodiment, the YOLO model described above can be the YOLOv8 model. The two-dimensional convolutional neural network described above can be the DenseNet-201 network.

[0076] In another example, the method further includes the following steps: converting the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, inputting the CWT spectrogram into a YOLO model to mark the frequency band region in the CWT spectrogram that is related to the heart rate of the user being tested, so as to obtain a third estimate of the heartbeat signal.

[0077] Further, determining the final estimate of the heartbeat signal based on the first estimate and the second estimate of the heartbeat signal includes:

[0078] The final estimate of the heartbeat signal is determined based on the first estimate, the second estimate, and the third estimate of the heartbeat signal.

[0079] It should be noted that, due to YOLO's powerful positioning capabilities, it can accurately locate the frequency band most relevant to heart rate. Spectrum brightening enhances the spectrum located by YOLO (preserving effective information while appropriately reducing out-of-band noise). Therefore, the enhanced spectrum, when input into the two-dimensional convolutional neural network, significantly improves the accuracy of heart rate estimation. Consequently, the second estimate has a more critical impact on the accuracy of the final heart rate estimation than the third estimate. Subsequent experimental comparisons also revealed that the accuracy of the final heart rate estimate determined based on the first and second estimates is better than that determined based on the first and third estimates. Furthermore, the scheme based on the first, second, and third estimates is slightly better than the scheme combining the first and second estimates; therefore, this scheme can ultimately be considered as a further optimization.

[0080] Step S103: Extract multi-timescale features of the user's breathing and heartbeat based on the reconstructed breathing and heartbeat signals;

[0081] Step S104: Real-time stress detection is performed on the user under test based on the multi-timescale features of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.

[0082] Furthermore, since the YOLO model can be reused in determining the second and third estimates in the embodiments of this application, the third estimate may be described first and the second estimate later in subsequent embodiments of this application, taking into account the order in which the YOLO model appears. The above order adjustment is easy to understand for those skilled in the art and will not be explained in detail thereafter.

[0083] Specifically, as shown in Figure 2, the method provided in this application is based on the ZeroMQ framework and relies on four processes to carry out real-time millimeter-wave data acquisition, millimeter-wave data processing, vital sign monitoring, and stress state detection:

[0084] In one example, the above processes can be implemented using the following modules:

[0085] The first process, the millimeter-wave data acquisition module, acquires millimeter-wave echo data through a UDP (user datagram protocol) port;

[0086] The second process, the millimeter-wave data processing module, extracts the raw beat signal from the millimeter-wave echo signal;

[0087] The third process is the vital signs monitoring module, which monitors breathing and heartbeat.

[0088] The fourth process, the pressure state detection module, identifies pressure states and non-pressure states.

[0089] In one possible embodiment, the system framework provided in this application can be as shown in Figure 3: First, the detection system transmits low-power millimeter waves to the user to be detected via a millimeter-wave module and detects the echo signal generated by the reflection of the signal from the human body (e.g., chest cavity); second, after clutter suppression and echo selection of the echo signal, the original beat signal is extracted; then, based on motion artifact suppression, the reconstructed respiratory and heartbeat signals are obtained, and respiratory and heartbeat frequency estimation is performed. Combining one-dimensional and two-dimensional neural networks further improves the accuracy of heart rate estimation; finally, using multimodal and multi-timescale fusion based on a deep learning network, high-precision pressure detection is achieved. Specifically, it includes four basic modules:

[0090] (1) Millimeter-wave data acquisition module. This module captures and receives millimeter-wave data via a UDP port and transmits the captured data to the host computer. First, it initializes key parameters such as the detection range and accuracy of the millimeter-wave radar. Second, it monitors the UDP port in real time to capture UDP packets via a Socket interface; simultaneously, it converts the binary data into int16 format. Finally, the converted data enters a 30-second queue for subsequent data processing.

[0091] During system initialization, when data is less than 30 seconds old, all data (regardless of length) is directly transmitted to subsequent processing modules. Once data reaches 30 seconds, data processing and transmission are framed in 30-second queues to ensure continuous operation of the real-time system and efficient data processing, as shown in Figure 4. After subsequent processing is complete, the system immediately retrieves the latest data from the queue and performs analysis. This seamless cycle effectively balances real-time data response speed and processing efficiency, ensuring stable system operation under varying loads.

[0092] (2) Millimeter-wave data processing module. This module processes individual chirps, suppresses clutter, and finally extracts the original beat signal. This module consists of three units:

[0093] (2-1) Single chirp processing unit. First, each frame of received signal contains one chirp; signals from multiple frames are received sequentially to form an echo matrix. Then, FFT (Fast Fourier Transform) is used...

[0094] The Lyotransform generates a range-time (distance-time) graph, using a matrix. As shown in Figure 5, in this range-time graph, vertical bright bars represent possible periodic vital signs, and the position of the bright bars on the horizontal axis from left to right indicates the distance units from near to far.

[0095] (2-2) Clutter Suppression Unit. Two main types of clutter exist in radar applications: static clutter generated by stationary objects and non-static clutter generated by moving objects. This system employs adaptive background subtraction to eliminate static clutter and singular value decomposition to eliminate non-static clutter. The range-time graph after clutter suppression is denoted as matrix Q.

[0096] (2-3) Extraction of beat signal units. After clutter suppression, the distance between the detected target and the radar is confirmed, and the original vital signs signal is extracted from the range unit where the detected target is located. Specifically, firstly, the detected target is identified by locating the range unit with the highest energy, and its corresponding range unit index is marked as n. max Subsequently, at a distance of unit n max Extract the original signal Q(:,n) max The phase of this signal is extracted to represent the original beat signal. The beat signal is a mixture of breathing, heartbeat, and noise signals, and its extraction method is as follows: Here, the arctan operation is used to determine the phase, while the unwrap operation is used to avoid phase entanglement and ensure phase continuity.

[0097] (3) Vital Signs Monitoring Module. This module is responsible for monitoring respiration and heart rate. Given that chest vibrations caused by respiration are stronger than those caused by heartbeats, heart rate monitoring presents more challenges than respiration monitoring. Therefore, this application employs motion artifact suppression and time-spectrum fusion techniques for heart rate estimation; while respiration monitoring uses traditional motion artifact suppression and partial multi-level filtering, as shown in Figure 3 (highlighted underlined). This module comprises two units:

[0098] (3-1) Motion Artifact Suppression Unit. As shown in Figure 3, this application suppresses motion noise through a combination of a traditional motion artifact suppression algorithm, a multi-stage filtering algorithm, and an improved motion artifact suppression algorithm. Since the improved motion artifact suppression algorithm relies on real-time acquisition of triaxial acceleration data, and considering the potential loss of this data in real-world applications due to equipment limitations, the improved motion artifact suppression algorithm is considered an optional rather than mandatory component in this application. The impact of employing or not employing the improved motion artifact suppression algorithm on the accuracy of heart rate estimation will be discussed further later in this application.

[0099] It should be noted that reconstructing the respiratory and heartbeat signals of the user under test based on the echo signals includes: extracting the corresponding beat signals from the echo signals; the beat signals carrying the respiratory and heartbeat signals of the user under test; using a normalized least mean squares (NLMS) algorithm based on polynomial fitting to initially suppress motion artifacts in the beat signals; and employing continuous wavelet transform. The transform (CWT) algorithm performs multi-stage filtering on the beat signal after initial suppression of motion artifacts. The multi-stage filtering includes: performing a first-stage bandpass filtering on the beat signal in the wavelet domain to obtain first-stage filtered respiratory and heartbeat signals; then combining peak-and-valley search algorithms and spectral evaluation to obtain first-stage respiratory and heartbeat signal estimates; performing a second-stage bandpass filtering on the first-stage filtered respiratory and heartbeat signals to obtain second-stage filtered respiratory and heartbeat signals; then combining peak-and-valley search algorithms and spectral evaluation to obtain second-stage respiratory and heartbeat signal estimates; and performing a third-stage bandpass filtering on the second-stage filtered heartbeat signal to obtain a third-stage filtered heartbeat signal; finally, combining peak-and-valley search algorithms and spectral evaluation to obtain a third-stage heartbeat signal estimate. The center frequency of the next-stage bandpass filtering is set with reference to the respiratory and heartbeat signal estimates of the previous stage.

[0100] The respiratory signal obtained after secondary filtering and the heartbeat signal obtained after tertiary filtering are used as the reconstructed respiratory and heartbeat signals of the user under test.

[0101] The second-level respiratory signal estimate obtained during the reconstruction of the user's respiratory signal is used as the estimated value of the respiratory signal;

[0102] (3-1-1) Traditional Motion Artifact Suppression Algorithm. This application first introduces a Normalized Least Mean Square Error (NLMS) algorithm based on polynomial fitting, as shown in Figure 6, to initially suppress motion artifacts. This algorithm does not require prior knowledge of the statistical characteristics of the input signal and noise; it uses polynomial fitting of the noise and dynamic iterative adjustment of parameters using NLMS. Specifically, the key to this algorithm lies in the selection of the reference signal, and this application uses polynomial fitting of the noise d(m):

[0103] Where, m i p is the variable m raised to the power of i. i (where i = 0, 1, ..., γ) are constant coefficients. In this embodiment, the 30s window waveform data is divided into three 10s windows for polynomial fitting; within each 10s window, γ is 50. Then, NLMS is used to calculate the relevant output signal y(m) from the filter coefficients w(m) and the reference signal d(m): y(m) = w(m) T d(m)

[0104] The estimation error e(m) can be expressed as:

[0105] According to the least mean square criterion, the error gradually converges. The update formula for the filter weight w(m) is:

[0106] Where μ is the step size of the adaptive filtering algorithm, and δ is a constant. Finally, This is a combination of respiratory and heartbeat signals after noise suppression.

[0107] (3-1-2) Multi-stage filtering algorithm. A three-stage filtering process using Continuous Wavelet Transform (CWT) is employed, as shown in Figure 3. First, a wide-bandwidth filter is used. Based on the initial frequency estimation, a narrow-band filter is then used, centered on that frequency, to further eliminate noise. For ease of explanation, a heartbeat signal is used as an example.

[0108] First, a primary filtering stage is performed. The heartbeat signal is then reconstructed using CWT. That is, bandpass filtering is performed in the wavelet domain, followed by inverse CWT operation. The frequency range of this filter is [0.7, 2.5] Hz, the center frequency is 1.6 Hz, and the half-bandwidth is 0.9 Hz. Next, heart rate estimation is performed by combining traditional peak-and-valley finding algorithms and spectral evaluation methods, denoted as hr1. The corresponding heart rate calculation formula is f. hr1 =hr1 / 60.

[0109] Next, a second-level filter is performed. The heartbeat signal is then further reconstructed using CWT. Compared to the initial filter, this filter has a narrower bandwidth, with a half-bandwidth HBW2 ≤ 0.9 Hz and a frequency range of [f 12 ,f 22 ]Hz:

[0110] In this application, HBW2 is set to a default value of 0.4Hz. Next, heart rate estimation, denoted as hr2, is performed using a combination of traditional peak-and-valley finding algorithms and spectral evaluation methods. The corresponding heart rate calculation formula is f. hr2 =hr2 / 60.

[0111] Finally, a three-stage filtering process is performed. The heartbeat signal is then reconstructed using CWT. Similar to the intermediate filter, this filter has a narrow bandwidth; the default value for HBW3 is 0.4Hz, and the bandpass filter cutoff frequency is:

[0112] Next, heart rate estimation is performed by combining traditional peak-and-valley finding algorithms and spectral evaluation methods, denoted as hr3; the corresponding formula for calculating heart rate is f. hr3 =hr3 / 60.

[0113] It should be noted that, as shown in Figure 3, if real-time triaxial acceleration data is unavailable, then f hr2 It is obtained directly from the second-level filter; however, if real-time triaxial acceleration data is available, then f hr2 Obtained by an improved motion artifact suppression algorithm. In the improved motion artifact suppression algorithm, x h2 (m) and f hr2 These are the input waveform data and the output heart rate value, respectively.

[0114] Furthermore, when the triaxial acceleration data of the user under test can be acquired in real time, the multi-level filtering performed on the heartbeat signal reconstructed based on the echo signal is as follows: First-level bandpass filtering is performed on the beat signal in the wavelet domain to obtain the first-level filtered heartbeat signal. Then, a peak-and-valley finding algorithm and spectral evaluation are used to obtain the first-level heartbeat signal estimate. Second-level bandpass filtering is performed on the first-level filtered heartbeat signal to obtain the second-level filtered heartbeat signal. Then, combined with the triaxial acceleration data of the user under test, the Q-PPG deep learning algorithm is used to estimate the heart rate of the second-level filtered heartbeat signal to obtain the second-level heartbeat signal estimate. Third-level bandpass filtering is performed on the second-level filtered heartbeat signal to obtain the third-level filtered heartbeat signal. Then, a peak-and-valley finding algorithm and spectral evaluation are used to obtain the third-level heartbeat signal estimate. The heartbeat signal obtained after the third-level filtering is used as the reconstructed heartbeat signal of the user under test.

[0115] (3-1-3) Improved Motion Artifact Suppression Algorithm. When triaxial acceleration data is available in real time, this application employs a simplified version of the Q-PPG (Quantized-PPG) deep learning algorithm to further suppress motion artifacts, as shown in Figure 7. This algorithm consists of two stages:

[0116] First, in the initialization phase, a simplified framework is generated using a seed network. Specifically, the seed network is a slightly modified version of TEMPONet, consisting of three convolutional blocks and three fully connected layers, as shown in Figure 7. Each convolutional block contains three convolutional layers. To expand the search space during subsequent architecture exploration, the seed network is modified from the original TEMPONet: the dilation parameter of all convolutional layers is set to 1, while the filter size is increased to maintain the same receptive field. As shown in Figure 7, the initial loss function is defined as L... task +λ*L cost L task and L cost These represent the task loss term and the regularization loss term, respectively. The task loss measures the prediction error for a specific task (such as heart rate estimation) to ensure model accuracy. The regularization loss imposes complexity constraints by controlling the mask value; the coefficient λ represents the regularization strength and is used to adjust the relative importance of the two loss terms to maintain a balance between heart rate tracking error and model complexity. In this application, the default value of λ is 10. -6 .

[0117] Secondly, in the optimization phase, two neural architecture search (NAS) tools, MorphNet and Pruning-In-Time (PIT), are used to optimize the network architecture. Specifically, MorphNet introduces mask parameters (i.e., α masks) into convolutional and fully connected layers, gradually removing weights with small mask values ​​and their associated output channels, thereby optimizing the number of output channels. On the other hand, PIT introduces mask parameters (i.e., β masks) at the time step of convolutional layers, adjusting the dilation rate of convolutional filters to affect the receptive field, thereby reducing computational load while maintaining accuracy.

[0118] (3-2) Time-frequency fusion estimation unit. Compared with respiratory signals, heartbeat signals are relatively weak. Therefore, this application adopts a neural network hybrid model that combines one-dimensional waveform estimation and two-dimensional spectrum estimation, as shown in Figure 8, which further improves the accuracy of heart rate estimation based on the fusion of time and frequency domains.

[0119] (3-2-1) In a one-dimensional neural network model, the reconstructed heartbeat signal x h3 (m) is first normalized to the range [-1, 1], and then processed by a one-dimensional neural network to estimate the heart rate hr. 1d (i.e., the first estimate of the heartbeat signal).

[0120] Specifically, the one-dimensional neural network model based on 1D waveform estimation fully leverages the advantages of convolutional operations and the Transformer. It extracts local features through convolutional layers to effectively capture detailed information, and extracts global features through the Transformer to effectively capture contextual relationships. This not only compensates for the shortcomings of convolution in long-range dependency modeling but also reduces the computational burden on the Transformer. The model consists of four parts: a Convolutional Neural Network (CNN) backbone, a pyramid input block, a Transformer encoding block, and a Dense connection block, as shown in Figure 8.

[0121] First, the CNN backbone. At the network's bottom layer, to achieve a good balance between model complexity and performance, four consecutive CNN backbone blocks are proposed for local feature compression extraction. These backbone blocks aim to effectively extract local features while significantly reducing the computational burden on subsequent steps. Each backbone block consists of 1D convolutions, batch normalization (BN), a rectified linear unit (ReLU) activation function, average pooling, and dropout. The 1D convolutions will use four different kernel sizes: 1×15, 1×11, 1×5, and 1×3. The kernel size will gradually decrease as the input changes to capture features at different scales. Similarly, average pooling downsamples features with kernel sizes of 5, 5, 2, and 2, effectively reducing size while preserving important feature information. Dropout aims to further improve generalization ability; its default value is 0.2, effectively preventing overfitting.

[0122] Second, the pyramid input block. This module aims to improve model performance by introducing feature learning at different scales, fusing multi-scale heartbeat signals, obtaining multi-layered receptive fields, and acquiring more raw signal information. Specifically, this module uses nearest-neighbor interpolation to downsample the input signal to different degrees: 1 / 15, 1 / 5, and 1 / 2. These downsampled signals are then combined with the corresponding CNN backbone blocks to participate in subsequent convolutional operations.

[0123] Third, the Transformer encoding block. Since heartbeat signals are still affected by minor limb movements or breathing noise even at rest, directly using local features learned by CNNs for heart rate estimation often yields unsatisfactory results. Furthermore, heartbeat signals are essentially time-series signals, and extracting only local features is insufficient. To address this issue, a Transformer multi-head attention mechanism is proposed, which effectively learns global features and captures the dependencies between heartbeats, thus providing a powerful complement to CNNs. Specifically, the processing flow of this module is as follows: the input heartbeat signal is compressed, and the positions of each dimension are swapped; after 6 layers of Transformer encoding, long-range dependencies in the signal are captured through a self-attention mechanism; the average value of the multi-channel dimensions is calculated and output, extracting the most representative features. The hyperparameters of this module are set as follows: d_model (model dimension) is 12, nhead (number of heads in multi-head attention) is 6, dim_feedforward (dimension of the feedforward network) is 24, and dropout is 0.2.

[0124] Fourth, Dense Connector Blocks. Two dense connectors will be used to predict the final heart rate estimate.

[0125] The hyperparameters of the 1D neural network are set as follows: epochs = 1000, learning rate = 0.001, batch size = 32; the optimizer is Adam, with hyperparameters set as beta_1 (exponential decay rate of the first moment estimate of the gradient) = 0.9, beta_2 (exponential decay rate of the second moment estimate of the gradient) = 0.999, and epsilon (numerical stability) = 1e-08.

[0126] (3-2-2) In the two-dimensional neural network model, the reconstructed heartbeat signal x h3 (m) is transformed into a two-dimensional CWT spectrogram, and heart rate estimation is performed using a cascaded network composed of YoLov8 and DenseNet-201. Specifically,

[0127] On one hand, the CWT spectrogram is input into the YOLOv8 pre-trained model. The architecture of the YOLOv8 pre-trained model is shown in Figure 9. The core frequency band region related to heart rate is outlined, and the estimated heart rate value hr is output. 2dY (i.e., the third estimate of the heartbeat signal). The hyperparameters of the network are set as follows: batch size is 10, epochs are 200, and learning rate is 0.01.

[0128] On the other hand, a linear transformation is used to adjust the pixel values ​​within the frequency band framed by YOLOv8 to improve the contrast of the CWT spectrogram. pixel new pixel =clip(α·old) pixel +β,0,255)

[0129] In this model, α controls contrast (>1 indicates increased contrast, <1 indicates decreased contrast), β controls brightness (>0 indicates increased brightness, <0 indicates decreased brightness), and clip limits the result to the range [0, 255] to avoid overflow. In this embodiment, α = 1.2 and β = 30. The enhanced CWT spectrum is then input into the fine-tuned DenseNet-201 network, the architecture of which is shown in Figure 10. During fine-tuning, the bottom layer of the pre-trained DenseNet-201 model is retained, while the top layer is retrained using the two-dimensional CWT spectrum from this application, outputting a heart rate estimate hr. 2dD (i.e., the second estimate of the heartbeat signal). The hyperparameters of the network are set as follows: batch size of 16, epochs of 800, and learning rate of 0.0025. It should be noted that the main reason why this application uses CWT to generate the two-dimensional spectrogram is that, although CWT, short-time Fourier transform (STFT), and smoothed pseudo-Wigner–Ville distribution (SPWVD) are all suitable for analyzing time-varying signals, CWT has higher accuracy and more efficient processing capabilities compared to STFT and SPWVD.

[0130] (3-2-3) Preferably, the final estimated value of the heartbeat signal is determined based on the multi-level heartbeat signal estimates, the first estimated value of the heartbeat signal, the second estimated value of the heartbeat signal, and the third estimated value of the heartbeat signal obtained in the multi-level filtering process.

[0131] For example, heart rate estimates (hr) from one-dimensional and two-dimensional networks 1d ,hr 2dY ,hr 2dD The hr1 to hr3 values ​​obtained from multi-level filtering during motion artifact suppression are input into the ridge regression ensemble model to generate the final heart rate estimate hr (i.e., the final estimate of the heartbeat signal).

[0132] (4) Stress State Detection Module. This module aims to identify stress states based on multimodal and multi-scale fusion. This application employs two deep learning networks: a residual-temporal convolution network (Res-TCN) and a self-supervised learning network (SSL), which process the reconstructed respiratory and heart rate signals using a long detection window of 30 seconds and a short detection window of 10 seconds, respectively. After analysis by the two networks, the 30-second and 10-second multi-timescale features are fused, and the respiratory and heart rate features are fused, improving the accuracy of stress recognition. This module includes three units:

[0133] (4-1) Res-TCN Unit. The Res-TCN consists of four residual blocks: one-dimensional convolution, a causal analyzer, dilated convolution, and residual connections, as shown in Figure 11. The one-dimensional convolution captures the temporal variation characteristics of the time-series data; the causal analyzer ensures that the prediction results are consistent with the real-world scenario by limiting the convolution process to past and present data; the dilated convolution facilitates downsampling and provides a larger receptive field; and the residual connections accelerate convergence by allowing gradient propagation without the need for activation functions.

[0134] Specifically, each residual block in the network consists of three residual sub-blocks, and each sub-block contains three layers, representing a batch normalization layer, an activation function layer, and a convolutional layer, respectively. The parameters of the convolutional layer are represented by Res-U(F,K,S,D), where F represents the number of filters, K represents the filter size, S represents the stride, and D represents the dilation rate.

[0135] (4-2) SSL Unit. SSL consists of two parts: a self-supervised learning part and a stress detection part, as shown in Figure 12.

[0136] First, in the self-supervised learning part, a multi-task CNN model is trained using automatically generated labels to extract high-level abstract features from respiratory or heartbeat time series. Specifically, this module uses the recognition of reconstructed respiratory and heartbeat signals and their six signal transformations as a pre-training task, enabling the network to learn spatiotemporal features and thus improve its generalization ability. The weights obtained during training are then transferred to the stress detection part.

[0137] Next, in the stress detection part, the convolutional layers remain frozen, and the weights are consistent with those obtained in the self-supervised learning part, as shown in Figure 12; while the fully connected layers are fine-tuned to perform stress detection.

[0138] (4-3) Classifier Unit. As shown in Figure 3, this application uses a classifier for stress identification based on multimodal and multiscale fusion.

[0139] (4-3-1) First, features from multiple time scales of the respiratory and heartbeat signals are aggregated, including: (i) Res-TCN features extracted from a 30-second detection window and (ii) SSL features extracted from a 10-second detection window. Specifically, after aggregation, a total of 8 features are obtained in each 30-second detection window, including 4 respiratory-related features and 4 heartbeat-related features. Taking the respiratory features as an example, 1 feature comes from Res-TCN and 3 features come from SSL.

[0140] (4-3-2) Then, these 8 features, along with the estimated values ​​of the respiratory signal and the final estimated values ​​of the heartbeat signal, are input into a random forest classifier to classify the state of the user under test into two categories: non-stress and stress, and the results are visualized as shown in Figure 13. The identification results are updated once per second.

[0141] (5) To verify the reliability of the non-contact real-time stress detection method and system proposed in this application, this embodiment recruited 15 subjects to participate in two tests, each lasting 5 minutes (one test under non-stress conditions and the other under stress conditions induced by a mental arithmetic experiment; the two tests were 2 minutes apart). After each test, the subjects completed a simplified version of the State-Trait Anxiety Inventory (STAI). This questionnaire contains 6 questions to assess whether the subjects are under stress: a score greater than 15 points is considered a stress condition, and a score less than or equal to 15 points is considered a non-stress condition. During the experiment, an IWR1642BOOST+DCA1000EVM millimeter-wave radar device was used to collect millimeter-wave data for vital sign detection and stress identification. At the same time, a wearable ECG device MAX-ECG-MONITOR was used to collect ECG data as a benchmark for millimeter-wave heart rate estimation; a Shimmer3 IMU device was used to collect triaxial acceleration data to suppress motion artifacts. The visualization effect of this embodiment is shown in Figure 13, and the experimental results are shown in Tables 1 and 2.

[0142] (5-1) The heart rate estimation results of this embodiment are shown in Table 1. This embodiment used three metrics for performance evaluation: MAE (mean absolute error), RMSE (root mean square error), and Corr (correlation coefficient). As can be seen from Table 1, in both the motion artifact suppression and heart rate estimation stages, the method proposed in this embodiment exhibits superior performance in both scenarios with and without real-time three-axis accelerometer signals (wACCs).

[0143] Table 1. Experimental results of motion artifact suppression and heart rate estimation.

[0144] Explanation: w ACCs and wo ACCs represent scenarios with and without real-time triaxial acceleration signals, respectively; CorNET indicates heart rate estimation using the CorNET one-dimensional network; 1D DCNN indicates heart rate estimation using 1D DCNN; 1D(CNN+Pyramid+Transformer) indicates heart rate estimation using the one-dimensional neural network model proposed in this embodiment; VGG16+STFT indicates using STFT to generate a spectrogram and inputting it into the VGG16 network for heart rate estimation; DenseNet+STFT indicates using STFT to generate a spectrogram and inputting it into the DenseNet network for heart rate estimation; ResNet+CWT indicates using CWT to generate a spectrogram and inputting it into the ResNet network for heart rate estimation; DenseNet+CWT indicates using CWT to generate a spectrogram and inputting it into the DenseNet network for heart rate estimation; 2D YoLo This indicates that a spectrogram is generated using CWT and input into the YOLO model for heart rate estimation; 2D YoLo+频谱提亮+DenseNet This indicates that the spectrogram generated using CWT is input into the YOLO model for spectrogram enhancement, and then input into the DenseNet network for heart rate estimation; 2D(2D) YoLo +2D YoLo+频谱提亮+DenseNet This indicates that a spectrogram is generated using CWT and input into the YOLO model for heart rate estimation. Then, the spectrogram is enhanced and input into the DenseNet network for further heart rate estimation. The results of the two heart rate estimations are then integrated to obtain the final heart rate estimate; 1D+2D YoLoThis indicates that 1D is used for heart rate estimation, followed by 2D. YoLo Heart rate estimation is performed, and the results of two heart rate estimations are integrated to obtain the final heart rate estimate; 1D+2D YoLo+频谱提亮 +DenseNet This indicates that 1D is used for heart rate estimation, followed by 2D. YoLo+频谱提亮+DenseNet Heart rate estimation is performed, and the results of two heart rate estimations are integrated to obtain the final heart rate estimate; 1D+2D+hr 1~3 This means that heart rate estimation is performed using 1D, then 2D, and finally the three heart rate estimation results obtained from the multi-stage filtering process are integrated to obtain the final heart rate estimate.

[0145] (5-1-1) Regarding the WO ACCs case, firstly, in the motion artifact suppression stage, experimental results show that the method proposed in this embodiment can effectively suppress motion noise compared to the six leading traditional algorithms. Specifically, the MAE of the method proposed in this embodiment is 4.501 bpm (beats per minute), which is better than the best-performing multinomial fitting-based LMS algorithm among the six leading comparison algorithms, with an MAE of 4.545 bpm. Subsequently, in the heart rate estimation stage, the 1D+2D+hr proposed in this embodiment... 1~3 The proposed method significantly outperforms six leading-edge comparison algorithms. The six leading-edge comparison algorithms have the lowest MAE (4.320 bpm) for 1D DCNN, while the proposed method in this embodiment has an MAE of 3.780 bpm. Furthermore, Table 1 also demonstrates the following points:

[0146] 1) The 1D (CNN + Pyramid + Transformer) method proposed in this embodiment (MAE of 4.059 bpm) outperforms two traditional 1D DCNN methods: CorNET (MAE of 5.642 bpm) and 1D DCNN (MAE of 4.320 bpm). The reason why the 1D method proposed in this embodiment is significantly better than the traditional methods is that this embodiment uses CNN to extract local features, uses the pyramid structure to extract multi-scale features, and uses Transformer to extract global features, effectively making up for the shortcomings of traditional DCNN networks in multi-scale and long-range dependency modeling.

[0147] 2) YOLO outperforms traditional 2D DCNN networks. In this embodiment, YOLO's MAE is 4.366 bpm, while the lowest MAE of traditional 2D methods is 4.411 bpm (DenseNet+CWT). The reason why the YOLO method used in this embodiment is significantly better than traditional 2D DCNN networks is that, for YOLO, heart rate estimation is regarded as a spectral localization problem. This perspective can make full use of YOLO's high-precision localization capabilities, thereby improving the accuracy of heart rate estimation.

[0148] 3) The combination of YOLO, spectrum enhancement, and DenseNet (MAE 4.105 bpm) outperforms YOLO alone (MAE 4.366 bpm) or DenseNet alone (MAE 4.411 bpm). This is because YOLO's powerful localization capabilities accurately pinpoint the frequency band most relevant to heart rate. Spectrum enhancement enhances the spectrum located by YOLO (preserving effective information while appropriately reducing out-of-band noise). Therefore, the enhanced spectrum, when input into DenseNet, significantly improves the accuracy of heart rate estimation.

[0149] 4) The 1D+2D method (MAE 3.815 bpm) outperforms the standalone 1D (MAE 4.059 bpm) or 2D (MAE 4.070 bpm) methods. This is because the fusion of 1D time-domain features and 2D frequency-domain features improves the accuracy of heart rate estimation.

[0150] 5) 1D+2D+hr 1~3 (MAE is 3.780 bpm) This method outperforms all other methods. The reason is that this method not only integrates the time-domain and frequency-domain features extracted by the deep learning network, but also makes full use of the time-domain features of the three previous filtered signals, thus achieving the highest heart rate recognition accuracy.

[0151] (5-1-2) Regarding wACCs, firstly, in the motion artifact suppression stage, experimental results show that the DL method proposed in this embodiment has significant advantages compared to three cutting-edge deep learning (DL) algorithms. Specifically, the best MAE of the three cutting-edge methods is 3.779 bpm (Q-PPG), while the MAE of the method proposed in this embodiment is 3.603 bpm. On the one hand, the performance of this embodiment is better than the three cutting-edge DL comparison methods; on the other hand, compared with traditional motion artifact suppression methods, this embodiment is better than the best-performing LMS based on polynomial fitting among the six cutting-edge comparison algorithms, with an MAE of 4.545 bpm. This result shows that combining the traditional method (NLMS based on polynomial fitting) with the DL framework (Q-PPG) can further enhance the motion artifact suppression performance. Secondly, in the heart rate estimation stage, experimental results again confirm that the 1D+2D+hr method proposed in this embodiment has significant advantages compared to the three cutting-edge deep learning (DL) algorithms. 1~3 The effectiveness of the method is demonstrated. The best MAE among the six leading-edge comparison algorithms is 3.281 bpm (DenseNet+CWT), while the MAE of the method proposed in this embodiment is 2.824 bpm. Furthermore, the relevant results further validate the five conclusions drawn from the WO ACCs data in Table 1.

[0152] In summary, the method proposed in this embodiment fully considers practical application scenarios, including w0 and w0 ACCs, and provides a practical solution. The proposed method demonstrates superior performance in both motion artifact suppression and heart rate estimation, showcasing its potential in vital sign monitoring systems.

[0153] (5-2) The stress detection results of this embodiment are shown in Table 2. In order to evaluate the stress detection performance, this embodiment uses five-fold cross-validation and leave-one-out cross-validation on a self-built dataset: on the one hand, five-fold cross-validation ignores the differences of individual subjects and aims to determine the generalization ability of the algorithm on the entire dataset; on the other hand, leave-one-out cross-validation focuses on the learning of irrelevant subjects and aims to evaluate the generalization ability of the algorithm among different subjects.

[0154] Table 2 Results of the 50% and Leave-one-out cross-validation experiments

[0155] Note: B and H represent respiratory and heart rate data, respectively.

[0156] (5-2-1) As shown in Table 2, this embodiment used four metrics for performance evaluation: accuracy, precision, F1 score, and area under the curve (AUC). Experimental results show that for five-fold cross-validation, the values ​​of the four relevant metrics in the last row of Table 2 all exceed 0.940, indicating that this embodiment can accurately identify stress; similarly, for leave-one-out cross-validation, the values ​​of the four relevant metrics in the last row of Table 2 are all greater than 0.953, indicating that this embodiment can accurately identify the stress state of unknown subjects based on data from known individuals.

[0157] (5-2-2) Table 2 also shows a comparison between the proposed method in this embodiment and four cutting-edge methods. In general, these four cutting-edge methods can be divided into two categories: traditional feature extraction methods (handcrafted) such as EQ-Radio, and deep learning methods (DL) such as Stressalyzer, Res-TCN, and SSL. The experimental results in Table 2 show that the proposed method in this embodiment significantly outperforms these four cutting-edge methods. For example, for five-fold cross-validation, the stress recognition accuracy of the proposed method in this embodiment is 0.940, significantly higher than the traditional feature extraction method EQ-Radio (see the third row of Table 2) and the best-performing deep learning algorithm Res-TCN (see the seventh row of Table 2), whose accuracies are 0.713 and 0.820, respectively.

[0158] (5-2-3) Furthermore, the data in the last three rows of Table 2 show that the multimodal multiscale fusion method proposed in this embodiment performs significantly better than the two single-data methods. For example, for five-fold cross-validation, the accuracy of the method proposed in this embodiment is 0.940, which is significantly higher than the accuracy of the other two single-data methods, which are 0.793 and 0.873, respectively.

[0159] Figure 14 is a diagram of the architecture of a non-contact real-time pressure detection system based on millimeter-wave radar provided in an embodiment of this application; as shown in Figure 14, it includes:

[0160] The signal reconstruction module 1410 is used to acquire the echo signal of the millimeter-wave radar reflected from the user under test in real time, and reconstruct the respiratory signal and heartbeat signal of the user under test based on the echo signal, and obtain the estimated value of the respiratory signal.

[0161] The heart rate final estimation module 1420 is used to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain a first estimate of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional CWT spectrogram, input the CWT spectrogram into a YOLO model to mark the frequency band regions in the CWT spectrogram that are related to the heart rate of the user being tested; then highlight the marked frequency band regions, and input the spectrum-highlighted CWT spectrogram into a two-dimensional convolutional neural network to obtain a second estimate of the heartbeat signal; and determine the final estimate of the heartbeat signal based on the first estimate and the second estimate of the heartbeat signal.

[0162] The stress detection module 1430 is used to extract multi-timescale features of the user's breathing and heartbeat based on the reconstructed breathing and heartbeat signals; and to perform real-time stress detection on the user based on the multi-timescale features of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.

[0163] It should be understood that the above system is used to execute the methods in the above embodiments. The corresponding program modules in the system are similar in implementation principle and technical effect to those described in the above methods. The working process of the system can be referred to the corresponding process in the above methods, and will not be repeated here.

[0164] Based on the methods in the above embodiments, this application provides an electronic device, as shown in FIG15. The electronic device may include a processor 1510, a communication interface 1520, a memory 1530, and a communication bus 1540, wherein the processor 1510, the communication interface 1520, and the memory 1530 communicate with each other via the communication bus 1540. The processor 1510 may call logical instructions in the memory 1530 to execute the methods in the above embodiments.

[0165] Furthermore, the logical instructions in the aforementioned memory 1530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0166] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0167] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0168] It is understood that the processor in the embodiments of this application can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0169] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A contactless real-time pressure detection method based on millimeter wave radar, characterized by, include: The system acquires the echo signal of the millimeter-wave radar reflected from the user under test in real time, and reconstructs the respiratory signal and heartbeat signal of the user under test based on the echo signal, and obtains the estimated value of the respiratory signal. The reconstructed heartbeat signal is input into a one-dimensional neural network to obtain the first estimate of the heartbeat signal; the reconstructed heartbeat signal is converted into a two-dimensional continuous wavelet transform (CWT) spectrum, and the CWT spectrum is input into the YOLO model to mark the frequency band region in the CWT spectrum that is related to the heart rate of the user being tested; The marked frequency band region is then highlighted, and the highlighted CWT spectrum is input into a two-dimensional convolutional neural network to obtain the second estimate of the heartbeat signal. The final estimate of the heartbeat signal is determined based on the first estimate and the second estimate of the heartbeat signal. Multi-timescale features of the user's breathing and heartbeat were extracted based on the reconstructed breathing and heartbeat signals. Based on the multi-timescale characteristics of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal, the user under test is subjected to real-time stress detection.

2. The method of claim 1, wherein, The one-dimensional neural network includes: multiple convolutional neural network backbone blocks, pyramid input blocks, Transformer encoding blocks, and Dense connection blocks; The multiple convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales through different convolutional kernels; The pyramid input module is used to perform feature learning at different scales on the signals input to each convolutional neural network backbone block, and then combine the signals obtained from feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block can capture local features at the corresponding scale. The Transformer encoding block is used to combine a multi-head attention mechanism to learn the global features corresponding to the local features of the heartbeat signals output by the multiple convolutional neural networks at different scales, and to capture the most representative features of the heartbeat signals among the local features at different scales. The Dense connection block is used to combine the output signal of the Transformer coding block to obtain a first estimate of the heartbeat signal.

3. The method of claim 1, wherein, Based on the echo signal, the respiratory and heartbeat signals of the user under test are reconstructed, and an estimated value of the respiratory signal is obtained, including: The corresponding beat signal is extracted from the echo signal; the beat signal carries the respiratory signal and heartbeat signal of the user being tested. The normalized minimum mean square error (NLMS) algorithm based on polynomial fitting is used to initially suppress motion artifacts in the beat signal. Continuous wavelet transform (CWT) is used to perform multi-stage filtering on the beat signal after initial suppression of motion artifacts. The multi-stage filtering includes: performing a first-stage bandpass filtering on the beat signal in the wavelet domain to obtain first-stage filtered respiratory and heartbeat signals; then combining peak-and-valley finding algorithms and spectral evaluation to obtain first-stage respiratory and heartbeat signal estimates; performing a second-stage bandpass filtering on the first-stage filtered respiratory and heartbeat signals to obtain second-stage filtered respiratory and heartbeat signals; then combining peak-and-valley finding algorithms and spectral evaluation to obtain second-stage respiratory and heartbeat signal estimates; and performing a third-stage bandpass filtering on the second-stage filtered heartbeat signal to obtain a third-stage filtered heartbeat signal; finally, combining peak-and-valley finding algorithms and spectral evaluation to obtain a third-stage heartbeat signal estimate. The center frequency of each subsequent bandpass filter is set with reference to the corresponding respiratory and heartbeat signal estimates from the previous stage. The respiratory signal obtained after secondary filtering and the heartbeat signal obtained after tertiary filtering are used as the reconstructed respiratory signal and heartbeat signal of the user under test; The second-level respiratory signal estimate obtained during the reconstruction of the user's respiratory signal is used as the estimated value of the respiratory signal.

4. The method of claim 3, wherein, When the triaxial acceleration data of the user under test can be acquired in real time, the multi-level filtering performed to reconstruct the user's heartbeat signal based on the echo signal is as follows: First-level bandpass filtering is performed on the beat signal in the wavelet domain to obtain the first-level filtered heartbeat signal. Then, a peak-and-valley finding algorithm and spectral evaluation are used to obtain the first-level heartbeat signal estimate. Second-level bandpass filtering is performed on the first-level filtered heartbeat signal to obtain the second-level filtered heartbeat signal. Then, combined with the user's triaxial acceleration data, the Q-PPG deep learning algorithm is used to estimate the heart rate of the second-level filtered heartbeat signal to obtain the second-level heartbeat signal estimate. Third-level bandpass filtering is performed on the second-level filtered heartbeat signal to obtain the third-level filtered heartbeat signal. Then, a peak-and-valley finding algorithm and spectral evaluation are used to obtain the third-level heartbeat signal estimate. The heartbeat signal obtained after the third-level filtering is used as the reconstructed heartbeat signal of the user under test.

5. The method according to claim 3 or 4, characterized in that, Also includes: The reconstructed heartbeat signal is converted into a two-dimensional continuous wavelet transform (CWT) spectrum. The CWT spectrum is then input into the YOLO model to mark the frequency band region in the CWT spectrum that is related to the heart rate of the user being tested, so as to obtain a third estimate of the heartbeat signal. The final estimate of the heartbeat signal is determined based on the multi-level heartbeat signal estimates, the first estimate of the heartbeat signal, the second estimate of the heartbeat signal, and the third estimate of the heartbeat signal obtained in the multi-level filtering process.

6. The method of claim 1, wherein, The extraction of multi-timescale features of the user's breathing and heartbeat includes: Based on the reconstructed respiratory and heartbeat signals, a first deep learning network is used to extract the respiratory and heartbeat features of the user under test in the first time span using a detection window based on the first time span. Based on the reconstructed respiratory and heartbeat signals, a second deep learning network is used to extract the respiratory and heartbeat features of the user under test in the second time span using a detection window based on the second time span. The respiratory and heart rate characteristics of the test users at different time spans were used as multi-time-scale features.

7. The method of claim 6, wherein, The first deep learning network includes: a one-dimensional convolution, a causal analyzer, dilated convolution, and residual connections; the one-dimensional convolution is used to capture the temporal variation features of the time series data corresponding to respiratory and heart rate signals within a first time span; the causal analyzer is used to restrict the convolution process to past and present data, ensuring that the prediction results are consistent with the real situation; the dilated convolution is used to assist downsampling and provide a larger receptive field; the residual connections accelerate convergence by allowing gradient propagation without the need for activation functions, thereby extracting respiratory and heart rate features within the corresponding time span; The second deep learning network is a self-supervised learning network, used to learn and extract the spatiotemporal features of respiratory and heartbeat signals within the corresponding time span.

8. A contactless real-time pressure detection system based on millimeter wave radar, characterized by, include: The signal reconstruction module is used to acquire the echo signal of the millimeter-wave radar reflected from the user under test in real time, and reconstruct the respiratory signal and heartbeat signal of the user under test based on the echo signal, and obtain the estimated value of the respiratory signal. The heart rate final estimation module is used to input the reconstructed heartbeat signal into a one-dimensional neural network to obtain the first estimate of the heartbeat signal; convert the reconstructed heartbeat signal into a two-dimensional continuous wavelet transform (CWT) spectrum, and input the CWT spectrum into the YOLO model to mark the frequency band region in the CWT spectrum that is related to the heart rate of the user being tested; The marked frequency band region is then highlighted, and the highlighted CWT spectrum is input into a two-dimensional convolutional neural network to obtain the second estimate of the heartbeat signal. The final estimate of the heartbeat signal is determined based on the first estimate and the second estimate of the heartbeat signal. The pressure detection module is used to extract multi-timescale features of the user's breathing and heartbeat based on the reconstructed breathing and heartbeat signals; The system performs real-time stress detection on the user under test based on the multi-timescale features of breathing and heartbeat, the estimated value of the breathing signal, and the final estimated value of the heartbeat signal.

9. The system of claim 8, wherein, The one-dimensional neural network used in the final heartbeat estimation module includes: multiple convolutional neural network backbone blocks, a pyramid input block, a Transformer encoding block, and a Dense connection block. The multiple convolutional neural network backbone blocks are used to capture local features of the reconstructed heartbeat signal at different scales using different convolutional kernels. The pyramid input block is used to learn features at different scales from the signals input to each convolutional neural network backbone block, and then combine the signals obtained from the feature learning with each convolutional neural network backbone block so that each convolutional neural network backbone block captures local features at the corresponding scale. The Transformer encoding block is used to learn global features corresponding to the local features of the heartbeat signal output by the multiple convolutional neural networks at different scales using a multi-head attention mechanism, and captures the most representative heartbeat signal features among the local features at different scales. The Dense connection block is used to combine the output signal of the Transformer encoding block to obtain a first estimate of the heartbeat signal.

10. An electronic device, comprising: include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-7.