Non-contact sleep monitoring method based on cross-modal compensation
Through multimodal fusion and feature optimization algorithms of radar, video and audio sensors, the problem of insufficient signal continuity and accuracy in contactless sleep monitoring is solved, and high-precision sleep staging is achieved in complex environments.
Patent Information
- Application Number
- CN202510398815.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing contactless sleep monitoring technologies are mostly single mode or dual mode, with poor signal continuity and accuracy, and the physiological signals are not rich enough. They lack interactive verification between data, making it difficult to fully and accurately reflect the real vital signs of the human body.
Radar sensors, video sensors and audio sensors are used for multimodal fusion, feature optimization is combined with ReliefF algorithms and machine learning algorithms, decision-making level fusion is used for use with Naive Bayes classifiers, and multi-sensor data fusion system is built to achieve cross-modal compensation.
It improves the high-precision extraction of physiological signals in complex environments, enhances the accuracy and robustness of sleep staging, and reduces the discomfort and measurement errors caused by traditional contact sensors.
Smart Images

Figure CN120240971A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of non-contact sleep monitoring, and specifically is a non-contact sleep monitoring method based on video, audio, and radio frequency multi-modal fusion and cross-modal compensation. Technical Background
[0002] Sleep monitoring technology has always been an important means in sleep science and medical research, and is also one of the key technologies for evaluating sleep quality and diagnosing sleep disorders. Traditional sleep monitoring technology usually relies on a variety of physiological signals, such as electroencephalogram (EEG), electrooculogram (EOG), electromyogram (EMG), respiratory monitoring and body temperature monitoring. Through these signals, sleep staging and monitoring of respiratory abnormalities can be carried out to comprehensively evaluate sleep quality and diagnose sleep disorders.
[0003] This traditional sleep monitoring technology based on polysomnography performs well under ideal conditions, but they are mostly wearable and patch-type monitoring, which inevitably affects the user's sleep, so non-contact vital sign measurement is booming. However, existing non-contact sleep monitoring is mostly single-mode or dual-mode, the continuity and accuracy of the acquired signals are poor, and the physiological signals are not rich enough. The data of each sensor is analyzed independently, and there is a lack of interactive verification between data, making it difficult to fully and accurately reflect the true vital sign status of the human body.
[0004] To solve this problem, the present invention proposes a non-contact sleep monitoring system with cross-modal compensation, which enables physiological signals to complement each other. When a certain modality data deviates due to environmental interference or its own limitations, other modality data can quickly fill the information gap. Summary of the invention
[0005] The present invention proposes a non-contact sleep monitoring technology based on multi-sensor data fusion. The system obtains body movement, breathing and heartbeat signals of the human body through radar sensors, obtains heartbeat signals of the human body through video sensors, and obtains breathing and snoring signals through audio sensors, avoiding the discomfort and measurement errors caused by traditional contact sensors.
[0006] Compared with existing non-contact sleep monitoring, the present invention is more robust in complex environments and can ensure high-precision extraction of physiological signals even in undesirable environments.
[0007] Specifically, the main innovations of the present invention are embodied in the following aspects:
[0008] 1. Non-contact sleep monitoring: Traditional sleep monitoring methods usually require contact equipment (such as EEG, ECG, etc.), while the method proposed in this paper realizes non-contact monitoring by fusion of three modalities through radar sensor, video sensor and audio sensor, reducing the discomfort and measurement errors caused by traditional contact sensors.
[0009] 2. Light-Intensity-Based Heartbeat Signal Optimization Algorithm: Aiming at the problem of poor detection performance of heartbeat signals by video sensors in low-light environments, a light-intensity-based heartbeat signal optimization algorithm is proposed. This algorithm selects a more accurate source of heartbeat signals according to different light intensities, effectively improving the detection accuracy of heartbeat signals in low-light environments.
[0010] 3. Sleep Staging Based on Multi-Sensor Data Fusion: By combining the data of radar, video, and audio sensors, extracting signals and characteristic parameters, the accuracy and robustness of sleep staging are improved. The synchronization of multiple sensors and the accuracy of data are ensured through the construction of experimental scenarios.
[0011] 4. Multi-Sensor Feature-Level Fusion Model: The present invention uses the ReliefF algorithm to optimize the characteristic parameters and combines machine learning algorithms (such as Subspace KNN, Bagged Trees, etc.) for classification.
[0012] 5. Multi-Sensor Decision-Level Fusion Model: The present invention constructs a fusion system framework based on a naive Bayes classifier and improves the performance of the model by optimizing classifier parameters (such as the number of learners). BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flowchart of the present invention for processing signal monitoring sleep.
[0014] Figure 2 It is the result of the light-intensity-based heartbeat signal optimization algorithm.
[0015] Figure 3 It is the characteristic parameters of the radar sensor.
[0016] Figure 4 It is the characteristic parameters of the audio sensor.
[0017] Figure 5 It is a diagram of the feature-level fusion framework.
[0018] Figure 6 It is a diagram of the decision-level fusion framework. DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention proposes a multi-modal non-contact sleep monitoring method through radar, video, and audio sensors. The flowchart is as Figure 1 shown. According to the flowchart, the present invention can be divided into five steps: signal acquisition and processing, signal optimization, characteristic parameter extraction, data fusion, and result comparison and verification. The specific description is as follows:
[0020] 1. Signal Acquisition and Processing
[0021] The 2.4GHz digital intermediate frequency continuous wave radar emits electromagnetic waves and receives the reflected echo signals to extract the vital sign signals of the human body (such as respiration, heartbeat, body movement, etc.). A near-infrared camera is used as a video sensor to capture the RGB signals in the face area through the camera, and the photoplethysmography (PPG) principle is used to extract the heartbeat signal. The adaptive threshold method is used to detect respiration events and snoring events. The voiced and unvoiced segments are distinguished by the endpoint detection method (such as the double-threshold method, spectral entropy method, etc.). The frequency range of the respiration signal is 0.13 - 0.4Hz.
[0022] 2. Signal Optimization
[0023] Under different lighting conditions, compare the accuracy of the heartbeat signals of the radar and the video sensor; select the signal source with higher accuracy according to the light intensity; use the decision tree model to judge the light intensity and select the optimal signal source.
[0024] Calculation of light intensity: I = 0.299×R + 0.587×G + 0.114×B, where (R, G, B) are the three-channel values of the face area of each frame of video image.
[0025] The optimization results are as Figure 2 , and the heartbeat signal optimization algorithm based on light intensity can effectively improve the accuracy of heart rate measurement, eliminate the influence of light intensity on heartbeat signal detection, has higher measurement accuracy compared with the single-modal system, can adapt to different environments, and has stronger universality and robustness.
[0026] 3. Feature Parameter Extraction
[0027] 3.1 Feature Parameter Extraction of Radar Sensor
[0028] According to the characteristics of the human physiological signals in each sleep stage, a total of 11 feature parameters used to divide the sleep stages are extracted from the radar signals, including respiration, heartbeat, and other feature parameters. The feature parameters are extracted at intervals of every 30 seconds. As Figure 3 shown. RPM_ADA (accumulated difference in amplitude of respiration signal): Calculate the accumulated difference in amplitude between adjacent peaks of the respiration signal, which reflects the change in respiration amplitude.
[0029]
[0030] Among them, A(n) is the amplitude of the respiration signal
[0031] RPM_MOVE (body movement parameter in respiration signal): Detect body movement events in the respiration signal by setting an amplitude threshold.
[0032]
[0033] Where A(n) is the amplitude of the respiratory signal
[0034] 3.2 Feature Parameter Extraction of Video Sensor
[0035] According to the characteristics of human physiological signals in each sleep stage, the characteristic parameters extracted from the video signal for dividing the sleep stages are mainly heartbeat signals, which supplement other modalities.
[0036] Heart rate variability (HRV): It shows specific characteristics in different stages of sleep, is closely related to sleep structure, and effectively reflects sleep quality.
[0037]
[0038] Low frequency (LF) and high frequency (HF) power: reflects the balance of the autonomic nervous system and effectively reflects the quality of sleep
[0039]
[0040] 3.3 Feature Parameter Extraction of Audio Sensor
[0041] According to the characteristics of human physiological signals in each sleep stage, 23 characteristic parameters are initially extracted from the audio signal to divide the sleep stages, which are divided into two categories: sleep breathing related parameters and linear and nonlinear parameters of snoring. Figure 4 As shown in the figure, LLE (Lyapunov exponent): used to measure the divergence speed of trajectories in a chaotic system.
[0042]
[0043] Where δ(t) is the separation distance of the trajectories.
[0044] ApEn (Approximate Entropy): Used to measure the complexity and regularity of time series.
[0045] ApEn=Φm(r)-Φm+ 1 (r)
[0046] Among them, Φm(r) is the probability of similar patterns.
[0047] 4. Data Fusion
[0048] 4.1 Feature Level Fusion
[0049] The feature data of radar, video and audio sensors are integrated and classified using machine learning algorithms. The framework diagram is as follows: Figure 5 . Use the ReliefF algorithm to optimize the feature parameters and eliminate unimportant features.
[0050] Weight update formula of ReliefF algorithm:
[0051]
[0052] Among them, diff(A, R1, R2) represents the difference between samples R1 and R2 on feature A.
[0053] Use classifiers such as Subspace KNN and Bagged Trees to classify the feature data.
[0054] Distance calculation formula of Subspace KNN classifier: Decision tree construction formula of Bagged Trees classifier:
[0055] 4.2 Decision-level fusion
[0056] The decision-level fusion model fuses the classification results of radar, video, and audio sensors, and uses a naive Bayes classifier for decision-making. The framework diagram is as Figure 6 . The basic principle of the naive Bayes classifier is based on Bayes' theorem to calculate the conditional probabilities of each class.
[0057] Bayes' theorem:
[0058] Conditional probability calculation:
[0059] Input the remaining radar segment data and radar, video, and audio segment data into the Subspace KNN and Bagged Trees classifiers respectively to obtain classification results. Through experiments, for the Subspace KNN classifier: the value of the parameter subspaceDimension is set to 6, and the value of the parameter Numbers oflearners is set to 50. For the BaggedTrees classifier: the value of the parameter Numbers oflearners is set to 5.
[0060] Use the naive Bayes classifier to fuse the classification results to obtain the final sleep staging results for sleep analysis.
[0061] Decision formula of the naive Bayes classifier:
[0062] 5. Result comparison and verification
[0063] By comparing with the sleep monitoring results of PSG, the feasibility and accuracy of multi-sensor data fusion technology in non-contact sleep staging and sleep monitoring are verified.
Claims
1. A non-contact sleep monitoring method based on cross-modal compensation, characterized by the following steps: S1. Non-contact acquisition of human vital sign signals through a radar sensor, a video sensor, and an audio sensor. S2. Optimize the acquired signals. When the light intensity is lower than the preset threshold, preferentially select the heartbeat signal of the radar sensor; When the light intensity is higher than the preset threshold, preferentially select the heartbeat signal of the video sensor. S3. Extract feature parameters from the optimized signals. S4. Perform multi-modal data fusion on the extracted feature parameters, including feature-level fusion: use the ReliefF algorithm to optimize and select the feature parameters, and combine with the Subspace KNN or Bagged Trees classifier for classification; decision-level fusion: fuse the classification results of each sensor through a naive Bayes classifier. S5. Compare the multi-sensor fusion results with the sleep staging results of a polysomnogram (PSG) to verify the accuracy of sleep staging, specifically by calculating the accuracy rate, recall rate, and F1 value to complete the verification.
2. In the non-contact sleep monitoring method according to claim 1, for the acquisition of vital sign signals by the three sensors in step S1, specifically, the radar sensor is a 2.4 GHz digital intermediate frequency continuous wave radar, and the vital sign signals are extracted by transmitting electromagnetic waves and receiving the reflected echo signals; the video sensors are a mobile phone and a near-infrared camera, and the pulse wave can still be extracted through facial features when the light in the sleep environment is insufficient, so as to extract the heartbeat signal; the audio sensor uses a double-threshold method for detecting breathing and snoring events.
3. In the non-contact sleep monitoring method according to claim 1, for the determination of the light intensity in step S2, it further includes implementing signal source selection through a decision tree model to ensure the robustness of heartbeat signal detection in a low-light environment.
4. In the non-contact sleep monitoring method according to claim 1, for the extraction of feature parameters in step S3, specifically, extract the I / Q two-channel signals through digital down-conversion and quadrature decomposition, and use high-pass filters and low-pass filters to separate breathing, heartbeat, and body movement signals; automatically perform facial recognition on video frames to obtain the ROI, record the time axis, restore the video frames with an accurate time axis through linear interpolation, and filter the extracted pulse wave signals with a 0.5 - 4 Hz band-pass filter; distinguish breathing and snoring signals through energy spectrum detection and mean square error method, and extract non-linear parameters including LLE, ApEn, and detrended fluctuation analysis coefficients.
5. In the non-contact sleep monitoring method according to claim 1, for the feature-level fusion model in step S4, use the Subspace KNN classifier for the remaining radar segment data (when there is no audio signal); use the Bagged Trees classifier for the radar + audio fusion segment data; use the Bagged Trees classifier for the video. Optimize the classification results through temporal features, and the temporal features reflect the continuity law of the sleep cycle.
6. In the non-contact sleep monitoring method according to claim 1, for the decision-level fusion model in step S4, Subspace KNN and Bagged Trees classifiers are independently trained for radar, video, and audio data respectively; probability fusion is performed on the classification results based on the Naive Bayes algorithm. The parameter settings of the Subspace KNN classifier are: the subspace dimension is 6, and the number of learners is 50; the number of learners of the Bagged Trees classifier is 5.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the non-contact sleep monitoring method according to claim 1.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to claim 1.
Citation Information
Cited By
Sleep state detection method, sleep state detection equipment and storage medium
CN120514341A
Sleep state detection method, sleep state detection device, and storage medium
CN120514341B
Sleep awakening detection and evaluation method and device based on radar and PPG signals
CN121533701A
Sleep-wake detection and assessment methods and equipment based on radar and PPG signals
CN121533701B