Multimodal optical non-contact health monitoring system and method
By using a multimodal optical non-contact health monitoring system, which combines multiple sensors and data processing algorithms, the problems of multimodal fusion, motion interference, and accuracy limitations in existing technologies have been solved. This system enables the detection of multiple physiological indicators in complex environments, thereby improving the reliability and accuracy of the detection.
Patent Information
- Application Number
- CN202510680447.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing non-contact physiological testing technologies suffer from the lack of multimodal fusion technology, the conflict between motion interference and privacy protection, and the limited accuracy of blood flow parameter detection, making it difficult to meet the needs of multi-parameter joint analysis in complex clinical scenarios.
A multimodal optical non-contact health monitoring system is employed. This system acquires facial video images, TOF depth maps, and IMU data through a sensing and measurement module. Combined with a micro-motion stabilization and correction module, a multimodal data processing module, and a physiological indicator calculation module, it achieves joint detection of multiple physiological parameters. The system includes an RGB camera, a TOF sensor, an IMU, and a tunable multi-wavelength light source. Data processing and correction are performed using the SURF algorithm, a lightweight visual transforerifier, a U-Net blood flow segmentation network, a diffuse optical spectral mapping network, and a temporal feature fusion network.
It significantly improves the reliability and accuracy of detection in complex environments with motion and changing lighting conditions, solves the problem of multimodal data coupling interference, and simultaneously acquires multidimensional physiological indicators such as blood pressure, heart rate, blood oxygen, and body temperature, meeting the multi-parameter analysis needs in complex clinical scenarios and improving the predictive accuracy and stability of detection.
Smart Images

Figure CN120392025B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application is directed to obtaining human physiological indicators multi-modal information of human face by optical imaging method, and relates to a physiological parameter detection device, in particular discloses a multi-modal optical non-contact health monitoring system and method, belonging to the technical field of measurement and testing. BACKGROUND
[0002] As an important development direction in the field of biomedical engineering, non-contact physiological detection technology has shown broad application prospects in the fields of health monitoring and clinical diagnosis in recent years. The existing technology mainly realizes physiological parameter detection based on optical, video analysis, Laser Speckle Contrast Analysis (LASCA), Optical Coherence Tomography (OCT) and other principles, but there are still significant technical bottlenecks in practical application:
[0003] 1. Lack of multi-modal fusion technology
[0004] Current mainstream devices generally rely on a single detection mode. For example, although optical detection technology based on photoplethysmography (PPG) can realize heart rate monitoring, it is difficult to simultaneously obtain multi-dimensional physiological indicators such as blood oxygen and blood pressure. Video photoplethysmography (VPG) can non-contactly obtain hemodynamic information, but it is easily disturbed by environmental light, skin pigment differences and other factors, resulting in insufficient stability of the detection results. This single mode limitation makes it difficult for the device to meet the demand of multi-parameter joint analysis in complex clinical scenarios.
[0005] 2. Motion interference and privacy protection conflict
[0006] Video analysis technology faces serious motion artifact problems in dynamic monitoring. The existing solution such as the privacy protection method based on target recognition and dynamic frame processing proposed in Chinese patent CN112545339A can protect personal information through blurring processing, but it needs to extract features and track targets frame by frame for video stream, resulting in a sharp increase in system computational complexity and difficulty in ensuring real-time performance. This contradiction restricts the large-scale application of this technology in wearable devices and other scenarios.
[0007] 3. Blood flow parameter detection precision limitation
[0008] Traditional blood flow detection technology has the problem that resolution and dynamics are difficult to balance: although laser speckle technology can realize high-resolution imaging of blood flow velocity, it cannot synchronously obtain dynamic parameters such as blood pressure; OCT technology can provide microscale resolution of blood vessel structure information, but it is difficult to realize real-time monitoring of hemodynamic parameters. The single dimension of this detection makes the existing equipment unable to meet the clinical needs such as early diagnosis of microcirculation disorders.
[0009] In summary, the existing non-contact physiological detection technology has significant technical gaps in multi-modal fusion algorithm, anti-motion interference mechanism, high-precision blood flow parameter detection, etc., and urgently needs to break through based on a multi-modal optical non-contact health monitoring system. SUMMARY
[0010] The purpose of the present application is to overcome the shortcomings of the above background art, and to provide a multi-modal optical non-contact health monitoring system and method. The multi-modal optical data is obtained by optical processing of the face image after anti-motion interference processing. The multi-physiological parameters are jointly detected by the data processing method of mapping the human physiological indicators by multi-modal optical data. The technical problems of the existing non-contact physiological detection technology, such as single detection mode, susceptibility to motion interference, and limited blood flow parameter monitoring accuracy, are solved. The purpose of the present application is to realize non-contact joint detection of human physiological parameters based on multi-modal optical data.
[0011] The present application adopts the following technical solutions to achieve the above-mentioned purpose of the present application:
[0012] The multi-modal optical non-contact health monitoring system comprises a perception measurement module, a data synchronous acquisition module, a micro-motion stabilization and correction module, a three-dimensional trajectory reconstruction module, a multi-modal data processing module, and a physiological index calculation module. The perception measurement module is used to acquire facial video images, TOF depth maps, and IMU data of a human body. The data synchronous acquisition module is used to synchronously acquire facial video images, TOF depth maps, and IMU data of a human body. The micro-motion stabilization and correction module is used to stabilize and correct the synchronously acquired facial video images, output stable facial video frames, final local micro-motion fields, a main frequency signal of a heart rate physiological frequency band, and a main frequency signal of a breathing frequency physiological frequency band. The three-dimensional trajectory reconstruction module is used to reconstruct a local micro-motion three-dimensional trajectory according to two-dimensional displacement time series data in the final local micro-motion field. The multi-modal data processing module is used to perform optical processing on the stable facial video frames, acquire multi-modal optical data, and align the multi-modal optical data according to a three-dimensional space-time reference frame of the local micro-motion three-dimensional trajectory. The physiological index calculation module is used to map blood flow parameters to preliminary blood pressure values, blood oxygen monitoring values, and body temperature monitoring values based on the three-dimensional space-time aligned multi-modal optical data through a physical model driving, extract heart rate and breathing frequency related local micro-motion time series features from the main frequency signals of the heart rate and breathing frequency physiological frequency bands, acquire heart rate and breathing frequency monitoring values, dynamically weightedly fuse the preliminary blood pressure values and blood pressure values calculated from the heart rate related local micro-motion time series features, acquire blood pressure monitoring values, and output the heart rate, blood pressure, blood oxygen, body temperature, and breathing frequency monitoring values.
[0013] As a further optimization scheme of the multi-modal optical non-contact health monitoring system, the perception measurement module comprises an RGB camera for acquiring facial video images of a human body, a time-of-flight sensor for acquiring TOF depth maps of a human body, and an inertial measurement unit for acquiring IMU data. The data synchronous acquisition module acquires facial video images, TOF depth maps, and IMU data of a human body through a hardware synchronization technology.
[0014] As a still further optimization scheme of the multi-modal optical non-contact health monitoring system, the multi-modal optical module adopts a tunable multi-wavelength light source or a multi-wavelength laser for acquiring multi-modal optical data including but not limited to laser speckle images, hyperspectral data, and infrared thermal imaging.
[0015] As a further optimization scheme of the multi-modal optical non-contact health monitoring system, the micro-motion stabilization and correction module comprises: a stable image unit and an artifact correction unit; the stable image unit corrects the face video frame image through a SURF algorithm correction model, and outputs a stable face video frame; the artifact correction unit performs pixel-level two-dimensional motion vector prediction on the face video frame image through a lightweight visual TRANSFORMER, generates a preliminary dense motion vector field, generates a preliminary local dense motion vector field through an improved pyramid Lucas-Kanade optical flow method, corrects the preliminary local dense motion vector field to obtain a corrected local micro-motion field, generates a complete motion field according to the corrected local micro-motion field and performs global motion component correction, and finally obtains the final local micro-motion field from the final local micro-motion field. The main frequency signals of the physiological frequency bands of heart rate and respiratory rate are extracted.
[0016] As a further optimization scheme of the multi-modal optical non-contact health monitoring system, the physiological index calculation module comprises: a U-Net blood flow segmentation network, a diffuse optical spectrum mapping network and a time sequence feature fusion network; the U-Net blood flow segmentation network is used for multi-scale blood flow region feature extraction of multi-modal optical data, reconstruction of a blood flow region mask, generation of a binary blood flow region segmentation result, and extraction of blood flow parameters of each blood flow region, including but not limited to blood flow velocity, blood oxygen saturation and hemoglobin concentration; the diffuse optical spectrum mapping network is used to establish a mathematical mapping relationship for mapping blood flow parameters to preliminary physiological indicators, and output preliminary blood pressure values, blood oxygen monitoring values and body temperature monitoring values; the time sequence feature fusion network extracts heart rate and respiratory rate related local micro-motion time sequence features from the main frequency signals of the physiological frequency bands of heart rate and respiratory rate through the time sequence feature fusion network, dynamically weights and fuses the preliminary blood pressure values and the blood pressure values calculated from the heart rate related local micro-motion time sequence features through the spatio-temporal attention mechanism, and obtains blood pressure monitoring values.
[0017] As a further optimization scheme of the multi-modal optical non-contact health monitoring system, the system further comprises an environment adaptive module, which adjusts the motion of the perception measurement module through a feedback control algorithm; or adjusts the calculation parameters of the micro-motion stabilization and correction module; or adjusts the calculation parameters of the three-dimensional trajectory reconstruction module; or adjusts the intensity of the multi-wavelength light source; the physiological index calculation module further comprises a temperature calibration sub-module, which compensates for the temperature drift error in the blood flow parameters through Kalman filtering based on the temperature data of the infrared thermal imaging.
[0018] The multi-modal optical non-contact health monitoring method is realized based on the above system, comprising the following steps:
[0019] S1, synchronously collecting a human face video image, a TOF depth map and IMU data of a perception measurement module;
[0020] S2, performing stabilization processing and motion artifact correction on the collected face video image to obtain a stable face video frame, a final local micro-motion field, a main frequency signal of a heart rate physiological frequency band and a main frequency signal of a respiration frequency;
[0021] S3, reconstructing a local micro-motion three-dimensional trajectory according to two-dimensional displacement time series data in the final local micro-motion field;
[0022] S4, performing laser speckle blood flow imaging, hyperspectral imaging and infrared thermal imaging on the stabilized face video image to obtain multi-modal optical data, and aligning the multi-modal optical data according to a three-dimensional space-time reference framework of the local micro-motion three-dimensional trajectory;
[0023] S5, mapping blood flow parameters to preliminary blood pressure values and blood oxygen monitoring values and body temperature monitoring values through a physical model driven based on the multi-modal optical data aligned in three-dimensional space-time, extracting heart rate and respiration frequency related local micro-motion time series features from the main frequency signal of the heart rate physiological frequency band and the main frequency signal of the respiration frequency physiological frequency band respectively through time series feature fusion technology to obtain heart rate and respiration frequency monitoring values, and dynamically weighting and fusing the preliminary blood pressure values and the blood pressure values calculated from the heart rate related local micro-motion time series features through a space-time attention mechanism to obtain blood pressure monitoring values.
[0024] As a further optimization method of the multi-modal optical non-contact health monitoring method, S2 specifically comprises the following steps:
[0025] S2A, performing alignment processing on the input face video frame image by affine transformation registration of adjacent face video images;
[0026] S2B, detecting global feature points on the face video image processed by S2A using a SURF algorithm, obtaining a global motion component by calculating a global affine transformation matrix, performing geometric inverse transformation on the input face video frame image according to the global affine transformation matrix, and outputting a stable face video frame;
[0027] S2C, performing pixel-level two-dimensional motion vector prediction on the face video image processed by S2A through a lightweight visual TRANSFORMER to generate a preliminary dense motion vector field;
[0028] S2D, filtering global feature points on the face video image corrected by the global motion in S2B through a Hessian matrix, and dividing local regions with the filtered global feature points as anchor points;
[0029] S2E employs an improved pyramid Lucas-Kanade optical flow method to process the initial dense motion vector field, generating an initial local dense motion vector field;
[0030] S2F uses the RANSAC algorithm to detect outliers in the initial local dense motion vector field, removes abnormal motions, and obtains the corrected local micro-motion field.
[0031] S2G superimposes the global motion components with the corrected local micro-motion fields to generate a complete motion vector field;
[0032] S2H uses the complete motion vector field to verify whether there are residual anomalies in the global motion components. If residual anomalies exist, they are removed by the RANSAC algorithm to obtain the corrected global motion components.
[0033] S2I is obtained by subtracting the corrected global motion components from the complete motion field to obtain the final local micro-motion field.
[0034] S2J preprocesses the final local micro-motion field, extracts the two-dimensional displacement time series data of the physiological frequency band of the preprocessed final local micro-motion field, and uses wavelet transform for multi-scale decomposition and hard threshold denoising. It retains the physiological frequency band coefficients of heart rate and respiratory rate, and after reconstructing the physiological frequency band coefficients of heart rate and respiratory rate, it extracts the main frequency of the physiological frequency bands of heart rate and respiratory rate respectively.
[0035] As a further optimization of the multimodal optical non-contact health monitoring method, S2E specifically includes the following steps:
[0036] S2E1 performs multi-scale downsampling on each local region to construct a pyramid for optimizing optical flow estimation layer by layer from low resolution to high resolution;
[0037] S2E2 uses the initial dense motion vector field generated by S2C as the initial value to quickly estimate the motion range of the local area and optimize the low-resolution layer.
[0038] S2E3, based on the assumption of constant brightness, solves the optical flow equation to iteratively optimize the two-dimensional motion vector of each pixel in the local region, and performs high-resolution layer optimization.
[0039] As a further optimization of the multimodal optical non-contact health monitoring method, S3 reconstructs the three-dimensional trajectory of local micro-motion based on the two-dimensional displacement time-series data in the final local micro-motion field, specifically:
[0040] S3A acquires synchronously collected TOF depth maps, human facial video images, and IMU data;
[0041] S3B, using statistical outlier removal and radius filtering to eliminate noise points in the TOF depth map data, generate a depth map of the human face region, combine the face video image, use Mask R-CNN to segment the depth map of the human face region, and obtain the depth map of the focused physiological micro-movement;
[0042] S3C, by iteratively registering the point cloud data of the focused physiological micro-movement depth map, separating the head whole motion from the local micro-movement, determining the target region from the separated local micro-movement point cloud data, and establishing the mapping relationship between the target region depth map and the face video image three-dimensional coordinates according to the TOF sensor internal and external parameter calibration, generating the local micro-movement three-dimensional trajectory.
[0043] The technical scheme of the present application has the following beneficial effects:
[0044] (1) The present application uses SURF algorithm to correct the model for global motion compensation of continuous video frames, stabilizes and corrects motion artifacts, and has an environment adaptive module, which significantly improves the reliability and accuracy of non-contact health monitoring in complex environments such as motion and light changes.
[0045] (2) The present application uses U-Net network to process optical imaging data first, and then combines with video micro-movement features to solve the problem of multi-modal data coupling interference in traditional methods. By establishing a mathematical mapping relationship between blood flow parameters and physiological indicators through a diffuse optical spectrum mapping network, the segmented blood flow parameters are mapped to preliminary physiological indicators such as blood pressure, blood oxygen, and body temperature. Through a time series feature fusion network, local micro-movement time series features related to heart rate and respiration rate are extracted to solve the time lag problem of single modal data. Then, through the spatio-temporal attention mechanism, the weight of the blood pressure value calculated by combining the local micro-movement time series features related to heart rate and the preliminary blood pressure mapping value is dynamically adjusted to enhance the robustness in complex environments and improve the prediction accuracy and stability of the non-contact health monitoring system.
[0046] (3) The present application combines laser speckle imaging, hyperspectral data, infrared thermal imaging and other multi-modal optical data to synchronously obtain multi-dimensional physiological indicators such as blood pressure, heart rate and oxygen saturation, overcoming the limitations of traditional single modal such as PPG or VPG, meeting the demand of multi-parameter joint analysis in complex clinical scenarios, and improving the comprehensiveness of detection dimension.
[0047] (4) The present application maps two-dimensional optical flow to three-dimensional displacement through a kinematic model, combines Kalman filtering to suppress noise, provides spatial dynamic information, significantly improves the analysis accuracy and robustness of physiological signals such as skin vibration and chest and abdominal fluctuation, and provides a unified spatio-temporal reference frame for multi-modal data fusion.
[0048] (5) The present application introduces the Beer-Lambert law and other physical constraints in the diffuse optical spectral mapping network, constrains the relationship between the blood oxygen saturation prediction value and the optical reflectance intensity, enhances the physical interpretability, ensures that the output result conforms to the medical law, and improves the trust degree of the doctor to the system output.
[0049] (6) The present application has an environment adaptive module, which dynamically adjusts the intensity of the multi-wavelength light source through feedback control, automatically adapts to different light conditions, reduces the influence of external interference on the optical signal quality, and ensures the detection stability.
[0050] (7) The present application separates global motion and local micro-motion through SURF algorithm, RANSAC outlier rejection, Kalman filter and other multi-level correction, and can still accurately extract physiological signals such as heart rate and respiratory rate in the scene of slight movement of the patient or shaking of the camera. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 is a system module schematic diagram in example 1.
[0052] Figure 2 is a monitoring method flowchart in example 2.
[0053] Figure 3 is a flowchart of S2 in example 2 using lightweight visual TRANSFORMER and SURF algorithm to correct motion artifacts to obtain stable face video frames.
[0054] Figure 4 is a flowchart of S5 in example 2 using an improved pyramid Lucas-Kanade optical flow method to generate a refined local dense motion vector field. DETAILED DESCRIPTION
[0055] The technical solutions of the present application will be further described in detail below in combination with specific embodiments, but the present embodiments are not used to limit the present application, and any similar structure and its similar changes that adopt the present application shall be included in the protection scope of the present application, the English letters in the present application are case-sensitive.
[0056] Example 1
[0057] As shown in Figure 1 , in one embodiment of the present application, a multi-modal optical non-contact health monitoring system is provided, which comprises a perception measurement module, a data synchronous acquisition module, a micro-motion stabilization and correction module, a three-dimensional trajectory reconstruction module, a multi-modal data processing module and a physiological index calculation module.
[0058] The perception measurement module comprises a red, green, and blue (RGB) camera, a time of flight (TOF) sensor, and an inertial measurement unit (IMU). The RGB camera is used to acquire a face video image of a human body, and the face video image contains texture and color information of a face of the human body. The RGB camera is one of a high-frame-rate camera, a high-speed CMOS camera array, and a PTZ camera. The high-speed CMOS camera array separates different wavelengths of reflected light to a dedicated sensor by a light splitting prism to achieve multi-spectral synchronous acquisition, and the frame rate is greater than or equal to 200 fps to avoid crosstalk. The TOF sensor is used to acquire a TOF depth map of the face of the human body, and the TOF depth map reflects depth information of the face of the human body. The IMU is used to collect IMU data of the perception measurement module to record a motion state of the perception measurement module, such as acceleration and angular velocity.
[0059] The data synchronous acquisition module is used to synchronously acquire the face video image, the TOF depth map, and the IMU data of the human body. Specifically, hardware synchronization such as a GPIO trigger is used to ensure the spatiotemporal consistency of the acquired data.
[0060] The micro-motion stabilization and correction module is used to perform stabilization processing and motion artifact correction on the acquired face video image, and output a stabilized face video frame, a final local micro-motion field, a dominant frequency signal of a heart rate physiological frequency band, and a dominant frequency signal of a respiration frequency physiological frequency band. The micro-motion stabilization and correction module comprises a lightweight visual TRANSFORMER and a SURF algorithm correction model. The lightweight visual TRANSFORMER adopts an encoder-decoder structure to predict a two-dimensional motion vector of each pixel in the face video frame to generate a preliminary dense motion vector field , which captures subtle motions such as skin vibration and chest and abdomen fluctuation. Then, an improved pyramid Lucas-Kanade optical flow method is used to improve the motion estimation accuracy of a local area to generate a preliminary local dense motion vector field. The preliminary local dense motion vector field provides a high-resolution input for subsequent motion correction and signal analysis. The preliminary local dense motion vector field is corrected to obtain a corrected local micro-motion field. A complete motion field is generated according to the corrected local micro-motion field, and a global motion component correction is performed. A final local micro-motion field is obtained from the corrected global motion component. The final local micro-motion field comprises a skin vibration micro-motion field and a chest and abdomen fluctuation micro-motion field. The dominant frequency signal of the heart rate physiological frequency band and the dominant frequency signal of the respiration frequency physiological frequency band are extracted from the final local micro-motion field. The SURF algorithm correction model depends on the global motion component to achieve accurate correction. This design combines the stability of traditional feature point detection and the data-driven advantage of deep learning, and is one of the core technical breakthroughs of the non-contact health monitoring system.
[0061] a three-dimensional trajectory reconstruction module configured to reconstruct a local micro-motion three-dimensional trajectory according to the two-dimensional displacement time-series data in the final local micro-motion field.
[0062] a multi-modal data processing module configured to perform optical processing on the stable facial video frames to obtain multi-modal optical data, wherein the multi-modal optical data comprises laser speckle images, hyperspectral data, and infrared thermal imaging, the multi-modal optical data is aligned according to a three-dimensional spatiotemporal reference frame of the local micro-motion three-dimensional trajectory to ensure consistency of the multi-modal optical data in space and time, and preferably, the multi-modal optical module comprises a tunable multi-wavelength light source or a multi-wavelength laser, the tunable multi-wavelength light source integrates RGB, near-infrared, and far-infrared light sources, the RGB light source can be a three-color light source with optional wavelengths of 660 nm, 530 nm, and 450 nm, the near-infrared light source can be a near-infrared light source with optional wavelengths of 850 nm and 940 nm, the far-infrared light source can be a far-infrared light source with an optional wavelength of 10-14 μm, and the tunable multi-wavelength light source can be provided by a tunable laser, the multi-wavelength laser comprises laser light sources with wavelengths of 660 nm, 850 nm, and 940 nm, and the multi-wavelength laser can be provided by a multi-laser array, the multi-modal tunable laser light source composed of the multi-wavelength laser or the tunable multi-wavelength light source can perform laser speckle blood flow imaging, hyperspectral imaging, and infrared thermal imaging on the stable facial video output by the micro-motion stabilization and correction module to obtain multi-modal optical data comprising laser speckle images, hyperspectral data, and infrared thermal imaging, the laser speckle images are used to capture dynamic changes in micro-artery blood flow velocity, the hyperspectral data are used to capture dynamic changes in micro-artery blood oxygen saturation and hemoglobin concentration, and the infrared thermal imaging is used to capture changes in facial skin temperature to assist in blood flow hemodynamic analysis.
[0063] a physiological index calculation module configured to map blood flow parameters to preliminary blood pressure values and blood oxygen monitoring values and body temperature monitoring values by driving a physical model based on the multi-modal optical data aligned in the three-dimensional space-time, extract local micro-motion time-series features related to heart rate and respiration frequency from main frequency signals in physiological frequency bands of the heart rate and respiration frequency, obtain heart rate and respiration frequency monitoring values by a time-series feature fusion network, obtain dynamically weighted and fused blood pressure monitoring values by dynamically weighting and fusing the preliminary blood pressure values and blood pressure values calculated from the local micro-motion time-series features related to the heart rate through a spatiotemporal attention mechanism, and output the heart rate, blood pressure, blood oxygen, body temperature, and respiration frequency monitoring values, and the physical model comprises a U-Net blood flow segmentation network, a Diffuse Optical Spectroscopy (DOS) mapping network, and a Long Short-Term Memory (LSTM) time-series feature fusion network.
[0064] The U-Net blood flow segmentation network comprises an encoder and a decoder, multi-modal optical data is extracted through the encoder to obtain multi-scale blood flow region features, and the blood flow region mask is reconstructed through the decoder; the encoder is a 5-layer convolution block, each layer comprising Conv2D+BN+ReLU, and the output size is halved layer by layer, and the output sizes are 1024, 512, 256, 128 and 64 in sequence; the decoder is a 5-layer deconvolution block, and the convolution block comprises a transposed convolution and a jump connection, the low-layer texture and high-layer semantic information are fused, and the blood flow region mask with the same resolution as the input is output.
[0065] The U-Net blood flow segmentation network further comprises an input layer and an output layer, the input layer is used for receiving the laser speckle image and the hyperspectral data, the resolution is greater than or equal to 1024*768 pixels, the number of channels is a multi-wavelength spectral dimension, and the multi-wavelength spectrum comprises light with wavelengths of 660 nm, 850 nm and 940 nm; the output layer generates a binary blood flow region segmentation result through a Sigmoid activation function, the binary blood flow region segmentation result is used to distinguish microarteries, veins and capillaries, and blood flow parameters of each blood flow region are extracted and output to the diffusion mapping network optical spectrum mapping network, wherein the blood flow parameters comprise blood flow velocity, blood oxygen saturation SpO2 and hemoglobin concentration.
[0066] The diffusion optical spectrum mapping network comprises three layers of fully connected networks and ReLU activation layers, the three layers of fully connected layers have a neuron structure of 512->256->128, preliminary physiological indexes such as blood pressure, blood oxygen and body temperature are calculated based on a physical model of multi-modal optical data such as the laser speckle image, the hyperspectral data and the infrared thermal image, a loss function with a physical constraint such as a Beer-Lambert law correction term is combined, the blood flow region segmentation accuracy and the physical rationality of physiological index prediction are simultaneously optimized, a mathematical mapping relationship between the blood flow parameters and the blood pressure, the blood oxygen and the body temperature is established, and the blood flow parameters of each blood flow region after segmentation are mapped to preliminary blood pressure monitoring values, blood oxygen monitoring values and body temperature monitoring values; through the loss function with the physical constraint, the blood flow region segmentation accuracy is improved, and the blood oxygen saturation value conforming to the physical law is output, thereby providing a double guarantee for the reliability and the explainability of the medical AI system.
[0067] The time sequence feature fusion network comprises a bidirectional LSTM layer, the LSTM layer has a structure of 128 units, the periodicity of the physiological frequency band main frequency signals of the heart rate and the respiratory frequency is analyzed based on the local micro motion field such as the skin vibration motion field, the local micro motion time sequence features related to the heart rate and the respiratory frequency are respectively extracted, and the heart rate monitoring value and the respiratory frequency monitoring value are obtained; through a space-time attention mechanism, the weight of the blood pressure value calculated from the local micro motion time sequence features related to the heart rate and the preliminary blood pressure monitoring value is dynamically adjusted, and the robustness in a complex environment is enhanced.
[0068] The physiological index calculation module further comprises a temperature calibration submodule, which compensates for the temperature drift error in the blood flow parameter based on the temperature data collected by the infrared thermal imaging module. The infrared thermal imaging module collects skin temperature data in real time through an infrared camera. The infrared thermal imaging submodule collects infrared rays emitted by an object through an infrared lens, focuses them on an infrared detector, converts the infrared radiation energy into an electrical signal, and then processes the electrical signal through a signal processing circuit for amplification, filtering, etc. Finally, the image processing module converts the electrical signal into a visual infrared thermal image, thereby reflecting the temperature distribution of the object surface. A TinyML light model is embedded in the edge device to directly generate preliminary health warnings such as abnormal heart rate and blood pressure fluctuations. The original video desensitization is completed on the local device, such as face blurring. By embedding a hardware encryption chip such as TPM 2.0 in the video acquisition module, differential privacy noise is introduced in the desensitization stage: adding Laplace noise to the encrypted feature vector adding Laplace noise to the encrypted feature vector satisfies wherein, is a noise scale parameter; ensures that the data is irreversible; only uploads the noise feature vector, avoiding the cloud interaction dependence of federated learning. The differential privacy noise desensitizes the encrypted feature vector, only uploads the noise data, avoids sensitive information leakage, complies with medical data privacy protection regulations such as HIPAA, and reduces the dependence on cloud federated learning.
[0069] The system further comprises an environment adaptive module, which adjusts the motion of the perception measurement module through a feedback control algorithm; or adjusts the calculation parameters of the micro-motion stabilization and correction module; or adjusts the calculation parameters of the three-dimensional trajectory reconstruction module; or adjusts the intensity of the multi-wavelength light source, automatically adapts to different lighting conditions, reduces the influence of external interference on the optical signal quality, ensures the detection stability, and the calculation formula of the intensity of the multi-wavelength light source is:
[0070] ;
[0071] wherein, is an adjustment coefficient, which is optimized in real time through Kalman filtering; is the adjusted light source intensity; is the light source intensity before adjustment.
[0072] Embodiment 2
[0073] As Figure 2 shown, in an embodiment of the present application, a multi-modal optical non-contact health monitoring method is provided, which comprises:
[0074] S1, synchronously collect human face video images, TOF depth maps and IMU data of the perception measurement module;
[0075] S2, perform stabilization processing and motion artifact correction on the collected face video images to obtain stable face video frames, final local micro-motion fields, a main frequency signal of a heart rate physiological frequency band and a main frequency signal of a respiration frequency;
[0076] S3, reconstruct a local micro-motion three-dimensional trajectory according to two-dimensional displacement time series data in the final local micro-motion field;
[0077] S4, perform laser speckle blood flow imaging, hyperspectral imaging and infrared thermal imaging on the stabilized face video images to obtain multi-modal optical data, and align the multi-modal optical data according to a three-dimensional space-time reference framework of the local micro-motion three-dimensional trajectory;
[0078] S5, map blood flow parameters to preliminary blood pressure values and blood oxygen monitoring values and body temperature monitoring values through a physical model driven based on the multi-modal optical data aligned in the three-dimensional space-time, extract local micro-motion time series features related to a heart rate and a respiration frequency from the main frequency signals of the heart rate physiological frequency band and the respiration frequency physiological frequency band through a time series feature fusion network to obtain heart rate monitoring values and respiration frequency monitoring values, and dynamically weight and fuse the preliminary blood pressure values and blood pressure values calculated from the local micro-motion time series features related to the heart rate through a space-time attention mechanism to obtain blood pressure monitoring values;
[0079] S2 performs stabilization and correction of motion artifacts on the collected face video images through a micro-motion stabilization and correction module, as shown in FIG. 2, and the specific execution of S2 includes S2A-S2J. Figure 3
[0080] S2A, input an aligned image pair: register adjacent frame face video images through affine transformation to ensure spatial consistency of the two frame face video images and eliminate the influence of global motion such as head translation and rotation.
[0081] S2B, global motion correction: detect global feature points such as eye corners and nose tips on the face video images processed through S2A using a SURF algorithm, obtain a global motion component by calculating a global affine transformation matrix to eliminate camera jitter or patient overall displacement. Optionally, histogram equalization or adaptive gamma correction is performed on each frame of face video images to reduce the interference of light changes on feature extraction; wherein the parameter of the gamma correction is dynamically adjusted, and the pixel value after gamma correction is , wherein I raw is an original pixel value.
[0082] S2B includes the specific steps of S2B1-S2B5:
[0083] S2B1, detect global feature points of the adjacent frame face video images by SURF algorithm, obtain matching global feature point pairs, the global feature points include head contour and key points of five organs, such as eye corner, nose tip and other facial landmark points, estimate the overall motion between the adjacent frame face video images, such as head translation, rotation or camera shaking, wherein the saliency of the key points is measured by the determinant of Hessian matrix, the greater the absolute value of the determinant, the more significant the image structure at the key point, such as high-contrast spots or corners;
[0084] S2B2, use RANSAC algorithm to eliminate abnormal matching global feature point pairs;
[0085] S2B3, score the confidence of the matching global feature point pairs remaining after S2B2 processing by MobileNetV2, retain high-confidence matching global feature point pairs, such as retaining matching global feature point pairs with a confidence score ≥0.8, output a stable sparse matching global feature point pair set;
[0086] S2B4, use least squares method or SVD decomposition, for the sparse matching global feature point pair set output by S2B3, based on multiple matching global feature point pairs in the adjacent frame face video images, calculate the global affine transformation matrix :
[0087] , wherein, is the rotation, scaling and shearing parameters; is the translation component, i.e. displacement along the x and y axes; the third row is fixed as , for homogeneous coordinate transformation;
[0088] S2B5, use the global affine transformation matrix to predict the global motion component of all global feature points : ;
[0089] wherein, for each global feature point / pixel point in the current face video image , the predicted motion vector of each global feature point / pixel point is obtained by conversion , wherein, is the coordinate of the global feature point / pixel point after the global affine transformation matrix transformation.
[0090] For example, input: matching global feature point pairs in adjacent frame face video images ;
[0091] Calculation: Solve by least squares , such that ;
[0092] Output: Predicted global motion component , the predicted motion vector of each global feature point is .
[0093] S2C, generate initial dense motion vector field: perform pixel-level motion prediction on the aligned face video image after S2A by lightweight visual TRANSFORMER, output two-dimensional motion vector of each pixel , form a preliminary dense motion vector field , covering all image pixels; specifically, the encoder uses MobileViT as the backbone network, divides the aligned face video image into image blocks, extracts global motion features through self-attention mechanism, adds a lightweight upsampling module such as transpose convolution or sub-pixel convolution in the decoder, gradually restores the spatial resolution, and outputs a dense motion vector field with the same resolution as the input; finally, through 1x1 convolution, the global motion features extracted by the decoder are mapped to two-dimensional motion vectors , generate a preliminary dense motion vector field ; embed window attention (Window Attention) and shifted window (Shifted Window) in MobileViT to capture long-distance motion dependencies while reducing computational complexity, model the spatio-temporal correlation between adjacent frames through cross-frame attention (Cross-frame Attention), and improve the accuracy of motion vector prediction.
[0094] S2D, local region division: filter the global feature points on the face video image after global motion correction by S2B through Hessian matrix, and divide the local region such as 16x16 pixel block with the filtered global feature points as anchor points.
[0095] Filter the global feature points on the face video image after global motion correction by S2B through Hessian matrix, which includes S2D1~S2D5:
[0096] S2D1, construct scale space: convolve the face video image after global motion correction by S2B with different standard deviations of Gaussian kernel to generate multi-scale images;
[0097] S2D2, calculate Hessian matrix: ; wherein represents the face video image after global motion correction by S2B along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the second-order Gaussian derivative convolution result along the direction of the standard deviation of the Gaussian kernel;
[0098] S2D3, computing response value: the determinant det(H) of the Hessian matrix computed by S2D2 evaluates the keypoint saliency;
[0099] S2D4, non-maximum suppression: screening local extreme points in the scale space and neighborhood constructed in S2D1, i.e., screening global feature points with response value higher than threshold τ;
[0100] S2D5, global feature point positioning: accurately positioning to sub-pixel level, eliminating low-contrast or edge response points from the global feature points screened in S2D4.
[0101] S2E, optical flow method refinement: in each local region, the improved pyramid Lucas-Kanade optical flow method is used to process the preliminary dense motion vector field to generate a refined local dense motion vector field.
[0102] As shown in Figure 4 , S2E specifically includes S2E1-S2E3:
[0103] S2E1, pyramid construction: multi-scale down-sampling is performed on each local region to construct a pyramid for layer-by-layer optimization of optical flow estimation from coarse to fine; for example, a 3-layer pyramid is constructed to optimize the two-dimensional motion vector of the local region layer by layer from the low-resolution layer to the high-resolution layer, and the calculation complexity is reduced through coarse-fine layer-by-layer optimization;
[0104] S2E2, low-resolution layer optimization: the two-dimensional motion vector predicted by the lightweight visual TRANSFORMER is used as the initial value to quickly estimate the large-scale motion of the local region;
[0105] S2E3, high-resolution layer optimization: based on the brightness constancy assumption, the optical flow equation is solved to iteratively optimize the two-dimensional motion vector of each pixel in the local region to refine the micro-motions of each local region, such as skin vibration and chest and abdomen fluctuation; wherein, is the spatial gradient, is the time gradient, is the horizontal and vertical displacement of the pixel.
[0106] The improved pyramid Lucas-Kanade optical flow method fuses the lightweight visual TRANSFORMER output, predicts the global motion trend through the encoder-decoder structure of the FlowNet type, reduces the number of iterations of the optical flow method in the complex motion scene, and then uses the preliminary dense motion vector field generated by the lightweight visual TRANSFORMER as the initial value of the optical flow method, instead of the zero initialization of the traditional method, to improve the motion estimation accuracy of the local area and generate a refined local dense motion vector field, i.e., a preliminary local dense motion vector field .
[0107] S2F, abnormal motion elimination: using the RANSAC algorithm to detect outliers in the preliminary local dense motion vector field, eliminate abnormal motion such as random noise or non-physiological motion, and retain physiological signal-related micro-motions such as skin vibration and chest and abdominal fluctuation, to obtain a corrected local micro-motion field ; wherein the number of RANSAC iterations and the inlier error threshold need to be set based on experiments, such as 500 iterations and 1 pixel inlier error threshold;
[0108] S2G, superimposition to generate a complete motion field: superimpose the global motion component and the corrected local micro-motion field to generate a complete motion vector field , the formula is:
[0109] ;
[0110] S2H, correction of the global motion component: use the complete motion vector field to verify whether the global motion component has residual abnormalities, and if there are residual abnormalities, eliminate them through the RANSAC algorithm to obtain the corrected global motion component Further eliminate residual artifacts; Residual outliers can also reduce the impact of outliers on model parameters by combining the squared loss for normal data and the absolute loss for outliers; Or by calculating the Z-score of the motion vector, removing outliers that deviate from the mean by more than a threshold, such as removing outliers with Z > 3; Or compare the optical flow vectors of adjacent frames, remove outliers with significantly inconsistent direction or amplitude; Or based on density clustering, identify isolated points that do not belong to any dense cluster as outliers; Or based on the verification of the motion model by fitting the residual to detect outliers; Or by training a convolutional neural network (CNN) or autoencoder to distinguish between normal and abnormal motion patterns; RANSAC algorithm can also be combined with one or more of the above outlier detection methods to further improve the accuracy and adaptability of global motion component correction.
[0111] S2I, separate local motion field: subtract the corrected global motion component from the complete motion field , get the final local micro-motion field , Only physiological related motion such as skin vibration, chest and abdominal fluctuation is retained;
[0112] ;
[0113] For example, the global motion is a slight right shift of the patient's head, and the local micro-motion is the slight tremor of the facial skin due to heartbeat. The SURF SURF algorithm correction model detects the global feature points: the head as a whole moves 2 pixels to the right and rotates 1°;
[0114] The global motion component is obtained by affine transformation matrix of global feature points For simplicity, the global affine transformation matrix only contains horizontal translation is: ;
[0115] ;
[0116] If the actual motion of a pixel is the complete motion vector field , assuming that the global motion component does not need to be corrected, the final local micro-motion field reflects the skin vibration.
[0117] S2J, motion field post-processing: the final local micro-motion field Preprocessing including Kalman filtering and Gaussian filtering is performed to achieve time smoothing and space smoothing and suppress high-frequency noise; wherein the state equation of Kalman filtering is a uniform motion model, and the convolution kernel size of Gaussian filtering can be 3*3; then the two-dimensional displacement time sequence data of the final local micro-motion field physiological frequency band 0.1-4 Hz after preprocessing is extracted through FFT, and the physiological frequency band to be extracted covers the physiological frequency bands of heart rate and breathing frequency, wherein the physiological frequency band of heart rate is 0.8-3 Hz, and the physiological frequency band of breathing frequency is 0.1-0.5 Hz; and multi-scale decomposition and hard threshold denoising are performed by using wavelet transform, and the rate physiological frequency band coefficient and the breathing frequency physiological frequency band coefficient are retained; and the rate physiological frequency band coefficient and the breathing frequency physiological frequency band coefficient are reconstructed, and the main frequency signals of the rate physiological frequency band and the breathing frequency physiological frequency band are extracted through FFT or power spectral density analysis.
[0118] The specific steps of multi-scale decomposition and hard threshold denoising by using wavelet transform are as follows:
[0119] S2J1, signal preprocessing: the two-dimensional displacement data is subjected to first-order difference or polynomial fitting to eliminate baseline drift;
[0120] S2J2, discrete wavelet decomposition: the two-dimensional displacement time sequence data of the final local micro-motion field after preprocessing is subjected to multi-scale decomposition to separate high-frequency noise, physiological frequency band signals and low-frequency artifacts, and each scale corresponds to a different frequency range; for example: high-frequency scale decomposition, obtaining : corresponding to transient noise, i.e. high-frequency noise > 4 Hz, : the proxy of high-frequency noise, the global noise level is estimated by statistics of the median of the absolute value; medium-frequency scale decomposition, obtaining : covering the physiological frequency band 0.5-4 Hz; low-frequency scale decomposition, obtaining : corresponding to baseline drift, i.e. low-frequency artifact < 0.5 Hz.
[0121] S2J3, physiological frequency band extraction: retaining the physiological frequency band wavelet coefficients, i.e. the wavelet coefficients of 0.5-4 Hz, corresponding to the heart rate and the breathing frequency , and performing hard threshold processing on the non-physiological frequency bands, such as high-frequency noise D1, D2 and low-frequency motion artifact A J ; wherein the hard threshold is adaptively set according to noise, and the general hard threshold is , and the hard threshold in the application is ; the wavelet coefficients of the non-physiological frequency bands are set to zero, and only the wavelet coefficients of the medium-frequency scale, i.e. the heart rate physiological frequency band and the breathing frequency physiological frequency band, are retained;
[0122] S2J4, signal reconstruction: the retained wavelet coefficients of the heart rate physiological frequency band and the breathing frequency physiological frequency band are reconstructed into pure signals;
[0123] S2J5, physiological signal analysis: the main frequency of the pure signal is extracted by FFT or power spectral density (PSD), such as the main frequency of the heart rate physiological frequency band and the main frequency of the respiratory frequency physiological frequency band. The main frequency of the heart rate physiological frequency band is 1.2 Hz. The peak value of the pulse wave is identified by the adaptive threshold method, and the heart rate variability is calculated.
[0124] In addition, the original face video frame image can be deformed in reverse according to the inverse global affine transformation matrix, and the output of the face video frame after eliminating the camera camera jitter or overall displacement is stabilized The reverse deformation realizes the spatial alignment of the image sequence through inverse geometric transformation, while retaining the local micro-motion information. The reverse deformation and wavelet transformation have clear division of labor and jointly improve the system robustness.
[0125] Among them, the reverse deformation is a key step to eliminate global motion artifacts through inverse geometric transformation. The reverse deformation not only ensures the spatial alignment of the image sequence, but also retains the local micro-motion information, providing a reliable data basis for high-precision physiological monitoring.
[0126] In S3, the two-dimensional displacement time series data in the final local micro-motion field are mapped to three-dimensional motion using a kinematic model such as a perspective projection model or a rigid body motion model, reflecting the three-dimensional dynamic changes of the human body surface or organs. Then, Kalman filtering is used to smooth the three-dimensional trajectory in time sequence, suppressing high-frequency noise.
[0127] The formula for three-dimensional trajectory reconstruction is:
[0128]
[0129] Among them, is the three-dimensional displacement vector of the local micro-motion, is the horizontal and vertical displacement of the target point cloud in the camera coordinate system; is the depth change of the target point cloud along the optical axis direction, which needs to be constrained or assumed, such as weak perspective projection; is the camera intrinsic matrix; is the two-dimensional displacement vector in the final local micro-motion field ;
[0130] Mapping through a normalized coordinate system:
[0131] ;
[0132] Assuming the depth , i.e. planar motion or known depth , the three-dimensional displacement of local micro-motion is: ,
[0133] If the depth is unknown, the depth variation of the target region, such as the face, is directly obtained by integrating a millimeter-level precision TOF camera The specific steps are: mapping relationship between the depth map and the three-dimensional coordinates of the camera is established through TOF sensor internal parameter calibration and external parameter calibration, the TOF sensor internal parameter calibration includes focal length and distortion parameters, and the TOF sensor external parameter calibration includes the positional relationship with the RGB camera.
[0134] S3 specifically includes S3A to S3C.
[0135] S3A: Multi-modal data synchronous acquisition: TOF depth map, facial video image and IMU data are synchronously acquired.
[0136] S3B: Point cloud preprocessing: statistical outlier removal and radius filtering are used to eliminate noise points in the TOF depth map data, and the depth map of the human face region is generated; combined with the facial video image, the depth map of the human face region is segmented using Mask R-CNN, and the depth map of the focused physiological micro-motion is obtained.
[0137] S3C: Motion compensation and trajectory generation: the point cloud data of the focused physiological micro-motion depth map is registered by the Iterative Closest Point (ICP) algorithm, the global motion of the head such as rigid transformation and the local micro-motion such as skin vibration are separated, and the target region is determined from the separated local micro-motion point cloud data; the mapping relationship between the target region depth map and the three-dimensional coordinates of the camera is established according to the TOF sensor internal parameter calibration and external parameter calibration, and the three-dimensional trajectory of the local micro-motion is generated.
[0138] Target region depth variation is much smaller than its distance to the camera , which is approximately a planar motion,
[0139] At this time, the three-dimensional displacement is simplified as: ;
[0140] Target region depth variation is much smaller than its distance to the camera , which is approximately a planar motion, under typical scenarios, the three-dimensional displacement is simplified as: ;
[0141] wherein, is the average depth measured by the TOF sensor, is the camera intrinsic matrix, is the two-dimensional displacement vector in the corrected local micro-motion field; the relative error of this simplified model is less than 0.1%, which meets the medical-level precision requirement.
[0142] The application maps two-dimensional optical flow into three-dimensional displacement through a kinematic model, combines Kalman filtering to time sequence smooth the three-dimensional trajectory of local micro-movement after separating the overall movement of the head, suppresses high-frequency noise, provides spatial dynamic information, significantly improves the analysis accuracy and robustness of physiological signals such as skin vibration and chest and abdominal fluctuation, and provides a unified space-time reference framework for multi-modal data fusion.
[0143] The process of mapping blood flow parameters into preliminary blood pressure values, blood oxygen monitoring values and body temperature monitoring values based on multi-modal optical data aligned in three-dimensional space-time in S5 through a physical model is as follows:
[0144] Data input: blood flow parameters of each blood flow region extracted from the U-Net blood flow segmentation network, blood flow parameter samples obtained after standardization processing, and blood flow parameter samples input into a 3-layer fully connected network;
[0145] Model output: 3-layer fully connected network outputs preliminary physiological indicators such as systolic pressure , diastolic pressure , blood oxygen , and body temperature ;
[0146] Loss calculation:
[0147] Data loss: the mean square error between the physiological indicator prediction value obtained according to the i th blood flow parameter sample and the true label is calculated as
[0148] ;
[0149] Physical constraint loss: based on the Beer-Lambert law, the blood oxygen saturation prediction value is constrained to be consistent with the red light / near-infrared light reflectance intensity ratio , and the physical constraint loss function is calculated based on the Beer-Lambert law as
[0150]
[0151] wherein, and are calibration parameters, which force the network output to comply with the optical absorption law; N represents the total number of data samples;
[0152] Parameter optimization: for the mathematical mapping relationship of mapping blood flow parameters into blood oxygen saturation, the total loss Parameter fine-tuning is performed, and for the mathematical mapping relationship of mapping blood flow parameters to blood pressure and body temperature, parameter fine-tuning is performed according to data loss, to ensure that the diffusion optical spectrum mapping network output conforms to the physical model.
[0153] The physical model based on the laser speckle image maps the blood flow velocity , pulse wave transmission time to the preliminary blood pressure value BP:
[0154] ;
[0155] wherein, are regression coefficients calibrated by clinical data.
[0156] The physical model based on hyperspectral data maps the red / near-infrared light reflection intensity ratio to the blood oxygen saturation monitoring value :
[0157] ;
[0158] wherein, are parameters determined by the calibration curve.
[0159] The physical model based on infrared thermal imaging maps the blood flow velocity , facial skin temperature to the body temperature monitoring value :
[0160] ;
[0161] wherein, is the facial skin temperature, is the correction coefficient, is the reference blood flow velocity.
[0162] In S5, the timing feature fusion network extracts heart rate and respiration rate related local micro-motion timing features from the main frequency signals of the heart rate physiological frequency band and the respiration rate physiological frequency band, respectively, to obtain heart rate and respiration rate monitoring values, specifically:
[0163] The main frequency of the heart rate physiological frequency band and the main frequency of the respiration rate physiological frequency band obtained from the final local micro-motion field are input into the LSTM, and the LSTM is used for prediction to extract the heart rate and the respiration rate :
[0164] ;
[0165] ;
[0166] wherein the main frequency of the respiratory frequency physiological frequency band is the two-dimensional motion displacement in the final local micro-motion field Wavelet transform is performed to obtain the main frequency in the 0.1-0.5 Hz frequency band.
[0167] In S5, the preliminary blood pressure value and the blood pressure value calculated from the heart rate related local micro-motion time sequence feature are dynamically weighted and fused by the spatio-temporal attention mechanism to obtain a blood pressure monitoring value, specifically:
[0168] The weights of the blood pressure value calculated from the heart rate related local micro-motion time sequence feature and the preliminary blood pressure value are dynamically adjusted by the spatio-temporal attention mechanism 、 to enhance the robustness in complex environments; the spatio-temporal attention mechanism analyzes the time sequence changes of physiological indicators in the time dimension, dynamically weights the confidence of different modal data through the attention mechanism for multi-modal data fusion, realizes accurate analysis of the time sequence changes of physiological indicators, and combines adaptive weighting of the preliminary physiological indicators to finally improve the prediction accuracy and stability of the non-contact health monitoring system; the time sequence changes of physiological indicators are analyzed in the time dimension; specifically: periodic capture: modeling the rhythmic changes of heartbeat and respiration; trend tracking: identifying long-term fluctuations of blood pressure and blood oxygen; noise suppression: reducing the weight of interference time steps. The reliability and accuracy of non-contact health monitoring in complex environments such as exercise and light changes are significantly improved.
[0169] According to the dynamically adjusted weight , the systolic blood pressure value calculated from the heart rate related local micro-motion time sequence feature is weighted and the preliminary systolic blood pressure value , the blood pressure estimation value is optimized by Kalman filtering to obtain a real-time prediction value of systolic blood pressure, i.e. a systolic blood pressure monitoring value . The dynamic weighting fusion method of diastolic blood pressure is the same as that of systolic blood pressure, and the present application will not be repeated.
[0170] ;
[0171] wherein the pulse wave transmission time is obtained according to the main frequency signal of the heart rate physiological frequency band.
[0172] At the first use, the user measures the baseline value by a medical device , the system dynamically adjusts the weight according to the difference between the blood pressure real-time prediction value and the baseline value , and outputs a more accurate final blood pressure :
[0173] ;
[0174] wherein, represents the final blood pressure value after error compensation; represents the user individualized weight coefficient; represents the blood pressure value directly predicted by an algorithm such as a diffuse optical spectroscopy mapping network; represents the user baseline blood pressure value, i.e., the initial calibration value; the present application introduces physical constraints such as the Beer-Lambert law in the diffuse optical spectroscopy mapping network, constrains the relationship between the blood oxygen saturation prediction value and the optical reflection intensity, enhances the physical interpretability of the model, ensures that the output result conforms to the medical law, and improves the trust degree of the doctor on the system output. Clinical verification: data are collected synchronously with the gold standard equipment, the regression coefficient is optimized through cross-validation, and the heart rate error is ≤±2 BPM, the blood pressure error is ≤±5 mmHg, and the SpO2 error is ≤±2%.
[0175] In order to verify the effect of the present application, the blood pressure measurement error performance of the traditional method and the present application in two dynamic scenes is compared, and the specific data is shown in Table 1.
[0176] Table 1
[0177] Scenario Conventional method error (mmHg) Invention error (mmHg) Head level movement ±8.2 ±3.5 Light abrupt change (2000 lux) ±12.1 ±4.8
[0178] As can be seen from Table 1, in the head horizontal movement scene, the error of the traditional method is ±8.2 mmHg, and the error of the present application is only ±3.5 mmHg; in the light mutation (2000 lux) scene, the error of the traditional method reaches ±12.1 mmHg, and the error of the present application is ±4.8 mmHg. As can be seen, in these two dynamic scenes, the present application has obvious advantages in error control compared with the traditional method, and the measurement result is more accurate.
[0179] The present application combines laser speckle imaging, hyperspectral data, infrared imaging data and other multi-modal optical data, synchronously acquires blood pressure, heart rate, blood oxygen saturation and other multi-dimensional physiological indicators, overcomes the limitations of traditional single modal such as only PPG or VPG, meets the demand of multi-parameter joint analysis in complex clinical scenes, and improves the comprehensiveness of detection dimension. Through multi-level correction such as SURF algorithm, RANSAC abnormal rejection and Kalman filtering, global motion and local micro-motion are separated, and in the scene of slight movement of the patient or shaking of the camera, physiological signals such as heart rate and respiratory rate can still be accurately extracted.
[0180] The application solves the problem of multi-modal data coupling interference in the traditional method by preferentially processing optical imaging data through the U-Net network and then fusing with the video micro-motion features. The mathematical mapping relationship between the blood flow parameters and the physiological indexes is established through the diffuse optical spectrum mapping network, and the segmented blood flow parameters such as blood flow velocity, blood oxygen saturation SpO2 and hemoglobin concentration are mapped to the preliminary physiological indexes such as blood pressure, blood oxygen and body temperature. The time sequence features such as heart rate and respiratory rate extracted by the time sequence feature fusion network solve the time lag problem of single modal data. Then, the weight of the blood pressure value calculated according to the heart rate related local micro-motion time sequence features and the preliminary blood pressure mapping value is dynamically adjusted through the space-time attention mechanism, the robustness in complex environment is enhanced, and the prediction accuracy and stability of the non-contact health monitoring system are improved.
[0181] While the preferred embodiments of the application have been described, additional alternatives, modifications, and variations would become apparent to one of ordinary skill in the art after having the benefit of this disclosure. Therefore, the appended claims shall be construed to include all such alternatives, modifications, and variations as falling within the scope of the application.
Claims
1. A multimodal optical non-contact health monitoring system, characterized in that, include: The perception and measurement module is used to acquire facial video images, TOF depth maps, and IMU data of the human body; The data synchronization acquisition module is used to simultaneously acquire facial video images, TOF depth maps, and IMU data of the human body; The micro-motion stabilization and correction module is used to stabilize and correct motion artifacts in synchronously acquired facial video images, and output stable facial video frames, the final local micro-motion field, the main frequency signal of the heart rate physiological frequency band, and the main frequency signal of the respiratory rate physiological frequency band. The three-dimensional trajectory reconstruction module is used to reconstruct the three-dimensional trajectory of local micro-motion based on the two-dimensional displacement time-series data in the final local micro-motion field. A multimodal data processing module is used to perform optical processing on the stable facial video frames, acquire multimodal optical data including but not limited to laser speckle images, hyperspectral data, and infrared thermal imaging through a multimodal optical module, and align the multimodal optical data according to a three-dimensional spatiotemporal reference frame of the local micro-motion three-dimensional trajectory; and, The physiological index calculation module is used to map blood flow parameters into preliminary blood pressure, blood oxygen, and body temperature values based on three-dimensional spatiotemporally aligned multimodal optical data through a physical model. It extracts local micro-motion temporal features related to heart rate and respiratory rate from the main frequency signals of the physiological frequency bands of heart rate and respiratory rate, respectively, to obtain heart rate and respiratory rate monitoring values. It dynamically weights and fuses the preliminary blood pressure value with the blood pressure value calculated from the local micro-motion temporal features related to heart rate to obtain the blood pressure monitoring value, and outputs the monitoring values of heart rate, blood pressure, blood oxygen, body temperature, and respiratory rate.
2. The multimodal optical non-contact health monitoring system according to claim 1, characterized in that, The perception and measurement module includes: an RGB camera for acquiring video images of the human face, a time-of-flight sensor for acquiring TOF depth maps of the human face, and an inertial measurement unit for acquiring IMU data; the data synchronization acquisition module acquires video images of the human face, TOF depth maps, and IMU data through hardware synchronization technology.
3. The multimodal optical non-contact health monitoring system according to claim 2, characterized in that, The multimodal optical module uses a tunable multi-wavelength light source.
4. The multimodal optical non-contact health monitoring system according to claim 3, characterized in that, The micro-motion stabilization and correction module includes: A stable image unit corrects facial video frame images using the SURF algorithm correction model, outputting stable facial video frames; and... The artifact correction unit performs pixel-level two-dimensional motion vector prediction on facial video frame images using a lightweight visual transformer to generate a preliminary dense motion vector field. It then generates a preliminary local dense motion vector field using an improved pyramid Lucas-Kanade optical flow method. The preliminary local dense motion vector field is corrected to obtain a corrected local micro-motion field. Based on the corrected local micro-motion field, a complete motion field is generated and global motion component correction is performed. Finally, the corrected global motion components are used to obtain the final local micro-motion field. The main frequency signals of the physiological frequency bands of heart rate and respiratory rate are extracted from the final local micro-motion field.
5. The multimodal optical non-contact health monitoring system according to claim 4, characterized in that, The physiological index calculation module includes: The U-Net blood flow segmentation network is used to extract multi-scale blood flow region features from the multimodal optical data, reconstruct blood flow region masks, generate binarized blood flow region segmentation results, and extract blood flow parameters for each blood flow region, including but not limited to blood flow velocity, blood oxygen saturation, and hemoglobin concentration. A diffuse optical spectral mapping network is used to establish a mathematical mapping relationship between blood flow parameters and preliminary physiological indicators, outputting preliminary blood pressure, blood oxygen monitoring values, and body temperature monitoring values; and, The temporal feature fusion network extracts local micro-motion temporal features related to heart rate and respiratory rate from the main frequency signals of the physiological frequency bands of heart rate and respiratory rate, respectively. The preliminary blood pressure value is dynamically weighted and fused with the blood pressure value calculated from the local micro-motion temporal features related to heart rate through a spatiotemporal attention mechanism to obtain the blood pressure monitoring value.
6. The multimodal optical non-contact health monitoring system according to claim 5, characterized in that, The system also includes an environment adaptation module, which adjusts the motion of the sensing and measurement module through a feedback control algorithm; or adjusts the calculation parameters of the micro-motion stabilization and correction module; or adjusts the calculation parameters of the three-dimensional trajectory reconstruction module; or adjusts the intensity of the multi-wavelength light source. The physiological index calculation module also includes a temperature calibration submodule, which compensates for temperature drift errors in blood flow parameters by using Kalman filtering based on infrared thermal imaging temperature data.
7. A multimodal optical non-contact health monitoring method, characterized in that, The system implementation based on any one of claims 1 to 6 includes the following steps: S1, synchronously acquires human facial video images, TOF depth maps, and IMU data from the perception measurement module; S2, performs stabilization processing and motion artifact correction on the acquired facial video images to obtain stable facial video frames, the final local micro-motion field, the main frequency signal of the heart rate physiological frequency band, and the main frequency signal of the respiratory rate; S3, Reconstruct the three-dimensional trajectory of local micro-motion based on the two-dimensional displacement time series data in the final local micro-motion field; S4. Laser speckle blood flow imaging, hyperspectral imaging, and infrared thermal imaging are performed on the stabilized facial video images to acquire multimodal optical data. The multimodal optical data are aligned with the three-dimensional spatiotemporal reference frame of the local micro-motion three-dimensional trajectory. S5, based on three-dimensional spatiotemporally aligned multimodal optical data, maps blood flow parameters to preliminary blood pressure, blood oxygen, and body temperature values through a physical model. It extracts local micro-motion temporal features related to heart rate and respiratory rate from the dominant frequency signals of the physiological frequency bands of heart rate and respiratory rate, respectively, through temporal feature fusion technology, and obtains heart rate and respiratory rate monitoring values. It then dynamically weights and fuses the preliminary blood pressure value with the blood pressure value calculated from the local micro-motion temporal features related to heart rate through a spatiotemporal attention mechanism to obtain the blood pressure monitoring value.
8. The multimodal optical non-contact health monitoring method according to claim 7, characterized in that, S2 specifically includes the following steps: S2A aligns the input facial video frame images by registering adjacent frames through affine transformation. S2B: For the facial video image after S2A alignment, the SURF algorithm is used to detect global feature points, and the global motion components are obtained by calculating the global affine transformation matrix. The geometric inverse transformation is performed on the input facial video frame image according to the global affine transformation matrix to output a stable facial video frame. S2C uses a lightweight visual transformer to perform pixel-level two-dimensional motion vector prediction on facial video images aligned by S2A, generating a preliminary dense motion vector field. S2D uses the Hessian matrix to filter global feature points on the facial video image after global motion correction by S2B, and uses the selected global feature points as anchor points to divide local regions. S2E employs an improved pyramid Lucas-Kanade optical flow method to process the initial dense motion vector field, generating an initial local dense motion vector field; S2F uses the RANSAC algorithm to detect outliers in the initial local dense motion vector field, removes abnormal motions, and obtains the corrected local micro-motion field. S2G superimposes the global motion components with the corrected local micro-motion fields to generate a complete motion vector field; S2H uses the complete motion vector field to verify whether there are residual anomalies in the global motion components. If residual anomalies exist, they are removed by the RANSAC algorithm to obtain the corrected global motion components. S2I is obtained by subtracting the corrected global motion components from the complete motion field to obtain the final local micro-motion field. S2J preprocesses the final local micro-motion field, extracts the two-dimensional displacement time series data of the physiological frequency band of the preprocessed final local micro-motion field, and uses wavelet transform for multi-scale decomposition and hard threshold denoising. It retains the physiological frequency band coefficients of heart rate and respiratory rate, and after reconstructing the physiological frequency band coefficients of heart rate and respiratory rate, it extracts the main frequency of the physiological frequency bands of heart rate and respiratory rate respectively.
9. The multimodal optical non-contact health monitoring method according to claim 8, characterized in that, The S2E specifically includes the following steps: S2E1 performs multi-scale downsampling on each local region to construct a pyramid for optimizing optical flow estimation layer by layer from low resolution to high resolution; S2E2 uses the initial dense motion vector field generated by S2C as the initial value to quickly estimate the motion range of the local area and optimize the low-resolution layer. S2E3, based on the assumption of constant brightness, solves the optical flow equation to iteratively optimize the two-dimensional motion vector of each pixel in the local region, and performs high-resolution layer optimization.
10. The multimodal optical non-contact health monitoring method according to claim 9, characterized in that, S3 reconstructs the three-dimensional trajectory of the local micro-motion based on the two-dimensional displacement time-series data in the final local micro-motion field, specifically as follows: S3A acquires synchronously collected TOF depth maps, human facial video images, and IMU data; S3B employs statistical outlier removal and radius filtering to eliminate noise points in TOF depth map data, generating a depth map of the human face region. Combined with facial video images, Mask R-CNN is used to segment the depth map of the human face region to obtain a depth map focusing on physiological micro-movements. S3C uses an iterative nearest-point algorithm to register and focus on the point cloud data of the depth map of physiological micro-movements, separating the overall head movement from local micro-movements. The target area is determined by the separated local micro-movement point cloud data, and the local micro-movement 3D trajectory is generated based on the mapping relationship between the target area depth map and the 3D coordinates of the facial video image established by the intrinsic and extrinsic parameter calibration of the TOF sensor.
Citation Information
Patent Citations
Bathtub with multiple massage functions
CN112545339A
Electrophysiology and hemodynamics multi-mode intelligent wearable system and data evaluation method
CN119837506A
Vision-Based Cardiorespiratory Monitoring
US20250025052A1