Non-contact multi-parameter monitoring method and system for physical and mental health analysis
By combining visible light and thermal infrared video data, and utilizing a deentangled representation learning network and a deep physiological parameter estimation network, the problem of detection instability in existing health monitoring systems under complex scenarios is solved, achieving high-precision and robust non-contact multi-parameter monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WEST CHINA HOSPITAL SICHUAN UNIV
- Filing Date
- 2023-03-31
- Publication Date
- 2026-07-14
AI Technical Summary
Existing health monitoring systems suffer from issues such as contact monitoring methods causing user inconvenience and discomfort, and non-contact monitoring solutions exhibit unstable detection performance and poor adaptability in complex scenarios such as changes in lighting conditions and significant head movements.
By acquiring data from visible light facial video and thermal infrared facial video, and using a deentanglement representation learning network and a deep physiological parameter estimation network to deentangle physiological and non-physiological features, a non-contact multi-parameter monitoring method and system are constructed to improve detection accuracy and robustness.
In complex scenarios such as changes in lighting conditions and significant head movements, the accuracy and robustness of physiological parameter detection are improved, providing highly comfortable and integrated multi-physiological parameter monitoring services.
Smart Images

Figure CN116403734B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data monitoring technology, and in particular to a non-contact multi-parameter monitoring method and system for physical and mental health analysis. Background Technology
[0002] Existing health monitoring systems include contact and non-contact solutions. Most existing contact-based solutions are based on wearable devices such as finger-clip pulse oximeters, wristbands, and portable blood pressure monitors. These contact-based methods can cause physiological or psychological discomfort to users and are not suitable for use in special populations such as newborns and patients with large skin wounds. Most existing non-contact health monitoring solutions use traditional signal processing methods to calculate physiological parameters. These solutions have poor data adaptability, and their performance fluctuates significantly due to noise interference in complex scenarios such as large changes in lighting conditions and significant head movements. Therefore, existing health monitoring systems generally suffer from poor monitoring effectiveness. Summary of the Invention
[0003] To address the aforementioned technical problems, this application provides a non-contact multi-parameter monitoring method and system for physical and mental health analysis.
[0004] In a first aspect, embodiments of this application provide a non-contact multi-parameter monitoring method for physical and mental health analysis, the method comprising:
[0005] Collect visible light facial video and thermal infrared facial video of the user;
[0006] Based on the visible light facial video and the thermal infrared facial video, obtain m first region of interest frame sequences and m second region of interest frame sequences;
[0007] A first spatiotemporal graph is constructed based on m first region of interest frame sequences, and a second spatiotemporal graph is constructed based on m second region of interest frame sequences;
[0008] The first and second spatiotemporal graphs are de-entangled using a pre-constructed de-entanglement representation learning network to obtain multiple physiological feature data by de-entanglement of physiological and non-physiological features.
[0009] Multiple physiological parameters of the user are obtained by calculating multiple physiological feature data through a pre-constructed deep physiological parameter estimation network.
[0010] Secondly, embodiments of this application provide a non-contact multi-parameter monitoring system for physical and mental health analysis, the system comprising:
[0011] The video capture module is used to capture the user's visible light facial video and thermal infrared facial video;
[0012] The acquisition module is used to acquire m first region of interest frame sequences and m second region of interest frame sequences based on the visible light facial video and the thermal infrared facial video;
[0013] The construction module is used to construct a first spatiotemporal graph based on m first region of interest frame sequences and to construct a second spatiotemporal graph based on m second region of interest frame sequences.
[0014] The processing module is used to perform physiological and non-physiological feature unentanglement processing on the first spatiotemporal graph and the second spatiotemporal graph through a pre-constructed unentangled representation learning network to obtain multiple physiological feature data.
[0015] The calculation module is used to calculate multiple physiological parameters of the user by using a pre-built deep physiological parameter estimation network to calculate multiple physiological feature data.
[0016] The non-contact multi-parameter monitoring method and system for physical and mental health analysis provided in this application acquires visible light facial videos and thermal infrared facial videos of users; obtains m first region of interest frame sequences and m second region of interest frame sequences based on the visible light facial videos and the thermal infrared facial videos; constructs a first spatiotemporal map based on the m first region of interest frame sequences and a second spatiotemporal map based on the m second region of interest frame sequences; performs physiological and non-physiological feature decoupling processing on the first and second spatiotemporal maps using a pre-constructed deentanglement representation learning network to obtain multiple physiological feature data; and calculates multiple physiological parameters of the user using a pre-constructed deep physiological parameter estimation network. Thus, by using a deentanglement representation learning network and a deep physiological parameter estimation network for non-contact calculation of multiple physiological parameters, the accuracy and robustness of physiological parameter detection are improved in complex scenarios such as large changes in lighting conditions and significant head movements, providing users with a highly comfortable, highly integrated, and highly robust multi-physiological parameter monitoring service. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be considered as a limitation on the scope of protection of this application. In the various drawings, similar components are numbered similarly.
[0018] Figure 1 This illustration shows one of the flowcharts of a non-contact multi-parameter monitoring method for physical and mental health analysis provided in an embodiment of this application;
[0019] Figure 2This illustration shows one of the structural schematic diagrams of the unentangled representation learning network provided in an embodiment of this application;
[0020] Figure 3 This illustration shows one of the structural schematic diagrams of the deep physiological parameter estimation network provided in an embodiment of this application;
[0021] Figure 4 A schematic diagram of one of the respiratory waveforms provided in an embodiment of this application is shown;
[0022] Figure 5 One of the schematic diagrams of heart rate waveforms provided in the embodiments of this application is shown;
[0023] Figure 6 This illustration shows one of the structural schematic diagrams of a non-contact multi-parameter monitoring system for physical and mental health analysis provided in an embodiment of this application.
[0024] Icons: 600 - Non-contact multi-parameter monitoring system for physical and mental health analysis; 601 - Video acquisition module; 602 - Acquisition module; 603 - Construction module; 604 - Processing module; 605 - Calculation module. Detailed Implementation
[0025] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0026] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0027] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0028] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0029] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0030] Example 1
[0031] This application provides a non-contact multi-parameter monitoring method for physical and mental health analysis.
[0032] See Figure 1 The non-contact multi-parameter monitoring method for physical and mental health analysis includes steps S101 to S105. This non-contact multi-parameter monitoring method for physical and mental health analysis can be applied to a non-contact multi-parameter monitoring system for physical and mental health analysis. The steps are described below.
[0033] Step S101: Acquire the user's visible light facial video and thermal infrared facial video.
[0034] In this embodiment, the non-contact multi-parameter monitoring system for physical and mental health analysis includes a video acquisition module, which consists of a binocular vision module and a thermal infrared module, and acquires visible light facial video and thermal infrared facial video of the user in real time.
[0035] Step S102: Obtain m first region of interest frame sequences and m second region of interest frame sequences based on the visible light facial video and the thermal infrared facial video.
[0036] In this embodiment, since the forehead, cheek, and nose regions of the face contain more physiological information and are less involved in facial movements, the forehead, left cheek, right cheek, and nose regions in the visible light facial video are selected as the four first regions of interest (ROIs) for physiological parameter monitoring, respectively denoted as the first forehead ROI1, the first nose ROI2, the first left cheek ROI3, and the first right cheek ROI4. Frame sequences of the first forehead ROI1, the first nose ROI4, the first left cheek ROI3, and the first right cheek ROI4 are obtained from the visible light facial video.
[0037] Four regions of interest (ROIs) were selected from the thermal infrared facial video: the forehead, left cheek, right cheek, and nose. These ROIs were designated as ROI5 (forehead), ROI6 (nose), ROI7 (left cheek), and ROI8 (right cheek). Frame sequences for each ROI were then obtained from the thermal infrared facial video.
[0038] In one embodiment, step S102 includes:
[0039] A first visible light facial video frame is determined from the visible light facial video, and a first infrared facial video frame is determined from the thermal infrared facial video;
[0040] The first visible light facial video frame and the first infrared facial video frame are registered to obtain the first visible light facial video registration frame and the first infrared facial video registration frame.
[0041] The face detection algorithm is used to detect s facial feature points in the first visible light face video registration frame;
[0042] Based on the s facial feature points, determine m first regions of interest in the first visible light facial video registration frame;
[0043] Based on the m first regions of interest, determine m second regions of interest from the first infrared facial video registration frame;
[0044] The m first regions of interest in subsequent frames of the visible light facial video and the m second regions of interest in subsequent frames of the thermal infrared facial video are tracked respectively to obtain a sequence of m first regions of interest frames and a sequence of m second regions of interest frames.
[0045] As an example, the first visible light facial video frame can be the first visible light video frame of a visible light facial video, and the first infrared facial video frame can be the first thermal infrared video frame of a thermal infrared facial video. The Roberts operator is used to detect the facial edges of the first visible light facial video frame and the first infrared facial video frame, and affine transformation is used to register the facial regions in the two frames to reduce the spatial deviation between them.
[0046] For the first visible light facial video frame, 81 facial feature points were detected using the Dlib facial detection algorithm. The 31st detected feature point was used as the reference point, and the distances to the left and right sides of the 31st feature point were calculated on the same horizontal line. The positions of the two points are respectively determined as the upper left and upper right vertices of the nose region. Then, using the upper left and upper right vertices of the nose region as references, the positions at a distance h below on the same vertical line are respectively determined as the lower left and lower right vertices of the nose region. Thus, the position of the first nose region of interest (ROI2) is determined based on the upper left, upper right, upper left, and upper right vertices of the nose region. Similarly, the position of the first forehead region of interest (ROI1) is determined using the midpoint of the line connecting feature point 71 and feature point 24 as the reference point; the position of the first left cheek region of interest (ROI3) is determined using the midpoint of the line connecting feature point 2 and feature point 32 as the reference point; and the position of the first right cheek region of interest (ROI4) is determined using the midpoint of the line connecting feature point 36 and feature point 17 as the reference point.
[0047] After obtaining the regions of interest (ROIs) in the first visible light facial video frame, linear coordinate mapping is used to map the points in the first visible light facial video frame's first forehead ROI1, first nose ROI2, first left cheek ROI3, and first right cheek ROI4 to the first infrared facial video frame, thus determining the second forehead ROI5, second nose ROI6, second left cheek ROI7, and second right cheek ROI8 in the thermal infrared video. Then, a fully convolutional Siamese network (SiamFC) is used to track the corresponding ROIs in subsequent frames of the visible light facial video, resulting in the first forehead ROI1 frame sequence, the first nose ROI4 frame sequence, the first left cheek ROI3 frame sequence, and the first right cheek ROI4 frame sequence.
[0048] A fully convolutional twin network was used to track the corresponding regions of interest in subsequent frames of the thermal infrared facial video, resulting in the following frame sequences: second forehead region of interest (ROI) 5-frame sequence, second nose region of interest (ROI) 6-frame sequence, second left cheek region of interest (ROI) 7-frame sequence, and second right cheek region of interest (ROI) 8-frame sequence.
[0049] Step S103: Construct a first spatiotemporal map based on m first region of interest frame sequences, and construct a second spatiotemporal map based on m second region of interest frame sequences.
[0050] If the m first region of interest frame sequences are the first forehead region of interest ROI1 frame sequence, the first nose region of interest ROI2 frame sequence, the first left cheek region of interest ROI3 frame sequence, and the first right cheek region of interest ROI4 frame sequence, then the first spatiotemporal map ST1 is determined based on the first forehead region of interest ROI1 frame sequence, the first nose region of interest ROI4 frame sequence, the first left cheek region of interest ROI3 frame sequence, and the first right cheek region of interest ROI4 frame sequence.
[0051] If the m second region of interest frame sequences are the second forehead region of interest ROI 5 frame sequence, the second nose region of interest ROI 6 frame sequence, the second left cheek region of interest ROI 7 frame sequence, and the second right cheek region of interest ROI 8 frame sequence, then the second spatiotemporal map ST2 is determined based on the second forehead region of interest ROI 5 frame sequence, the second nose region of interest ROI 6 frame sequence, the second left cheek region of interest ROI 7 frame sequence, and the second right cheek region of interest ROI 8 frame sequence.
[0052] In one embodiment, step S103 includes:
[0053] Each first region of interest frame sequence is divided into n first sub-regions, and the first denoised temporal signal of each first sub-region in color space is obtained; a first spatiotemporal map is constructed based on multiple first denoised temporal signals.
[0054] Each second region of interest frame sequence is divided into n second sub-regions, and the second denoised temporal signal of each second sub-region in color space is obtained; a second spatiotemporal map is constructed based on multiple second denoised temporal signals.
[0055] As an example, suppose there are T frames of first region of interest (ROI) or T frames of second region of interest (ROI) in a first region of interest (ROI) or second region of interest (ROI). Divide each ROI into n first sub-regions or n second sub-regions. Then, average the pixel values of each first sub-region or each second sub-region. For the m-th ROI, assume P... R (m,i,T) represents the average pixel value of the i-th first or second sub-region in the red channel of the T-th frame. The temporal signal of the i-th sub-region in the RGB color space is determined by the following formula:
[0056] R mi ={P R (m,i,1),…,P R (m,i,t),…,P R (m,i,T)}
[0057] G mi ={P G (m,i,1),…,P G (m,i,t),…,P G (m,i,T)}
[0058] B mi ={P B (m,i,1),…,PB (m,i,t),…,P B (m,i,T)}
[0059] Various denoising operations can be used to obtain the first denoised time domain signal or the second denoised time domain signal of each first sub-region or each second sub-region.
[0060] In one embodiment, acquiring the first denoised temporal signal of each of the first sub-regions in the color space includes:
[0061] The first initial temporal signal of each first sub-region is determined based on the average pixel value of each first sub-region;
[0062] Perform a Fourier transform on each of the first initial time-domain signals to obtain a first frequency-domain signal; select a second frequency-domain signal within a preset frequency range from each of the first frequency-domain signals; perform an inverse Fourier transform on each of the second frequency-domain signals to obtain each of the first denoised time-domain signals;
[0063] The step of obtaining the second denoised temporal signal of each of the second sub-regions in the color space includes:
[0064] The second initial temporal signal of each second sub-region is determined based on the average pixel value of each second sub-region;
[0065] Perform a Fourier transform on each of the second initial time-domain signals to obtain a third frequency-domain signal; select a fourth frequency-domain signal within a preset frequency range from each of the third frequency-domain signals, and perform an inverse Fourier transform on each of the fourth frequency-domain signals to obtain each of the second denoised time-domain signals.
[0066] In this embodiment, in order to eliminate noise in the time-domain signals of each first sub-region or each second sub-region, Fourier transform (FFT) is used to transform the first initial time-domain signal or the second initial time-domain signal R of each first sub-region or each second sub-region. mi G mi B mi The signal is converted to either the first or second frequency domain signal, and signals outside 0.75-3Hz are set to zero. Then, the first or second frequency domain signal is converted back to the time domain using an inverse Fourier transform (IFFT).
[0067] To maximize the use of visual vital signs signals, four denoised ROI frames—the first forehead ROI1 frame sequence, the first nose ROI2 frame sequence, the first left cheek ROI3 frame sequence, and the first right cheek ROI4 frame sequence—were used. mi G mi B miUsing temporal information, a first spatiotemporal map ST1 of size 4n×T×3 is constructed. Four denoised R... mi G mi B mi Temporal information is used to construct a second spatiotemporal graph ST2 with a size of 4n×T×3. The first spatiotemporal graph ST1 and the second spatiotemporal graph ST2 serve as input data for the deep physiological parameter estimation network.
[0068] Step S104: The first spatiotemporal graph and the second spatiotemporal graph are processed by untangling physiological and non-physiological features through a pre-constructed unentangled representation learning network to obtain multiple physiological feature data.
[0069] See Figure 2 The unentangled representation learning network includes a first unentangled encoder E. p1 First noise encoder E n1 The second unentangled encoder E p2 Second noise encoder E n2 First encoder D1, second encoder D2, third encoder D3, fourth encoder D4, and the first unentangled encoder E of the network learning the unentangled representation. p1 First noise encoder E n1 Input the first spatiotemporal graph ST1, and input the second unentangled encoder E. p2 Second noise encoder E n2 The input and the second spatiotemporal graph ST2 are used to calculate the first spatiotemporal graph ST1 and the second spatiotemporal graph ST2 through a deentanglement representation learning network, which can yield multiple physiological feature data.
[0070] In one embodiment, the plurality of physiological characteristic data includes a first physiological characteristic, a second physiological characteristic, a third physiological characteristic, and a fourth physiological characteristic;
[0071] The process of untangling physiological and non-physiological features of the first and second spatiotemporal graphs using a pre-constructed unentangled representation learning network includes:
[0072] The first physiological feature is obtained from the first spatiotemporal map by the first unentangled encoder;
[0073] The first non-physiological feature is obtained from the first spatiotemporal map by the first noise encoder;
[0074] The second physiological feature is obtained from the second spatiotemporal map by a second unentangled encoder;
[0075] The second non-physiological feature is obtained from the second spatiotemporal map by a second noise encoder;
[0076] A first pseudo-spatiotemporal graph is constructed by the first encoder based on the first physiological feature and the second non-physiological feature;
[0077] A second pseudo-spatiotemporal map is constructed by the second encoder based on the first non-physiological feature and the second physiological feature;
[0078] The third physiological feature is obtained from the first pseudo-spatiotemporal map by the first unentangled encoder;
[0079] The fourth physiological feature is obtained from the second pseudo-spatiotemporal graph by a second unentangled encoder.
[0080] In this embodiment, in order to obtain more accurate and richer physiological features, a first pseudo-spatiotemporal map and a second pseudo-spatiotemporal map are generated using a first encoder or a second encoder based on physiological features and non-physiological features from different spatiotemporal maps.
[0081] Please see again Figure 2 Through the first unentangled encoder E p1 The first physiological feature f is obtained from the first spatiotemporal map ST1. p1 ; via the first noise encoder E n1 The first non-physiological feature f is obtained from the first spatiotemporal map ST1. n1 ; through the second unentangled encoder E p2 The second physiological characteristic f is obtained from the second spatiotemporal diagram ST2. p2 ; via the second noise encoder E n2 The second non-physiological feature f is obtained from the second spatiotemporal diagram ST2. n2 ; based on the first physiological feature f by the first encoder D1 p1 and the second non-physiological characteristic f n2 Constructing the first pseudo-spacetime graph ST re_ ; based on the first non-physiological feature f by the second encoder D2 n1 and the second physiological characteristic f p2 Constructing the second pseudo-spacetime graph ST re_ ; through the first unentangled encoder E p1 From the first pseudo-spacetime map ST re_ Obtain the third physiological characteristic f re_ ; through the second unentangled encoder E p2 From the second pseudo-spacetime map ST re_ Obtain the fourth physiological characteristic f re_ .
[0082] In one embodiment, the non-contact multi-parameter monitoring method for physical and mental health analysis further includes:
[0083] A first reconstructed spatiotemporal map is constructed by a third encoder based on the first physiological feature and the first non-physiological feature;
[0084] The second reconstructed spatiotemporal map is constructed by the fourth encoder based on the second physiological feature and the second non-physiological feature.
[0085] Please refer to it again. Figure 2 Based on the first physiological feature f, the third encoder D3... p1 and the first non-physiological feature f n1 Construct the first reconstructed spatiotemporal map ST3. Then, based on the second physiological characteristic f, use the fourth encoder D4. p2 and the second non-physiological characteristic f n2 Construct the second reconstruction spatiotemporal graph ST4.
[0086] Thus, using the first physiological feature f from the same video p1 and the first non-physiological characteristic f n1 The first reconstructed spatiotemporal graph TS3 was rebuilt using a second physiological feature f from the same video. p2 Second non-physiological characteristic f n2 The second reconstruction spatiotemporal graph TS4 is reconstructed to ensure that the decoder can effectively reconstruct the spatiotemporal graph.
[0087] In one embodiment, the non-contact multi-parameter monitoring method for physical and mental health analysis further includes:
[0088] The third non-physiological feature is obtained from the first pseudo-spatiotemporal map by the first noise encoder;
[0089] via the second noise encoder E n2 The fourth non-physiological feature is obtained from the second pseudo-spatiotemporal map.
[0090] Please see again Figure 2 Through the first noise encoder E n1 From the first pseudo-spacetime map ST re_ Obtain the third non-physiological feature f re_ ; via the second noise encoder E n2 From the second pseudo-spacetime map ST re_ Obtain the fourth non-physiological feature f re_ .
[0091] In summary, utilizing the first physiological characteristic f p1 and second non-physiological characteristics f n2 Generate the first pseudo-spacetime graph ST re_Utilizing the second physiological characteristic f p2 and the first non-physiological characteristic f n1 Generate the second pseudo-spacetime graph ST re_ For the first pseudo-spacetime graph ST generated re_ The second pseudo-spacetime graph ST re_ Using the first unentangled encoder E p1 First noise encoder E n1 The second unentangled encoder E p2 Second noise encoder E n2 Extract the third physiological feature f re_1 Fourth physiological characteristic f re_2 Third non-physiological characteristic f re_ Fourth non-physiological characteristic f re_ This achieves the untangling of physiological and non-physiological features in the pseudo-spatiotemporal graph. Finally, the first physiological feature f is... p1 Second physiological characteristics f p2 Third physiological characteristic f re_ Fourth physiological characteristic f re_2 All of these will be fed into a subsequent deep physiological parameter estimation network for the extraction of heart rate and respiratory rate.
[0092] Step S105: Calculate multiple physiological parameters of the user by using a pre-built deep physiological parameter estimation network to calculate multiple physiological feature data.
[0093] In this embodiment, the deep physiological parameter estimation network mainly uses convolutional neural networks such as ResNet, VGGNet, MobileNet, and DenseNet. The network mainly consists of two branches, and the input of each branch is the first physiological feature f after unwinding. p1 Second physiological characteristics f p2 Third physiological characteristic f re_1 Fourth physiological characteristic f re_2 Features are shared in each branch to monitor respiratory rate and heart rate.
[0094] See Figure 3 The deep physiological parameter estimation network includes a shared computation submodule 301, a first channel attention module 302, a second channel attention module 303, a first computation submodule 304, and a second computation submodule 305, which inputs the first physiological feature f p1 Second physiological characteristics f p2 Third physiological characteristic f re_ and the fourth physiological characteristic f re_2By inputting a deep physiological parameter estimation network for calculation, the user's respiratory rate and heart rate can be obtained. The shared computation submodule 301 includes four convolutional layers 3011 and four first residual blocks 3012. Each convolutional layer 3011 is connected to a first residual block 3012. The four convolutional layers 3011 respectively receive the first physiological feature f. p1 Second physiological characteristics f p2 Third physiological characteristic f re_ Fourth physiological characteristic f re_2 The convolutional layer 3011 and the corresponding first residual block 3012 collaboratively calculate the feature maps of each physiological feature. The first calculation submodule 304 includes two second residual blocks 3042 and one first fully connected layer 3041; the second calculation submodule 305 includes two third residual blocks 3052 and one second fully connected layer 3051.
[0095] Since similar features from two channels can mutually reinforce each other, a channel attention mechanism is employed to perform weighted fusion of features across different channels based on their correlation. Two parallel channel attention modules, 302 and 303, respectively capture the correlation between features in the channel dimension and spatial dimension (a dimension corresponding to the spatiotemporal information of the spatiotemporal graph) of the feature map output by the first residual block 3012 using the channel attention mechanism. After weighted fusion using the attention mechanism, the weighted fused features are fed into different channels, where the first computation submodule 304 and the second computation submodule 305 respectively monitor respiratory rate and heart rate.
[0096] In one embodiment, the deep physiological parameter estimation network includes a shared computation submodule, a first extraction network branch, and a second extraction network branch. The calculation of multiple physiological feature data using the pre-constructed deep physiological parameter estimation network includes:
[0097] The shared computing submodule calculates the feature maps of each of the physiological characteristic data;
[0098] The first channel attention module of the first extraction network branch performs weighted fusion on multiple feature maps to obtain a first fused feature map;
[0099] The user's respiratory rate is obtained by calculating the first fused feature map through the first calculation module of the first extraction network branch;
[0100] The second channel attention module of the second extraction network branch performs weighted fusion on the multiple feature maps to obtain a second fused feature map;
[0101] The user's heart rate is obtained by calculating the second fused feature map through the second calculation module of the second extraction network branch.
[0102] Please see Figure 4 , Figure 4 The image shown is one of the schematic diagrams of respiratory waveforms. Figure 4 You can see the trend of changes in a user's breathing. Please refer to [link / reference]. Figure 5 , Figure 5 The image shown is one of the schematic diagrams of heart rate waveforms. Figure 5 It can reveal the trend of changes in a user's heart rate.
[0103] Thus, this embodiment provides a non-contact multi-parameter monitoring scheme for physical and mental health analysis. It constructs a deentangled representation learning network based on physiological and non-physiological feature deentanglement strategies, studies a deep physiological parameter estimation network based on a multi-task network, and realizes synchronous monitoring of heart rate and respiratory rate.
[0104] This embodiment provides a non-contact multi-parameter monitoring method for physical and mental health analysis, which collects visible light facial videos and thermal infrared facial videos of users. Based on the visible light and thermal infrared facial videos, m first regions of interest (ROI) frame sequences and m second ROI frame sequences are obtained. A first spatiotemporal map is constructed based on the m first ROI frame sequences, and a second spatiotemporal map is constructed based on the m second ROI frame sequences. A pre-constructed deentanglement representation learning network is used to deentangle the first and second spatiotemporal maps using physiological and non-physiological features, resulting in multiple physiological feature data. A pre-constructed deep physiological parameter estimation network is used to calculate the multiple physiological feature data to obtain multiple physiological parameters of the user. Thus, by using a deentanglement representation learning network and a deep physiological parameter estimation network for non-contact calculation of multiple physiological parameters, the accuracy and robustness of physiological parameter detection are improved in complex scenarios such as significant changes in lighting conditions and obvious head movements, providing users with a highly comfortable, highly integrated, and highly robust multi-physiological parameter monitoring service.
[0105] Example 2
[0106] Furthermore, embodiments of this application provide a non-contact multi-parameter monitoring system for physical and mental health analysis.
[0107] Specifically, such as Figure 6 As shown, the non-contact multi-parameter monitoring system 600 for physical and mental health analysis includes:
[0108] The video acquisition module 601 is used to acquire visible light facial video and thermal infrared facial video of the user;
[0109] The acquisition module 602 is used to acquire m first region of interest frame sequences and m second region of interest frame sequences based on the visible light facial video and the thermal infrared facial video;
[0110] Construction module 603 is used to construct a first spatiotemporal map based on m first region of interest frame sequences and to construct a second spatiotemporal map based on m second region of interest frame sequences;
[0111] The processing module 604 is used to perform physiological and non-physiological feature unentanglement processing on the first spatiotemporal graph and the second spatiotemporal graph through a pre-constructed unentangled representation learning network to obtain multiple physiological feature data.
[0112] The calculation module 605 is used to calculate multiple physiological parameters of the user by using a pre-built deep physiological parameter estimation network to calculate multiple physiological feature data.
[0113] In one embodiment, the deep physiological parameter estimation network includes a first extraction network branch and a second extraction network branch, and the computing module 605 includes:
[0114] The first calculation submodule is used to obtain the first feature map of each of the physiological feature data through the first residual block of the first extraction network branch;
[0115] The first feature map is obtained by weighted fusing multiple first feature maps through the first channel attention module of the first extraction network branch;
[0116] The user's respiratory rate is obtained by calculating the first fused feature map through the first calculation module of the first extraction network branch;
[0117] The second calculation submodule is used to obtain the second feature map of each of the physiological feature data through the second residual block of the second extraction network branch;
[0118] The second channel attention module of the second extraction network branch performs weighted fusion on multiple second feature maps to obtain a second fused feature map;
[0119] The user's heart rate is obtained by calculating the second fused feature map through the second calculation module of the second extraction network branch.
[0120] In one embodiment, the acquisition module 602 is further configured to determine a first visible light facial video frame from the visible light facial video and a first infrared facial video frame from the thermal infrared facial video.
[0121] The first visible light facial video frame and the first infrared facial video frame are registered to obtain the first visible light facial video registration frame and the first infrared facial video registration frame.
[0122] The face detection algorithm is used to detect s facial feature points in the first visible light face video registration frame;
[0123] Based on the s facial feature points, determine m first regions of interest in the first visible light facial video registration frame;
[0124] Based on the m first regions of interest, determine m second regions of interest from the first infrared facial video registration frame;
[0125] The m first regions of interest in subsequent frames of the visible light facial video and the m second regions of interest in subsequent frames of the thermal infrared facial video are tracked respectively to obtain a sequence of m first regions of interest frames and a sequence of m second regions of interest frames.
[0126] In one embodiment, the construction module 603 is further configured to divide each of the first regions of interest in each frame sequence of the first region of interest into n first sub-regions, obtain a first denoised temporal signal of each first sub-region in the color space, and construct a first spatiotemporal map based on the multiple first denoised temporal signals;
[0127] Each second region of interest frame sequence is divided into n second sub-regions, and the second denoised temporal signal of each second sub-region in color space is obtained; a second spatiotemporal map is constructed based on multiple second denoised temporal signals.
[0128] In one embodiment, the construction module 603 is further configured to determine a first initial temporal signal for each of the first sub-regions based on the average pixel value of each of the first sub-regions;
[0129] Perform a Fourier transform on each of the first initial time-domain signals to obtain a first frequency-domain signal; select a second frequency-domain signal within a preset frequency range from each of the first frequency-domain signals; perform an inverse Fourier transform on each of the second frequency-domain signals to obtain each of the first denoised time-domain signals;
[0130] The second initial temporal signal of each second sub-region is determined based on the average pixel value of each second sub-region;
[0131] Perform a Fourier transform on each of the second initial time-domain signals to obtain a third frequency-domain signal; select a fourth frequency-domain signal within a preset frequency range from each of the third frequency-domain signals, and perform an inverse Fourier transform on each of the fourth frequency-domain signals to obtain each of the second denoised time-domain signals.
[0132] In one embodiment, the plurality of physiological characteristic data includes a first physiological characteristic, a second physiological characteristic, a third physiological characteristic, and a fourth physiological characteristic;
[0133] Processing module 604 is used to obtain the first physiological feature from the first spatiotemporal map through the first unentangled encoder;
[0134] The first non-physiological feature is obtained from the first spatiotemporal map by the first noise encoder;
[0135] Through the second deentanglement encoder E p2 The second physiological characteristic is obtained from the second spatiotemporal diagram ST2;
[0136] The second non-physiological feature is obtained from the second spatiotemporal map by a second noise encoder;
[0137] A first pseudo-spatiotemporal graph is constructed by the first encoder based on the first physiological feature and the second non-physiological feature;
[0138] A second pseudo-spatiotemporal map is constructed by the second encoder based on the first non-physiological feature and the second physiological feature;
[0139] The third physiological feature is obtained from the first pseudo-spatiotemporal map by the first unentangled encoder;
[0140] The fourth physiological feature is obtained from the second pseudo-spatiotemporal graph by a second unentangled encoder.
[0141] In one embodiment, the processing module 604 is further configured to construct a first reconstructed spatiotemporal map based on the first physiological feature and the first non-physiological feature using a third encoder;
[0142] The second reconstructed spatiotemporal map is constructed by the fourth encoder based on the second physiological feature and the second non-physiological feature.
[0143] In one embodiment, the processing module 604 is further configured to obtain a third non-physiological feature from the first pseudo-spatiotemporal map through the first noise encoder;
[0144] The fourth non-physiological feature is obtained from the second pseudo-spatiotemporal map by the second noise encoder.
[0145] The non-contact multi-parameter monitoring system 600 for physical and mental health analysis provided in this embodiment can implement the non-contact multi-parameter monitoring method for physical and mental health analysis provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0146] It should be noted that the non-contact multi-parameter monitoring system for physical and mental health analysis may also include modules such as a controller, interface display module, storage module, transmission module, and power supply module. The controller is electrically connected to all other modules and is used to control the operation of the entire non-contact multi-parameter monitoring system for physical and mental health analysis. A switch button can be electrically connected to the controller, and the controller can be started using the switch button to send commands to control the non-contact multi-parameter monitoring system for physical and mental health analysis to start working. The video acquisition module consists of a binocular vision module and a thermal infrared module, which acquires visible light facial video and thermal infrared facial video of the user in real time. The visible light facial video, thermal infrared facial video, m first region of interest frame sequences and m second region of interest frame sequences, and multiple physiological parameters can be transmitted to the storage module and the interface display module. After receiving the visible light and thermal infrared facial video, the acquisition module 602, the construction module 603, the processing module 604, and the calculation module 605 process and calculate the received data to obtain the user's physiological data such as heart rate and respiratory rate. The calculation process can be referred to the previous description and will not be repeated here.
[0147] Multiple physiological parameters are transmitted to the storage module and the interface display module. The interface display module presents the system functions and physiological parameter calculation results to the user through a front-end interface. For example, it can display visible light images, thermal infrared images, and real-time respiratory wave remote photoplethysmography (rPPG) signals. Exemplarily, the left half of the interface displays the user's visible light and thermal infrared facial images, while the right half displays the real-time calculation results of physiological parameters such as heart rate and respiratory rate, as well as respiratory wave and pulse wave waveforms. The storage module temporarily stores the visible light and thermal infrared facial videos acquired by the video acquisition module, as well as the calculated heart rate, respiratory rate, and other physiological parameters. It also periodically transmits the data to the system server for long-term storage via the transmission module. The power module is electrically connected to an external power source via a power adapter to power the non-contact multi-parameter monitoring system used for physical and mental health analysis.
[0148] The computer-readable storage medium provided in this embodiment can implement the non-contact multi-parameter monitoring method for physical and mental health analysis provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0149] This embodiment provides a non-contact multi-parameter monitoring system for physical and mental health analysis. It collects visible light and thermal infrared facial videos of the user; obtains m first regions of interest (ROI) frame sequences and m second ROI frame sequences based on the visible light and thermal infrared facial videos; constructs a first spatiotemporal map based on the m first ROI frame sequences and a second spatiotemporal map based on the m second ROI frame sequences; performs physiological and non-physiological feature decoupling processing on the first and second spatiotemporal maps using a pre-constructed deentanglement representation learning network to obtain multiple physiological feature data; and calculates multiple physiological parameters of the user using a pre-constructed deep physiological parameter estimation network. Thus, by employing a deentanglement representation learning network and a deep physiological parameter estimation network for non-contact calculation of multiple physiological parameters, the system improves the accuracy and robustness of physiological parameter detection in complex scenarios such as significant changes in lighting conditions and obvious head movements, providing users with a highly comfortable, highly integrated, and highly robust multi-physiological parameter monitoring service.
[0150] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0151] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0152] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A non-contact multi-parameter monitoring method for physical and mental health analysis, characterized in that, The method includes: Collect visible light facial video and thermal infrared facial video of the user; Based on the visible light facial video and the thermal infrared facial video, obtain m first region of interest frame sequences and m second region of interest frame sequences; A first spatiotemporal graph is constructed based on m first region of interest frame sequences, and a second spatiotemporal graph is constructed based on m second region of interest frame sequences; The first and second spatiotemporal graphs are de-entangled using a pre-constructed de-entanglement representation learning network to obtain multiple physiological feature data by de-entanglement of physiological and non-physiological features. Multiple physiological parameters of the user are obtained by calculating multiple physiological feature data through a pre-constructed deep physiological parameter estimation network; The method involves obtaining m first impressions based on the visible light facial video and the thermal infrared facial video. The region of interest frame sequence and m second region of interest frame sequences include: A first visible light facial video frame is determined from the visible light facial video, and a first infrared facial video frame is determined from the thermal infrared facial video; The first visible light facial video frame and the first infrared facial video frame are registered to obtain the first visible light facial video registration frame and the first infrared facial video registration frame. The face detection algorithm is used to detect s facial feature points in the first visible light face video registration frame; Based on the s facial feature points, determine m first regions of interest in the first visible light facial video registration frame; Based on the m first regions of interest, determine m second regions of interest from the first infrared facial video registration frame; The m first regions of interest in subsequent frames of the visible light facial video and the m second regions of interest in subsequent frames of the thermal infrared facial video are tracked respectively to obtain a sequence of m first regions of interest frames and a sequence of m second regions of interest frames. The deep physiological parameter estimation network includes a shared computation submodule, a first extraction network branch, and a second extraction network branch. The calculation of multiple physiological feature data using the pre-constructed deep physiological parameter estimation network includes: The shared computing submodule calculates the feature maps of each of the physiological characteristic data; The first channel attention module of the first extraction network branch performs weighted fusion on multiple feature maps to obtain a first fused feature map; The user's respiratory rate is obtained by calculating the first fused feature map through the first calculation module of the first extraction network branch; The second channel attention module of the second extraction network branch performs weighted fusion on the multiple feature maps to obtain a second fused feature map; The user's heart rate is obtained by calculating the second fused feature map through the second calculation module of the second extraction network branch.
2. The method according to claim 1, characterized in that, The step of constructing a first spatiotemporal graph based on m first region of interest frame sequences and constructing a second spatiotemporal graph based on m second region of interest frame sequences includes: Each first region of interest frame sequence is divided into n first sub-regions, and the first denoised temporal signal of each first sub-region in color space is obtained; a first spatiotemporal map is constructed based on multiple first denoised temporal signals. Each second region of interest frame sequence is divided into n second sub-regions, and the second denoised temporal signal of each second sub-region in color space is obtained; a second spatiotemporal map is constructed based on multiple second denoised temporal signals.
3. The method according to claim 2, characterized in that, The step of obtaining the first denoised temporal signal of each of the first sub-regions in the color space includes: The first initial temporal signal of each first sub-region is determined based on the average pixel value of each first sub-region; Perform a Fourier transform on each of the first initial time-domain signals to obtain a first frequency-domain signal; select a second frequency-domain signal within a preset frequency range from each of the first frequency-domain signals; perform an inverse Fourier transform on each of the second frequency-domain signals to obtain each of the first denoised time-domain signals; The step of obtaining the second denoised temporal signal of each of the second sub-regions in the color space includes: The second initial temporal signal of each second sub-region is determined based on the average pixel value of each second sub-region; Perform a Fourier transform on each of the second initial time-domain signals to obtain a third frequency-domain signal; select a fourth frequency-domain signal within a preset frequency range from each of the third frequency-domain signals, and perform an inverse Fourier transform on each of the fourth frequency-domain signals to obtain each of the second denoised time-domain signals.
4. The method according to claim 1, characterized in that, The multiple physiological characteristic data include a first physiological characteristic, a second physiological characteristic, a third physiological characteristic, and a fourth physiological characteristic; The process of untangling physiological and non-physiological features of the first and second spatiotemporal graphs using a pre-constructed unentangled representation learning network includes: The first physiological feature is obtained from the first spatiotemporal map by the first unentangled encoder; The first non-physiological feature is obtained from the first spatiotemporal map by the first noise encoder; The second physiological feature is obtained from the second spatiotemporal map by a second unentangled encoder; The second non-physiological feature is obtained from the second spatiotemporal map by a second noise encoder; A first pseudo-spatiotemporal graph is constructed by the first encoder based on the first physiological feature and the second non-physiological feature; A second pseudo-spatiotemporal map is constructed by the second encoder based on the first non-physiological feature and the second physiological feature; The third physiological feature is obtained from the first pseudo-spatiotemporal map by the first unentangled encoder; The fourth physiological feature is obtained from the second pseudo-spatiotemporal graph by a second unentangled encoder.
5. The method according to claim 4, characterized in that, The method further includes: A first reconstructed spatiotemporal map is constructed by a third encoder based on the first physiological feature and the first non-physiological feature; The second reconstructed spatiotemporal map is constructed by the fourth encoder based on the second physiological feature and the second non-physiological feature.
6. The method according to claim 4, characterized in that, The method further includes: The third non-physiological feature is obtained from the first pseudo-spatiotemporal map by the first noise encoder; The fourth non-physiological feature is obtained from the second pseudo-spatiotemporal map by the second noise encoder.
7. A non-contact multi-parameter monitoring system for physical and mental health analysis, characterized in that, The system includes: The video capture module is used to capture the user's visible light facial video and thermal infrared facial video; The acquisition module is used to acquire m first region of interest frame sequences and m second region of interest frame sequences based on the visible light facial video and the thermal infrared facial video; The construction module is used to construct a first spatiotemporal graph based on m first region of interest frame sequences and to construct a second spatiotemporal graph based on m second region of interest frame sequences. The processing module is used to perform physiological and non-physiological feature unentanglement processing on the first spatiotemporal graph and the second spatiotemporal graph through a pre-constructed unentangled representation learning network to obtain multiple physiological feature data. The calculation module is used to calculate multiple physiological parameters of the user by using a pre-built deep physiological parameter estimation network to calculate multiple physiological feature data. The acquisition module is further configured to determine a first visible light facial video frame from the visible light facial video and a first infrared facial video frame from the thermal infrared facial video. The first visible light facial video frame and the first infrared facial video frame are registered to obtain the first visible light facial video registration frame and the first infrared facial video registration frame. The face detection algorithm is used to detect s facial feature points in the first visible light face video registration frame; Based on the s facial feature points, determine m first regions of interest in the first visible light facial video registration frame; Based on the m first regions of interest, determine m second regions of interest from the first infrared facial video registration frame; The m first regions of interest in subsequent frames of the visible light facial video and the m second regions of interest in subsequent frames of the thermal infrared facial video are tracked respectively to obtain a sequence of m first regions of interest frames and a sequence of m second regions of interest frames. The deep physiological parameter estimation network includes a first extraction network branch and a second extraction network branch, and the computation module includes: The first calculation submodule is used to obtain the first feature map of each of the physiological feature data through the first residual block of the first extraction network branch; The first feature map is obtained by weighted fusing multiple first feature maps through the first channel attention module of the first extraction network branch; The user's respiratory rate is obtained by calculating the first fused feature map through the first calculation module of the first extraction network branch; The second calculation submodule is used to obtain the second feature map of each of the physiological feature data through the second residual block of the second extraction network branch; The second channel attention module of the second extraction network branch performs weighted fusion on multiple second feature maps to obtain a second fused feature map; The user's heart rate is obtained by calculating the second fused feature map through the second calculation module of the second extraction network branch.