Health detection method and system based on face image
By performing face detection and initial state assessment on facial video streams, and collecting recovery period image sequences after identifying suspected agitated states, the time series of heart rate and skin perfusion index are extracted for modeling. This solves the problem of health detection being easily interfered with in existing technologies and achieves more accurate health assessment.
Patent Information
- Application Number
- CN202512028176.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-02-13
AI Technical Summary
Existing health detection methods based on facial images are easily affected by the user's transient physiological state, making it difficult to effectively distinguish between transient physiological disturbances and chronic pathological features, resulting in insufficient reliability and accuracy of the assessment results.
By performing face detection and initial state assessment on facial video streams, a health mode is entered after identifying suspected agitation. Facial image sequences during the recovery period are collected, and time series of heart rate and skin perfusion index are extracted. Dynamic modeling and feature parameterization of the recovery curve are performed, and finally, a health status assessment is conducted.
It significantly improves the accuracy and reliability of health assessments, effectively distinguishing between transient physiological disturbances and potential chronic pathological features, and ensuring the timeliness and accuracy of detection.
Smart Images

Figure CN121528547A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of health detection, and more specifically, to a health detection method and system based on facial images. Background Technology
[0002] With the rapid development of information technology and the increasing public awareness of health, preventative and routine health monitoring is becoming an important part of modern life. Traditional health monitoring methods, such as blood pressure monitors and electrocardiographs, while providing accurate physiological data, usually require contact-based specialized equipment. Their operational limitations and invasiveness restrict their application in daily, continuous monitoring scenarios. Therefore, developing a convenient and non-invasive health monitoring technology is of significant practical importance. In recent years, remote photoplethysmography (rPPG) technology based on computer vision has emerged. This technology can non-contactly extract key physiological indicators such as heart rate and blood oxygen saturation by analyzing subtle changes in skin color caused by heartbeats in facial video streams captured by ordinary cameras. Due to its low cost and high convenience, it shows broad application prospects in smart homes, telemedicine, and daily health management.
[0003] However, most existing health detection solutions based on facial images employ a "snapshot" measurement model, which involves capturing a short video clip and immediately calculating and evaluating physiological parameters. A key limitation of this approach is that its measurement results are highly susceptible to interference from the user's current transient physiological state. The human physiological system is highly dynamic; when an individual experiences physical exercise, emotional fluctuations, or external environmental stimuli, indicators such as heart rate and skin blood perfusion can undergo temporary and drastic changes. These deviations in physiological parameters caused by immediate stress responses may morphologically be confused with the long-term signs of certain chronic diseases (such as hypertension and cardiovascular dysfunction). Therefore, traditional static single-point measurement methods struggle to effectively distinguish between transient physiological disturbances and stable pathological features, easily misinterpreting temporary physiological peaks as potential health risks. This significantly reduces the reliability and accuracy of the assessment results, thus limiting their application in serious health assessment scenarios.
[0004] Therefore, there is an urgent need for an optimized health detection method and system based on facial images. Summary of the Invention
[0005] This application is made in order to solve the above-mentioned technical problems.
[0006] According to one aspect of this application, a health detection method based on facial images is provided, comprising: Face detection and initial state determination are performed on the acquired target face video stream to obtain an aligned facial sequence and initial state flags; In response to the initial state flag indicating suspected agitation, the system enters health mode and acquires facial image sequences during the recovery period; Time series extraction of multidimensional physiological signals during the recovery period was performed on facial image sequences during the recovery period to obtain heart rate time series and skin perfusion index time series. Recovery curve dynamics modeling and feature parameterization were performed on heart rate time series and skin perfusion index time series to obtain recovery feature vectors; A health status assessment is performed based on the recovery feature vector to obtain a health assessment report.
[0007] According to another aspect of this application, a health detection system based on facial images is provided, comprising: The face detection and initial state assessment module is used to perform face detection and initial state assessment on the acquired target face video stream to obtain an aligned face sequence and initial state flags. The health trigger and acquisition module is used to respond to the initial state flag being suspected agitation, when the system enters health mode and acquires facial image sequences during the recovery period; The physiological signal time series extraction module is used to extract multidimensional physiological signal time series from facial image sequences during the recovery period to obtain heart rate time series and skin perfusion index time series. The modeling and feature parameterization module is used to perform recovery curve dynamics modeling and feature parameterization on heart rate time series and skin perfusion index time series to obtain recovery feature vectors; The health status assessment module is used to assess health status based on recovery feature vectors to obtain a health assessment report.
[0008] Compared with existing technologies, this application provides a health detection method and system based on facial images. First, it performs preliminary state recognition on real-time facial video streams. When it determines that the user is in a state of excitement or post-exercise stress, it triggers recovery period monitoring. Subsequently, the system simultaneously extracts time series of multidimensional physiological signals such as heart rate and skin perfusion index from the facial image sequence during the recovery period and performs dynamic modeling of its recovery trajectory. By parameterizing the morphology, rate, and dynamic coupling relationship between multiple signals of the recovery curve, a recovery feature vector that profoundly reflects an individual's cardiovascular autonomic nervous system regulation function is obtained. Finally, a comprehensive health assessment is performed based on this vector. This effectively distinguishes between transient physiological disturbances and potential chronic pathological features, significantly improving the accuracy and reliability of health assessment. Attached Figure Description
[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 This is a flowchart of a health detection method based on face images according to an embodiment of this application.
[0011] Figure 2 This is a data flow diagram of a health detection method based on face images according to an embodiment of this application.
[0012] Figure 3 This is a flowchart of sub-step S1 of the health detection method based on face images according to an embodiment of this application.
[0013] Figure 4 This is a flowchart of sub-step S13 of the health detection method based on face images according to an embodiment of this application.
[0014] Figure 5 This is a flowchart of sub-step S3 of the health detection method based on face images according to an embodiment of this application.
[0015] Figure 6 This is a flowchart of sub-step S32 of the health detection method based on face images according to an embodiment of this application.
[0016] Figure 7 This is a flowchart of sub-step S4 of the health detection method based on face images according to an embodiment of this application.
[0017] Figure 8 This is a block diagram of a health detection system based on face images according to an embodiment of this application. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] To address the problems mentioned above in the background technology, this application proposes a health detection method based on face images. Figure 1 This is a flowchart of a health detection method based on face images according to an embodiment of this application. Figure 2 This is a data flow diagram of a health detection method based on face images according to an embodiment of this application. For example... Figure 1 and Figure 2 As shown, the health detection method based on face images includes the following steps: S1, performing face detection and initial state judgment on the acquired target face video stream to obtain an aligned facial sequence and an initial state marker; S2, in response to the initial state marker indicating suspected excitation, the system enters a health mode and acquires a recovery period facial image sequence; S3, performing recovery period multidimensional physiological signal time-series extraction on the recovery period facial image sequence to obtain a heart rate time series and a skin perfusion index time series; S4, performing recovery curve dynamic modeling and feature parameterization on the heart rate time series and skin perfusion index time series to obtain a recovery feature vector; S5, performing a health status assessment based on the recovery feature vector to obtain a health assessment report.
[0020] In the aforementioned health detection method based on facial images, step S1 involves performing face detection and initial state assessment on the acquired target face video stream to obtain an aligned facial sequence and initial state markers. It should be understood that the acquired target face video stream contains interference such as head movement, posture changes, and lighting fluctuations, and the user's current physiological state cannot be directly determined, resulting in low accuracy of subsequent physiological signal extraction and difficulty in determining whether recovery period monitoring needs to be initiated. Therefore, this application performs face detection on the target face video stream to locate the face region and performs initial state assessment to identify the user's physiological state, thereby obtaining an aligned facial sequence and initial state markers. This eliminates the impact of facial position changes on signal extraction and provides a basis for determining whether to enter health mode, ensuring the accuracy of subsequent health assessments.
[0021] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S1 of the health detection method based on face images according to an embodiment of this application. Figure 3 As shown, step S1 includes: S11, registering and normalizing the temporal face region of the target face video stream to obtain an aligned face sequence and a stable keypoint sequence; S12, extracting the original photoelectric signal of the region of interest from the aligned face sequence based on the stable keypoint sequence to obtain the average RGB temporal signal of the region of interest; S13, performing state discrimination based on short time window signal features on the average RGB temporal signal of the region of interest to obtain the initial state flag.
[0022] Specifically, step S11 involves temporally registering and normalizing the face regions of the target face video stream to obtain aligned facial sequences and stable keypoint sequences. It should be understood that, due to the translation, rotation, and scaling of the face in the target face video stream over time, the position and scale of the face regions in different frames are inconsistent, and facial keypoints are easily affected by motion interference, causing jitter and affecting the accuracy of subsequent region of interest localization. Therefore, this application performs temporal registration of the face regions of the target face video stream to unify the spatial position of the face regions in each frame, and simultaneously performs normalization processing to eliminate scale differences, thereby obtaining aligned facial sequences and stable keypoint sequences. This ensures that physiological signals are extracted from the same spatial location subsequently, avoiding the impact of motion artifacts on signal quality and laying the foundation for high-quality original photoelectric signal extraction.
[0023] Specifically, in one possible embodiment, step S11 is implemented as follows: First, for each frame of the target face video stream, MTCNN is used to detect the face region and determine the face bounding box of each frame. Then, based on the dlib facial keypoint detection model, the coordinates of 68 key feature points of the face in each frame are extracted. Next, the key points of the face in the first frame are selected as a reference template, and the affine transformation matrix between the key points of subsequent frames and the reference template is calculated. The face region of each frame is then subjected to affine transformation using this matrix to achieve temporal registration of the face region. After that, the registered face region is scaled and normalized to adjust the face region of all frames to a preset size to obtain an aligned face sequence. Finally, the Kalman filter algorithm is used to smooth the temporal keypoint coordinates, filtering out keypoint jitter noise to obtain a stable keypoint sequence.
[0024] Specifically, in step S12, the original photoelectric signal of the region of interest (ROI) is extracted from the aligned facial sequence based on the stable keypoint sequence to obtain the average RGB time-series signal of the ROI. It should be understood that, since the physiological signal intensity varies in different regions of the aligned facial sequence, and non-ROI regions (such as hair and background) do not contain effective physiological information, directly extracting the entire face signal would introduce a large amount of noise, reducing the accuracy of subsequent physiological parameter calculations. Therefore, this application further locates the ROI rich in physiological signals based on the stable keypoint sequence, extracts the original photoelectric signal within this region, and performs spatial averaging to obtain the average RGB time-series signal of the ROI. This allows focusing on the effective physiological signal region, reducing noise interference, and obtaining an RGB time-series signal with a high signal-to-noise ratio, providing a reliable data source for subsequent extraction of physiological parameters such as heart rate and skin perfusion index.
[0025] Specifically, in one possible embodiment, step S12 is implemented as follows: First, based on the stable keypoint sequence, the region of interest (ROI) is defined: the forehead ROI is determined using the forehead keypoints (numbers 19-24) as boundaries. The bilateral cheek ROIs are determined using the cheek keypoints (numbers 2-4, 32, 49 and 14-16, 35, 53) as boundaries. Then, each frame of the aligned facial sequence is traversed, and pixel data of the forehead and cheek ROIs are extracted from each frame according to the above definitions. Next, the average pixel intensity of the R, G, and B channels is calculated for the pixel data of each ROI. Finally, the RGB average values of the ROIs in each frame are arranged in chronological order to form the average RGB temporal signal of the ROI.
[0026] Specifically, in step S13, the average RGB time-series signal of the region of interest is subjected to state discrimination based on short-time window signal features to obtain the initial state flag. It should be understood that while the average RGB time-series signal of the region of interest contains dynamic changes in heart rate and skin color intensity, long-time window analysis tends to lag behind the user's immediate physiological state and cannot quickly capture stress responses. Therefore, this application divides the average RGB time-series signal of the region of interest into short-time windows, extracts features within the windows, and performs state discrimination to determine the initial state flag. This allows for real-time identification of whether the user is in a suspected agitated state, providing immediate evidence for whether to initiate recovery period monitoring, avoiding health assessment bias caused by delayed state judgment, and ensuring the timeliness and accuracy of the detection.
[0027] In particular, in one specific embodiment, Figure 4 This is a flowchart of sub-step S13 of the health detection method based on face images according to an embodiment of this application. Figure 4 As shown, step S13 includes: S131, extracting heart rate features from the average RGB time-series signal of the region of interest to obtain a temporary heart rate; S132, extracting skin color intensity features from the average RGB time-series signal of the region of interest to obtain a temporary color intensity; S133, inputting the temporary heart rate and temporary color intensity into the state decision engine to obtain the initial state flag.
[0028] More specifically, step S131 involves extracting heart rate features from the average RGB time-series signal of the region of interest to obtain a temporary heart rate. It should be understood that, because skin blood flow pulsation causes periodic fluctuations in the RGB channel intensity within the average RGB time-series signal of the region of interest, these fluctuations are directly related to heart rate. However, feature extraction is required to convert this implicit heart rate information into a quantifiable indicator. Therefore, this application further extracts heart rate features from the average RGB time-series signal of the region of interest to obtain a temporary heart rate. This allows for the acquisition of the user's current real-time heart rate data, serving as a core physiological basis for determining whether the user is in an excited state, providing quantitative support for subsequent state discrimination, and avoiding misjudgments due to a lack of heart rate indicators.
[0029] Specifically, in one possible embodiment, step S131 is implemented as follows: First, the G channel signal with a high signal-to-noise ratio is selected from the average RGB time-series signal of the region of interest. Next, a bandpass filter of 0.7-3.0 Hz is applied to the G channel signal to filter out low-frequency noise and high-frequency interference caused by changes in illumination. Then, a fast Fourier transform is performed on the filtered signal to obtain its frequency domain power spectrum. Finally, the peak frequency is located in the power spectrum, and the peak frequency is multiplied by 60 to convert it into heart rate per minute, thus obtaining the temporary heart rate.
[0030] More specifically, step S132 involves extracting skin tone intensity features from the average RGB time-series signal of the region of interest to obtain temporary color intensity. It should be understood that when a user is excited, facial vasodilation causes skin tone to flush, which is reflected in the intensity distribution of the average RGB time-series signal of the region of interest. However, feature extraction is needed to convert this skin tone change into a discriminative indicator. Therefore, this application further extracts skin tone intensity features from the average RGB time-series signal of the region of interest to obtain temporary color intensity. This allows for the acquisition of dynamic skin tone change data of the user, serving as a supplementary basis for heart rate features, enabling multi-dimensional state discrimination, and improving the comprehensiveness and accuracy of excited state identification.
[0031] Specifically, in one possible embodiment, step S132 is implemented as follows: First, the RGB values of each frame of the average RGB time-series signal of the region of interest are converted to the CIELAB color space. Next, the α component related to skin redness is extracted from the converted signal; this component is most sensitive to changes in skin blood flow. Then, the mean of the α component within a short time window is calculated to eliminate the influence of minor inter-frame fluctuations. Finally, this mean is defined as the temporary color intensity, completing the skin intensity feature extraction process.
[0032] More specifically, in step S133, the temporary heart rate and temporary color intensity are input into the state decision engine to obtain the initial state flag. It should be understood that since both temporary heart rate and temporary color intensity indicators have limitations—relying solely on heart rate may misjudge resting heart rate fluctuations, and relying solely on skin color intensity may misjudge skin color changes caused by light exposure—both types of indicators need to be combined to improve the reliability of the judgment. Therefore, this application further inputs temporary heart rate and temporary color intensity into the state decision engine to obtain the initial state flag. In this way, through multi-indicator collaborative judgment, the risk of misjudgment by a single indicator can be effectively eliminated, accurately identifying suspected agitation or suspected resting states, providing a precise decision-making basis for whether to initiate recovery period monitoring.
[0033] Specifically, in one possible embodiment, step S133 is implemented as follows: First, preset threshold parameters are loaded within the state decision engine. For example, for adult users, the heart rate threshold can be set to 100 beats / minute, which is generally considered the critical point for tachycardia at rest. The skin color intensity threshold can be set to be 15% or one standard deviation higher than the mean of the a* component in the user's historical resting state to capture significant facial flushing caused by vasodilation. Next, the engine compares the real-time calculated temporary heart rate with the heart rate threshold (e.g., 100 beats / minute) and the temporary color intensity with the skin color intensity threshold (e.g., 15% above the baseline value). Finally, if at least one of the indicators, temporary heart rate or temporary color intensity, exceeds its corresponding threshold, the engine outputs a suspected agitated initial state flag; otherwise, it outputs a suspected resting initial state flag.
[0034] In the aforementioned health detection method based on facial images, step S2 involves the system entering a health mode and acquiring a facial image sequence during the recovery period in response to the initial state flag indicating suspected agitation. It should be understood that when the initial state flag indicates suspected agitation, the user's physiological indicators are experiencing transient fluctuations under stress. Directly performing a health assessment at this time would misjudge these transient disturbances as chronic pathological features, leading to distorted assessment results. Therefore, this application initiates a health mode upon triggering the suspected agitation flag, simultaneously acquiring a facial image sequence during the recovery period to obtain complete data on the physiological indicators returning to a stable state from the stress state. This provides foundational data for subsequent extraction of recovery trajectory features and establishment of a dynamic model, effectively distinguishing between transient physiological disturbances and potential pathological features, and significantly improving the accuracy and reliability of health assessment.
[0035] Specifically, in one possible embodiment, step S2 is implemented as follows: First, after the system detects that the initial state flag is suspected agitation, it immediately triggers a health mode switching command and starts the recovery period data acquisition module. Next, the acquisition parameters are configured, for example, setting the camera frame rate to 30 frames / second and adjusting the image resolution to 640x480 pixels to ensure signal temporal resolution and computational efficiency. Simultaneously, an upper limit for the acquisition duration is set, typically 180 seconds, which is sufficient to cover the recovery process from an agitated state to a resting state for most individuals. Then, the system receives facial images transmitted from the camera in real time and performs quality checks on each frame, discarding invalid frames due to motion blur or facial occlusion. Finally, valid images are stored in chronological order to form a continuous sequence of recovery period facial images until physiological indicators such as heart rate are continuously stable within the resting range for 15 seconds, or the upper limit of the acquisition duration of 180 seconds is reached, at which point the system automatically stops acquisition.
[0036] In the aforementioned health detection method based on facial images, step S3 involves extracting multidimensional physiological signals from the recovery period facial image sequence to obtain heart rate and skin perfusion index time series. It should be understood that while the recovery period facial image sequence contains physiological signals reflecting cardiovascular regulatory function, these signals are implicitly present in the image as pixel intensity and cannot be directly used for recovery trajectory modeling and health assessment. Therefore, this application further extracts multidimensional physiological signals from the recovery period facial image sequence to obtain heart rate and skin perfusion index time series, thereby providing quantitative data support for subsequent dynamic modeling. This transforms implicit physiological information in the image into analyzable time-series data, accurately capturing the recovery process of physiological indicators, laying the foundation for distinguishing between transient physiological disturbances and potential pathological features, and ensuring the accuracy of health assessment.
[0037] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S3 of the health detection method based on face images according to an embodiment of this application. Figure 5 As shown, step S3 includes: S31, acquiring and converting the original photoelectric time-series signal and color space of the facial image sequence during the recovery period to obtain the original RGB time-series signal and the Lab color space a* component time-series signal; S32, purifying the blood volume pulse signal of the original RGB time-series signal to obtain the PPG signal; S33, performing a short-time Fourier transform on the PPG signal to obtain the heart rate time series; S34, performing low-pass filtering and downsampling on the Lab color space a* component time-series signal to obtain the skin perfusion index time series.
[0038] Specifically, step S31 involves acquiring raw photoelectric time-series signals and converting the color space of the recovery period facial image sequence to obtain raw RGB time-series signals and Lab color space a* component time-series signals. It should be understood that each frame of the recovery period facial image sequence contains pixel intensity information, which carries the raw photoelectric signals generated by changes in skin blood flow. However, RGB signals need to be directly acquired to extract blood volume information, while the a* component of the Lab color space is more sensitive to changes in skin redness and needs to be converted to obtain perfusion-related signals. Therefore, this application further acquires raw photoelectric time-series signals and converts the color space to obtain two types of raw time-series signals. This allows for the separation of raw data directly related to heart rate and skin perfusion from the images, eliminating interference from irrelevant pixel information, and providing a high-quality initial data source for subsequent signal purification and index calculation.
[0039] Specifically, in one possible embodiment, step S31 is implemented as follows: First, each frame of the facial image sequence during the recovery period is traversed, and the regions of interest (ROIs) for the forehead and cheeks are determined based on stable facial key points. Next, the average intensity values of the RGB three channels of all pixels within each ROI are acquired and stored in chronological order to form the original RGB temporal signals. Then, the RGB values of each frame are converted to the CIELAB color space, and the average pixel value of the a-component within the ROI is extracted. Finally, the average a-component values of each frame are integrated in chronological order to form the Lab color space a* component temporal signal, thus completing the acquisition of the two types of original temporal signals.
[0040] Specifically, step S32 involves purifying the original RGB time-series signal by extracting the blood volume pulse signal to obtain the PPG signal. It should be understood that the original RGB time-series signal contains noise such as light fluctuations and minute head movements, and the blood volume pulse component is masked by noise. Directly using it for heart rate calculation would lead to significant errors. The PPG signal is the core signal reflecting changes in blood volume and needs to be purified. Therefore, this application further purifies the original RGB time-series signal by extracting the blood volume pulse signal to obtain a pure PPG signal. This eliminates the interference of noise on the blood volume signal, obtains a PPG signal that accurately reflects the heartbeat cycle, provides a reliable basis for the accurate calculation of the subsequent heart rate time series, and avoids heart rate calculation errors caused by noise.
[0041] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S32 of the health detection method based on face images according to an embodiment of this application. Figure 6As shown, step S32 includes: S321, performing time normalization on the original RGB time-series signal to obtain a normalized RGB time-series signal; S322, performing signal projection on the normalized RGB time-series signal based on the POS algorithm to obtain an orthogonal first projection signal and a second projection signal; S323, performing adaptive signal fusion on the first projection signal and the second projection signal to obtain a PPG signal.
[0042] More specifically, step S321 involves timing normalization of the original RGB timing signal to obtain a normalized RGB timing signal. It should be understood that because each channel (R, G, B) in the original RGB timing signal has different DC components, and is affected by changes in light intensity, the signal baseline will drift, resulting in incomparable amplitudes for each channel. Direct signal projection would reduce the efficiency of useful signal extraction. Therefore, this application further performs timing normalization on the original RGB timing signal to obtain a normalized RGB timing signal. This eliminates the differences in DC components between channels and the baseline drift caused by light, ensuring that the signals of each channel fluctuate within the same amplitude range. This provides standardized data for subsequent signal projection based on the POS algorithm, improving the separation accuracy of the useful signal.
[0043] Specifically, in one possible embodiment, step S321 is implemented as follows: First, each channel (R, G, B) of the original RGB timing signal is traversed, and the signal mean of each channel over the entire timing period is calculated. Next, for each timing data point of each channel, the signal mean of that channel is subtracted to eliminate the DC component. Then, the amplitude of each channel signal is scaled to unify the standard deviation of each channel signal to a preset value. Finally, the processed R, G, and B channel signals are integrated in chronological order to form a normalized RGB timing signal.
[0044] More specifically, step S322 involves projecting the normalized RGB time-series signal using a POS algorithm to obtain orthogonal first and second projected signals. It should be understood that because residual noise such as motion artifacts still exists in the normalized RGB time-series signal, and the blood volume pulse signal is unevenly distributed across the RGB channels, orthogonal projection is necessary to separate the useful physiological signal from the noise. The POS algorithm can effectively separate signals by utilizing inter-channel correlation. Therefore, this application further projects the normalized RGB time-series signal using the POS algorithm to obtain two orthogonal projected signals. This allows the useful signal related to the blood volume pulse to be projected to a specific direction, while noise is projected to an orthogonal direction, achieving preliminary separation of the useful signal and noise, laying the foundation for subsequent signal fusion.
[0045] Specifically, in one possible embodiment, step S322 is implemented as follows: First, based on the absorption and reflection characteristics of skin tissue to RGB three-channel light, i.e., green light is most sensitive to changes in blood volume, while red and blue light are easily interfered with, the projection coefficients of the POS algorithm are determined. It is clarified that the first projection signal focuses on capturing the pulsation differences between the G and B channels, while the second projection signal focuses on canceling the common interference of the R, G, and B channels. Next, the normalized R, G, and B channel time-series signals (each frame corresponds to one time-series data point) are substituted into the preset projection calculation logic, and the first and second projection signals are calculated point-by-point. Finally, the inner product of the two projection signals over the entire time period is calculated to verify their orthogonality, ensuring that the inner product result approaches zero, i.e., there is no correlation. If the orthogonality requirement is not met, the projection coefficients are readjusted and the projection calculation is performed again until the two projection signals achieve effective orthogonality, completing the signal projection process.
[0046] More specifically, step S323 involves adaptively fusing the first projection signal and the second projection signal to obtain the PPG signal. It should be understood that since the first projection signal and the second projection signal carry different proportions of blood volume pulse information and noise, neither signal alone can provide a PPG signal with the optimal signal-to-noise ratio; the fusion weights need to be adaptively adjusted based on the signal quality of both. Therefore, this application further performs adaptive signal fusion on the two projection signals to obtain a pure PPG signal. In a specific example of this application, step S323 includes: adaptively fusing the first projection signal and the second projection signal using the following formula:
[0047]
[0048] in, and These are the first projection signal and the second projection signal, respectively. To calculate the standard deviation function, For adaptive fusion coefficients, For time variables, This represents the timing value of the PPG signal. This approach fully leverages the advantages of both projection signals, suppresses their respective noise, significantly improves the signal-to-noise ratio of the PPG signal, ensures the accuracy of subsequent heart rate calculations, and provides high-quality physiological data support for health status assessment.
[0049] Specifically, step S33 involves performing a short-time Fourier transform on the PPG signal to obtain a heart rate time series. It should be understood that since the PPG signal is a periodic pulse signal, its frequency characteristics directly correspond to heart rate. However, this signal only reflects pulse changes and cannot directly obtain the heart rate value at each time point. Furthermore, a short-time Fourier transform is needed to correlate the frequency domain with the time domain to obtain the instantaneous heart rate. Therefore, this application further performs a short-time Fourier transform on the PPG signal to obtain a heart rate time series. This allows the frequency characteristics of the PPG signal to be converted into continuous heart rate data, accurately capturing the dynamic trend of heart rate changes during the recovery period, providing continuous quantitative evidence of heart rate for dynamic modeling of the recovery curve, and reflecting the recovery process of cardiovascular regulatory function.
[0050] Specifically, in one possible embodiment, step S33 is implemented as follows: First, a bandpass filter of 0.7-3.0 Hz is applied to the PPG signal to filter out noise outside the heart rate frequency band. Next, the filtered PPG signal is divided into preset short-time windows (e.g., 8 seconds), and a window function is set to reduce spectral leakage. Then, a short-time Fourier transform is performed on the signal within each short-time window to obtain the frequency domain power spectrum within that window. Subsequently, the peak frequency in the power spectrum is located, converted into the corresponding heart rate value, and associated with the center time of the short-time window. Finally, the heart rate values at each time point are integrated in chronological order to form a heart rate time series.
[0051] Specifically, in step S34, the time-series signal of the a* component in the Lab color space is low-pass filtered and downsampled to obtain the skin perfusion index time series. It should be understood that the a component in the Lab color space is directly related to skin redness, which is a direct reflection of blood flow in the epidermis and dermis. Therefore, the dynamic changes in the a* component can serve as an effective physiological proxy indicator for measuring changes in skin perfusion levels. However, the original a* component time-series signal contains high-frequency noise, and its sampling rate is much higher than the slow rate of change in skin perfusion. Therefore, this application further performs low-pass filtering and downsampling on the signal to extract a smooth, noise-reduced time-series signal, which is defined as the skin perfusion index time series. This filters out high-frequency noise, smooths the signal, reduces data redundancy, matches the rate of change in perfusion, accurately reflects the dynamic recovery process of skin perfusion during the recovery period, and provides stable perfusion quantification data for multi-signal coupling analysis.
[0052] Specifically, in one possible embodiment, step S34 is implemented as follows: First, to extract an effective sequence representing the skin perfusion level from the a* component time-series signal, a fourth-order Butterworth low-pass filter is used to filter it, and the cutoff frequency is set to 0.5Hz. This cutoff frequency is chosen because it effectively filters out high-frequency interference caused by camera sensor noise or minor light fluctuations, while completely preserving the skin blood flow perfusion information regulated by the autonomic nervous system and exhibiting a relatively slow rate of change. Next, considering that the physiological rate of change in skin perfusion is much lower than the video frame rate, the filtered signal is downsampled, reducing its sampling rate from the original video frame rate (e.g., 30Hz) to 2Hz. This reduces data redundancy while being sufficient to capture the dynamic process of perfusion recovery. Then, the downsampled signal undergoes consistency verification, for example, by using the 3-sigma principle to remove outliers caused by instantaneous interference. Finally, the verified signal data is stored in chronological order, and the resulting sequence is used as the skin perfusion index time series, ensuring that it smoothly and stably reflects the recovery trajectory of skin perfusion.
[0053] In the aforementioned health detection method based on facial images, step S4 involves performing recovery curve dynamics modeling and feature parameterization on the heart rate time series and skin perfusion index time series to obtain recovery feature vectors. It should be understood that since heart rate time series and skin perfusion index time series are continuous dynamic data, direct use cannot quantify the recovery rate, amplitude, and baseline level of physiological indicators, and it is difficult to extract core information that can distinguish between transient disturbances and pathological features. Therefore, this application further performs recovery curve dynamics modeling and feature parameterization on the two types of time series data to obtain recovery feature vectors. This transforms the dynamic recovery process into quantifiable feature parameters, profoundly reflecting the individual's cardiovascular autonomic nervous system regulation function, providing accurate quantitative evidence for subsequent health assessments, and effectively improving the accuracy and reliability of health assessments.
[0054] In particular, in one specific embodiment, Figure 7 This is a flowchart of sub-step S4 of the health detection method based on face images according to an embodiment of this application. Figure 7 As shown, step S4 includes: S41, performing parameter optimization fitting on the heart rate time series and skin perfusion index time series based on nonlinear least squares method to obtain heart rate fitting parameters and skin perfusion index fitting parameters; S42, integrating the heart rate fitting parameters and skin perfusion index fitting parameters into feature vectors to obtain recovery feature vectors, wherein the recovery feature vectors include heart rate recovery amplitude, heart rate recovery time constant, post-recovery heart rate baseline, skin perfusion amplitude, skin perfusion recovery time constant, and post-recovery skin perfusion baseline.
[0055] Specifically, in step S41, the heart rate time series and skin perfusion index time series are subjected to parameter optimization fitting based on nonlinear least squares method to obtain heart rate fitting parameters and skin perfusion index fitting parameters. It should be understood that since the recovery process of heart rate and skin perfusion index follows a nonlinear exponential decay law, linear fitting methods cannot accurately fit this dynamic process, and noise in the time series data will further increase the fitting error, leading to inaccurate parameters. Therefore, this application further employs nonlinear least squares method to perform parameter optimization fitting on the two types of time series data to obtain accurate heart rate fitting parameters and skin perfusion index fitting parameters. This allows for precise matching of the nonlinear recovery trajectory of physiological indicators, effectively suppressing noise interference, ensuring the reliability of the fitting parameters, and providing a high-quality parameter foundation for the subsequent construction of recovery feature vectors.
[0056] Specifically, in one possible embodiment, step S41 is implemented as follows: First, the heart rate time series and skin perfusion index time series are cleaned to remove outliers and perform trend correction. Next, an exponential decay model is defined as the fitting function, and the parameters to be optimized are identified, including amplitude, time constant, and baseline. Then, with the goal of minimizing the sum of squared residuals, the Levenberg-Marquardt algorithm is used to perform nonlinear least squares optimization, iteratively adjusting the parameters until the residuals meet the convergence condition. Finally, the heart rate fitting parameters for the heart rate time series and the skin perfusion index fitting parameters for the skin perfusion index time series are output respectively, and the validity of the parameters is verified.
[0057] Specifically, in step S42, the heart rate fitting parameters and skin perfusion index fitting parameters are integrated using feature vectors to obtain a recovery feature vector. This recovery feature vector includes the heart rate recovery amplitude, heart rate recovery time constant, post-recovery heart rate baseline, skin perfusion amplitude, skin perfusion recovery time constant, and post-recovery skin perfusion baseline. It should be understood that because the heart rate fitting parameters and skin perfusion index fitting parameters are scattered and independent, using them alone cannot comprehensively reflect the overall recovery function of an individual's cardiovascular system, and they lack a unified quantitative carrier, making it difficult to support subsequent comprehensive health assessments. Therefore, this application further integrates the feature vectors of the two types of fitting parameters to obtain a recovery feature vector containing six key parameters. This systematically integrates the scattered parameters to form a unified quantitative index that comprehensively reflects the recovery characteristics of heart rate and skin perfusion, providing complete and systematic feature inputs for health status assessment and ensuring the comprehensiveness and accuracy of the assessment results.
[0058] Specifically, in one possible embodiment, step S42 is implemented as follows: First, from the heart rate fitting parameters obtained through previous optimization using the nonlinear least squares method, three core parameters—heart rate recovery amplitude, heart rate recovery time constant, and post-recovery heart rate baseline—are precisely extracted. Simultaneously, from the skin perfusion index fitting parameters, three core parameters—skin perfusion amplitude, skin perfusion recovery time constant, and post-recovery skin perfusion baseline—are extracted to ensure that the extraction accuracy of each parameter is consistent with the previous fitting results. Next, the parameters are arranged in an ordered manner according to a preset parameter sorting logic. This logic prioritizes heart rate-related parameters over skin perfusion-related parameters, and within each parameter category, the order of recovery amplitude, recovery time constant, and post-recovery baseline is followed to ensure the consistency and interpretability of the feature vector structure. Finally, the sorted six types of parameters are sequentially embedded into a preset six-dimensional vector structure. Each dimension forms a unique correspondence with a specific parameter, ultimately forming a recovery feature vector with fixed dimensions and clear parameter positions and physical meanings. This ensures that the subsequent health status assessment module can directly read the information from each dimension of the vector and perform feature analysis and health status determination based on a preset algorithm.
[0059] Specifically, when modeling and parameterizing the recovery curve dynamics of the physiological recovery process, it is necessary to consider that the physiological recovery process is not dominated by a single, homogeneous physiological mechanism, but rather a multi-system, multi-rate binary process. That is, heart rate recovery comprises two main parts: a rapid recovery phase, primarily driven by the rapid reactivation of the parasympathetic nervous system (vagus nerve), which quickly lowers the heart rate; and a slow recovery phase, mainly related to the slow withdrawal of sympathetic activity, a decrease in body temperature, and the clearance of metabolites (such as lactic acid) from the blood. The aforementioned POS algorithm already provides clues for separating these different processes, namely, designing orthogonal signals... and ,in It is the signal most sensitive to blood volume pulsation (i.e., the core of the PPG signal) and can more directly reflect macroscopic changes in the cardiac cycle. It is a signal that is more sensitive to motion artifacts and changes in specular reflection, and It is not merely noise; as a non-pulsating optical change, it may also carry physiological information, such as changes in skin microcirculation tension and muscle tremors under the influence of sympathetic nerve activity. The recovery rate of these processes may differ from the macroscopic recovery rate of heart rate.
[0060] In another preferred embodiment, step S4 includes: fitting the heart rate time series and the skin perfusion index time series using a double exponential decay model, wherein the double exponential decay model includes a fast recovery component and a slow recovery component; by fitting the heart rate time series and the skin perfusion index time series, obtaining the fast recovery amplitude and fast recovery time constant representing the fast recovery component, the slow recovery amplitude and slow recovery time constant representing the slow recovery component, and the stable baseline after recovery; wherein the recovery feature vector includes the fast recovery amplitude, the fast recovery time constant, the slow recovery amplitude, the slow recovery time constant, and the stable baseline after recovery.
[0061] In other words, when modeling the recovery curve dynamics, a double exponential decay model can be used to fit the heart rate time series and the skin perfusion exponential time series. This double exponential decay model includes a fast recovery component and a slow recovery component. Here, the mathematical form of the double exponential decay model is:
[0062] in, and These are the amplitude and time constant of the fast recovery component, respectively. and These are the amplitude and time constant of the slow recovery component, respectively. It is the stable baseline after recovery. These are the fitted time-series values of heart rate or skin perfusion index. Therefore, this model containing 5 parameters is needed to fit the heart rate time series and the skin perfusion index time series respectively. That is, by fitting the heart rate time series and the skin perfusion index time series, we obtain the rapid recovery amplitude and rapid recovery time constant representing the rapid recovery component, the slow recovery amplitude and slow recovery time constant representing the slow recovery component, and the stable baseline after recovery. The recovery feature vector includes the rapid recovery amplitude, rapid recovery time constant, slow recovery amplitude, slow recovery time constant, and the stable baseline after recovery.
[0063] First, due to the increased complexity of the model, good initial values are crucial for the optimization algorithm to converge to the global optimum. A stripping method can be used to estimate the initial fitting values. Specifically, in this preferred embodiment, before fitting the data using the double exponential decay model, the stripping method is used to estimate the initial fitting values for the parameters of the double exponential decay model. This includes: fitting a single exponential model to the latter half of the time series data to obtain initial estimates of the slow recovery amplitude, slow recovery time constant, and stable baseline; fitting a slow component curve based on the initial estimates, and subtracting the slow component curve from the complete time series to obtain a residual series; fitting a single exponential model to the residual series to obtain initial estimates of the fast recovery amplitude and fast recovery time constant; and using these initial estimates as the initial values for the nonlinear least squares optimization fitting.
[0064] Specifically, we first estimate the slow component, assuming it occurs in the latter half of the recovery process (e.g., t > 3*). The fast component has largely decayed, and the data is now dominated by the slow component. Therefore, a single exponential model can be fitted to the latter half of the time series data to obtain initial estimates of the slow recovery magnitude, slow recovery time constant, and stable baseline.
[0065] get , and The initial estimates are used to avoid fitting failures caused by improper initial values, ensure the stability and convergence of the double exponential model optimization, reduce coupling interference between parameters, and improve the accuracy of the initial parameters.
[0066] Then, a slow component curve is fitted based on the initial estimate, and the slow component curve is subtracted from the complete time series to obtain the residual series:
[0067] in, These are the time-series values of the original heart rate time series. This is the residual time-series value between the heart rate time series and the fitted values of the slow component. This effectively separates the fast and slow components of heart rate recovery, eliminates mutual interference between the two components, provides a clean data carrier for the initial estimation of the fast component, and ensures the accuracy of subsequent fast component parameter estimation.
[0068] Then, estimate the rapid components, at this point It should be approximately equal to the fast component, that is:
[0069] A single-exponential model is fitted to the residual sequence to obtain initial estimates of the rapid recovery amplitude and rapid recovery time constant. This allows for the accurate extraction of the kinetic parameters of the rapid component, providing all the initial parameters required for the double-exponential model, and offering a comprehensive starting point for subsequent nonlinear optimization, ensuring that the model fully reflects the fast and slow processes of heart rate recovery.
[0070] Then, the initial estimate is used as the initial value for the nonlinear least squares optimization fitting, so as to simultaneously optimize the five parameters. , , , and The optimization process still aims to minimize the sum of squared residuals. The eigenvectors are integrated to obtain the recovered eigenvectors, which fully considers the mutual influence between components, significantly improves the model's fitting accuracy to the actual data, and ensures that the optimized parameters accurately reflect cardiovascular regulatory capacity, providing reliable feature parameters for health assessment.
[0071] Thus, because the bi-exponential model has higher degrees of freedom, it can more accurately fit the real, non-monotonic exponential recovery curve, thereby significantly reducing the sum of squared residuals and improving the model's fitting accuracy. Furthermore, it can obtain features with a higher depth of physiological information analysis, thereby decoupling the recovery capabilities of different physiological systems, for example... It can more purely reflect the rapid regulatory capacity of the vagus nerve, which is a very valuable indicator in exercise physiology and cardiology. It can reflect the efficiency of sympathetic nerve activity withdrawal and metabolite clearance. For example, in individuals with low training levels or under stress, the sympathetic nervous system may remain excited after exercise, leading to... Significantly increased. And the amplitude was greater than The contribution of the rapid recovery mechanism to the overall heart rate recovery process can be quantified; for example, well-trained athletes may have a higher contribution ratio to rapid recovery.
[0072] Furthermore, the specificity and robustness of health assessments have been improved; for example, the model can learn to distinguish between two different types of poor recovery. It's big, but This is normal; it may indicate vagus nerve dysfunction. Normal, but A large number could indicate overtraining or over-excitation of the sympathetic nervous system.
[0073] In the aforementioned health detection method based on facial images, step S5 involves assessing health status based on the recovered feature vector to obtain a health assessment report. It should be understood that while the recovered feature vector contains core quantitative parameters such as heart rate and skin perfusion recovery, these parameters are in professional numerical form and cannot directly reflect health status. Furthermore, they need to be combined with physiological mechanisms and population benchmarks to correlate health risks in order to form a valuable assessment conclusion. Therefore, this application further assesses health status based on the recovered feature vector to obtain a health assessment report. This transforms professional parameters into easily understandable health conclusions, accurately correlates with cardiovascular autonomic nervous system regulation, effectively distinguishes between transient physiological disturbances and potential chronic pathological features, provides users with comprehensive and interpretable health feedback, and significantly enhances the practical value and guiding significance of health detection.
[0074] Specifically, in one possible embodiment, step S5 is implemented as follows: First, the recovery feature vector is standardized to eliminate the impact of parameter magnitude differences on the evaluation model. Health assessment models use preset parameter dimension adaptation logic and feature importance weights to uniformly transform recovery features of different structures into the input format required for evaluation, ensuring that the recovery feature vectors in both embodiments can be accurately parsed. Next, a pre-trained health assessment model is loaded. The training process of this model begins with large-scale clinical data collection, i.e., recruiting subjects with different health levels to perform standardized cardiovascular stress tests, such as the 3-minute stair climb test. During the recovery period after the test, their facial videos are simultaneously recorded using a camera, and their heart rate, heart rate variability (HRV), and other physiological indicators are simultaneously measured using standard medical equipment such as an electrocardiograph (ECG) or a finger-clip pulse oximeter. Subsequently, steps S1 to S4 are performed on the collected facial videos to extract recovery feature vectors for each subject as input features for training the model; while the simultaneously recorded standard physiological indicators, such as the heart rate decrease value 1 minute after recovery or the RMSSD value in HRV, serve as output labels that the model needs to learn and predict. Using these feature-label data pairs, a multi-objective regression model (such as Gradient Boosting Tree (XGBoost) or a small fully connected neural network) is selected for supervised learning. The dataset is divided into training and validation sets, and the model parameters are optimized to minimize the error between the model's predicted values and the true labels until its performance on the validation set reaches a preset standard. Once the trained and validated model is deployed to the system, it can be used for actual health assessments. During detection, the system first standardizes the user's recovery feature vector and then inputs it into the model. The model outputs a set of quantitative assessment indicators, such as a cardiopulmonary function score (1-100 points) and an autonomic balance index. Finally, the system integrates these assessment indicators, the user's historical trend analysis, and corresponding health recommendations to generate a comprehensive and easy-to-understand health assessment report.
[0075] In summary, the health detection method based on facial images according to the embodiments of this application is explained. It first performs preliminary state recognition on real-time facial video streams, triggering recovery period monitoring when the user is determined to be in a state of excitement or post-exercise stress. Subsequently, the system simultaneously extracts time series of multidimensional physiological signals such as heart rate and skin perfusion index from the facial image sequence during the recovery period, and performs dynamic modeling of its recovery trajectory. By parameterizing the morphology, rate, and dynamic coupling relationship between multiple signals of the recovery curve, a recovery feature vector that profoundly reflects an individual's cardiovascular autonomic nervous system regulation function is obtained. Finally, a comprehensive health assessment is performed based on this vector. This effectively distinguishes between instantaneous physiological disturbances and potential chronic pathological features, significantly improving the accuracy and reliability of health assessment.
[0076] Figure 8 This is a block diagram of a health detection system based on face images according to an embodiment of this application. Figure 8 As shown, the health detection system 100 based on face images according to an embodiment of this application includes: a face detection and initial state judgment module 110, used to perform face detection and initial state judgment on the acquired target face video stream to obtain an aligned facial sequence and an initial state flag; a health triggering and acquisition module 120, used to enter a health mode and acquire a recovery period facial image sequence in response to the initial state flag being suspected agitation; a physiological signal time series extraction module 130, used to perform recovery period multidimensional physiological signal time series extraction on the recovery period facial image sequence to obtain a heart rate time series and a skin perfusion index time series; a modeling and feature parameterization module 140, used to perform recovery curve dynamics modeling and feature parameterization on the heart rate time series and skin perfusion index time series to obtain a recovery feature vector; and a health status assessment module 150, used to perform a health status assessment based on the recovery feature vector to obtain a health assessment report.
[0077] As described above, the face image-based health detection system 100 according to the embodiments of this application can be implemented in various wireless terminals, such as servers with face image-based health detection algorithms. In one possible implementation, the face image-based health detection system 100 according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the face image-based health detection system 100 can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the face image-based health detection system 100 can also be one of many hardware modules of the wireless terminal.
[0078] Alternatively, in another example, the face image-based health detection system 100 and the wireless terminal can also be separate devices, and the face image-based health detection system 100 can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.
[0079] Here, those skilled in the art will understand that the specific operations of each step in the above-described health detection system based on facial images have been referenced above. Figures 1 to 7 The description of the health detection method based on face images is detailed here, and therefore, its repeated description will be omitted.
Claims
1. A health detection method based on facial images, characterized in that, include: Face detection and initial state determination are performed on the acquired target face video stream to obtain an aligned facial sequence and initial state flags; In response to the initial state flag indicating suspected agitation, the system enters health mode and acquires facial image sequences during the recovery period; Time series extraction of multidimensional physiological signals during the recovery period was performed on facial image sequences during the recovery period to obtain heart rate time series and skin perfusion index time series. Recovery curve dynamics modeling and feature parameterization were performed on heart rate time series and skin perfusion index time series to obtain recovery feature vectors; A health status assessment is performed based on the recovery feature vector to obtain a health assessment report.
2. The health detection method based on face images according to claim 1, characterized in that, The acquired target face video stream is subjected to face detection and initial state determination to obtain an aligned facial sequence and initial state flags, including: Temporal registration and normalization of the target face video stream are performed to obtain aligned facial sequences and stable keypoint sequences. Based on the stable keypoint sequence, the original photoelectric signal of the region of interest is extracted from the aligned facial sequence to obtain the average RGB time-series signal of the region of interest; The initial state flag is obtained by performing state discrimination based on the short time window signal characteristics on the average RGB time-series signal of the region of interest.
3. The health detection method based on face images according to claim 2, characterized in that, The initial state flag is obtained by performing state discrimination based on short-time-window signal characteristics on the average RGB time-series signal of the region of interest, including: Heart rate features are extracted from the average RGB time-series signal of the region of interest to obtain a temporary heart rate. Skin color intensity features are extracted from the average RGB time-series signal of the region of interest to obtain temporary color intensity; The temporary heart rate and temporary color intensity are input into the state decision engine to obtain the initial state flag.
4. The health detection method based on face images according to claim 1, characterized in that, Time series extraction of multidimensional physiological signals during the recovery period was performed on facial image sequences to obtain heart rate time series and skin perfusion index time series, including: The original photoelectric time-series signal was acquired and the color space was converted from the facial image sequence during the recovery period to obtain the original RGB time-series signal and the Lab color space a* component time-series signal; The original RGB timing signal was purified by blood volume pulse signal extraction to obtain the PPG signal; Short-time Fourier transform of PPG signal to obtain heart rate time series; Low-pass filtering and downsampling were performed on the time series signal of the a* component in the Lab color space to obtain the skin perfusion index time series.
5. The health detection method based on face images according to claim 4, characterized in that, The original RGB timing signal is purified by blood volume pulse signal extraction to obtain the PPG signal, including: The original RGB timing signal is time-normalized to obtain a normalized RGB timing signal; Normalized RGB time-series signals are projected using a POS algorithm to obtain orthogonal first and second projected signals. Adaptive signal fusion is performed on the first projection signal and the second projection signal to obtain the PPG signal.
6. The health detection method based on face images according to claim 5, characterized in that, Adaptive signal fusion of the first projection signal and the second projection signal to obtain the PPG signal includes: adaptively fusing the first projection signal and the second projection signal using the following formula, wherein the formula is: ; ;in, and These are the first projection signal and the second projection signal, respectively. To calculate the standard deviation function, For adaptive fusion coefficients, For time variables, This represents the timing value of the PPG signal.
7. The health detection method based on face images according to claim 1, characterized in that, Recovery curve dynamics modeling and feature parameterization were performed on heart rate time series and skin perfusion index time series to obtain recovery feature vectors, including: The heart rate time series and skin perfusion index time series were optimized and fitted using nonlinear least squares method to obtain the heart rate fitting parameters and skin perfusion index fitting parameters. The heart rate fitting parameters and skin perfusion index fitting parameters are integrated into a feature vector to obtain a recovery feature vector. The recovery feature vector includes heart rate recovery amplitude, heart rate recovery time constant, post-recovery heart rate baseline, skin perfusion amplitude, skin perfusion recovery time constant, and post-recovery skin perfusion baseline.
8. The health detection method based on face images according to claim 1, characterized in that, Recovery curve dynamics modeling and feature parameterization were performed on heart rate time series and skin perfusion index time series to obtain recovery feature vectors, including: The heart rate time series and skin perfusion index time series were fitted using a double exponential decay model, which includes a fast recovery component and a slow recovery component. By fitting the heart rate time series and the skin perfusion index time series, the rapid recovery amplitude and rapid recovery time constant representing the rapid recovery component, the slow recovery amplitude and slow recovery time constant representing the slow recovery component, and the stable baseline after recovery are obtained, respectively. The recovery feature vector includes the fast recovery amplitude, fast recovery time constant, slow recovery amplitude, slow recovery time constant, and the stable baseline after recovery.
9. The health detection method based on face images according to claim 8, characterized in that, Before fitting the model using the double exponential decay model, the method further includes estimating initial fitting values for the parameters of the double exponential decay model using a stripping method, specifically including: A single exponential model is fitted to the second half of the time series data to obtain initial estimates of the slow recovery magnitude, slow recovery time constant, and stable baseline. The slow component curve is fitted based on the initial estimate, and the slow component curve is subtracted from the complete time series to obtain the residual series; A single exponential model is fitted to the residual sequence to obtain initial estimates of the fast recovery amplitude and the fast recovery time constant; The initial estimate is used as the initial value for the nonlinear least squares optimization fitting.
10. A health detection system based on facial images, characterized in that, include: The face detection and initial state assessment module is used to perform face detection and initial state assessment on the acquired target face video stream to obtain an aligned face sequence and initial state flags. The health trigger and acquisition module is used to respond to the initial state flag being suspected agitation, when the system enters health mode and acquires facial image sequences during the recovery period; The physiological signal time series extraction module is used to extract multidimensional physiological signal time series from facial image sequences during the recovery period to obtain heart rate time series and skin perfusion index time series. The modeling and feature parameterization module is used to perform recovery curve dynamics modeling and feature parameterization on heart rate time series and skin perfusion index time series to obtain recovery feature vectors; The health status assessment module is used to assess health status based on recovery feature vectors to obtain a health assessment report.