Indoor Mid-to-Long Distance Camera Heart Rate Detection Method and System with Environmental Interference Resistance

By constructing a distance-focal length correspondence table and using multi-view fusion technology, the problems of weak signal, high noise, and poor stability in heart rate measurement at medium and long distances were solved, enabling low-power, low-latency heart rate detection in indoor environments and improving signal stability and robustness.

CN122074935APending Publication Date: 2026-05-26XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-01-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies cannot reliably perform non-contact heart rate measurement at medium to long distances, especially in indoor environments where they suffer from weak signals, high noise, poor stability, and poor real-time performance. Furthermore, in high magnetic environments, contact devices or close-range non-contact measurement methods cannot be used.

Method used

By constructing a preset correspondence table between distance and available focal length, combining face detection and ROI localization, rigid and non-rigid noise are filtered out, Kalman filtering is used for posture correction, and through multi-threaded real-time scheduling and quality closed-loop linkage, adaptive adjustment of camera focal length and multi-view fusion are achieved, thereby improving signal stability and robustness.

Benefits of technology

Stable measurement of heart rate signals was achieved at medium to long distances of 2 meters and above, improving the strength and measurement range of the analytical signal, reducing noise interference, and ensuring low power consumption and long-term stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074935A_ABST
    Figure CN122074935A_ABST
Patent Text Reader

Abstract

This invention discloses an indoor mid-to-long-range camera heart rate detection method and system with resistance to environmental interference. The host computer acquires one video stream from each camera to determine the video frame sequence of the current optimal viewing angle. Face detection, tracking, and ROI localization are performed on the video frame sequence to obtain the face ROI region for each video frame. Rigid motion noise filtering is applied to the video frame sequence to obtain an aligned frame sequence and a reference frame sequence. The face ROI region of each aligned frame is determined based on the face ROI region of the reference frame. The ROI signal is determined based on the face ROI region of the aligned frame, and non-rigid noise is filtered out to obtain a pulse wave signal. The heart rate value corresponding to the current sliding window is determined based on the pulse wave signal, and a corresponding quality score is determined. This, combined with a preset correspondence between distance and available focal length, allows for adaptive adjustment of the camera focal length. This invention exhibits high robustness, accuracy, and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of non-contact detection technology, specifically relating to an indoor mid-to-long-range camera heart rate detection method and system that is resistant to environmental interference. Background Technology

[0002] Heart rate is an important physiological indicator reflecting the state of the human circulatory system. Traditional heart rate monitoring methods mostly use electrocardiogram (ECG) or contact photoplethysmography (PPG), which require electrode attachment or sensor wearing, and may cause discomfort, skin irritation or poor compliance. For special groups such as newborns, burn victims, and disabled elderly, contact methods are even more inconvenient.

[0003] Camera-based remote photoplethysmography (PPG) technology analyzes subtle color fluctuations in skin reflection caused by changes in blood volume to achieve non-contact heart rate measurement. It offers advantages such as being non-invasive, easy to deploy, and capable of continuous remote monitoring. The academic community has proposed various signal extraction algorithms, including independent component analysis methods based on blind source separation, algorithms based on color space projection, methods based on orthogonal plane projection of skin tone, and enhancement strategies combining signal amplification (e.g., Euler video amplification) with quality assessment.

[0004] Most existing technologies cannot perform non-contact measurements at medium to long distances, and the measurement systems lack stability for long-term monitoring. Furthermore, existing solutions typically require the subject to remain still, which is restrictive in everyday scenarios such as health monitoring of the elderly. In high-magnetic environments, strong electromagnetic interference at close range prevents the use of contact devices or close-range non-contact measurement methods. Therefore, non-contact physiological parameter acquisition at medium to long distances is essential. Combining medium to long-distance health monitoring with indoor cameras to achieve full scene coverage and enable unobtrusive, natural monitoring has broad application scenarios (such as monitoring vital signs of infants in hospital incubators and monitoring the physiological state of bank tellers).

[0005] However, in indoor medium-to-long-distance scenarios, rPPG faces more stringent comprehensive challenges: the farther away the distance, the fewer the facial ROI pixels and the higher the noise ratio; changes in indoor lighting and flickering will superimpose on skin color changes; the natural movement of the test subject will cause ROI drift, occlusion and deformation; at the same time, engineering deployment requires the device to have low power consumption, long-term operation and easy maintenance.

[0006] In summary, current technical solutions for non-contact video heart rate measurement suffer from problems such as weak signal, high noise, poor stability, and poor real-time performance due to factors such as distance attenuation and human motion artifacts. Summary of the Invention

[0007] To address the aforementioned problems in the existing technology, this invention provides an indoor mid-to-long-range camera heart rate detection method and system that is resistant to environmental interference.

[0008] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides an indoor mid-to-long-range camera heart rate detection method with resistance to environmental interference, applied to a host computer. The method includes: Determine the operating parameters of at least one camera, as well as a preset correspondence table between distance and available focal length; Acquire video stream data from each camera and obtain the optimal viewpoint video frame sequence within the current sliding window based on the video stream data. The video frame sequence contains K consecutive video frames, where K is a positive integer greater than or equal to 2. Each sliding window corresponds to a preset time period. By performing face detection, inter-frame tracking, and face ROI localization on the K consecutive video frames, the face ROI region of each video frame is obtained. By sequentially performing rigid motion noise filtering on the K consecutive video frames, the K consecutive video frames are transformed into K consecutive aligned frames that correspond one-to-one, and the K consecutive video frames are used as K reference frames. Using the face ROI region of each of the K reference frames as a reference, the face ROI region of each of the K consecutive alignment frames is determined. The ROI signal is determined based on the face ROI region of each of the K consecutive alignment frames, and the ROI signal is subjected to non-rigid noise filtering to obtain the pulse wave signal. Based on the pulse wave signal, the heart rate value corresponding to the current sliding window is determined. Based on the pulse wave signal and the face detection results of each video frame in the K consecutive video frames, the quality score corresponding to the current sliding window is determined to perform adaptive adjustment of the camera focal length.

[0009] This invention also provides an indoor mid-to-long-range camera heart rate detection system resistant to environmental interference, comprising: At least one camera, each camera is used to collect one video stream of data and transmit it to the host computer; A host computer is connected to each of the at least one camera and is used to execute the above-described method for detecting heart rate in indoor mid-to-long-range cameras with resistance to environmental interference.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: 1) This invention improves the stability of the face ROI region during the measurement process by pre-constructing a correspondence table between distance and available focal length (e.g., pre-constructing it through offline calibration) and combining it with the real-time calculated quality score Qt to perform adaptive focal length adjustment of the camera. This allows the system to maintain a resolvable BVP signal even in scenarios of 2 meters and above, thereby increasing the strength of the resolved BVP signal and expanding the measurement stability and measurement range of the system.

[0011] 2) This invention parses the RTCP SR from each RTSP stream and maps the RTP timestamp to a unified NTP time axis to complete the alignment. Furthermore, based on phase consistency, it calculates the quality score Qt of each stream in real time to determine the optimal viewing angle, thereby improving the continuity and robustness under occlusion, head turning and walking conditions.

[0012] 3) This invention sequentially processes rigid motion, non-rigid disturbance, and distance attenuation compensation, and completes multi-threaded real-time scheduling and quality closed-loop linkage through the host computer, realizing low power consumption, low latency, and long-term stable heart rate output, which greatly improves the environmental robustness of the system. Furthermore, it combines the low power consumption and easy deployment characteristics of embedded host computers.

[0013] 4) This invention estimates the pose angle of the face based on the face detection results and performs temporal smoothing through Kalman filtering to construct a compensation matrix and perform affine alignment, thereby reducing the inconsistency of ROI space and motion artifacts caused by head movement (within 30°) and reducing signal noise.

[0014] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the indoor mid-to-long-range camera heart rate detection method with resistance to environmental interference provided in an embodiment of the present invention. Figure 2 This is an exemplary preset correspondence table between distance and available focal length provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the principle of a sliding window; Figure 4 This is a schematic diagram of the face ROI region provided in an embodiment of the present invention; Figure 5 This is a schematic diagram comparing the effects of POS enhancement filtering before and after according to an embodiment of the present invention; Figure 6 This is a schematic diagram of selecting the main peak from the heart rate frequency band of a frequency domain signal according to an embodiment of the present invention; Figure 7This is a schematic diagram of the architecture of an indoor mid-to-long-range camera heart rate detection system with resistance to environmental interference provided in an embodiment of the present invention; Figure 8 This is a schematic diagram showing the effect of storing some log data on the cloud platform provided in this embodiment of the invention. Detailed Implementation

[0016] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0017] The purpose of this invention is to provide a method and system for stable heart rate detection at medium to long distances indoors at distances of 2 meters and above. Through hardware-algorithm-structure collaboration, proactive measures are taken to significantly improve the ability to resist environmental interference and engineering usability. Specific objectives include: (1) Maintaining sufficient pixels and clarity of the face ROI at medium to long distances using a camera with optical zoom capability and gimbal control, and providing distance adaptive compensation, breaking through the distance limitations of traditional rPPG measurement systems; (2) Stabilizing imaging brightness and suppressing signal distortion caused by sudden changes in illumination and flicker through a hysteresis control strategy; (3) Improving continuity in head-turning and walking situations and avoiding single-view loss through redundant deployment of multiple cameras and multi-view fusion indoors. The algorithm link of posture correction, layered denoising and signal quality assessment is used to suppress rigid motion, non-rigid motion and distance noise respectively, stably extracting BVP heart rate signals, and combining confidence level for quality gating to achieve a window ratio of more than 80% that can continuously output heart rate at 3m; (4) Combining the portable and low-power characteristics of embedded devices, ensuring the portability and long-term stability of the system.

[0018] This invention provides an indoor mid-to-long-range camera heart rate detection method that resists environmental interference. The method is applied to a host computer, and the host computer is connected to at least one camera. Figure 1 This is a flowchart illustrating an indoor mid-to-long-range camera heart rate detection method with environmental interference resistance provided by an embodiment of the present invention. Figure 1 As shown, the method includes: S101. Determine the operating parameters of at least one camera, and a preset correspondence table between distance and available focal length.

[0019] For example, the operating parameters of each camera include: exposure parameters, focal length, camera intrinsic parameter matrix, and distortion coefficients.

[0020] A preset mapping table of distances to available focal lengths contains at least one camera focal length corresponding to each distance. For example, when each camera has multiple different focal length settings, and each focal length setting corresponds to a focal length parameter, the preset mapping table of distances to available focal lengths contains at least one focal length setting for each distance, and the focal length corresponding to that focal length setting. For example, Figure 2 The table shown is a preset correspondence between distance and available focal length. For example, a distance of 1.5m corresponds to two focal length settings, namely focal length setting 2 and focal length setting 4, and focal lengths 9.3715 and 12 correspond to focal lengths respectively.

[0021] S102. Obtain one video stream data from each camera and obtain the video frame sequence of the optimal viewpoint within the current sliding window based on the video stream data. The video frame sequence contains K consecutive video frames, where K is a positive integer greater than or equal to 2. Each sliding window corresponds to a preset time period.

[0022] S103. By performing face detection, inter-frame tracking, and face ROI localization on K consecutive video frames, the face ROI region of each video frame is obtained.

[0023] S104. By sequentially performing rigid motion noise filtering on K consecutive video frames, the K consecutive video frames are transformed into K consecutive aligned frames that correspond one-to-one, and the K consecutive video frames are used as K reference frames.

[0024] S105. Using the face ROI region of each of the K reference frames as a reference, determine the face ROI region of each of the K consecutive aligned frames.

[0025] S106. Determine the ROI signal based on the face ROI region of each of the K consecutive aligned frames, and perform non-rigid noise filtering on the ROI signal to obtain the pulse wave signal.

[0026] S107. Determine the heart rate value corresponding to the current sliding window based on the pulse wave signal. Based on the pulse wave signal and the face detection results of each video frame in K consecutive video frames, determine the quality score Q corresponding to the current sliding window in order to perform adaptive adjustment of the camera focal length.

[0027] In some embodiments, S102 is implemented by steps S1021-S1022 or S1023-S1029: S1021. When at least one camera is a camera and the camera captures a video stream, acquire one video stream captured by the camera and perform data packet parsing to obtain an original video frame sequence received in the current preset time period. Using the camera intrinsic parameter matrix and distortion coefficients contained in the camera's working parameters, perform distortion removal processing on each video frame in the original video frame sequence to obtain the video frame sequence of the optimal viewpoint within the current sliding window corresponding to the current preset time period.

[0028] It should be noted that each sliding window corresponds to a preset time period, meaning the window length of each sliding window is the preset time period. Furthermore, each time a new sliding window is generated, it is generated using a preset sliding step size, which determines the overlap rate between two adjacent sliding windows. For example, Figure 3 This is a schematic diagram illustrating the principle of a sliding window. (For example...) Figure 3 As shown, when the window length T of each sliding window w =10s (seconds), sliding step size T s =5s, when the overlap rate between two adjacent sliding windows is 50%, there are 5 sliding windows (referred to as sliding windows) within the period from 0s to 30s.

[0029] It should be noted that how to use the camera intrinsic parameter matrix and distortion coefficients to perform distortion correction on each video frame is a well-known technique, and this invention does not limit it.

[0030] S1022. Calculate the sampling frequency of the original video frame sequence after distortion correction, and perform signal interpolation processing on the original video frame sequence after distortion correction according to the calculated sampling frequency to obtain the video frame sequence of the optimal viewpoint within the current sliding window corresponding to the current preset time period.

[0031] Since the acquisition time of video frames may be variable (e.g., frame rate fluctuations, frame drops), the core assumption for subsequent frequency domain transformation (FFT) is that the sampling interval is constant. Non-uniform sampling can lead to frequency leakage and inaccurate heart rate calculation. Therefore, it is necessary to convert non-uniform signals into uniform signals through interpolation.

[0032] Specifically, the original video frame sequence after distortion correction is a segment of original non-uniformly sampled signal (i.e., a segment of original non-uniformly sampled RGB signal). The sampling frequency of this segment of original non-uniformly sampled signal is calculated as follows: based on the time window length (dur) and the number of samples (N) of this segment of original non-uniformly sampled signal, the original sampling frequency (N / dur) is calculated. After calculating the original sampling frequency, it is adjusted to obtain the target uniform sampling frequency (fs), which is within the range of 15Hz to 30Hz. This frequency range is suitable for the Nyquist sampling requirements of heart rate analysis and is also compatible with common video frame rate characteristics. Next, signal interpolation processing is performed on this segment of original non-uniformly sampled signal using the target uniform sampling frequency fs. Specifically, taking the start time of the time window of this original non-uniform sampling signal as the reference, a uniform time point sequence (tt) covering the length of the time window of this original non-uniform sampling signal is generated according to the time interval (i.e., 1 / fs) of the target uniform sampling frequency fs. The uniform time point sequence (tt) consists of M uniform time points, where M is obtained by rounding fs × dur, and M ≥ 16 (to ensure the effective execution of the subsequent FFT). Next, for each uniform time point, the corresponding RGB value is calculated from this original non-uniform sampling signal using linear interpolation. Specifically, for each uniform time point, if the uniform time point is less than or equal to the earliest time point of this original non-uniform sampling signal, the first RGB value in this original non-uniform sampling signal is directly used as the signal value corresponding to that uniform time point; if the uniform time point is greater than the latest time point of this original non-uniform sampling signal, the last RGB value in this original non-uniform sampling signal is directly used as the signal value corresponding to that uniform time point; if the uniform time point is greater than... If the earliest time point of this original non-uniformly sampled signal is less than the latest time point of this original non-uniformly sampled signal, then the adjacent time points of this uniform time point in this original non-uniformly sampled signal are located using the bisection method and denoted as t1 and t2. The weight coefficient α = (t-t1) / (t2-t1) of this uniform time point in the interval t1~t2 is calculated. Then, the RGB value of this uniform time point is obtained by linear interpolation using the formula RGB = RGB1 × (1-α) + RGB2 × α, where RGB1 and RGB2 are the RGB values ​​corresponding to time points t1 and t2, respectively, and t represents the uniform time point. Finally, after the RGB values ​​of all uniform time points in the uniform time point sequence (tt) are filled, a segment of equally spaced sampled signal composed of the uniform time axis sequence (tt) and the corresponding RGB values ​​is output. This segment of equally spaced sampled signal is the video frame sequence of the optimal viewpoint within the current sliding window. The following example further illustrates this.For example, if the starting time of the time window of an original non-uniform sampling signal is 0.0s and the length of the time window dur is 3.0s, then the target uniform sampling frequency fs = 30Hz. Therefore, the time interval Δt between two adjacent uniform time points in the uniform time point sequence (tt) is 1 / 30 ≈ 0.0333 seconds, and the number of uniform time points in the uniform time point sequence (tt) is M = round(30×3) = 90 (satisfying M≥16). Therefore, the process of generating these 90 uniform time points is as follows: When i=0, that is, the first uniform time point: t0=0.0+0×0.0333=0.0 seconds; When i=1, that is, the second uniform time point: t1=0.0+1×0.0333=0.0333 seconds; When i=2, that is, the third uniform time point: t2=0.0+2×0.0333=0.0666 seconds; ... When i=89, that is, the 90th uniform time point: t 89 =0.0 + 89 × 0.0333 ≈ 2.9637 seconds (covering the original 3-second window). This ultimately generates 90 equally spaced time points with a fixed time interval of 0.0333 seconds, forming a complete uniform time point sequence (tt).

[0033] S1023. When at least one camera is two or more cameras, and M cameras acquire M video streams, acquire the M video streams and parse the data packets of each, to obtain the M original video frame sequences received in the current preset time period and the NTP timestamp of each original video frame in each original video frame sequence. and RTP timestamp .

[0034] S1024, NTP-based timestamp and RTP timestamp The acquisition times of the M original video frame sequences are converted into sampling times on the same NTP time axis to obtain the standard acquisition times of each original video frame in the M original video frame sequences.

[0035] In this invention, the host computer acts as a Network Time Protocol (NTP) server. Each camera periodically synchronizes its time with the host computer, ensuring that the time bases of all cameras are as consistent as possible, providing a common reference for subsequent cross-camera timeline mapping. Each video stream is an RTSP stream. The host computer parses the RTCP Sender Report (SR packet) from each RTSP stream. The SR packet contains not only video frames but also NTP timestamps. (Absolute time) and corresponding RTP timestamp (Counting Clock). For each original video frame (e.g., referred to as video frame k) in each of the M original video frame sequences, since RTP timestamps are 32-bit counts and wraparound exists, the relative count difference of video frame k is first calculated. ,Right now , Indicates modulo, Represents the NTP timestamp of video frame k Then, based on the RTP clock frequency (90 kHz is commonly used for video), and the relative count difference of video frame k. and NTP timestamp Calculate the standard acquisition time of video frame k. ,in, In this way, the RTP clock frequency can be utilized. The counting clock is converted into a time increment to obtain the acquisition time (i.e., the standard acquisition time) of video frame k on the NTP time axis. By using this method, the RTP counts of each parsed video frame are converted into sampling times on the same NTP time axis, so that the video frames can be compared and aligned under a unified time reference.

[0036] S1025. Sort each original video frame sequence in ascending order according to the standard acquisition time to obtain M time-rearranged original video frame sequences.

[0037] S1026. Calculate the sampling frequency of each time-rearranged original video frame sequence, and perform signal interpolation processing on the time-rearranged original video frame sequence according to the calculated sampling frequency to obtain M video frame sequences with uniform sampling intervals.

[0038] The specific principle of this step is the same as that of step S1022 above, and will not be repeated here.

[0039] S1027. Using the camera intrinsic parameter matrix and distortion coefficients included in the camera's working parameters, perform distortion removal processing on each video frame in each video frame sequence with uniform sampling intervals to obtain M distortion-removed video frame sequences.

[0040] S1028. Perform face detection on each distortion-corrected video frame sequence, and calculate the quality score of the relevant face ROI region for each distortion-corrected video frame sequence based on the face detection results, to obtain M quality scores of the relevant face ROI regions.

[0041] Specifically, for each distortion-corrected video frame sequence, face detection is performed on the sequence to obtain facial landmarks in each frame. For example, the dlib library can be used for face detection. dlib is a commonly used computer vision library that provides mature tools for face detection and facial landmark localization. The face bounding box is obtained using dlib's frontal face detector, and within this bounding box, dlib's 68-point landmark model is used to locate landmarks such as the corners of the eyes, the tip of the nose, the corners of the mouth, and the jawline. Next, based on the facial keypoints in each video frame of the distorted video frame sequence, the normalized face scale and face sharpness of each video frame in the distorted video frame sequence are calculated. The normalized face scale is the ratio of the height of the face region to the height of the original video frame, which can be directly calculated using face detection results and video frame size. Face sharpness can be calculated using the variance method and the average gradient method. Then, the mean of the normalized face scale and the mean of the face sharpness of all video frames in the distorted video frame sequence are calculated. Finally, the mean of the normalized face scale and the mean of the face sharpness of the distorted video frame sequence are weighted and summed to obtain the quality score of the relevant face ROI region in the distorted video frame sequence. It should be noted that before performing a weighted summation of the mean of the normalized face scale and the mean of the face sharpness, the mean of the normalized face scale and the mean of the face sharpness are first standardized to map them to the same numerical space before the weighted summation is performed.

[0042] It should be noted that the weights corresponding to the mean of normalized face scale and the mean of face sharpness can be determined according to actual needs, and this invention does not impose any limitations on this.

[0043] S1029. Take the video frame sequence with the highest quality score for the relevant face ROI region after distortion correction as the video frame sequence of the optimal viewpoint within the current sliding window corresponding to the current preset time period.

[0044] Here, selecting the view with the highest score as the output view can ensure phase consistency and stability of the main peak of the spectrum when switching between multiple views.

[0045] In some embodiments, the specific implementation principle of S103 is as follows: By using the dlib library to perform face detection on K consecutive video frames, and simultaneously using KCF face bounding box tracking and KLT face keypoint tracking to perform inter-frame tracking on the K consecutive video frames, it is possible to obtain whether a target face exists in each of the K consecutive video frames and the face keypoints of the target face. Then, for video frames containing a target face, based on the face keypoints (e.g., 68 keypoints) of the target face in the video frame, all face ROI regions in the video frame are determined. For example, Figure 4 This is a schematic diagram of the ROI (Region of Interest) of a face. For example... Figure 4 As shown, the facial ROI region can include two cheek ROI regions and one forehead ROI region. In some embodiments, the facial ROI region can also include other ROI regions of the face.

[0046] It should be noted that other facial landmark models (such as MediaPipe Face Mesh) and other trackers (KLT / optical flow / correlation filtering) can also be used for target tracking, and this invention does not limit them.

[0047] In some embodiments, S104 is implemented by S1041 to S1044: S1041. Using the PnP method, the face detection results of each video frame in K consecutive video frames are used to solve the pose angle of the face in each video frame in K consecutive video frames, so as to obtain the observed pose angle of each video frame in K consecutive video frames. At the same time, Kalman filtering is used to smooth the prediction of the pose angle of the face in each video frame in K consecutive video frames, so as to obtain the predicted pose angle of each video frame in K consecutive video frames.

[0048] Here, the Chinese name for the PnP method is the n-point perspective problem method, used to solve for pose. Specifically, the PnP method uses the face detection results (i.e., 68 facial keypoints) of each of K consecutive video frames to calculate the pose angle of the face in each video frame. It should be noted that during the solution process, the translation of each of the K consecutive video frames can be calculated simultaneously. Since the PnP method is an existing method, this invention will not elaborate on the specific solution principle. Furthermore, based on the face detection results of each video frame, the interpupillary distance ratio of the faces in each video frame is also calculated. The attitude angles of a human face include three Euler angles: yaw, pitch, and roll.

[0049] Here, the Kalman filter smoothing process can be divided into two steps: prediction and update. Prediction involves assuming a uniform change in angle or displacement over a short period of time to estimate the state of the next frame. If the state at a certain moment is... ,in It refers to the face's attitude angles (i.e., the three Euler angles of yaw, pitch, and roll, as well as the three-dimensional translation vector). It is the rate of change. Therefore, the system prediction for the next moment is: The prediction is updated after prediction, using the observations of the current frame to correct the predicted value. If the observation noise is low, the weights are adjusted accordingly. Larger output values ​​more closely resemble observed values. If the observation noise is high, with abrupt changes or rigid motion noise, then the weights will be adjusted. The smaller the output, the closer it is to the prediction, resulting in a smoother output. It should be noted that Kalman filtering smoothing is an existing method, and the specific principles will not be elaborated upon in this invention.

[0050] S1042. Based on the magnitude of the observed attitude angle and the predicted attitude angle of each video frame in K consecutive video frames, determine the weight of the observed attitude angle and the weight of the predicted attitude angle respectively. Based on the determined weights, sum the observed attitude angle and the predicted attitude angle of each video frame in K consecutive video frames to obtain the final attitude angle of each video frame in K consecutive video frames.

[0051] Specifically, for any video frame among K consecutive video frames (e.g., denoted as the k-th video frame), if the difference (i.e., the residual) between the observed attitude angle and the predicted attitude angle of the k-th video frame exceeds a set threshold, it indicates that there is significant noise or abrupt change in the detection of the k-th video frame. Therefore, the weight of the observed attitude angle of the k-th video frame is reduced, while the weight of the predicted attitude angle of the k-th video frame is increased. The predicted attitude angle of the k-th video frame is based on the (k-1)-th video frame. The attitude angles of the k-th video frame are predicted from the observed attitude angles. Conversely, if the residual is small, it indicates that the observation quality of the k-th video frame is reliable. In this case, the weight of the observed attitude angle of the k-th video frame is increased to ensure the real-time performance and accuracy of attitude tracking, while the weight of the predicted attitude angle of the k-th video frame is decreased. Then, the observed attitude angle and the predicted attitude angle of the k-th video frame are weighted and summed using dynamically adjusted weights (the sum of their weights is 1) to obtain the final smoothed attitude angle of the k-th video frame. It should be noted that this invention does not limit the magnitude of the adjustment or the base value used for adjustment. The attitude angles of the face include three Euler angles: yaw, pitch, and roll.

[0052] S1043. Using the weighted attitude angle of each video frame in K consecutive video frames, construct an affine compensation matrix to obtain K affine compensation matrices that correspond one-to-one with the K consecutive video frames, and determine the inverse matrix of each of the K affine compensation matrices to obtain K alignment matrices.

[0053] For example, for any one of K consecutive video frames (e.g., denoted as the k-th video frame), based on the roll angle in the final attitude angle of the k-th video frame... and the translation amount of the kth video frame. and the interpupillary distance ratio of the k-th video frame. The expression for the corresponding affine compensation matrix is ​​as follows: Affine compensation matrix The inverse matrix is ​​an alignment matrix. .

[0054] S1044. Using K alignment matrices, perform affine transformations on K consecutive video frames that correspond one-to-one to obtain K consecutive aligned frames, and use the K consecutive video frames as K reference frames.

[0055] Continuing with the example above, using a matrix By performing an affine transformation on the k-th video frame, the k-th aligned frame can be obtained. ,Right now ,in, This represents the affine transformation operation. This represents the k-th video frame. In this way, we can obtain K consecutive aligned frames.

[0056] In some embodiments, the specific implementation principle of S105 is as follows: for each alignment frame, based on the two cheek ROI regions and one forehead ROI region in the reference frame corresponding to the alignment frame, two cheek ROI regions and one forehead ROI region are cropped from the alignment frame to obtain the face ROI region of the alignment frame.

[0057] In some embodiments, the above-mentioned S106 is implemented by S1061 to S1065: S1061. Calculate the average pixel values ​​of the R, G, and B channels of the face ROI region in each of the K consecutive aligned frames to obtain a three-channel discrete pixel mean sequence composed of the average pixel values ​​of the R, G, and B channels of the K consecutive aligned frames, and use the three-channel discrete pixel mean sequence as the ROI signal.

[0058] Specifically, for each alignment frame (e.g., denoted as the k-th alignment frame), the average pixel values ​​of the R, G, and B channels for the two cheek ROI regions and one forehead ROI region are calculated. For example, the average pixel value of any ROI region in the k-th alignment frame is calculated as follows: Average pixel value of each color channel The expression is: , This represents the number of pixels within any given ROI region, where any ROI region can be either a cheek ROI region or a forehead ROI region. Thus, the average pixel values ​​of the R, G, and B channels for the two cheek ROI regions and the average pixel values ​​of the R, G, and B channels for the forehead ROI region in the k-th aligned frame constitute a set of pixel average values. In this way, the K sets of pixel average values ​​corresponding to K consecutive aligned frames constitute a three-channel discrete pixel mean sequence, which is a segment of the ROI signal.

[0059] S1062. Using a preset time length as the segmentation benchmark, the ROI signal is segmented according to a preset overlap rate to obtain multiple segmented ROI signals. After normalizing each segmented ROI signal, multiple normalized segmented ROI signals are obtained.

[0060] Here, the preset time length can be set according to actual needs, for example, it can be 1.5s or 2s, and the preset overlap rate can also be set according to actual needs, for example, it can be 50%. This invention does not make specific limitations on these.

[0061] It should be noted that any existing normalization method can be used, and this invention does not limit it.

[0062] S1063. After sequentially splicing multiple normalized segmented ROI signals into a single signal, a third-order Butterworth bandpass filter (0.75–3.0Hz) is used to filter the single signal to suppress non-heart rate frequency band noise, resulting in a denoised single signal.

[0063] S1064. The POS method is used to process a segment of the denoised signal to obtain a pulse wave signal.

[0064] Here, the Chinese name for the POS method is Plane Orthogonal to Skin Color Method, which is an existing method. The POS method linearly combines the RGB channels of a denoised signal segment, highlighting the blood volume pulse wave (BVP) component and suppressing common-mode components introduced by facial expressions and local shadows, thus obtaining a BVP signal segment. For example, Figure 5 This refers to the raw signal captured by the camera when the distance between the camera and the person is 2.5 meters. Figure 5The yellow curve shown above) and the signal after POS enhancement filtering of the original signal (i.e., the pulse wave signal extracted after non-rigid noise filtering) Figure 5 A comparison diagram showing the yellow curve shown below. (For example...) Figure 5 As shown, trend and abrupt noise in the signal are suppressed, the signal-to-noise ratio of the rPPG signal is improved, and Figure 5 The vertical axis values ​​displayed are -1.00, -0.50, 0.00, 0.50, and 1.00, respectively. Figure 5 The horizontal axis displays coordinate values ​​of 0.0s, 2.0s, 4.0s, 6.0s, 8.0s, and 10.0s, respectively.

[0065] In some embodiments, the specific principle of determining the heart rate value corresponding to the current sliding window based on the pulse wave signal in S107 above is as follows: A Fast Fourier Transform (FFT) is performed on the obtained pulse wave signal to obtain a frequency domain signal; the main peak f* is selected within the heart rate frequency band of the frequency domain signal; and the main peak f* is then transformed in the frequency domain to obtain the heart rate value. ,Right now In some embodiments, the obtained heart rate value can be constrained by the difference between the maximum confidence output and historical values ​​(less than 15 bpm) to suppress aberrant jumps. For example, Figure 6 This is a schematic diagram of selecting the main peak f* from the heart rate frequency band of the frequency domain signal. As shown in the figure, the light blue area represents the heart rate frequency band of the frequency domain signal.

[0066] In some embodiments, the determination of the quality score Qt corresponding to the current sliding window based on the pulse wave signal and the face detection results of each video frame in K consecutive video frames in S107 above is implemented through S1071~S1075: S1071. Calculate the average signal-to-noise ratio (SNR) of the pulse wave signal.

[0067] Specifically, the method for calculating the average signal-to-noise ratio (SNR) of the pulse wave signal is an existing method, and this invention will not elaborate on it.

[0068] S1072. Based on the face detection results of each of the K consecutive video frames, calculate the average motion of the K consecutive video frames.

[0069] Here, the motion of each video frame in K consecutive video frames is calculated, and then the average motion of the K consecutive video frames is obtained. The motion of each video frame indicates the amplitude of the subject's movement during the measurement process and can be used to determine whether rigid motion compensation is needed. Specifically, for each video frame in K consecutive video frames (e.g., denoted as the k-th video frame), the motion of the k-th video frame is...k The calculation formula is: ; ; ; ; in, and It is the average of the x and y coordinates of the 68 facial keypoints in the k-th video frame, representing the center coordinates of the keypoints in the k-th video frame. and It is the average of the x and y coordinates of the 68 facial keypoints in the (k-1)th video frame, representing the center coordinates of the keypoints in the (k-1)th video frame. It is the roll angle in the pose angle of the face in the k-th video frame. k The roll angle in the pose angle of the face in the (k-1)th video frame k-1 The angular difference between them. and It is the 68 facial key points in the k-th video frame. The x and y coordinates of facial landmarks and It is the first of the 68 facial key points in the (k-1)th video frame. The x and y coordinates of key facial features. Indicates the first Euclidean displacement of key facial features. It is the mean of the Euclidean displacements of 68 facial key points. It represents the standard deviation of Euclidean displacement. These are empirical weighting coefficients, for example, with values ​​of 0.4, 0.4, and 0.2, respectively, to focus on the detection of rigid head movements. This indicates normalization processing. In other words, when processing... , and Before weighted summation, the three parameters need to be normalized to map them to the same value range (e.g., [0,1]) so that they can be compared during weighted fusion.

[0070] S1073. The average signal-to-noise ratio (SNR), average motion, and the mean of normalized face scale and the mean of face sharpness of the video frame sequence within the current sliding window are weighted and summed to obtain the sum value.

[0071] Specifically, the weighted sum of the average signal-to-noise ratio (SNR), average motion, mean normalized face scale, and mean face sharpness is equal to 1. Furthermore, the weights of these parameters can be set according to actual needs, and this invention does not impose any limitations on them. It should also be noted that the average SNR, average motion, mean normalized face scale, and mean face sharpness all need to be normalized before weighted summation to map their values ​​to the same numerical space.

[0072] S1074, Calculate the preset coefficients The product of the average motion and the average motion is used to obtain the product value.

[0073] here, The value can be set according to actual needs.

[0074] S1075. Calculate the difference between the sum and the product to obtain the quality score Q corresponding to the current sliding window.

[0075] For example, the expression for the quality score Q corresponding to the current sliding window is: ,in, The third of the four parameters represents the average signal-to-noise ratio (SNR), average motion, mean normalized face scale, and mean face sharpness. One parameter, The value of is between 1 and 4. Indicates the first The weights of each parameter, Indicates the first The normalized values ​​of the parameters This represents the normalized average amount of exercise.

[0076] In some embodiments, the above method further includes S108~S109: S108. Based on the relationship between the quality score Q corresponding to the current sliding window and the quality score threshold Q (min), determine whether to adjust the camera focal length.

[0077] Specifically, if the quality score Q corresponding to the current sliding window is greater than or equal to the quality score threshold Q (min), then there is no need to adjust the camera focal length; conversely, if the quality score Q corresponding to the current sliding window is less than the quality score threshold Q (min), then there is a need to adjust the camera focal length.

[0078] S109. When determining to adjust the camera focal length, the camera focal length is adjusted based on the preset correspondence table between distance and available focal length and the calculation of the quality score.

[0079] Specifically, when adjusting the camera focal length, you can select a focal length corresponding to the current distance from the preset correspondence table of distance and available focal lengths based on the distance between the current camera and the target (i.e., the person being tested), and then modify the focal length of all cameras to the selected focal length.

[0080] To suppress the significant jitter in rPPG signal measurement caused by frequent camera focus switching, a hysteresis and hold constraint are introduced for focus updates. Specifically, after changing the camera's focus, a waiting period is observed before continuing the process described in steps S102-S107 to obtain the quality score corresponding to the next sliding window. If the quality score corresponding to the next sliding window... If the quality score Q of the current sliding window is greater than the quality score Q of the current sliding window, after Z (e.g., 8) sliding windows, the quality score of the latest sliding window is calculated again, and the principle of S108 above is used to determine whether the focus needs to be adjusted. If the focus needs to be adjusted, S109 above is executed again. If the quality score of the next sliding window is greater than the quality score Q of the current sliding window, the quality score of the next sliding window is calculated. When the heart rate is also less than the quality score threshold Q(min), after passing through Z / 2 (i.e., 4) sliding windows, the quality score corresponding to the latest sliding window is calculated again, and the principle of S108 above is used to determine whether the focus needs to be adjusted. If the focus needs to be adjusted, S109 above is executed again. In this way, periodic heart rate detection and closed-loop control are achieved.

[0081] This invention also provides an indoor mid-to-long-range camera heart rate detection system resistant to environmental interference. The system includes a host computer and at least one camera, with the host computer and at least one camera respectively connected. Each of the at least one camera is used to collect one video stream of data and transmit it to the host computer. The host computer is used to execute the aforementioned indoor mid-to-long-range camera heart rate detection method resistant to environmental interference. It should be noted that optical zoom cameras with different magnifications, binocular / depth cameras, or near-infrared cameras can be used; this invention does not limit the scope of the invention.

[0082] In some embodiments, Figure 7 This is a schematic diagram of the system architecture. For example... Figure 7As shown, from a hardware and software perspective, this system includes a hardware support system module, a signal acquisition and tracking system module, a hierarchical noise filtering system module, and a heart rate calculation and closed-loop control module. The hardware support system module is hardware, while the signal acquisition and tracking system module, hierarchical noise filtering system module, and heart rate calculation and closed-loop control module are all software modules. The hardware support system module provides high-resolution video acquisition, zoom control, edge computing, and multi-camera coverage capabilities. The main function of the signal acquisition and tracking system module is to stably locate and continuously track the facial physiological region in the acquired real-time video stream, estimate the subject's head posture and perform geometric alignment, and, if necessary, perform multi-view selection or fusion to avoid occlusion. The hierarchical noise filtering system module performs hierarchical suppression according to the noise source; it first processes changes in rigid head posture, then processes non-rigid disturbances such as facial expressions / blinking, and finally compensates for distance attenuation and removes common-mode disturbances. The heart rate calculation and closed-loop control module uses a sliding window to perform signal normalization, detrending, enhancement, and spectrum estimation, outputting quality indicators such as heart rate and confidence level. A quality indicator feedback and distance adaptive adjustment mechanism ensures real-time and stable output. The hardware support system module includes a host computer and multiple zoom cameras. The signal acquisition and tracking system module, the layered noise filtering system module, and the heart rate calculation and closed-loop control module can be deployed on the host computer, enabling it to execute the aforementioned environmentally interference-resistant indoor mid-to-long-range camera heart rate detection method, thereby controlling the camera's working focal length and detecting the target object's heart rate. In some embodiments, the system further includes a host computer storage and display module for receiving heart rate results, quality indicators, and keyframe fragments output from the edge terminal (i.e., the host computer), providing real-time curve display, data storage, and cloud platform synchronization to support long-term health monitoring. For example, Figure 8 This is a display of some log data stored on the cloud platform.

[0083] Compared to existing single-point improvement solutions (such as relying solely on single-camera zoom and ROI stabilization, using only RGB-NIR dual-channel anti-lighting, or algorithms only targeting camera shake), this invention achieves comprehensive advantages through coordinated optimization across four levels: structure, hardware, algorithm, and closed-loop control. 1) The non-contact heart rate estimation of this invention has an adaptive compensation mechanism for different distances, which improves the accuracy in medium and long distance scenarios to a certain extent. By using optical zoom, the stability of the face ROI region during the measurement process is ensured. By using an offline calibrated set of available focal lengths, adaptive gain compensation is achieved at different distances. Even in scenarios of 2 meters and above, a resolvable BVP signal can still be maintained, thus broadening the measurement range of the system.

[0084] 2) This invention proposes a multi-camera redundancy and quality-driven multi-view switching method: the host computer acts as an NTP server to complete time synchronization; it parses the RTCP SR from each RTSP stream and maps the RTP timestamp to a unified NTP time axis to complete alignment; based on phase consistency, it calculates the quality score of each stream, selects / switches the optimal view or performs fusion according to the hysteresis rule, thereby improving the continuity and robustness under occlusion, head turning and walking conditions.

[0085] 3) The monitoring system provided by this invention has good stability and flexibility. Combining the low power consumption and easy deployment characteristics of the embedded host computer, a closed-loop control mechanism enables continuous monitoring of physiological indicators over a long period of time. The monitoring system can utilize a cloud platform to update users' historical measurement data in real time, maintaining system availability.

[0086] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0087] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0088] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0089] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for detecting heart rate using an indoor mid-to-long-range camera with resistance to environmental interference, characterized in that, Applied to a host computer, the method includes: Determine the operating parameters of at least one camera, as well as a preset correspondence table between distance and available focal length; Acquire video stream data from each camera and obtain the optimal viewpoint video frame sequence within the current sliding window based on the video stream data. The video frame sequence contains K consecutive video frames, where K is a positive integer greater than or equal to 2. Each sliding window corresponds to a preset time period. By performing face detection, inter-frame tracking, and face ROI localization on the K consecutive video frames, the face ROI region of each video frame is obtained. By sequentially performing rigid motion noise filtering on the K consecutive video frames, the K consecutive video frames are transformed into K consecutive aligned frames that correspond one-to-one, and the K consecutive video frames are used as K reference frames. Using the face ROI region of each of the K reference frames as a reference, the face ROI region of each of the K consecutive alignment frames is determined. The ROI signal is determined based on the face ROI region of each of the K consecutive alignment frames, and the ROI signal is subjected to non-rigid noise filtering to obtain the pulse wave signal. Based on the pulse wave signal, the heart rate value corresponding to the current sliding window is determined. Based on the pulse wave signal and the face detection results of each video frame in the K consecutive video frames, the quality score corresponding to the current sliding window is determined to perform adaptive adjustment of the camera focal length.

2. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 1, characterized in that, The step of acquiring one video stream data from each camera and obtaining the video frame sequence of the optimal viewpoint within the current sliding window based on the video stream data includes: When the at least one camera is a single camera and the camera captures a video stream, the video stream captured by the camera is obtained and the data packet is parsed to obtain an original video frame sequence received in the current preset time period. The camera intrinsic parameter matrix and distortion coefficient contained in the working parameters of the camera are used to perform distortion removal processing on each video frame in the original video frame sequence to obtain the original video frame sequence after distortion removal. Calculate the sampling frequency of the original video frame sequence after distortion correction, and perform signal interpolation processing on the original video frame sequence after distortion correction based on the calculated sampling frequency to obtain the video frame sequence of the optimal viewpoint within the current sliding window corresponding to the current preset time period.

3. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 2, characterized in that, The step of acquiring one video stream data from each camera and obtaining the video frame sequence within the current sliding window based on the video stream data includes: When there are two or more cameras, and M cameras capture M video streams, the M video streams are acquired and data packets are parsed to obtain M original video frame sequences received in the current preset time period and the NTP timestamp of each original video frame in each original video frame sequence. and RTP timestamp ; Based on NTP timestamps and RTP timestamp The acquisition times of the M original video frame sequences are converted into sampling times on the same NTP time axis to obtain the standard acquisition times of each original video frame in the M original video frame sequences. Each original video frame sequence is sorted in ascending order according to the standard acquisition time to obtain M time-rearranged original video frame sequences; Calculate the sampling frequency of each time-rearranged original video frame sequence, and perform signal interpolation processing on each time-rearranged original video frame sequence based on the calculated sampling frequency to obtain M video frame sequences with uniform sampling intervals. Using the camera intrinsic parameter matrix and distortion coefficients included in the camera's operating parameters, distortion removal processing is performed on each video frame in the M video frame sequences with uniform sampling intervals to obtain M distortion-removed video frame sequences. Face detection is performed on each distortion-corrected video frame sequence. Based on the face detection results, the quality scores of the relevant face ROI regions in each distortion-corrected video frame sequence are calculated to obtain M quality scores of the relevant face ROI regions. The video frame sequence with the highest quality score for the relevant face ROI region after distortion correction is taken as the video frame sequence of the optimal viewpoint within the current sliding window corresponding to the current preset time period.

4. The indoor mid-to-long-range camera heart rate detection method under environmental interference as described in claim 3, characterized in that, The NTP-based timestamp and RTP timestamp The acquisition times of the M original video frame sequences are converted into sampling times on the same NTP time axis to obtain the standard acquisition times of each original video frame in the M original video frame sequences, including: For each original video frame in each of the M original video frame sequences, based on the NTP timestamp of the original video frame... and RTP timestamp Calculate the relative count difference of the original video frames. ,in, , Indicates modulo, Indicates the NTP timestamp of the original video frame ; Based on RTP clock frequency and the relative count difference of the original video frames. and NTP timestamp Calculate the standard acquisition time of the original video frame. ,in, .

5. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 3, characterized in that, The process of performing face detection on each distortion-corrected original video frame sequence and calculating the quality score of the relevant face ROI region for each distortion-corrected original video frame sequence based on the face detection results includes: Face detection is performed on each distortion-free original video frame sequence to obtain the facial key points in each original video frame of each distortion-free original video frame sequence; Based on the facial key points in each original video frame in each distortion-free original video frame sequence, calculate the normalized face scale and face sharpness of each original video frame in each distortion-free original video frame sequence, where the normalized face scale is the ratio of the height of the face region to the height of the original video frame. Calculate the mean of normalized face scale and the mean of face sharpness for all original video frames in each distortion-free original video frame sequence. The mean of the normalized face scale and the mean of the face sharpness of each distortion-corrected original video frame sequence are weighted and summed to obtain the quality score of the relevant face ROI region for each distortion-corrected original video frame sequence.

6. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 1, characterized in that, The method further includes: Based on the relationship between the quality score Qt corresponding to the current sliding window and the quality score threshold, it is determined whether to adjust the camera focal length. When determining to adjust the camera focal length, the camera focal length is adjusted based on a preset correspondence table between distance and available focal length and the calculation of quality score.

7. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 1, characterized in that, The step of sequentially performing rigid motion noise filtering on the K consecutive video frames to transform them into K consecutive aligned frames, and using the K consecutive video frames as K reference frames, includes: The PnP method is used to solve the pose angle of the face in each of the K consecutive video frames using the face detection results of each video frame in the K consecutive video frames, so as to obtain the observed pose angle of each video frame in the K consecutive video frames. At the same time, Kalman filtering is used to smooth the face and predict the pose angle of the face in each of the K consecutive video frames, so as to obtain the predicted pose angle of each video frame in the K consecutive video frames. Based on the magnitude between the observed attitude angle and the predicted attitude angle of each video frame in the K consecutive video frames, the weights of the observed attitude angle and the predicted attitude angle are determined respectively. Based on the determined weights, the observed attitude angle and the predicted attitude angle of each video frame in the K consecutive video frames are weighted and summed to obtain the final attitude angle of each video frame in the K consecutive video frames. Using the final pose angle of each of the K consecutive video frames, an affine compensation matrix is ​​constructed to obtain K affine compensation matrices that correspond one-to-one with the K consecutive video frames. The inverse matrix of each of the K affine compensation matrices is then determined to obtain K alignment matrices. Using the K alignment matrices, affine transformations are performed on the K consecutive video frames that correspond one-to-one, resulting in K consecutive aligned frames, which are then used as K reference frames.

8. The indoor mid-to-long-range camera heart rate detection method under environmental interference according to claim 1, characterized in that, The process of determining the ROI signal based on the face ROI region of each of the K consecutive aligned frames, and performing non-rigid noise filtering on the ROI signal to obtain the pulse wave signal includes: The average pixel values ​​of the R, G, and B channels of the face ROI region in each of the K consecutive aligned frames are calculated to obtain a three-channel discrete pixel mean sequence composed of the average pixel values ​​of the R, G, and B channels of the K consecutive aligned frames, and the three-channel discrete pixel mean sequence is used as the ROI signal. Using a preset time length as the segmentation benchmark, the ROI signal is segmented according to a preset overlap rate to obtain multiple segmented ROI signals. After normalizing each segmented ROI signal, multiple normalized segmented ROI signals are obtained. After sequentially splicing the multiple normalized segmented ROI signals into a single signal, a third-order Butterworth bandpass filter is used to filter the single signal to suppress non-heart rate frequency band noise, resulting in a denoised single signal. The denoised segment of the signal is processed using the POS method to obtain a pulse wave signal.

9. The method for detecting heart rate in an indoor mid-to-long-range camera under environmental interference as described in claim 5, characterized in that, The process of determining the quality score Q corresponding to the current sliding window based on the pulse wave signal and the face detection results of each of the K consecutive video frames includes: Calculate the average signal-to-noise ratio (SNR) of the pulse wave signal; Based on the face detection results of each of the K consecutive video frames, calculate the average motion of the K consecutive video frames; The average signal-to-noise ratio (SNR), the average motion, and the mean of normalized face scale and the mean of face sharpness of the video frame sequence within the current sliding window are weighted and summed to obtain the sum value. Calculate the preset coefficients The product of the average motion value and the motion value is obtained. The difference between the sum and the product is calculated to obtain the quality score Qt corresponding to the current sliding window.

10. An indoor mid-to-long-range camera heart rate detection system resistant to environmental interference, characterized in that, include: At least one camera, each camera is used to collect one video stream of data and transmit it to the host computer; A host computer is connected to each of the at least one camera and is used to execute the indoor mid-to-long-range camera heart rate detection method with resistance to environmental interference as described in any one of claims 1 to 9.