Non-contact Respiratory Rate Monitoring Method and System Based on Dual-Spectrum Facial Videos
Through the dual-spectral facial video monitoring method, combined with thermal infrared and visible light video, nuclear-related filtering and affine transformation technology are used to solve the problem of environmental impact of a single channel data source, and high-precision respiratory rate detection and abnormal breathing rate detection are achieved.
Patent Information
- Application Number
- CN202311281289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-09-28
AI Technical Summary
In the existing non-contact breathing rate monitoring methods, a single channel data source is susceptible to ambient light and motion artifacts, resulting in poor quality of respiratory signals and low accuracy.
The dual-spectral facial video monitoring method is used, combined with thermal infrared video and visible light video, and the kernel-related filtering algorithm and affine transformation technology are used to locate and signal fusion of the region of interest in the nose, thereby improving the accuracy of respiratory frequency detection through signal processing.
The accuracy of respiratory rate detection is improved, and the visible light video can be used to map and signal supplement when thermal infrared video cannot locate the nose ROI area, realizing the detection of abnormal breathing rates.
Smart Images

Figure CN117357093B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of respiratory rate monitoring, and particularly to a non-contact respiratory rate monitoring method and system based on dual-spectrum facial videos. Background Art
[0002] With the development of information technology, in the prior art, the respiratory rate of a user is often extracted through thermal infrared videos or visible light videos. Specifically: based on the periodic temperature change in the nose area when the user breathes, the respiratory rate can thus be analyzed and obtained.
[0003] However, the above technical solutions are extremely susceptible to factors such as environmental light and motion artifacts, making the respiratory signals extracted through a single-channel data source (such as visible light videos) contain a large amount of noise, resulting in poor quality of the extracted respiratory signals, and thus leading to a low accuracy in respiratory rate monitoring.
[0004] Therefore, there is an urgent need for a non-contact respiratory rate monitoring method to solve the problem of low accuracy in non-contact respiratory rate monitoring in the prior art. Summary of the Invention
[0005] (I) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the prior art, the present invention provides a non-contact respiratory rate monitoring method and system based on dual-spectrum facial videos, which solves the problem of low accuracy in non-contact respiratory rate monitoring in the prior art.
[0007] (II) Technical Solutions
[0008] To achieve the above object, the present invention is realized through the following technical solutions:
[0009] In the first aspect of the present invention, a non-contact respiratory rate monitoring method based on dual-spectrum facial videos is provided. The method includes:
[0010] Simultaneously acquire a thermal infrared video and a visible light video containing the face of the user to be measured;
[0011] Based on the kernel correlation filtering algorithm, track and locate the region of interest of the nose of the user to be measured in the thermal infrared video, extract the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and use the thermal infrared image frames in the thermal infrared video where the region of interest of the nose is not tracked and located as the first thermal infrared image frames;
[0012] Extract the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal; and select the visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame;
[0013] Using the method of affine transformation, map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, obtain the region of interest of the nose of the user to be measured in the first visible light image;
[0014] Extract the respiratory signal corresponding to the second thermal infrared image as the second respiratory signal;
[0015] Fuse the first respiratory signal and the second respiratory signal, and use the signal after signal fusion as the third respiratory signal;
[0016] Based on a preset signal processing method, perform signal processing on the third respiratory signal to obtain a respiratory signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing;
[0017] Based on the respiratory signal curve, obtain the respiratory rate of the user to be measured.
[0018] Optionally, using the method of affine transformation to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image, includes:
[0019] Select a visible light image frame with the same timestamp as the second thermal infrared image frame in the visible light video as the second visible light image frame, and use the second thermal infrared image frame and the second visible light image with the same timestamp as a preset image pair; wherein, the second thermal infrared image frame represents the image frame containing the first thermal infrared image in the thermal infrared video;
[0020] Based on the mediapipe tool, obtain the coordinates of the upper left corner point A(x1, y1) and the lower right corner point B(x2, y2) of the region of interest of the nose in the second visible light image frame of the preset image pair; and based on the kernel correlation filtering algorithm, obtain the coordinates of the upper left corner point A′(u1, v1) and the lower right corner point B′(u2, v2) of the region of interest of the nose in the second thermal infrared image frame;
[0021] Based on the coordinates of the upper left corner points A, A′ and the lower right corner points B, B′ in the image pair, determine the affine transformation matrix of the same coordinates;
[0022] Based on the affine transformation matrix, perform affine transformation processing on the coordinates of the upper left corner point A and the lower right corner point B in the first visible light image frame to obtain the coordinates of the upper left corner point A′ and the lower right corner point B′ in the first thermal infrared image frame;
[0023] Based on the coordinates of the upper left corner point A' and the lower right corner point B' in the first thermal infrared image frame, extract the image of the nose region of interest in the first thermal infrared image frame as the second thermal infrared image.
[0024] Optionally, extracting the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal includes:
[0025] Convert the first thermal infrared image into a grayscale image as the first grayscale image;
[0026] Calculate the pixel average value of the first grayscale image as the first respiratory signal.
[0027] Optionally, performing signal processing on the third respiratory signal based on a preset signal processing method to obtain a respiratory signal curve includes:
[0028] Perform detrending processing on the third respiratory signal, and use the detrended third respiratory signal as the fourth respiratory signal;
[0029] Perform normalization processing on the fourth respiratory signal, and use the normalized fourth respiratory signal as the fifth respiratory signal;
[0030] Set the data form of the experimental data corresponding to the fifth respiratory signal to And perform fitting processing on the experimental data to obtain the fitted Gaussian function as the respiratory signal curve; where T i represents the abscissa time in the respiratory curve; represents the ordinate respiratory value in the respiratory curve;
[0031]
[0032] a, c, and b respectively represent the parameters to be solved, and their physical meanings are the peak height, peak position, and half-width information of the Gaussian curve; b0, b1, and b2 respectively represent the intermediate parameters in the process of solving the parameters a, c, and b to be solved.
[0033] Optionally, the method further includes:
[0034] When consecutive image frames of the thermal infrared video are missing and the missing image frames cannot be obtained through affine transformation, if the number of the missing image frames is lower than the first preset value, then perform Gaussian fitting processing on the respiratory signal curve;
[0035] If the number of the missing image frames is not lower than the first preset value, then perform translation processing on the respiratory signal curve to supplement the missing respiratory signals in the respiratory signal.
[0036] Optionally, the translation processing of the respiration signal curve includes:
[0037] Obtain the total number of frames of consecutive missing image frames, and the position of the first missing frame or the position of the last missing frame in the consecutive missing image frames;
[0038] Determine the time length of the continuously missing respiration signals based on the total number of frames as the first time length; and determine the starting time point of the continuously missing respiration signals based on the position of the first missing frame or determine the ending time point of the continuously missing respiration signals based on the position of the last missing frame;
[0039] Obtain the continuous respiration curve before the starting time point and with a time length of the first time length as the first curve; or obtain the continuous respiration curve after the starting time point and with a time length of the first time length as the second curve;
[0040] Translate the first curve backward by a length corresponding to the first time length according to the time sequence, or translate the second curve forward by a length corresponding to the first time length according to the time sequence to supplement the missing respiration signals in the respiration signal.
[0041] Optionally, obtaining the respiration rate of the user under test based on the respiration signal curve includes:
[0042] Obtain the number of wave peaks PN of the respiration signal curve; wherein, the first peak and the last peak are respectively marked as FP and LP, and the average distance between two adjacent peaks is ADP;
[0043] Calculate the respiration rate of the user under test based on a preset formula; wherein, the preset formula is:
[0044]
[0045] wherein, PN represents the total number of frames of the thermal infrared video obtained within one minute; RR represents the respiration rate of the user under test.
[0046] Optionally, the method further includes:
[0047] Set a time interval between two consecutive respiration rate monitors.
[0048] In a second aspect of the present invention, a non-contact respiration rate monitoring system based on a dual-spectrum facial video is provided, and the system includes:
[0049] A first acquisition module, configured to simultaneously acquire a thermal infrared video and a visible light video including the face of the user under test;
[0050] A tracking and positioning module, which is used to track and position the region of interest of the nose of the user to be measured in the thermal infrared video based on the kernel correlation filtering algorithm, extract the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and use the thermal infrared image frames in the thermal infrared video where the region of interest of the nose is not tracked and positioned as the first thermal infrared image frames;
[0051] A first extraction module, which is used to extract the respiration signal corresponding to the first thermal infrared image as the first respiration signal; and select the visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame;
[0052] An affine transformation module, which is used to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame by using the method of affine transformation, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, the region of interest of the nose of the user to be measured in the first visible light image is obtained;
[0053] A second extraction module, which is used to extract the respiration signal corresponding to the second thermal infrared image as the second respiration signal;
[0054] A signal fusion module, which is used to perform signal fusion on the first respiration signal and the second respiration signal, and use the signal after signal fusion as the third respiration signal;
[0055] A signal processing module, which is used to perform signal processing on the third respiration signal based on a preset signal processing method to obtain a respiration signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing;
[0056] A second acquisition module, which is used to obtain the respiration rate of the user to be measured based on the respiration signal curve.
[0057] (III) Advantageous Effects
[0058] The present invention provides a non-contact respiration rate monitoring method and system based on dual-spectrum facial videos. Compared with the prior art, it has the following advantageous effects:
[0059] The present invention provides a non-contact respiratory rate monitoring method and system based on dual-spectrum facial videos. The method includes: simultaneously acquiring a thermal infrared video and a visible light video containing the face of the user to be measured; tracking and positioning the region of interest of the nose of the user to be measured in the thermal infrared video based on the kernel correlation filtering algorithm, and extracting the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and using the thermal infrared image frames in the thermal infrared video where the region of interest of the nose is not tracked and positioned as the first thermal infrared image frames; extracting the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal; and selecting the visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame; using the method of affine transformation to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extracting the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, the region of interest of the nose of the user to be measured in the first visible light image is acquired; extracting the respiratory signal corresponding to the second thermal infrared image as the second respiratory signal; fusing the first respiratory signal and the second respiratory signal, and using the signal after signal fusion as the third respiratory signal; based on a preset signal processing method, performing signal processing on the third respiratory signal to obtain a respiratory signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing; based on the respiratory signal curve, obtaining the respiratory rate of the user to be measured. Based on the above processing, by performing dual-channel nose ROI (Region Of Interest) region detection on the facial video of the user to be measured, the ROI region detection becomes more robust. Specifically: when the ROI region is normally detected in the thermal infrared video, the respiratory signal can be normally extracted from the detected ROI region; when the ROI region is not located in the thermal infrared video, the ROI region of the visible light video can be accurately mapped to the thermal infrared video through dual-light registration, and then the respiratory signal is extracted, and the respiratory signal is spliced and fused with the respiratory signal normally extracted from the ROI region of the thermal infrared video, effectively improving the accuracy of the obtained respiratory rate.
[0060] Meanwhile, after detrending and normalizing the respiratory signal extracted from the ROI, the method of Gaussian fitting is used to smooth the data point set of the respiratory signal, obtaining a smooth and continuous fitting curve, and then through peak detection, the detection of abnormal respiratory rates is realized. Description of the Drawings
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0062] Figure 1 Flowchart of a non-contact respiratory rate monitoring method based on dual-spectrum facial video provided by an embodiment of the present invention;
[0063] Figure 2 Schematic diagram of the process of an affine transformation provided by an embodiment of the present invention;
[0064] Figure 3 Schematic diagram of the result of Gaussian fitting denoising provided by an embodiment of the present invention;
[0065] Figure 4 Schematic diagram of the process of non-contact respiratory rate monitoring provided by an embodiment of the present invention;
[0066] Figure 5 Architecture diagram of a non-contact respiratory rate monitoring system based on dual-spectrum facial video provided by an embodiment of the present invention;
[0067] Figure 6 Architecture diagram of a non-contact respiratory rate monitoring system provided by an embodiment of the present invention. Detailed implementation manners
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0069] The embodiments of the present invention provide a non-contact respiratory rate monitoring method and system based on dual-spectrum facial video, solving the problem of low accuracy in non-contact respiratory rate monitoring in the prior art, and achieving the accuracy of effectively obtaining the respiratory rate and detecting abnormal respiratory rates.
[0070] To better understand the above technical solutions, the following will elaborate on the above technical solutions in combination with the accompanying drawings of the specification and specific implementation manners.
[0071] The technical solutions in the embodiments of the present invention for solving the above technical problems have the following general idea:
[0072] There are mainly two current non-contact respiratory rate monitoring technical solutions:
[0073] 1. A non-contact respiratory rate monitoring method that directly analyzes the visible light video of the face and uses it as a single signal source.
[0074] 2. A non-contact respiratory rate monitoring method that directly analyzes the thermal infrared video of the face and uses it as a single signal source.
[0075] Since the above two mainstream technical solutions are extremely susceptible to environmental light and motion artifacts, the respiratory signal extracted through a single-channel data source (such as visible light video) may contain a large amount of noise, resulting in poor quality of the extracted respiratory signal and a decrease in the accuracy of respiratory rate estimation.
[0076] A few existing technical solutions use a dual-light registration method to assist in locating and segmenting the alar region of the thermal infrared video with the help of the visible light video, and then extract the respiratory signal. However, since the face detection algorithm may fail, it will lead to the absence of the respiratory signal and the accuracy of respiratory rate estimation is not high enough.
[0077] There is also a common defect in the three methods: almost all three technical solutions use the FFT method to process the respiratory signal. When the FFT method processes the signal, it will lose some respiratory information due to band-pass filtering. That is: when processing the signal, the fast Fourier transform FFT is usually used to denoise the signal within the normal frequency range [0.15, 0.40] Hz, but the signal outside this range is ignored.
[0078] The existing technical path for respiratory rate extraction has the problem of ignoring abnormal respiratory values, which makes there is a large room for improvement in the robustness of the existing technology.
[0079] To solve the above problems, the present invention provides a non-contact respiratory rate monitoring method based on dual-spectrum facial video. See Figure 1 , Figure 1 which is a schematic flowchart of a non-contact respiratory rate monitoring method based on dual-spectrum facial video provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0080] S1. Simultaneously obtain a thermal infrared video and a visible light video containing the face of the user to be measured;
[0081] S2. Based on the kernel correlation filtering algorithm, track and locate the region of interest of the nose of the user to be measured in the thermal infrared video, extract the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and use the thermal infrared image frame in the thermal infrared video where the region of interest of the nose is not tracked and located as the first thermal infrared image frame;
[0082] S3. Extract the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal; and select the visible light image frame with the same timestamp as that of the first thermal infrared image frame in the visible light video as the first visible light image frame.
[0083] S4. Use the method of affine transformation to map the region of interest of the nose of the user under test in the first visible light image frame to the first thermal infrared image frame, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, obtain the region of interest of the nose of the user under test in the first visible light image.
[0084] S5. Extract the respiratory signal corresponding to the second thermal infrared image as the second respiratory signal.
[0085] S6. Perform signal fusion on the first respiratory signal and the second respiratory signal, and use the signal after signal fusion as the third respiratory signal.
[0086] S7. Based on the preset signal processing method, perform signal processing on the third respiratory signal to obtain a respiratory signal curve; wherein, the preset signal processing method includes detrending processing, normalization processing, and Gaussian fitting processing.
[0087] S8. Based on the respiratory signal curve, obtain the respiratory rate of the user under test.
[0088] Based on the above processing, by performing dual-channel nose ROI (Region Of Interest) region detection on the facial video of the user under test, the ROI region detection becomes more robust. When the nose ROI region can be normally detected in the thermal infrared video, the respiratory signal can be normally extracted; when the nose ROI region is not located in the thermal infrared video, the ROI region of the visible light video can be accurately mapped to the thermal infrared video through dual-light registration, and then the respiratory signal is extracted and spliced and fused with the respiratory signal normally extracted from the ROI region of the thermal infrared video, effectively improving the accuracy of respiratory rate detection.
[0089] At the same time, after performing detrending and normalization processing on the extracted respiratory signal, the Gaussian fitting method is used to perform fitting processing on the data point set of the respiratory signal to obtain a smooth and continuous fitting curve, and then through peak detection, the detection of abnormal respiratory rates is realized.
[0090] Regarding step S1, the start times of the thermal infrared video and the visible light video are the same as the camera frame numbers, which can be understood that each frame image in the thermal infrared video and the visible light video corresponds to each other, and the timestamps or frame number sequences are the same.
[0091] Specifically, the user to be measured takes a picture at a distance of 0.5m - 1m from the visible light camera and the thermal infrared camera, so as to ensure that the above cameras can capture the head and shoulders of the user to be measured. After the shooting starts, the acquisition of the dual - light video will start simultaneously and continue, serving as the data source for subsequent respiratory rate detection.
[0092] For step S2, the KCF method (Kernel Correlation Filter) based on object tracking is used to track and locate the nose ROI region in the thermal infrared video to obtain the first thermal infrared image.
[0093] Specifically, the acquisition process of the above - mentioned first thermal infrared image can be divided into:
[0094] (1) Initialize the filter
[0095] The image at the initial moment in the thermal infrared video is recorded as the first frame. Based on the Mediapipe model, the face in the thermal infrared video is detected, which is used as the accurate nose ROI region positioning of the first frame, and the most ideal set response is Y1. Then, the first frame is sampled to obtain X1, and then mapped to a high - dimensional space to obtain φ(X1). The circulant matrix is used to achieve circular sampling to obtain C(φ(X i ))), and the initialization filter W1 is obtained through calculation according to Y1 and C(φ(X i ))). Among them, Y1 represents the response value of the first frame, X1 represents the positive sample obtained by sampling, and C(φ(X1)) represents the sample set obtained by circular sampling.
[0096] (2) Perform fast detection
[0097] After the second frame and subsequent frames in the thermal infrared video are input, the tracking position of the i - th frame is sampled to obtain X(φ(Z i ))), and the response value Y i-1 The one with the largest response value is i Located as the face ROI region of the i - th frame. Among them, C(φ(Z i )) represents the sample set obtained by circular sampling of the i - th frame, and Y i Represents the response value of the i - th frame.
[0098] (3) Tracker update
[0099] After the fast positioning of the i - th frame is achieved, circular sampling is performed on the i - th frame to obtain C(φ(X i ))), and the parameters of the filter are updated by Y i And C(φ(X i )) to generate a new tracker model.
[0100] The above process is cycled to achieve high-speed and accurate face tracking. Among them, C(φ(X i )) represents the sample set obtained by cyclic sampling of the i-th frame, and Y i represents the response value of the i-th frame.
[0101] In some embodiments, the process of obtaining the region of interest of the nose of the user to be measured in the visible light video based on the face detection algorithm and facial feature points includes:
[0102] First, in order to prevent the noise of the background from affecting the quality of the signal, the face in each frame of the visible light video is cropped, and the face is segmented from the background picture according to the face detection algorithm to obtain a clear face image. Specifically, there are a total of M feature points in the visible light face, and the abscissa of the i-th key feature point of the face is m i , and the ordinate is n i . The coordinates of the upper left corner point and the lower right corner point of the rectangular frame of the entire face can be determined as A(m min ,n min ), B(m max ,n max ), then
[0103] m min =min(M), n min =min(N)
[0104] m max =max(M), n max =max(N)
[0105] where M = (m1, m2,..., m j ), N = (n1, n2,..., n j ), m min represents the minimum abscissa value of the key feature point, min(M) represents taking the minimum abscissa value, n min represents the minimum ordinate value of the key feature point, min(N) represents taking the minimum ordinate value, m j represents the abscissa of the i-th key feature point, and n j represents the ordinate of the i-th key feature point.
[0106] Then, the facial feature points of the nose, eyes, mouth, chin, etc. of the face in the visible light video are accurately calibrated, so as to accurately locate the ROI region of interest of the face using the facial feature points, and the ROI regions of each frame in the visible light video are segmented to obtain an ROI image set.
[0107] In some embodiments, when the nose region of interest of the user to be measured in the thermal infrared video cannot be obtained based on the kernel correlation filtering algorithm, the nose region of interest of the user to be measured in the visible light video is mapped to the thermal infrared video by means of affine transformation to obtain a second thermal infrared image.
[0108] The above steps may include the following:
[0109] Step 1: Select a visible light image frame in the visible light video with the same timestamp as the second thermal infrared image frame as the second visible light image frame, and use the second thermal infrared image frame and the second visible light image with the same timestamp as a preset image pair; wherein, the second thermal infrared image frame represents the image frame containing the first thermal infrared image in the thermal infrared video.
[0110] Step 2: Based on the mediapipe tool, obtain the coordinates of the upper left corner point A(x1, y1) and the lower right corner point B(x2, y2) of the nose region of interest in the second visible light image frame in the preset image pair; and obtain the coordinates of the upper left corner point A′(u1, v1) and the lower right corner point B′(u2, v2) of the nose region of interest in the second thermal infrared image frame based on the kernel correlation filtering algorithm.
[0111] Step 3: Based on the coordinates of the upper left corner points A, A′ and the lower right corner points B, B′ in the image pair, determine the affine transformation matrix of the same coordinates.
[0112] Step 4: Based on the affine transformation matrix, perform affine transformation processing on the coordinates of the upper left corner point A and the lower right corner point B in the first visible light image frame to obtain the coordinates of the upper left corner point A′ and the lower right corner point B′ in the first thermal infrared image frame.
[0113] Step 5: Based on the coordinates of the upper left corner point A′ and the lower right corner point B′ in the first thermal infrared image frame, extract the image of the nose region of interest in the first thermal infrared image frame as the second thermal infrared image.
[0114] Since the segmented visible light ROI region cannot be directly mapped to the thermal infrared video, it is necessary to perform consistent conversion on the dual light (i.e., the visible light video and the thermal infrared video in the present invention) so that the feature points of the dual light faces can correspond one by one, and then locate the ROI region of the face in the thermal infrared video through the feature points of the visible light face, thereby ensuring the extraction quality of the breathing signal. Specifically, the face detected in the visible light video and the located feature points are accurately mapped to the thermal infrared video through affine transformation, so as to accurately locate the ROI region of the thermal infrared video.
[0115] In some embodiments, the above-mentioned affine transformation process can be divided into: First, use the mediapipe tool to perform face detection and nose ROI region positioning on the visible light videos existing in the training set, and record the coordinates of the upper left point A(x1, y1) and the lower right point B(x2, y2) of the nose ROI region in multiple frames. Then, based on the KCF object tracking algorithm, perform nose ROI region tracking and positioning on the corresponding thermal infrared videos of the same person existing in the training set, and record the coordinates of the upper left point A′(u1, v1) and the lower right point B′(u2, v2) of the nose ROI region. Similarly, the two-point coordinates of the dual-spectrum ROI regions in multiple corresponding frames can be obtained.
[0116] Then, calculate the affine matrix M.
[0117] Affine transformation (Affine Transformation or Affine Map) is a method of linearly transforming and mapping the two-dimensional coordinates (xi, yi) of visible light to the two-dimensional coordinates (ui, vi) of thermal infrared. The mathematical expression form is as follows:
[0118]
[0119] The corresponding homogeneous coordinate matrix representation form is:
[0120]
[0121] Among them, the affine transformation matrix
[0122] Calculate and fit M with multiple groups of visible light coordinate points (xi, yi), (xi+1, yi+1) and thermal infrared coordinate points (ui, vi), (ui+1, vi+1) obtained from the same frame, and obtain an accurate affine transformation matrix M. Among them, (xi, yi), (xi+1, yi+1) respectively represent visible light coordinate points; (ui, vi), (ui+1, vi+1) respectively represent thermal infrared coordinate points at the same position in the corresponding frame.
[0123] Finally, obtain the coordinates (x1, y1) of the upper left corner and the coordinates (x2, y2) of the lower right corner of the ROI region of the visible light video, and calculate the corresponding two-point coordinates (x1′, y1′) and (x2′, y2′) of the thermal infrared based on the affine transformation matrix M obtained by the above fitting, so as to locate the ROI region of the thermal infrared video and realize the registration of the dual-light ROI regions. See Figure 2 , Figure 2 which is a schematic flowchart of an affine transformation provided by an embodiment of the present invention.
[0124] For step S3, the extraction process of the respiration signal can be divided into:
[0125] Step Six: Convert the first thermal infrared image into a grayscale image, which serves as the first grayscale image;
[0126] Step Seven: Calculate the pixel average value of the first grayscale image, which serves as the first respiratory signal.
[0127] Similarly, the process of obtaining the second respiratory signal in step S5 is also as described in the content of Step Six and Step Seven.
[0128] Since during the breathing process of a person, there are obvious differences in the gas temperature near the nose, and the temperature change of the gas near the nose is periodic. Therefore, the image of the nose ROI area in the normally captured thermal infrared video can be converted into a grayscale image, and then the pixel average value of the thermal infrared nose area can be calculated to obtain the respiratory signal.
[0129] Assume that the image matrices of the R, G, and B channels of the nose region of interest in the thermal infrared video are R nose , G nose , B nose , respectively. The calculation method for extracting the respiratory signal in the i-th frame of the picture is:
[0130] Signal_nose i = mean(Gray)
[0131] Gray = 0.3R nose + 0.59G nose + 0.11B nose
[0132] Among them, R nose , G nose , B nose respectively represent the image matrices of the R, G, and B channels, Signal_nose i represents the extracted initial respiratory signal, mean(Gray) represents the function of taking the average value of the matrix Gray, and Gray indicates that Gray is the grayscale image matrix of the nose ROI area.
[0133] Assume that the respiratory signal extracted from the thermal infrared channel is TR, then:
[0134] TR = [SN1, SN2, …, SN i
[0135] Among them, TR represents the first set of initial respiratory signals, and SN i represents the extracted initial respiratory signal value.
[0136] In one implementation, the image frames in the thermal infrared video where no respiratory signal is extracted, that is, no ROI area is detected, can be marked with IndUN.
[0137] For the frames marked as IndUN, that is, the frames in the thermal infrared video where the ROI region cannot be normally located, the respiration signal is replaced by the respiration signal extracted by dual-light fusion using the corresponding frames of the visible light video. If the ROI region of the visible light is normally located, the dual-light fusion method is used to extract the respiration signal at this time. The feature points of the ROI region obtained from the visible light are mapped to the thermal infrared through the affine transformation matrix. After the ROI region of the thermal infrared is located, the image of the ROI region is converted into a grayscale image, and then the pixel average value of the nose region of the thermal infrared is calculated to obtain the respiration signal.
[0138] VR = [SN’1, SN’2, …, SN’ i
[0139] Wherein, VR represents the second initial respiration signal set, and SN’ i represents the extracted initial respiration signal value.
[0140] The respiration signals of the frames marked with Ind (denoted as the Ind1, Ind2, …, Ind i frames) are replaced by the values of the corresponding frames in VR:
[0141] UTR = [SN’ Ind1 , SN’ Ind2 , …, SN’ Indi
[0142] Wherein, UTR represents the respiration signal missing the first initial respiration signal extracted from the second initial respiration signal set, and SN’ Indi represents the extracted initial respiration signal value.
[0143] In some embodiments, the respiration signal fusion in step S6 may include the following:
[0144] The respiration signals of the marked frames extracted by dual-light fusion and the respiration signals normally extracted by thermal infrared are spliced and fused to form the respiration signal RR0.
[0145] RR0 = [TR, UTR].
[0146] In some embodiments, step S7 may include the following:
[0147] S701. Perform detrending processing on the third respiration signal, and use the detrended third respiration signal as the fourth respiration signal.
[0148] S702. Perform normalization processing on the fourth respiration signal, and use the normalized fourth respiration signal as the fifth respiration signal.
[0149] S703. Set the data format of the experimental data corresponding to the fifth respiration signal to and perform fitting processing on the experimental data to obtain the fitted Gaussian function as the respiration signal curve.
[0150] where T i represents the abscissa time in the respiration curve; represents the ordinate respiration value in the respiration curve;
[0151]
[0152] a, c, and b respectively represent the parameters to be solved, and the physical meanings they represent are the peak height, peak position, and half-width information of the Gaussian curve; b0, b1, and b2 respectively represent the intermediate parameters in the process of solving the parameters a, c, and b to be solved, and this solving problem can be simplified to a polynomial problem.
[0153] In one implementation, for step S701, since the initially collected respiration signal itself has a certain degree of oscillation, and the low-frequency component will also affect the nasal temperature signal; at the same time, due to the instability of the signal acquisition device and its susceptibility to interference from the surrounding environment, zero drift of the signal will occur, which often causes the signal to deviate from the baseline, and even the magnitude of the deviation from the baseline will change with time. The entire process of the deviation from the baseline changing with time is called the trend term of the signal. And the trend term will affect the quality and correctness of the signal, so the trend term needs to be eliminated.
[0154] Then, perform normalization processing on the respiration signal after detrending, as shown in the following formula:
[0155]
[0156] where σ represents the standard deviation, μ is the mean of the original signal, RR Detrend represents the signal before normalization processing, and RR SD is the signal after normalization processing.
[0157] Finally, perform Gaussian fitting on the processed signal data RR SD to use the Gaussian function to approximate the discrete data point set and achieve smoothing of the data to obtain a smooth respiration curve.
[0158] Specifically, the process of Gaussian fitting can be divided into the following content:
[0159] Assume that the normalized signal data is described by the Gaussian function:
[0160]
[0161] Among them, a, c, and b respectively represent the physical meanings of the peak height, peak position, and half-width information of the Gaussian curve.
[0162] Take the natural logarithm on both sides of the equation to transform it into:
[0163]
[0164]
[0165] Let
[0166]
[0167] We can obtain
[0168] Considering data and measurement errors, it is expressed in matrix form as follows
[0169]
[0170] Briefly recorded as:
[0171] Z n×1 = X n×3 B 3×1 + E n×1
[0172] Without considering the influence of the total range error E, according to the least squares principle, the generalized least squares solution of the matrix B composed of the fitting constants b0, b1, and b2 can be obtained as:
[0173] B = (X T X) -1 X T Z
[0174] The evaluation parameters a, b, and c can be obtained from the above formula:
[0175]
[0176] Substitute the obtained a, b, and c into the above formula, and the Gaussian function fitted from the experimental data can be obtained. Finally, the final respiratory signal RR is calculated through the Gaussian function obtained by fitting L . See Figure 3 , Figure 3 which is a schematic diagram of the result of Gaussian fitting denoising provided by the present invention.
[0177] Based on the above Gaussian fitting process, the obtained breathing curve can be made closer to the actual situation of the real user. For example, in cases of apnea or very rapid breathing, they can all be obtained through the above breathing curve analysis. However, existing filtering methods transform the time-domain signal to the frequency domain based on the fast Fourier transform (FFT), and then use a band-pass filter to retain the data with frequency values between 0.15 - 0.4 Hz and set the data values outside 0.15 - 0.4 Hz to zero. They can only detect the breathing values of normal people and cannot detect abnormal situations.
[0178] In some embodiments, when there are consecutive missing image frames in the first set of thermal infrared images and the missing image frames cannot be obtained through affine transformation, the method further includes the following steps:
[0179] Step a: If the number of the missing image frames is lower than a first preset value, then perform Gaussian fitting processing on the breathing signal curve. Among them, the first preset value is related to the frame rate of the visible light video or the thermal infrared video. The specific range can be set to 1 - 5.
[0180] Step b: If the number of the missing image frames is not lower than the first preset value, then perform translation processing on the breathing signal curve to supplement the missing breathing signals in the breathing signal.
[0181] In one implementation, step b may include the following content:
[0182] Step b01: Obtain the total number of frames of the continuously missing image frames, and the position of the first missing frame or the position of the last missing frame in the continuously missing image frames;
[0183] Step b02: Determine the time length of the continuously missing breathing signals based on the total number of frames as the first time length; and determine the starting time point of the continuously missing breathing signals based on the position of the first missing frame or determine the ending time point of the continuously missing breathing signals based on the position of the last missing frame;
[0184] Step b03: Obtain the continuous breathing curve with a time length of the first time length before the starting time point as the first curve; or obtain the continuous breathing curve with a time length of the first time length after the starting time point as the second curve;
[0185] Step b04: Translate the first curve backward by a length corresponding to the first time length in chronological order, or translate the second curve forward by a length corresponding to the first time length in chronological order to supplement the missing breathing signals in the breathing signal.
[0186] When the facial contour in the thermal infrared video is blurred and the key points of the region of interest cannot be located, or when the tracking target suddenly disappears and the target tracking algorithm cannot locate the key feature points of the region of interest of the thermal infrared face. In addition, when the face cannot be detected in the visible light video either, that is, the key feature points of the region of interest of the thermal infrared face cannot be located through the affine transformation matrix. In the above several situations, the loss or interruption of the normal breathing signal will occur.
[0187] Therefore, when these situations occur, the loss of the breathing signal in a few frames can be compensated by Gaussian fitting. When there is a loss of the breathing signal in more frames, record the subscript of the frame where the normal breathing signal loss starts as ind1, and the ending frame as ind2. When there is a loss of the breathing signal in more frames, the normal breathing signals before and after the abnormal segment (that is, the frames corresponding to the missing breathing signal in the present invention) can be translated to obtain the normal breathing signal waveforms of adjacent segments.
[0188] Specifically, assume that f(x ind1-t ) and f(x ind2+t ) represent the waveforms of [ind1 - t, ind1] and [ind2, ind2 + t] before and after the starting frame of the abnormal breathing signal. The waveforms of the front and rear parts of the abnormal breathing signal segment can be translated left or right by t units. Then, the translated waveforms can be represented by f(x ind1-t+t ) or f(x ind2+t-t ). Among them, ind1 - t represents the abscissa value of the left endpoint of the abnormal breathing signal segment, ind2 + t represents the abscissa value of the right endpoint of the abnormal breathing signal segment, t represents the translation unit value, and f(x ind1-t+t ), f(x ind2+t-t ) represent the translated waveforms.
[0189] In some embodiments, step S8 includes:
[0190] S801. Obtain the number of wave peaks PN of the breathing signal curve; among them, the first peak and the last peak are respectively marked as FP and LP, and the average distance between two adjacent peaks is ADP;
[0191] S802. Calculate the breathing frequency of the user to be measured based on a preset formula; among them, the preset formula is:
[0192]
[0193] where PN represents the total number of frames of the thermal infrared video obtained within one minute; RR represents the breathing frequency of the user to be measured. See Figure 4 , Figure 4A flowchart of non-contact respiratory rate monitoring provided by an embodiment of the present invention.
[0194] In some embodiments, a time interval is set between two consecutive respiratory rate monitors.
[0195] Among them, the time length of one detection of the respiratory rate can be set to one minute. To ensure the accuracy of the respiratory rate calculation result, the time interval between two detections needs to be more than 10 seconds. The time interval N (N can be a positive integer greater than 10) is affected by factors such as the hardware environment and the camera frame rate, and the specific value of N can be modified and updated.
[0196] The present invention also provides a non-contact respiratory rate monitoring system based on dual-spectrum facial videos. See Figure 5 , Figure 5 An architecture diagram of a non-contact respiratory rate monitoring system based on dual-spectrum facial videos provided by an embodiment of the present invention. As shown in Figure 5 shown, the system includes:
[0197] A first acquisition module 501, configured to simultaneously acquire a thermal infrared video and a visible light video including the face of the user to be measured;
[0198] A tracking and positioning module 502, configured to track and position the region of interest of the nose of the user to be measured in the thermal infrared video based on the kernel correlation filtering algorithm, extract the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and use the thermal infrared image frame in the thermal infrared video where the region of interest of the nose is not tracked and positioned as the first thermal infrared image frame;
[0199] A first extraction module 503, configured to extract the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal; and select a visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame;
[0200] An affine transformation module 504, configured to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame by using the method of affine transformation, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; among them, based on the face detection algorithm and facial feature points, the region of interest of the nose of the user to be measured in the first visible light image is obtained;
[0201] A second extraction module 505, configured to extract the respiratory signal corresponding to the second thermal infrared image as the second respiratory signal;
[0202] A signal fusion module 506, configured to perform signal fusion on the first respiratory signal and the second respiratory signal, and use the signal after signal fusion as the third respiratory signal;
[0203] A signal processing module 507, configured to perform signal processing on the third respiration signal based on a preset signal processing method to obtain a respiration signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing;
[0204] A second acquisition module 508, configured to acquire the respiration rate of the user to be measured based on the respiration signal curve.
[0205] It can be understood that the non-contact respiration rate monitoring system based on dual-spectrum facial video provided in the embodiments of the present invention corresponds to the above-mentioned non-contact respiration rate monitoring method based on dual-spectrum facial video. The explanations, examples, beneficial effects, etc. of the relevant content can refer to the corresponding content in the non-contact respiration rate monitoring method based on dual-spectrum facial video, and will not be elaborated here.
[0206] In one implementation, in a specific application scenario, a non-contact respiration rate monitoring system can be constructed to execute the above-mentioned non-contact respiration rate monitoring method based on dual-spectrum facial video. To improve the openness and maintainability of the above system, the system can be divided into several modules with high cohesion and low coupling. Among them, it includes an acquisition module, a monitoring module, a signal processing module, a display module, etc., which respectively execute functions such as video acquisition, respiration rate monitoring, respiration signal processing, and respiration rate display. See Figure 6 , Figure 6 is an architecture diagram of a non-contact respiration rate monitoring system provided in the embodiments of the present invention. With the upgrade of algorithms and the improvement of hardware performance, the monitoring system can be updated and iterated subsequently to meet higher requirements.
[0207] In summary, compared with the prior art, the following beneficial effects are achieved:
[0208] 1. The present invention proposes a non-contact respiration rate monitoring method and system based on dual-spectrum information fusion, which can make full use of dual-spectrum information. When the ROI area cannot be detected by the thermal infrared spectrum, the respiration signal extracted by the thermal infrared spectrum is supplemented according to the ROI area of the visible light video and using the method based on affine transformation and dual-light fusion, so as to achieve strong-robustness ROI area detection.
[0209] 2. After performing detrending, normalization, and Gaussian fitting processing on the extracted respiration signal, the waveform of the interrupted and missing part of the curve is translated, and then the peaks of the complete, smooth, and continuous respiration signal curve are detected, so as to achieve strong-robustness and high-precision respiration rate monitoring.
[0210] 3. Traditional respiratory rate detection methods rely excessively on professional detection equipment (such as respiratory belts and monitors). They not only have drawbacks such as high cost, low comfort, and poor portability, but also wearing them for a long time will affect people's normal work and life. In addition, the above-mentioned respiratory monitoring equipment cannot meet the respiratory monitoring needs in complex scenarios with limited medical conditions. Non-contact respiratory rate monitoring has advantages such as high precision, low interference, and can be worn for a long time, and is suitable for special scenarios such as patients with skin ulcers, newborns, and performing tasks. On this basis, the monitoring function of abnormal breathing makes this system also applicable to detecting extreme situations such as apnea and too rapid breathing.
[0211] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0212] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A non-contact breathing rate monitoring method based on dual-spectrum facial videos, characterized in that, The method includes: simultaneously acquiring a thermal infrared video and a visible light video containing the face of the user to be measured; tracking and positioning the region of interest of the nose of the user to be measured in the thermal infrared video based on the kernel correlation filtering algorithm, extracting the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and using the thermal infrared image frames in the thermal infrared video where the region of interest of the nose is not tracked and positioned as the first thermal infrared image frames; extracting the respiration signal corresponding to the first thermal infrared image as the first respiration signal; and selecting the visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame; using the method of affine transformation to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extracting the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, the region of interest of the nose of the user to be measured in the first visible light image is obtained; extracting the respiration signal corresponding to the second thermal infrared image as the second respiration signal; performing signal fusion on the first respiration signal and the second respiration signal, and using the signal after signal fusion as the third respiration signal; performing signal processing on the third respiration signal based on a preset signal processing method to obtain a respiration signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing; obtaining the respiration rate of the user to be measured based on the respiration signal curve.
2. The method according to claim 1, wherein Using the method of affine transformation to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extracting the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image, includes: selecting the visible light image frame with the same timestamp as the second thermal infrared image frame in the visible light video as the second visible light image frame, and using the second thermal infrared image frame and the second visible light image with the same timestamp as a preset image pair; wherein, the second thermal infrared image frame represents the image frame containing the first thermal infrared image in the thermal infrared video; based on the mediapipe tool, obtaining the coordinates of the upper left corner point A(x1, y1) and the lower right corner point B(x2, y2) of the region of interest of the nose in the second visible light image frame in the preset image pair; and obtaining the coordinates of the upper left corner point A′(u1, v1) and the lower right corner point B′(u2, v2) of the region of interest of the nose in the second thermal infrared image frame based on the kernel correlation filtering algorithm; determining the affine transformation matrix of the same coordinates based on the coordinates of the upper left corner points A, A′ and the lower right corner points B, B′ in the image pair; performing affine transformation processing on the coordinates of the upper left corner point A and the lower right corner point B in the first visible light image frame based on the affine transformation matrix to obtain the coordinates of the upper left corner point A′ and the lower right corner point B′ in the first thermal infrared image frame; Extract the image of the nose region of interest in the first thermal infrared image frame as the second thermal infrared image based on the coordinates of the upper left corner point A' and the lower right corner point B' in the first thermal infrared image frame.
3. The method according to claim 1, wherein Extract the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal, including: Convert the first thermal infrared image into a grayscale image as the first grayscale image; Calculate the pixel average value of the first grayscale image as the first respiratory signal.
4. The method according to claim 1, wherein Based on a preset signal processing method, perform signal processing on the third respiratory signal to obtain a respiratory signal curve, including: Perform detrending processing on the third respiratory signal, and use the detrended third respiratory signal as the fourth respiratory signal; Perform normalization processing on the fourth respiratory signal, and use the normalized fourth respiratory signal as the fifth respiratory signal; Set the data format of the experimental data corresponding to the fifth respiration signal to and perform fitting processing on the experimental data to obtain a fitted Gaussian function as the respiration signal curve; where T i represents the abscissa time in the respiration curve; represents the ordinate respiration value in the respiration curve; a, c, and b are respectively parameters to be solved, and their physical meanings are the peak height, peak position, and half-width information of the Gaussian curve; b0, b1, and b2 are respectively intermediate parameters in the process of solving the parameters a, c, and b to be solved.
5. The method according to claim 4, characterized in that, The method further includes: When consecutive image frames of the thermal infrared video are missing and the missing image frames cannot be obtained through affine transformation, if the number of missing image frames is lower than a first preset value, then perform Gaussian fitting processing on the respiratory signal curve; If the number of missing image frames is not lower than the first preset value, then perform translation processing on the respiratory signal curve to supplement the missing respiratory signal in the respiratory signal.
6. The method according to claim 5, characterized in that, The translation processing on the respiratory signal curve includes: Obtain the total number of consecutive missing image frames, and the position of the first missing frame or the last missing frame in the consecutive missing image frames; Determine the time length of the continuously missing respiratory signal based on the total number of frames as the first time length; and determine the starting time point of the continuously missing respiratory signal based on the position of the first missing frame or the ending time point of the continuously missing respiratory signal based on the position of the last missing frame; Obtain a continuous respiratory curve with a time length of the first time length before the starting time point as the first curve; or obtain a continuous respiratory curve with a time length of the first time length after the starting time point as the second curve; Translate the first curve backward by a length corresponding to the first time length in chronological order, or translate the second curve forward by a length corresponding to the first time length in chronological order to supplement the missing respiratory signal in the respiratory signal.
7. The method according to claim 1, characterized in that Based on the respiratory signal curve, obtain the respiratory rate of the user being measured, including: Obtain the number of wave peaks PN of the respiratory signal curve; where the first peak and the last peak are respectively marked as FP and LP, and the average distance between two adjacent peaks is ADP; Calculate the respiratory rate of the user being measured based on a preset formula; where the preset formula is: where PN represents the total number of frames of the thermal infrared video obtained within one minute; RR represents the respiratory rate of the user being measured.
8. The method according to claim 1, wherein The method further includes: Set a time interval between two consecutive respiratory rate monitors.
9. A non-contact respiratory rate monitoring system based on dual-spectrum facial videos, characterized in that, The system includes: A first acquisition module, configured to simultaneously acquire a thermal infrared video and a visible light video containing the face of the user to be measured; A tracking and positioning module, configured to track and position the region of interest of the nose of the user to be measured in the thermal infrared video based on the kernel correlation filtering algorithm, extract the image of the region of interest of the nose in the thermal infrared video as the first thermal infrared image; and use the thermal infrared image frame in the thermal infrared video where the region of interest of the nose is not tracked and positioned as the first thermal infrared image frame; A first extraction module, configured to extract the respiration signal corresponding to the first thermal infrared image as the first respiration signal; and select the visible light image frame with the same timestamp as the first thermal infrared image frame in the visible light video as the first visible light image frame; An affine transformation module, configured to use the method of affine transformation to map the region of interest of the nose of the user to be measured in the first visible light image frame to the first thermal infrared image frame, and extract the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; wherein, based on the face detection algorithm and facial feature points, the region of interest of the nose of the user to be measured in the first visible light image is obtained; A second extraction module, configured to extract the respiration signal corresponding to the second thermal infrared image as the second respiration signal; A signal fusion module, configured to perform signal fusion on the first respiration signal and the second respiration signal, and use the signal after signal fusion as the third respiration signal; A signal processing module, configured to perform signal processing on the third respiration signal based on a preset signal processing method to obtain a respiration signal curve; wherein, the preset signal processing method includes: detrending processing, normalization processing, and Gaussian fitting processing; A second acquisition module, configured to obtain the respiration rate of the user to be measured based on the respiration signal curve.
Citation Information
Patent Citations
Respiration rate calculation method and device
CN106539586A
Non-contact multi-parameter monitoring method and system for physical and psychological health analysis
CN116403734A