Non-contact Respiratory Rate Detection Method and System Based on Facial Thermal Imaging
Through deep learning methods based on facial thermal imaging, facial ROI tracking and respiratory signal noise reduction processing are performed, and the problem of insufficient robustness and generalization in the prior art is solved, and high-precision and stable respiratory rate detection are achieved.
Patent Information
- Application Number
- CN202411257478.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-09
AI Technical Summary
The existing non-contact respiratory rate detection technology is insufficient in many practical application scenarios, resulting in limited accuracy in breath rate estimation.
The facial thermal imaging method is used and combined with deep learning technology, facial ROI tracking is performed through thermal infrared video signal processing network, and the adaptive breathing signal noise reduction network is used to capture the signal key characteristics, obtain the noise reduction breathing wave signal, and finally calculate the breathing rate.
It significantly improves the quality of respiratory signals and the accuracy of respiratory rate estimation, enhances the generalization ability in a variety of practical application scenarios, and ensures the stability and high accuracy of respiratory rate detection.
Smart Images

Figure CN119184667B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of non-contact respiratory rate detection, and specifically to a non-contact respiratory rate detection method, system, storage medium and electronic device based on facial thermography. Background Art
[0002] In the body's metabolism, respiration is crucial. Regularly monitoring the respiratory condition can detect and prevent diseases in the respiratory system, cardiovascular and cerebrovascular systems, etc. as early as possible.
[0003] In related technologies, respiratory rate detection technologies are mainly divided into two categories: contact type and non-contact type. Among them, non-contact respiratory rate detection does not require the use of traditional contact devices. It can detect human body movements and physiological signals through the use of cameras, microwave radars or other sensors, and analyze to obtain the respiratory rate. For example, Patent CN117357093A discloses a non-contact respiratory rate monitoring method and system based on dual-spectrum facial videos, including: simultaneously acquiring thermal infrared videos and visible light videos; extracting the first thermal infrared image; extracting the respiratory signal corresponding to the first thermal infrared image as the first respiratory signal; and selecting the first visible light image frame; using the method of affine transformation to map the region of interest of the nose of the measured user in the first visible light image frame to the first thermal infrared image frame, and extracting the image of the mapped region in the first thermal infrared image frame as the second thermal infrared image; extracting the respiratory signal corresponding to the second thermal infrared image as the second respiratory signal; performing signal fusion on the first respiratory signal and the second respiratory signal, and using the signal after signal fusion as the third respiratory signal; obtaining a respiratory signal curve; and based on the respiratory signal curve, obtaining the respiratory frequency of the measured user.
[0004] However, the above solutions have deficiencies such as limited applicable scenarios and limited accuracy of respiratory rate estimation in variable actual scenarios. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the prior art, the present invention provides a non-contact respiratory rate detection method, system, storage medium and electronic device based on facial thermography, and solves the technical problems of insufficient robustness and generalization in various actual application scenarios.
[0007] (2) Technical Solutions
[0008] To achieve the above purposes, the present invention is realized through the following technical solutions:
[0009] A non-contact respiratory rate detection method based on facial thermal imaging, based on a robust non-contact respiratory rate detection framework, the robust non-contact respiratory rate detection framework includes a thermal infrared video signal processing network and an adaptive respiratory signal noise reduction network; the method includes:
[0010] Obtain the real-time thermal infrared video of the tested person;
[0011] For the image frame sequence of the real-time thermal infrared video, based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network, perform facial ROI tracking, and based on the change of the average pixel value of the tracked facial nose area, obtain the initial respiratory wave signal;
[0012] For the initial respiratory wave signal, based on the one-dimensional convolutional layer and the attention mechanism layer of the adaptive respiratory signal noise reduction network, capture the key features of the signal to obtain the denoised respiratory wave signal;
[0013] Detect the wave peaks of the denoised respiratory wave signal and calculate the respiratory rate of the tested person.
[0014] Preferably, the fully convolutional single-stage object detection neural network includes a ResNet50 backbone network, a feature pyramid, and a shared detection head; for the image frame sequence of the real-time thermal infrared video, perform facial ROI tracking based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network; including:
[0015] In the ResNet50 backbone network: Continuously pass the current image frame through five convolutional layers, and extract three image features C3, C4, and C5 with different scales through the last three convolutional layers respectively;
[0016] In the feature pyramid: Respectively pass C3, C4, and C5 through 2DCNNs with a convolution kernel of 1×1; use the processed C5 as D5, splice the processed C5 after upsampling and the processed C4 to get D4, and splice the processed C4 after upsampling and the processed C3 to get D3; input D5, D4, and D3 into the feature extraction layer, and respectively pass through 2DCNNs with a convolution kernel of 3×3 to form P3, P4, and P5, and separately pass P5 through two 2DCNNs with a convolution kernel of 3×3 to form P6 and P7 respectively;
[0017] In the detection head: Pass P3, P4, P5, P6, and P7 through four 2DCNN-GN-ReLU layers on different branches, and then pass through a convolutional layer with a convolution kernel of 3×3 to output the facial ROI category, ROI coordinates, and confidence on the current image frame respectively.
[0018] Preferably, obtaining the initial respiratory wave signal based on the change in the average pixel value of the tracked facial nose region includes:
[0019] Calculating the average pixel value of the facial nose region in the current image frame t:
[0020]
[0021] where pixel ij is the pixel value at the coordinate (x i , y j ), and (x i , y j ) is the planar coordinate of the pixel located in the i-th row and j-th column;
[0022] After traversing and calculating the pixel average value of the facial nose region in each image frame, an initial respiratory wave signal is formed:
[0023]
[0024] where m is the total number of frames of the real-time thermal infrared video.
[0025] Preferably, the adaptive respiratory signal denoising network specifically includes five 1DCNNs, two MaxPooling layers, two UPSAMPLE layers, and one Attention layer; for the initial respiratory wave signal, based on the one-dimensional convolutional layer and attention mechanism layer of the adaptive respiratory signal denoising network, capturing the key features of the signal to obtain a denoised respiratory wave signal includes:
[0026] At the encoder, input the initial respiratory wave signal into the first 1DCNN for convolutional calculation to obtain the first local feature, and then input the first local feature into the first MaxPooling layer for max-pooling operation to obtain the first downsampled feature;
[0027] And input the first downsampled feature into the second 1DCNN for convolutional calculation to obtain the second local feature, and then input the second local feature into the second MaxPooling layer for max-pooling operation to obtain the second downsampled feature;
[0028] Input the second downsampled feature into the attention mechanism layer to focus on the key information of the input signal by assigning different weights;
[0029] At the decoder, input the output feature of the attention mechanism layer into the third 1DCNN for convolutional calculation to obtain the third local feature, and then input the third local feature into the first UPSAMPLE layer for upsampling operation to obtain the first upsampled feature with the same dimension as the input signal;
[0030] Input the first upsampled feature into the fourth 1D CNN for convolutional calculation to obtain the fourth local feature, and then input the fourth local feature into the second UPSAMPLE layer for upsampling operation to obtain the second upsampled feature with the same dimension as the input signal;
[0031] And input the second upsampled feature into the fifth 1D CNN for convolutional calculation, obtain and output the fifth local feature, and use the fifth local feature as the denoised respiratory wave signal.
[0032] Preferably, the respiratory rate is expressed as:
[0033]
[0034] Where RR is the respiratory rate, PN is the number of detected wave peaks, the indices of the first peak and the last peak are marked as FP and LP respectively, the average distance between two adjacent peaks is ADP, and TF is the total number of real-time thermal infrared video frames obtained within one minute.
[0035] Preferably, the robust non-contact respiratory rate detection framework is trained using a diverse dataset during the training phase.
[0036] Preferably, the robust non-contact respiratory rate detection framework selects the following loss function during the training phase:
[0037]
[0038] Where x and y are the predicted respiratory wave signal and the real respiratory band signal respectively, and T is the signal length.
[0039] A non-contact respiratory rate detection system based on facial thermography, based on a robust non-contact respiratory rate detection framework, the robust non-contact respiratory rate detection framework includes a thermal infrared video signal processing network and an adaptive respiratory signal denoising network; the system includes:
[0040] An acquisition module, configured to acquire the real-time thermal infrared video of the tested person;
[0041] A processing module, configured to perform facial ROI tracking on the image frame sequence of the real-time thermal infrared video based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network, and obtain an initial respiratory wave signal based on the change of the average pixel value of the tracked facial nose area;
[0042] A denoising module, configured to capture key features of the signal for the initial respiratory wave signal based on the one-dimensional convolutional layer and the attention mechanism layer of the adaptive respiratory signal denoising network to obtain a denoised respiratory wave signal;
[0043] A detection module, configured to detect the peaks of the noise-reduced respiratory wave signals and calculate the respiratory rate of the subject.
[0044] A storage medium stores a computer program for non-contact respiratory rate detection based on facial thermography, wherein the computer program causes a computer to execute the non-contact respiratory rate detection method as described above.
[0045] An electronic device, comprising:
[0046] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the non-contact respiratory rate detection method as described above.
[0047] (III) Advantageous Effects
[0048] The present invention provides a non-contact respiratory rate detection method, system, storage medium and electronic device based on facial thermography. Compared with the prior art, the following advantageous effects are achieved:
[0049] The present invention uses thermal infrared videos and combines deep learning technology to accurately and stably extract high-quality respiratory signals from the thermal infrared spectrum. Meanwhile, by designing an adaptive respiratory signal noise reduction network based on the attention mechanism to perform adaptive automatic noise reduction processing on the initial respiratory wave signals, it is possible to deeply analyze the local details and global features in the respiratory signals, automatically distinguish and effectively suppress noise, retain key physiological information, and significantly improve the quality of the respiratory signals and the accuracy of respiratory rate estimation. Through the training of the deep learning model, this method can adapt to diverse respiratory characteristics of different people and environments, enhance the generalization ability in various practical application scenarios, and ensure the stability and high accuracy of respiratory rate detection. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0051] Figure 1 It is a block diagram of a non-contact respiratory rate detection method based on facial thermography provided by an embodiment of the present invention;
[0052] Figure 2 It is a framework diagram of a robust non-contact respiratory rate detection provided by an embodiment of the present invention;
[0053] Figure 3 A schematic structural diagram of a thermal infrared video signal processing network provided by an embodiment of the present invention;
[0054] Figure 4 A schematic structural diagram of a feature extraction layer provided by an embodiment of the present invention;
[0055] Figure 5 A schematic structural diagram of an adaptive breathing signal noise reduction network provided by an embodiment of the present invention. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] Embodiments of the present application provide a non-contact respiratory rate detection method, system, storage medium, and electronic device based on facial thermal imaging, which solve the technical problems of insufficient robustness and generalization in various practical application scenarios.
[0058] The general idea of the technical solutions in the embodiments of the present application to solve the above technical problems is as follows:
[0059] Embodiments of the present invention innovatively propose a non-contact respiratory rate detection method and system for monitoring physiological parameters of respiratory rate. The method uses thermal infrared information and combines deep learning technology to accurately and stably extract high-quality breathing signals from the thermal infrared spectrum. At the same time, by designing an adaptive breathing signal noise reduction network based on the attention mechanism, the initial breathing wave signal extracted is adaptively and automatically denoised, and the correlation between local and global information of the breathing signal is mined to achieve the generalization processing ability of the breathing signals extracted in different scenarios and populations and the accurate estimation of the respiratory rate, so as to realize the real-time analysis of the individual's physiological health status and provide objective and accurate physiological perception information for the screening of respiratory diseases.
[0060] In the signal acquisition stage, embodiments of the present invention abandon the traditional tracking method that relies on visible light face key point detection and instead adopt the thermal infrared video face ROI tracking technology based on the FCOS neural network. Through training with multi-scene face data, this technology can quickly, accurately, and stably locate the face ROI area in the thermal infrared video. Compared with the traditional method, the tracking technology of the present invention is not affected by scene changes such as the distance change and movement of the detector, showing higher robustness and being able to cope with the interference in the resting or moving state.
[0061] In the signal processing stage, traditional signal noise reduction methods rely on manual adjustment of algorithm parameters through signal preprocessing operations such as detrending, normalization, and fast Fourier transform, and are unable to robustly and adaptively process respiratory signals generated by different breathing patterns. The adaptive signal noise reduction method proposed in the embodiments of the present invention can deeply analyze local details and global features in respiratory signals by designing a neural network automatic noise reduction network, automatically distinguish and effectively suppress noise, retain key physiological information, and significantly improve the quality of respiratory signals and the accuracy of respiratory rate estimation. Through the training of deep learning models, this method can adapt to diverse respiratory characteristics of different people and environments, enhance the generalization ability in various practical application scenarios, and ensure the stability and high accuracy of respiratory rate detection.
[0062] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0063] Embodiment 1:
[0064] As Figure 1 shown, the embodiments of the present invention provide a non-contact respiratory rate detection method based on facial thermal imaging based on a robust non-contact respiratory rate detection framework as Figure 2 shown. The robust non-contact respiratory rate detection framework includes a thermal infrared video signal processing network and an adaptive respiratory signal noise reduction network; this method includes:
[0065] S1. Obtain the real-time thermal infrared video of the test subject;
[0066] S2. For the image frame sequence of the real-time thermal infrared video, perform facial ROI tracking based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network, and obtain the initial respiratory wave signal based on the change in the average pixel value of the tracked facial nose area;
[0067] S3. For the initial respiratory wave signal, capture the key features of the signal based on the one-dimensional convolutional layer and the attention mechanism layer of the adaptive respiratory signal noise reduction network to obtain the noise-reduced respiratory wave signal;
[0068] S4. Detect the wave peaks of the noise-reduced respiratory wave signal and calculate the respiratory rate of the test subject.
[0069] In an embodiment of the present invention, a tracking technique for the Region of Interest (ROI) of the face based on deep learning is proposed to achieve fast, accurate, and stable tracking of the facial ROI in a thermal infrared video, improving the quality and consistency of the initially extracted respiratory wave signals. At the same time, based on an adaptive respiratory signal denoising network, it can analyze the local and global features in the initial respiratory wave signals, automatically distinguish and suppress noise, and retain important physiological information to restore respiratory wave signals with better quality, thereby significantly improving the accuracy of respiratory rate estimation.
[0070] Next, each step of the above solution will be introduced in detail:
[0071] In step S1, a real-time thermal infrared video of the subject is obtained.
[0072] To ensure that the facial video of the subject can be captured, the thermal infrared camera can be kept 0.5m - 0.8m away from the subject. After clicking the start collection button, the collection device records the real-time thermal infrared video of the subject.
[0073] In step S2, for the image frame sequence of the real-time thermal infrared video, facial ROI tracking is performed based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network, and an initial respiratory wave signal is obtained based on the change in the average pixel value of the tracked facial nose region.
[0074] This step performs respiratory wave signal extraction operations based on the thermal infrared video signal processing network as shown in Figure 3 and involves facial ROI tracking and respiratory wave signal extraction, specifically including:
[0075] (1) Facial ROI tracking based on thermal infrared video
[0076] Precise tracking of the Region of Interest (ROI) of the face in a thermal infrared video is one of the key steps to obtain respiratory signals. Traditional ROI tracking methods that rely on affine transformation and feature point matching techniques are mainly based on RGB camera imaging and dual-spectrum matching of thermal infrared cameras. However, in complex scenarios with insufficient or excessive light conditions, traditional face detection methods are difficult to effectively track the ROI region, resulting in the failure of techniques such as affine transformation and feature point matching and the inability to achieve stable tracking of the facial ROI region. In contrast, methods based on thermal infrared video can capture thermal images generated by human activities and are less dependent on environmental light, so they are more reliable.
[0077] To this end, the embodiment of the present invention designs a thermal infrared video ROI tracking method based on a Fully Convolutional One-Stage Object Detection (FCOS) neural network to robustly obtain the coordinate position of the facial ROI region in the thermal infrared video, thereby ensuring the acquisition of stable and reliable respiratory wave signals from the thermal infrared video in complex and variable scenarios.
[0078] As Figure 3 shown, the FCOS neural network includes a ResNet50 backbone network, a feature pyramid, and a shared detection head. Correspondingly, the specific implementation steps for facial ROI tracking based on the thermal infrared video are as follows:
[0079] 1) In the ResNet50 backbone network: The current image frame is continuously passed through five convolutional layers. Among them, the last three convolutional layers extract three image features C3, C4, and C5 at different scales respectively. These feature information will be used for the classification and localization of the ROI later.
[0080] 2) In the feature pyramid: C3, C4, and C5 are respectively passed through a 2D CNN with a convolutional kernel of 1×1; the processed C5 is used as D5, the processed C5 is upsampled and then concatenated with the processed C4 to obtain D4, and the processed C4 is upsampled and then concatenated with the processed C3 to obtain D3; D5, D4, and D3 are input into the feature extraction layer as Figure 4 shown, and are respectively passed through a 2D CNN with a convolutional kernel of 3×3 to form P3, P4, and P5. And P5 is separately passed through two 2D CNNs with a convolutional kernel of 3×3 to form P6 and P7 respectively. Finally, the feature extraction layer extracts a total of five scales of features, namely P3, P4, P5, P6, and P7.
[0081] 3) In the detection head: The purpose of the detection head is to classify the objects, perform bounding box regression, and predict the centrality of the features extracted by the feature pyramid. And the above five scales of features share a detection head. P3, P4, P5, P6, and P7 are passed through four 2D CNN-GN-ReLU layers on different branches, and then passed through a convolutional layer with a convolutional kernel of 3×3 to respectively output the facial ROI category, ROI coordinates, and confidence on the current image frame.
[0082] It can be understood that the above FCOS network reflects good robustness and generalization through methods such as anchor-free design and multi-scale feature fusion; its fully convolutional structure and end-to-end training process further enhance the adaptability of the model on different datasets and tasks, enabling it to maintain high detection performance in various complex scenarios.
[0083] (2) Absorbing Wave Signal Extraction
[0084] During the human breathing process, the temperature change in the nose area is the most obvious. Therefore, in the embodiments of the present invention, an initial breathing wave signal is extracted based on the image sequence of the nose area in the thermal infrared facial video. During normal breathing, the gas exhaled from the mouth and nose is mainly carbon dioxide, which is produced by metabolism in the exhaled gas from the nose. The temperature of the exhaled gas is significantly higher than the temperature of the air around the nose, while the inhaled air is significantly lower than the temperature of the exhaled gas. This causes a slight change in the temperature of the air around the nose, which is manifested as a change in the average pixel value in the thermal infrared facial region of interest.
[0085] Correspondingly, the specific implementation steps of absorbing wave signal extraction are as follows:
[0086] Calculate the average pixel value (Average pixel, abbreviated as AP) of the facial nose area in the current image frame t:
[0087]
[0088] where pixel ij is the pixel value at the coordinate (x i , y j ), and (x i , y j ) is the planar coordinate of the pixel located in the i-th row and j-th column.
[0089] After traversing and calculating the pixel average value of the facial nose area in each image frame, an initial breathing wave signal is formed:
[0090]
[0091] where m is the total number of frames of the real-time thermal infrared video.
[0092] In step S3, for the initial breathing wave signal, based on the one-dimensional convolutional layer and attention mechanism layer of the adaptive breathing signal denoising network, the key features of the signal are captured to obtain a denoised breathing wave signal.
[0093] Since there is a certain amount of noise in the preliminarily processed breathing wave signal and it cannot be directly applied to the accurate estimation of the breathing rate, the obtained initial breathing wave signal needs to be denoised. However, the traditional breathing wave signal is basically achieved through normalization, detrending, band-pass filtering, etc. It relies on experience or a pre-set frequency range for manual configuration, and the fixed setting range cannot cope with the changing breathing patterns and scenarios, lacking adaptability and intelligent processing capabilities.
[0094] Therefore, to improve the anti-interference and robust performance of breathing signal processing, the embodiments of the present invention design as Figure 5The "1DCNN+ATTENTION" adaptive noise reduction network shown. This network aims to effectively capture the key features of the signal and extract the respiratory features in the physiological information to achieve automatic noise reduction of the initial respiratory wave signal.
[0095] In terms of the overall network architecture, as Figure 5 shown, the adaptive respiratory signal noise reduction network mainly includes three structures: an encoder, an attention layer, and a decoder. Among them:
[0096] The encoder converts the input data from a high-dimensional representation to a low-dimensional one, and extracts features at different levels through the stacking of multiple neural networks. The attention layer assigns different weights to the features obtained from different parts, enabling the model to dynamically focus on the key parts of the data and helping the model identify important segments in the signal. The decoder gradually generates the target output from the low-dimensional features generated by the encoder, restores the features compressed by the encoder to the same spatial dimension as the input signal, and finally generates the target output.
[0097] This neural network aims to significantly improve the signal noise reduction effect by efficiently capturing the key features in the local, global, and context of the signal. At the same time, this network can accurately extract the respiratory features in the physiological signal, thereby effectively restoring a higher-quality real respiratory signal.
[0098] Specifically, as Figure 5 shown, the adaptive respiratory signal noise reduction network includes five 1DCNNs, two MaxPooling layers, two UPSAMPLE layers, and one Attention layer. This structural design aims to improve the adaptability and practicality of the network.
[0099] It can be understood that the adaptive signal noise reduction module based on the "1DCNN+ATTENTION" neural network is a deep learning model, and through these network layers, it can effectively perform adaptive noise reduction processing on the respiratory signal. Since this network processes one-dimensional respiratory signals, all 5 convolutional layers are considered to use 1DCNNs. The one-dimensional convolutional layer is responsible for performing convolutional operations on the respiratory signal such as numbers. Through convolutional operations, the convolutional layer can extract the local features in the signal, and these local features can reflect the key features such as the frequency and waveform of the respiratory signal, which are very important for understanding the respiratory signal features. Specifically, the first two 1DCNN layers are mainly used to extract relevant features such as the initial respiratory wave signal of the input, and the last three 1DCNN layers are mainly used to restore the extracted high-dimensional features to a signal with the same dimension as the input signal, so as to achieve the purpose of adaptive signal noise reduction.
[0100] The MaxPooling layer divides the features extracted by the convolutional layer into several blocks of the same size and selects the maximum value of each block for output. This operation can reduce the spatial dimension of the data while retaining the most important features, thereby reducing the computational complexity of the network and making the network calculation more efficient. The Dropout layer randomly selects a part of the neurons and their connections during the training process, which helps to reduce the network's dependence on the training samples, reduce the risk of overfitting of the network, and improve the generalization ability of the network. And the Attention layer is a technique used to enhance the ability of the neural network to focus on key information. The attention mechanism can more effectively denoise the signal by assigning different weights to different parts of the input signal. In addition, the attention mechanism allows the network to dynamically adjust the focus according to the changes in the input signal, thereby realizing adaptive learning, which can help the network identify more signal categories and perform more effective signal denoising.
[0101] Correspondingly, the specific implementation steps for automatically denoising the initial respiratory wave signal are as follows:
[0102] At the encoder, the initial respiratory wave signal is input into the first 1DCNN for convolutional calculation to obtain the first local features. Then, the first local features are input into the first MaxPooling layer for max-pooling operation to obtain the first down-sampled features, so as to retain the most important features and reduce the spatial dimension of the data, improving the computational efficiency.
[0103] And the first down-sampled features are input into the second 1DCNN for convolutional calculation to obtain the second local features. Then, the second local features are input into the second MaxPooling layer for max-pooling operation to obtain the second down-sampled features, so as to further compress the feature data and improve the computational efficiency.
[0104] The second down-sampled features are input into the attention mechanism layer, and different weights are assigned to focus on the key information of the input signal, so as to improve the effect of signal denoising and realize adaptive learning.
[0105] At the decoder, the output features of the attention mechanism layer are input into the third 1DCNN for convolutional calculation to obtain the third local features. Then, the third local features are input into the first UPSAMPLE layer for up-sampling operation to obtain the first up-sampled features with the same dimension as the input signal.
[0106] The first up-sampled features are input into the fourth 1DCNN for convolutional calculation to obtain the fourth local features. Then, the fourth local features are input into the second UPSAMPLE layer for up-sampling operation to obtain the second up-sampled features with the same dimension as the input signal.
[0107] And input the second upsampled feature into the fifth 1D CNN for convolution calculation to obtain and output the fifth local feature, and use the fifth local feature as the denoised respiratory wave signal.
[0108] After passing through the adaptive signal denoising network, the output form of the denoised respiratory wave signal is as follows:
[0109]
[0110] In step S4, detect the peaks of the denoised respiratory wave signal and calculate the respiratory rate of the subject.
[0111] Detect the denoised respiratory wave signal The peaks of the signal. Assume that the number of detected peaks is called PN, the indices of the first peak and the last peak are marked as FP and LP respectively, and the average distance between two adjacent peaks is ADP. In addition, since the information in [0, FP] and [LP, TF] is ignored during the calculation of RR, the number of breaths in this part needs to be considered. The total number of real-time thermal infrared video frames obtained within one minute is TF, and the calculation formula for the respiratory rate RR is as follows:
[0112]
[0113] So far, the embodiment of the present invention has completed all the processes of non-contact respiratory rate detection.
[0114] Specifically, to improve the prediction effect, in the training stage of the robust non-contact respiratory rate detection framework, the embodiment of the present invention collects facial videos of different scenarios and subjects and records the real respiratory wave signals of the subjects through a respiratory belt to obtain a diverse dataset for training, so that it can learn more features related to different scenarios and subjects.
[0115] And, the robust non-contact respiratory rate detection framework specifically selects the following loss function in the training stage:
[0116]
[0117] Among them, x and y are the predicted respiratory wave signal and the real respiratory belt signal respectively, and T is the signal length.
[0118] Embodiment 2:
[0119] The embodiment of the present invention provides a non-contact respiratory rate detection system based on facial thermography, based on a robust non-contact respiratory rate detection framework, and the robust non-contact respiratory rate detection framework includes a thermal infrared video signal processing network and an adaptive respiratory signal denoising network; this system includes:
[0120] An acquisition module for acquiring the real-time thermal infrared video of the subject;
[0121] A processing module for performing facial ROI tracking on the image frame sequence of the real-time thermal infrared video based on the fully convolutional single-stage object detection neural network of the thermal infrared video signal processing network, and acquiring an initial respiratory wave signal based on the change of the average pixel value of the tracked facial nose region;
[0122] A noise reduction module for capturing key features of the signal based on the one-dimensional convolutional layer and the attention mechanism layer of the adaptive respiratory signal noise reduction network for the initial respiratory wave signal to obtain a noise-reduced respiratory wave signal;
[0123] A detection module for detecting the peaks of the noise-reduced respiratory wave signal and calculating the respiratory rate of the subject.
[0124] Embodiment 3:
[0125] An embodiment of the present invention provides a storage medium, characterized in that it stores a computer program for non-contact respiratory rate detection based on facial thermography, wherein the computer program enables a computer to execute the non-contact respiratory rate detection method as described in Embodiment 1.
[0126] Embodiment 4:
[0127] An embodiment of the present invention provides an electronic device, characterized by including:
[0128] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing non-contact respiratory rate detection as described in Embodiment 1.
[0129] It can be understood that the non-contact respiratory rate detection system, storage medium and electronic device provided by the embodiments of the present invention correspond to the non-contact respiratory rate detection method provided by the embodiments of the present invention. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding parts in the non-contact respiratory rate detection method, which will not be elaborated here.
[0130] In summary, compared with the prior art, the following beneficial effects are achieved:
[0131] 1. An embodiment of the present invention designs a method for tracking the facial ROI region of a thermal infrared video based on the FCOS neural network. Through training with multi-scene facial data, the proposed method can locate the facial region of interest in the thermal infrared video more quickly, accurately and stably, so as to improve the extraction quality of respiratory signals and lay a solid foundation for subsequent signal processing.
[0132] 2. An embodiment of the present invention proposes an adaptive breathing signal denoising method based on an attention mechanism. Through an end-to-end learning method, it deeply explores the internal connection between local and global information in the breathing signal, intelligently adjusts the denoising strategy, and effectively overcomes the problem of the decrease in the estimation accuracy of physiological parameters in the traditional signal preprocessing method within the abnormal range.
[0133] 3. When designing the deep learning network of the embodiment of the present invention, the diversity of different individuals and environments is fully considered. Through a large amount of data training, the adaptability and generalization ability of the model to new situations are enhanced, ensuring stability and reliability in different application scenarios. At the same time, in the design of the network structure, the embodiment of the present invention also makes innovations. Especially in the introduction of the attention mechanism, the network can focus more on the key information in the signal, improving the pertinence and effectiveness of signal processing.
[0134] 4. The embodiment of the present invention has achieved a significant improvement in comprehensive performance, including the accuracy of signal extraction, the robustness of denoising processing, and the generalization ability of the model.
[0135] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0136] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A non-contact respiratory rate detection method based on facial thermal imaging, characterized in that: Based on a robust non-contact breathing rate detection framework, the robust non-contact breathing rate detection framework includes a thermal infrared video signal processing network and an adaptive breathing signal denoising network; The method includes: Obtain real-time thermal infrared video of the subject; For the image frame sequence of the real-time thermal infrared video, a fully convolutional single-stage target detection neural network of the thermal infrared video signal processing network is used to track the facial ROI, and an initial respiratory wave signal is obtained based on the change in the average pixel value of the tracked facial nose area; For the initial respiratory wave signal, based on the one-dimensional convolution layer and the attention mechanism layer of the adaptive respiratory signal denoising network, the key features of the signal are captured to obtain a denoised respiratory wave signal; Detecting the peak of the noise-reduced respiratory wave signal and calculating the respiratory rate of the subject; The adaptive breathing signal denoising network specifically includes five 1DCNNs, two MaxPooling layers, two UPSAMPLE layers and one Attention layer; for the initial breathing wave signal, based on the one-dimensional convolution layer and the attention mechanism layer of the adaptive breathing signal denoising network, the key features of the signal are captured to obtain the denoised breathing wave signal; including: At the encoder, the initial respiratory wave signal is input into the first 1DCNN for convolution calculation to obtain the first local feature, and then the first local feature is input into the first MaxPooling layer for maximum pooling operation to obtain the first down-sampling feature; And inputting the first down-sampled feature into a second 1DCNN for convolution calculation to obtain a second local feature, and then inputting the second local feature into a second MaxPooling layer for maximum pooling operation to obtain a second down-sampled feature; Input the second downsampled features into the attention mechanism layer to focus on the key information of the input signal by assigning different weights; At the decoder, the output features of the attention mechanism layer are input into the third 1DCNN for convolution calculation to obtain the third local features, and then the third local features are input into the first UPSAMPLE layer for upsampling operation to obtain the first upsampling features with the same dimension as the input signal; Input the first up-sampled feature into the fourth 1DCNN for convolution calculation to obtain a fourth local feature, and then input the fourth local feature into the second UPSAMPLE layer for up-sampling operation to obtain a second up-sampled feature with the same dimension as the input signal; And input the second up-sampled feature into the fifth 1DCNN for convolution calculation, obtain and output the fifth local feature, and use the fifth local feature as the denoised respiratory wave signal.
2. The non-contact breathing rate detection method according to claim 1, characterized in that: The fully convolutional single-stage target detection neural network includes a ResNet50 backbone network, a feature pyramid, and a shared detection head; the fully convolutional single-stage target detection neural network based on the thermal infrared video signal processing network performs facial ROI tracking for the image frame sequence of the real-time thermal infrared video; including: In the ResNet50 backbone network: the current image frame is continuously subjected to five layers of convolution, wherein the last three layers of convolution respectively extract image features C3, C4, and C5 of three different scales; In the feature pyramid: C3, C4, and C5 are respectively processed by a 2DCNN with a convolution kernel of 1×1; the processed C5 is used as D5, and the processed C5 is upsampled and concatenated with the processed C4 to obtain D4, and the processed C4 is upsampled and concatenated with the processed C3 to obtain D3; D5, D4, and D3 are input into the feature extraction layer, and are respectively processed by a 2DCNN with a convolution kernel of 3×3 to form P3, P4, and P5, and P5 is separately processed by a 2DCNN with a convolution kernel of 3×3 twice to form P6 and P7 respectively; In the detection head: P3, P4, P5, P6, and P7 are passed through four 2DCNN-GN-ReLU layers on different branches, and then through corresponding convolution layers with a convolution kernel of 3×3, and the facial ROI category, ROI coordinates, and confidence on the current image frame are output respectively.
3. The non-contact breathing rate detection method according to claim 1, characterized in that: The method of obtaining an initial respiratory wave signal based on the change in the average pixel value of the nose area of the tracked face comprises: Calculate the average pixel value of the nose area of the face in the current image frame t: Among them, pixel ij is the coordinate (x i ,y i ), (x i ,y i ) is the plane coordinate of the pixel located at the i-th row and j-th column; After traversing and calculating the average pixel value of the nose area in each image frame, the initial breathing wave signal is formed: Where m is the total number of frames of the real-time thermal infrared video.
4. The non-contact breathing rate detection method according to claim 1, characterized in that: The respiration rate is expressed as: Among them, RR is the respiration rate, PN is the number of detected peaks, the indexes of the first peak and the last peak are marked as FP and LP respectively, the average distance between two adjacent peaks is ADP, and TF is the total number of frames of real-time thermal infrared video acquired within one minute.
5. The non-contact breathing rate detection method according to any one of claims 1 to 4, characterized in that: The robust contactless respiration rate detection framework is trained using a diverse dataset during the training phase.
6. The non-contact breathing rate detection method according to any one of claims 1 to 4, characterized in that: The robust contactless respiration rate detection framework selects the following loss function during the training phase: Among them, x and y are the predicted respiratory wave signal and the real respiratory band signal respectively, and T is the signal length.
7. A non-contact respiratory rate detection system based on facial thermal imaging, characterized in that: Based on a robust non-contact breathing rate detection framework, the robust non-contact breathing rate detection framework includes a thermal infrared video signal processing network and an adaptive breathing signal denoising network; The system is used to perform the non-contact breathing rate detection method as claimed in claim 1, comprising: An acquisition module is used to acquire real-time thermal infrared video of the subject; A processing module, for performing facial ROI tracking based on a fully convolutional single-stage target detection neural network of the thermal infrared video signal processing network for an image frame sequence of the real-time thermal infrared video, and obtaining an initial respiratory wave signal based on a change in an average pixel value of a nose region of the tracked face; A denoising module, for capturing key features of the signal based on the one-dimensional convolution layer and the attention mechanism layer of the adaptive respiratory signal denoising network for the initial respiratory wave signal, so as to obtain a denoised respiratory wave signal; The detection module is used to detect the peak of the noise-reduced respiratory wave signal and calculate the respiratory rate of the subject.
8. A storage medium, characterized in that: The computer program for non-contact respiratory rate detection based on facial thermal imaging is stored therein, wherein the computer program enables the computer to execute the non-contact respiratory rate detection method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs including instructions for executing the non-contact breathing rate detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Lightweight non-contact multi-parameter physiological detection method and system
CN119112107A