Real-time physiological parameter non-inductive monitoring method and system based on visual perception and intelligent calculation

Through the contactless physiological parameter monitoring method of visual perception and intelligent computing, machine vision and deep learning technology are used to achieve high-precision and real-time monitoring of heart rate, blood oxygen saturation, blood pressure and fatigue degree, solving the comfort and cost problems of traditional monitoring methods, and are suitable for special medical scenarios and home health management.

CN120236777APending Publication Date: 2025-07-01UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510249489.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing physiological parameter monitoring technology has limitations such as poor comfort in invasive detection methods, high cost of special medical equipment, and real-time lag in continuous monitoring, making it difficult to meet the needs of special medical scenarios and home health management.

Method used

Using a contactless physiological parameter monitoring method based on visual perception and intelligent calculation, a machine vision algorithm and deep learning model is used to achieve high-precision and real-time monitoring of heart rate, heart rate variability, blood oxygen saturation, blood pressure and fatigue degree, and a multimodal physiological signal analysis system is constructed.

Benefits of technology

It realizes physiological parameter monitoring with good comfort and high accuracy, meets the needs of special medical scenarios and home health management, and reduces equipment costs and operation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236777A_ABST
    Figure CN120236777A_ABST
Patent Text Reader

Abstract

The invention provides a real-time physiological parameter non-inductive monitoring method based on visual perception and intelligent calculation, belongs to the technical field of machine vision and deep learning, and aims at solving the problems that a traditional physiological parameter monitoring method is uncomfortable and high in cost. Video image acquisition equipment such as a common camera is used for acquiring video images of a face area of a subject, and the heart rate value, the heart rate variability, the oxyhemoglobin saturation and the blood pressure value of the subject are obtained; and establishing a fatigue degree evaluation network for the obtained physiological index parameters, and comprehensively evaluating the fatigue degree of the testee. The system has the advantages of good comfort, high precision and automatic monitoring function, can accurately monitor various physiological parameters such as heart rate, heart rate variability, blood oxygen saturation, blood pressure and fatigue degree of a subject, ensures the real-time performance of the monitoring process, and meets the application requirements of multiple fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of machine vision and deep learning, and particularly relates to a method and system for real-time non-intrusive monitoring of physiological parameters based on visual perception and intelligent computing. Background Art

[0002] The monitoring of human physiological parameters by detecting various physiological indicators of the human body in real time or periodically, such as heart rate, blood oxygen saturation, blood pressure, etc., can not only provide a basis for health management, help detect potential health problems early and carry out preventive interventions, but also play a key role in the process of disease diagnosis and treatment. Especially in the treatment process of chronic disease patients, regular monitoring of physiological parameters can help doctors adjust treatment plans in time, thus avoiding the deterioration of patients' conditions and helping patients recover.

[0003] The current physiological parameter monitoring technology faces three core challenges: (1) Traditional contact monitoring methods such as pulse oximeters, electrocardiogram patches, etc. need to be in direct contact with the human skin, which is likely to cause problems such as skin allergies and pressure sores, and there are application limitations in special medical scenarios such as the burn department and the neonatal intensive care unit; (2) Medical-grade monitoring devices such as multi-parameter monitors, ambulatory blood pressure monitors, etc. are costly to purchase and rely on professional personnel for operation, making it difficult to meet the inclusive needs of community medical care and home health management; (3) Existing non-contact monitoring technologies such as millimeter-wave radars have bottlenecks in terms of accuracy, real-time performance, etc., with insufficient dynamic monitoring accuracy and significant data latency, and there are relatively large application limitations.

[0004] The real-time non-intrusive monitoring technology of physiological parameters based on visual perception and intelligent computing has prominent advantages. By means of non-contact measurement, it abandons the limitations of wearing sensors or devices in traditional monitoring methods, and uses computer vision algorithms and deep learning models to accurately and real-time extract physiological parameters only with an ordinary camera. This non-intrusive monitoring not only eliminates the discomfort of the subject, but also ensures that the subject can carry out daily work in a natural state, avoiding interference with the subject's daily activities. Summary of the Invention

[0005] In view of the limitations of the existing physiological parameter detection technologies, such as poor comfort in invasive detection methods, high costs of dedicated medical equipment, and lag in real-time continuous monitoring, the present invention proposes a non-intrusive real-time physiological parameter monitoring method and system based on visual perception and intelligent computing. By deeply integrating machine vision and deep learning algorithms, the method and system construct a non-contact multi-modal physiological signal analysis system to achieve high-precision continuous monitoring of key physiological indicators such as heart rate (HR), heart rate variability (HRV), blood oxygen saturation (SpO2), dynamic blood pressure parameters (SBP / DBP), and fatigue index (FI), thereby ensuring the accuracy, system robustness, and real-time performance of the detection results. Particularly for special application scenarios such as the skin contact taboos of burn patients, the sensitive and fragile body surfaces of newborns, and the monitoring of the working status of drivers, this solution, through technological breakthroughs such as non-contact detection modes and lightweight deployment, significantly reduces equipment costs and operation complexity while improving monitoring efficiency, providing a universal solution for fields such as clinical monitoring, health management, and industrial safety.

[0006] The technical solution of the present invention is a non-intrusive real-time physiological parameter monitoring method based on visual perception and intelligent computing. The method includes the following steps:

[0007] Step 1: Build a video image acquisition system, set the camera device directly in front of the subject's face, and ensure stable, clear, and continuous capture of the subject's facial video image data;

[0008] Step 2: Perform image segmentation processing on the obtained facial video image of the subject, extract the region containing facial features therefrom, and obtain the facial region video image of the subject;

[0009] Step 3: Build a multi-scale rPPG signal extraction convolutional neural network based on channel attention mechanism to extract rPPG signals, and extract rPPG signals from the facial region video image of the subject obtained in Step 2;

[0010] Step 4: If the multi-scale rPPG signal extraction convolutional neural network based on channel attention mechanism has not been trained yet, execute Step 5 for network training and optimization; otherwise, if the training of the multi-scale rPPG signal extraction convolutional neural network based on channel attention mechanism has been completed and the performance meets the expectations, skip the training step and directly execute Step 6 for subsequent testing and evaluation;

[0011] Step 5: Use the open-source UBFC dataset to train the convolutional neural network for multi-scale rPPG signal extraction based on channel attention mechanism; divide the dataset into training set, validation set and test set to ensure data diversity during training and the accuracy of validation; in the training stage, optimize the model parameters through multiple rounds of iteration to gradually improve the performance of the network until the preset number of training rounds or performance metrics are reached, ensuring that the network can converge efficiently and prevent overfitting;

[0012] Step 6: Take the video image of the subject's facial area extracted in Step 2 as input and send it into the trained convolutional neural network for multi-scale rPPG signal extraction based on channel attention mechanism for processing. After network calculation, the rPPG signal of the subject is output;

[0013] Step 7: Filter the rPPG signal obtained in Step 6 using a Butterworth band-pass filter; set the passband range of the filter to 0.7 Hz to 3.0 Hz to ensure that only the frequency components related to heart rate are retained, filter out low-frequency noise and high-frequency interference, and obtain the filtered rPPG signal;

[0014] Step 8: Perform a fast Fourier transform on the filtered rPPG signal obtained in Step 7 to convert the time-domain rPPG signal into a frequency-domain representation and obtain the corresponding rPPG signal spectrum;

[0015] Step 9: Perform peak analysis on the rPPG signal spectrum obtained in Step 8 to identify the highest peak in the spectrum and its corresponding frequency. This peak frequency corresponds to the frequency of the heart rate. By calculation, multiply this frequency value by the time coefficient to obtain the actual heart rate value of the subject;

[0016] Step 10: Perform peak detection on the filtered rPPG signal obtained in Step 7, record the time points when each peak appears, and calculate the time interval between two adjacent peaks. This time interval represents the duration of each heartbeat of the heart, that is, the cardiac cycle;

[0017] Step 11: Calculate the absolute error between adjacent cardiac cycles for each cardiac cycle obtained in Step 10, and calculate the average value of the errors for all cardiac cycles to obtain the heart rate variability of the subject;

[0018] Step 12: Build a convolutional neural network for blood oxygen saturation extraction to extract the blood oxygen saturation value of the human body; if the convolutional neural network for blood oxygen saturation extraction has not been completed training, execute Step 13 for further training to ensure that the network parameters are fully adjusted and reach the expected performance level; otherwise, if the network has been completed training and performs stably on the validation set, directly jump to Step 14 for actual prediction and application of blood oxygen saturation;

[0019] Step 13: Use the open-source PURE dataset to train the convolutional neural network for blood oxygen saturation extraction; to ensure the diversity of training data and improve the generalization ability of the model, divide the dataset into a training set, a validation set, and a test set; during the training process, the model continuously optimizes the parameters through multiple rounds of iteration. In each round, it learns through the training set and is simultaneously verified in real-time on the validation set to monitor the training situation and validation accuracy of the model; as the training progresses, the network performance gradually improves until the preset number of training rounds or performance metrics are reached;

[0020] Step 14: Use the filtered rPPG signal obtained in Step 7 as the input and feed it into the trained convolutional neural network for blood oxygen saturation extraction. After network calculation, the blood oxygen saturation of the subject is obtained;

[0021] Step 15: Build a residual network for blood pressure extraction to extract the blood pressure value of the human body; if the network has not been completed training, execute Step 16 to ensure that the model can fully learn the data features and achieve the expected training effect; conversely, if the residual network for blood pressure extraction has been trained and is stable on the validation set, skip the training stage and directly execute Step 18 for blood pressure value prediction and application to achieve blood pressure monitoring;

[0022] Step 16: Collect the rPPG signals of at least 100 subjects and simultaneously collect the blood pressure data of the subjects, including systolic blood pressure and diastolic blood pressure, using a medical sphygmomanometer to construct a blood pressure dataset.

[0023] Step 17: Use the blood pressure dataset to train the residual network for blood pressure extraction. Divide it into a training set, a validation set, and a test set to perform iterative training on the network until the preset number of training rounds or performance metrics are reached, and the network training is completed;

[0024] Step 18: Input the filtered rPPG signal obtained in Step 7 into the trained residual network for blood pressure extraction. After network calculation, the blood pressure value of the subject, including systolic blood pressure and diastolic blood pressure, is obtained;

[0025] Step 19: Build a fatigue degree evaluation network to evaluate the fatigue degree of the subject; if the fatigue degree evaluation network has not been completed training, execute Step 20 for further training and optimization; conversely, if the fatigue degree evaluation network has been trained and has stable performance, directly execute Step 22 for the actual evaluation and application of the fatigue degree;

[0026] Step 20: Collect the heart rate value, heart rate variability, blood oxygen saturation, and blood pressure value of at least 100 subjects within one minute, and at the same time, have professional doctors judge the fatigue degree of the subjects to construct a fatigue degree dataset;

[0027] Step 21: Train the fatigue level evaluation network using the fatigue level dataset; the dataset is divided into a training set, a validation set, and a test set, and the network is optimized through multiple rounds of iterative training until the preset number of training rounds or performance metrics are reached, completing the training process of the network;

[0028] Step 22: Input the heart rate value of the subject obtained in Step 9, the heart rate variability of the subject obtained in Step 11, the blood oxygen saturation of the subject obtained in Step 14, the blood pressure value of the subject obtained in Step 18, and the corresponding fatigue level into the trained fatigue level evaluation network, and calculate the fatigue level of the subject through network calculation.

[0029] Further, the specific content of Step 2 is as follows:

[0030] Step 2.1: For the video image of the subject obtained in Step 1, use the open-source MediaPipe face detection model to perform face key point detection, locate 468 key points of the face, and obtain the specific coordinate information of each key point;

[0031] Step 2.2: Select the key points marking the face contour, and combine these key points to obtain a face contour mask;

[0032] Step 2.3: Perform a pixel-by-pixel multiplication operation on the face contour mask generated in Step 2.2 and the original video image of the subject obtained in Step 1, and use the mask processing technology to extract the features of the facial area, and finally obtain the facial video image of the subject, excluding the interference of the surrounding environment changes, so as to facilitate subsequent feature extraction.

[0033] Further, the specific content of Step 3 is as follows:

[0034] Step 3.1: Crop and reshape the facial video image of the subject obtained in Step 2 to obtain a video image with a length of n frames, a width and height of 72 pixels each, and three RGB channels, that is, a video image with a shape of n×72×72×3 as the input of the network;

[0035] Step 3.2: Shift the entire R channel of the video image forward by 1 frame, keep the G channel unchanged, and shift the entire B channel backward by 1 frame to obtain the video image after channel shifting; then use 2D convolution to extract features from the obtained video image after channel shifting, that is, complete one TSM convolution operation;

[0036] Step 3.3: Perform global information aggregation on the input feature map using global maximum pooling; then, generate the attention weights for each channel through a fully connected layer, and these weights are normalized to between 0 and 1 through the Sigmoid activation function; finally, apply the calculated channel attention weights to each channel of the input feature map to achieve dynamic weighting of important feature channels;

[0037] Step 3.4: Take the video image obtained in Step 3.1 as the first-layer video image of the Gaussian pyramid. Apply Gaussian blur to the first-layer video image of the Gaussian pyramid using a Gaussian convolution kernel, and perform a downsampling operation to obtain the second-layer video image of the Gaussian pyramid with a shape of n×36×36×3. Apply Gaussian blur to the second-layer video image of the Gaussian pyramid using a Gaussian convolution kernel, and perform a downsampling operation to obtain the third-layer video image of the Gaussian pyramid with a shape of n×18×18×3.

[0038] Step 3.5: Perform two TSM convolution module operations and one max pooling operation on the first-layer video image of the Gaussian pyramid obtained in Step 3.4. Its shape becomes n×36×36×32, and then stack it with the second-layer video image of the Gaussian pyramid that has undergone one TSM convolution module operation to obtain a feature map with a shape of n×36×36×64.

[0039] Step 3.6: Perform two TSM convolution module operations and one max pooling operation on the feature map obtained in Step 3.5. Its shape becomes n×18×18×64, and then stack it with the third-layer video image of the Gaussian pyramid that has undergone two TSM convolution module operations to obtain a feature map with a shape of n×18×18×128.

[0040] Step 3.7: Perform two TSM convolution module operations and one channel attention module operation on the feature map obtained in Step 3.6 to obtain a feature map with a shape of n×18×18×128.

[0041] Step 3.8: Flatten the feature map obtained in Step 3.7 after performing one max pooling operation, and use a fully connected layer for feature output to obtain an rPPG signal with a length of n as the network output.

[0042] Further, the specific content of Step 11 is as follows:

[0043] Step 11.1: In the analysis of the rPPG signal, the time interval between adjacent wave peaks represents the duration of a single cardiac cycle. By extracting the time parameters of the cardiac cycles within one minute from the filtered rPPG signal obtained in Step 7 and calculating the mean value, an index AVNN reflecting the heart rate characteristics can be obtained. Its mathematical expression is:

[0044]

[0045] where RR i is the duration of the i-th cardiac cycle, and n is the total number of cardiac cycles.

[0046] Further, the specific content of Step 12 is as follows:

[0047] Step 12.1: Use the filtered rPPG signal obtained in Step 7 as the input signal of the network. Its shape is a vector of length n. Normalize the input signal to eliminate baseline drift and high-frequency noise;

[0048] Step 12.2: Use a 1D convolutional layer to extract features in the time dimension of the input signal, and use the ReLU function for activation to obtain a time-domain feature map;

[0049] Step 12.3: Perform a fast Fourier transform on the input signal obtained in Step 12.1 to convert the time-domain signal into a frequency-domain signal, and extract harmonic features;

[0050] Step 12.4: Use a 1D convolutional layer to perform a convolution operation on the frequency-domain signal, and use the ReLU function for activation to obtain a frequency-domain feature map;

[0051] Step 12.5: Concatenate the feature map obtained in Step 12.2 and the feature map obtained in Step 12.4 to obtain a time-frequency hybrid feature map;

[0052] Step 12.6: Use a long short-term memory module (LSTM) on the time-frequency hybrid feature map obtained in Step 12.5 to capture the temporal dependence of the input signal and obtain a new feature map;

[0053] Step 12.7: Use a fully connected layer to output the feature map obtained in Step 12.6, and use the Sigmoid function as the activation function to limit the output range between 0 and 1 to obtain the percentage value of blood oxygen saturation.

[0054] Further, the specific content of Step 15 is as follows:

[0055] Step 15.1: Use the obtained rPPG signal as the input signal of the network. First, perform spatial feature extraction through a 1D convolutional layer, and then use the ReLU function as the activation function for non-linear transformation to obtain a preliminary feature map representation;

[0056] Step 15.2: Perform residual processing on the feature map obtained in Step 15.1. Adopt a 4-level residual module network architecture. Each level of the residual module contains two convolutional layers and skip connections. Gradually extract deep physiological features through hierarchical feature transformation to obtain a new feature map;

[0057] Step 15.3: Use 2 fully connected layers to output the feature map obtained in Step 15.2. The output of one fully connected layer points to the systolic blood pressure, and the output of the other fully connected layer points to the diastolic blood pressure. The systolic blood pressure and diastolic blood pressure are used as the two output values of the network.

[0058] Further, the specific content of Step 19 is as follows:

[0059] Step 19.1: Take a total of 5 parameters, namely the heart rate value, heart rate variability, blood oxygen saturation, systolic blood pressure, and diastolic blood pressure of the obtained subject, as the input layer of the network, and perform standardization processing on the input parameters to eliminate the dimension difference;

[0060] Step 19.2: Use two fully connected layers as the hidden layer, use Leaky ReLU as the activation function, and add a Dropout layer to prevent overfitting;

[0061] Step 19.3: Use one fully connected layer as the output layer, use the Sigmoid function as the activation function, limit the output within the range of 0 to 1, and the larger the value, the higher the degree of fatigue.

[0062] The present invention has good comfort, high precision, and an automatic monitoring function, can accurately monitor multiple physiological parameters such as the subject's heart rate, heart rate variability, blood oxygen saturation, blood pressure, and fatigue degree, and at the same time ensure the real-time nature of the monitoring process, meeting the application requirements of multiple fields. Brief Description of the Drawings

[0063] Figure 1 It is the overall flowchart of the physiological parameter monitoring system.

[0064] Figure 2 It is the convolutional neural network diagram for blood oxygen saturation extraction.

[0065] Figure 3 It is the residual network diagram for blood pressure extraction.

[0066] Figure 4 It is the network diagram for fatigue degree assessment.

[0067] Figure 5 It is the convolutional neural network diagram for multi-scale rPPG signal extraction based on channel attention mechanism.

[0068] Figure 6 It is the actual application diagram of the physiological parameter monitoring system.

[0069] Figure 7 It is the visualization interface diagram of the physiological parameter monitoring system.

[0070] Figure 8 It is the schematic diagram of the heart rate detection result of the present invention.

[0071] Figure 9 It is the schematic diagram of the AVNN detection result of the heart rate variability of the present invention.

[0072] Figure 10 It is the schematic diagram of the blood oxygen saturation detection result of the present invention. Detailed Implementation Manner

[0073] The following is a detailed description of a real-time physiological parameter non-intrusive monitoring method and system based on visual perception and intelligent computing according to the present invention with reference to the accompanying drawings:

[0074] Step 1: Build a video image acquisition system, set the camera device at a suitable position 30 cm to 50 cm directly in front of the subject's face, and ensure that the facial video image data of the subject is captured stably, clearly and continuously.

[0075] Step 2: Perform image segmentation processing on the obtained facial video image of the subject, extract the area containing facial features from it, and obtain the facial area video image of the subject.

[0076] Step 2.1: For the subject video image obtained in Step 1, use the open-source MediaPipe face detection model to perform face key point detection, locate 468 key points of the face, and obtain the specific coordinate information of each key point;

[0077] Step 2.2: Select the key points marking the face contour, combine these key points, and obtain the face contour mask;

[0078] Step 2.3: Perform pixel-by-pixel multiplication operation on the face contour mask generated in Step 2.2 and the original video image of the subject obtained in Step 1, and realize the feature extraction of the facial area through the mask processing technology. Finally, obtain the facial video image of the subject, excluding the interference of the surrounding environmental changes, so as to facilitate subsequent feature extraction.

[0079] Step 3: Build a multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism to extract the rPPG signal, and extract the rPPG signal from the facial area video image of the subject obtained in Step 2.

[0080] Step 3.1: Crop and reshape the facial video image of the subject obtained in Step 2 to obtain a video image with a length of n frames, a width and height of 72 pixels each, and having three RGB channels, that is, a video image with a shape of n×72×72×3 as the input of the network;

[0081] Step 3.2: Build a TSM convolutional module. Shift the entire R channel of the video image forward by 1 frame, keep the G channel unchanged, and shift the entire B channel backward by 1 frame to obtain the video image after channel translation. Then use 2D convolution to extract features from the obtained video image after channel translation, that is, complete a TSM convolution operation. Use the TSM convolutional module instead of the traditional 3D convolutional module. On the premise of being able to extract the spatio-temporal features of the video image, the calculation amount and parameters are greatly reduced, and the calculation speed of the algorithm is greatly improved to meet the requirements of real-time monitoring;

[0082] Step 3.3: Build a channel attention module. First, global information aggregation is performed on the input feature map using global max pooling. Then, a fully connected layer is used to generate the attention weights for each channel, and these weights are normalized between 0 and 1 through the Sigmoid activation function. Finally, the calculated channel attention weights are applied to each channel of the input feature map to achieve dynamic weighting of important feature channels, thereby enhancing the network's ability to focus on key features and strengthening the expressive power of feature representation;

[0083] Step 3.4: Build a Gaussian pyramid for video images. The video image obtained in Step 3.1 is used as the first-layer video image of the Gaussian pyramid. The first-layer video image of the Gaussian pyramid is blurred using a Gaussian convolution kernel and then downsampled once to obtain the second-layer video image of the Gaussian pyramid with a shape of n×36×36×3; the second-layer video image of the Gaussian pyramid is blurred using a Gaussian convolution kernel and then downsampled once to obtain the third-layer video image of the Gaussian pyramid with a shape of n×18×18×3;

[0084] Step 3.5: Perform two TSM convolution module operations and one max pooling operation on the first-layer video image of the Gaussian pyramid obtained in Step 3.4. Its shape becomes n×36×36×32, and then it is stacked with the second-layer video image of the Gaussian pyramid that has undergone one TSM convolution module operation to obtain a feature map with a shape of n×36×36×64;

[0085] Step 3.6: Perform two TSM convolution module operations and one max pooling operation on the feature map obtained in Step 3.5. Its shape becomes n×18×18×64, and then it is stacked with the third-layer video image of the Gaussian pyramid that has undergone two TSM convolution module operations to obtain a feature map with a shape of n×18×18×128;

[0086] Step 3.7: Perform two TSM convolution module operations and one channel attention module operation on the feature map obtained in Step 3.6 to obtain a feature map with a shape of n×18×18×128;

[0087] Step 3.8: Flatten the feature map obtained in Step 3.7 after performing one max pooling operation, and use a fully connected layer for feature output to obtain an rPPG signal feature map with a length of n as the network output.

[0088] Step 4: If the convolutional neural network for multi-scale rPPG signal extraction based on the channel attention mechanism has not been completed training, then execute Step 5 for network training and optimization; conversely, if the training of the convolutional neural network for multi-scale rPPG signal extraction based on the channel attention mechanism has been completed and the performance meets the expectations, then skip the training step and directly execute Step 6 for subsequent testing and evaluation.

[0089] Step 5: Use the open-source UBFC dataset to train a convolutional neural network for multi-scale rPPG signal extraction based on channel attention mechanism. Divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1 to ensure data diversity during training and the accuracy of validation. During the training phase, optimize the model parameters through multiple rounds of iteration to gradually improve the performance of the network until the preset number of training rounds or performance metrics are reached, ensuring that the network can converge efficiently and prevent overfitting.

[0090] Step 6: Take the video image of the subject's facial area extracted in Step 2 as input and feed it into the trained convolutional neural network for multi-scale rPPG signal extraction based on channel attention mechanism. After network calculation, output the rPPG signal of the subject.

[0091] Step 7: Filter the rPPG signal obtained in Step 6 using a Butterworth band-pass filter. Set the passband range of the filter to 0.7 Hz to 3.0 Hz to ensure that only the frequency components related to heart rate are retained, filter out low-frequency noise and high-frequency interference, and obtain the filtered rPPG signal.

[0092] Step 8: Perform a fast Fourier transform on the filtered rPPG signal obtained in Step 7 to convert the time-domain rPPG signal into a frequency-domain representation and obtain the corresponding rPPG signal spectrum.

[0093] Step 9: Perform peak analysis on the rPPG signal spectrum obtained in Step 8 to identify the highest peak in the spectrum and its corresponding frequency. This peak frequency corresponds to the frequency of the heart rate. Through further calculation, multiply this frequency value by the time coefficient to obtain the actual heart rate value of the subject.

[0094] Step 10: Detect the peaks of the filtered rPPG signal obtained in Step 7, record the time points at which each peak appears, and calculate the time interval between two adjacent peaks. This time interval represents the duration of each heartbeat of the heart, that is, the cardiac cycle.

[0095] Step 11: Calculate the absolute error between two adjacent cardiac cycles for each cardiac cycle obtained in Step 10, and calculate the average value of the errors for all cardiac cycles to obtain the heart rate variability of the subject.

[0096] Step 11.1: In rPPG signal analysis, the time interval between adjacent peaks characterizes the duration of a single cardiac cycle; by extracting the time parameters of the cardiac cycles within one minute from the filtered rPPG signal obtained in Step 7 and performing a mean calculation, an index AVNN reflecting heart rate characteristics can be obtained, and its mathematical expression is:

[0097]

[0098] In the formula, RR i is the duration of the i-th heartbeat cycle, and n is the total number of heartbeat cycles.

[0099] Step 12: Build a convolutional neural network for extracting blood oxygen saturation to extract the blood oxygen saturation value of the human body. If the convolutional neural network for extracting blood oxygen saturation has not been completed training, then execute Step 13 for further training to ensure that the network parameters are fully adjusted and reach the expected performance level; otherwise, if the network has completed training and performs stably on the validation set, then directly jump to Step 14 for actual prediction and application of blood oxygen saturation.

[0100] Step 12.1: Use the filtered rPPG signal obtained in Step 7 as the input signal of the network. Its shape is a vector with a length of n. Normalize the input signal to eliminate baseline drift and high-frequency noise;

[0101] Step 12.2: Use a 1D convolutional layer to extract the features in the time dimension of the input signal, and use the ReLU function for activation to obtain a time-domain feature map;

[0102] Step 12.3: Perform a fast Fourier transform on the input signal obtained in Step 12.1 to convert the time-domain signal into a frequency-domain signal to extract harmonic features;

[0103] Step 12.4: Use a 1D convolutional layer to perform a convolution operation on the frequency-domain signal, and use the ReLU function for activation to obtain a frequency-domain feature map;

[0104] Step 12.5: Concatenate the feature map obtained in Step 12.2 and the feature map obtained in Step 12.4 to obtain a time-frequency hybrid feature map;

[0105] Step 12.6: Use a long short-term memory module LSTM on the time-frequency hybrid feature map obtained in Step 12.5 to capture the time-dependent relationship of the input signal to obtain a new feature map;

[0106] Step 12.7: Use a fully connected layer to output the feature map obtained in Step 12.6, and use the Sigmoid function as the activation function to limit the output range between 0 and 1 to obtain the percentage value of blood oxygen saturation.

[0107] Step 13: Use the open-source PURE dataset to train the convolutional neural network for blood oxygen saturation extraction. To ensure the diversity of training data and improve the generalization ability of the model, the dataset is divided into a training set, a validation set, and a test set in the ratio of 7:2:1. During the training process, the model continuously optimizes its parameters through multiple rounds of iteration. In each round, it learns from the training set and simultaneously performs real-time validation on the validation set to monitor the training status and validation accuracy of the model. As the training progresses, the network performance gradually improves until the preset number of training rounds or performance metrics are reached.

[0108] Step 14: Take the filtered rPPG signal obtained in Step 7 as the input and feed it into the trained convolutional neural network for blood oxygen saturation extraction. After network calculation, the blood oxygen saturation of the subject is obtained.

[0109] Step 15: Build a residual network for blood pressure extraction to extract the blood pressure value of the human body. If the network has not been completed training, execute Step 16 to ensure that the model can fully learn the data features and achieve the expected training effect; otherwise, if the residual network for blood pressure extraction has been trained and is stable on the validation set, skip the training stage and directly execute Step 18 for blood pressure value prediction and application to achieve blood pressure monitoring.

[0110] Step 15.1: Take the obtained rPPG signal as the input signal of the network. First, perform spatial feature extraction through a 1D convolutional layer, and then use the ReLU function as the activation function for non-linear transformation to obtain a preliminary feature map representation.

[0111] Step 15.2: Perform residual processing on the feature map obtained in Step 15.1. Adopt a 4-level residual module network architecture. Each level of the residual module contains two convolutional layers and skip connections. Gradually extract deep physiological features through hierarchical feature transformation to obtain a new feature map.

[0112] Step 15.3: Use 2 fully connected layers to output the feature map obtained in Step 15.2. The output of one fully connected layer points to the systolic blood pressure, and the output of the other fully connected layer points to the diastolic blood pressure. The systolic blood pressure and diastolic blood pressure are used as the two output values of the network.

[0113] Step 16: Collect the rPPG signals of at least 100 subjects, and at the same time use a medical sphygmomanometer to collect the blood pressure data of the subjects, including systolic blood pressure and diastolic blood pressure, to construct a blood pressure dataset.

[0114] Step 17: Use the blood pressure dataset to train the residual network for blood pressure extraction. Divide it into a training set, a validation set, and a test set in the ratio of 7:2:1 for iterative training of the network until the preset number of training rounds or performance metrics are reached, and the network training is completed.

[0115] Step 18: Input the filtered rPPG signal obtained in Step 7 into the trained blood pressure extraction residual network. Through network calculation, the blood pressure values of the subject are obtained, including systolic blood pressure and diastolic blood pressure.

[0116] Step 19: Build a fatigue degree evaluation network to evaluate the fatigue degree of the subject. If the fatigue degree evaluation network has not been trained yet, execute Step 20 for further training and optimization; on the contrary, if the fatigue degree evaluation network has been trained and its performance is stable, directly execute Step 22 for the actual evaluation and application of the fatigue degree.

[0117] Step 19.1: Take a total of 5 parameters of the obtained heart rate value, heart rate variability, blood oxygen saturation, systolic blood pressure and diastolic blood pressure of the subject as the input layer of the network, and standardize the input parameters to eliminate the dimension difference.

[0118] Step 19.2: Use two fully connected layers as the hidden layer, use Leaky ReLU as the activation function, and add a Dropout layer to prevent overfitting.

[0119] Step 19.3: Use one fully connected layer as the output layer, use the Sigmoid function as the activation function, limit the output to the range of 0 to 1, and the larger the value, the higher the fatigue degree.

[0120] Step 20: Collect the heart rate value, heart rate variability, blood oxygen saturation, and blood pressure value within one minute of at least 100 subjects. At the same time, a professional doctor judges the fatigue degree of the subjects to construct a fatigue degree dataset.

[0121] Step 21: Use the fatigue degree dataset to train the fatigue degree evaluation network. The dataset is divided into a training set, a validation set and a test set according to the ratio of 7:2:1, and the network is optimized through multiple rounds of iterative training until the preset number of training rounds or performance indicators are reached, and the training process of the network is completed.

[0122] Step 22: Input the heart rate value of the subject obtained in Step 9, the heart rate variability of the subject obtained in Step 11, the blood oxygen saturation of the subject obtained in Step 14, the blood pressure value of the subject obtained in Step 18, and the corresponding fatigue degree into the trained fatigue degree evaluation network. Through network calculation, the fatigue degree of the subject is obtained.

[0123] The real-time physiological parameter non-invasive monitoring method based on visual perception and intelligent computing described in the present invention has proven the effectiveness of this method in many datasets.

[0124] Among them, for the heart rate value of the subject obtained in Step 9, it is verified on the UBFC-rPPG dataset, and the heart rate detection accuracy reaches 99%, as Figure 8As shown. According to the standards of heart rate-related medical devices, an instrument with an accuracy rate exceeding 95% is considered effective. The heart rate detection method proposed by the present invention clearly meets the relevant regulations.

[0125] Among them, for the heart rate variability AVNN of the subject obtained in step 11, it is verified on the UBFC-rPPG dataset, and the detection accuracy rate of the heart rate variability AVNN reaches 99%, specifically as Figure 9 shown.

[0126] Among them, for the blood oxygen saturation of the subject obtained in step 14, it is verified on the PURE dataset, and the MAPE value is 0.6, that is, the detection accuracy rate of the blood oxygen saturation is 99.4%. According to the standards of blood oxygen saturation-related medical devices, an instrument with an accuracy rate exceeding 98% is considered effective, as Figure 10 shown. The blood oxygen saturation detection method proposed by the present invention clearly meets the relevant regulations.

[0127] Among them, for the blood pressure value of the subject obtained in step 18, it is verified on the blood pressure dataset composed of 203 people. The proportion of data with an error between the predicted value and the true value within 15 mmHg is 99.43%, and the proportion of data with an error between the predicted value and the true value within 10 mmHg is 92.53%. According to the standards of blood pressure-related medical devices, an instrument with a proportion of data with an error between the predicted value and the true value within 10 mmHg exceeding 85% is considered effective. The blood pressure detection method proposed by the present invention clearly meets the relevant regulations.

Claims

1. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing, the method comprising the following steps: Step 1: Build a video image acquisition system and place the camera device in front of the subject's face to ensure stable, clear and continuous capture of the subject's facial video image data; Step 2: performing image segmentation processing on the acquired facial video image of the subject, extracting the area containing facial features therefrom, and obtaining a video image of the facial area of ​​the subject; Step 3: Build a multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism to extract rPPG signals, and extract rPPG signals from the video image of the subject's facial area obtained in step 2; Step 4: If the multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism has not been trained, execute step 5 to train and optimize the network; conversely, if the multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism has been trained and the performance has reached the expected level, skip the training step and directly execute step 6 for subsequent testing and evaluation; Step 5: Use the open source UBFC dataset to train the multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism; Divide the dataset into training, validation, and test sets to ensure data diversity during training and accuracy of validation; In the training phase, the model parameters are optimized through multiple rounds of iterations to gradually improve the performance of the network until the preset training rounds or performance indicators are reached, ensuring that the network can converge efficiently and prevent overfitting; Step 6: The video image of the subject's facial area extracted in step 2 is used as input and sent to the trained multi-scale rPPG signal extraction convolutional neural network based on the channel attention mechanism for processing. After network calculation, the subject's rPPG signal is output; Step 7: The rPPG signal obtained in step 6 is filtered using a Butterworth bandpass filter; the passband range of the filter is set to 0.7 Hz to 3.0 Hz to ensure that only the frequency components related to the heart rate are retained, and low-frequency noise and high-frequency interference are filtered out to obtain a filtered rPPG signal; Step 8: Perform a fast Fourier transform on the filtered rPPG signal obtained in step 7 to convert the time domain rPPG signal into a frequency domain representation to obtain the corresponding rPPG signal spectrum; Step 9: Perform peak analysis on the rPPG signal spectrum obtained in step 8 to identify the highest peak in the spectrum and its corresponding frequency. The peak frequency corresponds to the heart rate frequency. By calculation, the frequency value is multiplied by the time coefficient to obtain the actual heart rate value of the subject. Step 10: Perform peak detection on the filtered rPPG signal obtained in step 7, record the time point when each peak occurs, and calculate the time interval between two adjacent peaks. This time interval represents the duration of each heart beat, that is, the heartbeat cycle; Step 11: For each heartbeat cycle obtained in step 10, the absolute error between two adjacent heartbeat cycles is calculated, and the average value of the errors of all heartbeat cycles is calculated to obtain the heart rate variability of the subject; Step 12: Build a blood oxygen saturation extraction convolutional neural network to extract the blood oxygen saturation value of the human body; if the blood oxygen saturation extraction convolutional neural network has not been trained, execute step 13 for further training to ensure that the network parameters are fully adjusted and reach the expected performance level; on the contrary, if the network has been trained and performs stably on the validation set, jump directly to step 14 for actual prediction and application of blood oxygen saturation; Step 13: Use the open source PURE dataset to train the blood oxygen saturation extraction convolutional neural network; in order to ensure the diversity of training data and improve the generalization ability of the model, the dataset is divided into training set, validation set and test set; During the training process, the model continuously optimizes parameters through multiple rounds of iterations. Each round is learned through the training set, and real-time verification is performed on the verification set to monitor the training status and verification accuracy of the model. As the training progresses, the network performance gradually improves until the preset training rounds or performance indicators are reached. Step 14: The filtered rPPG signal obtained in step 7 is used as input and sent to the trained blood oxygen saturation extraction convolutional neural network. After network calculation, the blood oxygen saturation of the subject is obtained; Step 15: Build a blood pressure extraction residual network to extract human blood pressure values; if the network has not been trained, execute step 16 to ensure that the model can fully learn the data features and achieve the expected training effect; on the contrary, if the blood pressure extraction residual network has been trained and performs stably on the validation set, skip the training phase and directly execute step 18 to predict and apply blood pressure values ​​to achieve blood pressure monitoring; Step 16: Collect rPPG signals from at least 100 subjects, and use a medical sphygmomanometer to collect blood pressure data of the subjects, including systolic and diastolic pressure, to construct a blood pressure dataset. Step 17: Use the blood pressure data set to train the blood pressure extraction residual network, divide it into a training set, a validation set, and a test set to iteratively train the network until the preset training rounds or performance indicators are reached, and the network training is completed; Step 18: Input the filtered rPPG signal obtained in step 7 into the trained blood pressure extraction residual network, and calculate the blood pressure value of the subject through the network, including systolic pressure and diastolic pressure; Step 19: Build a fatigue assessment network to assess the fatigue level of the subject; if the fatigue assessment network has not been trained, execute step 20 for further training optimization; otherwise, if the fatigue assessment network has been trained and has stable performance, directly execute step 22 to conduct actual assessment and application of fatigue level; Step 20: Collect the heart rate, heart rate variability, blood oxygen saturation, and blood pressure values ​​of at least 100 subjects within one minute, and have professional doctors judge the fatigue level of the subjects to construct a fatigue level data set; Step 21: Use the fatigue level dataset to train the fatigue level assessment network; The data set is divided into training set, validation set and test set, and the network is optimized through multiple rounds of iterative training until the preset training rounds or performance indicators are reached, completing the network training process; Step 22: Send the heart rate value of the subject obtained in step 9, the heart rate variability of the subject obtained in step 11, the blood oxygen saturation of the subject obtained in step 14, the blood pressure value of the subject obtained in step 18 and the corresponding fatigue level to the trained fatigue level assessment network, and calculate the fatigue level of the subject through the network.

2. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: For the subject video image obtained in step 1, use the open source MediaPipe face detection model to detect facial key points, locate the 468 key points of the face, and obtain the specific coordinate information of each key point; Step 2.2: Select key points that mark the face contour, and combine these key points to obtain the face contour mask; Step 2.3: Perform pixel-by-pixel multiplication on the face contour mask generated in step 2.2 and the original video image of the subject obtained in step 1, and realize feature extraction of the facial area through mask processing technology, and finally obtain the facial video image of the subject, eliminating the interference of changes in the surrounding environment, so as to facilitate subsequent feature extraction.

3. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Crop and reshape the facial video image of the subject obtained in step 2 to obtain a video image with a length of n frames, a width and height of 72 pixels, and RGB three channels, that is, a video image with a shape of n×72×72×3 as the input of the network; Step 3.2: Shift the R channel of the video image forward by 1 frame, keep the G channel unchanged, and shift the B channel backward by 1 frame to obtain the video image after channel shifting; extract features from the obtained video image after channel shifting using 2D convolution, thus completing a TSM convolution operation; Step 3.3: Global maximum pooling is used to aggregate global information on the input feature map; then, the attention weights of each channel are generated through a fully connected layer, and these weights are normalized to between 0 and 1 through the Sigmoid activation function; finally, the calculated channel attention weights are applied to each channel of the input feature map to achieve dynamic weighting of important feature channels; Step 3.4: The video image obtained in step 3.1 is used as the first-layer video image of the Gaussian pyramid. The first-layer video image of the Gaussian pyramid is Gaussian blurred using a Gaussian convolution kernel, and a downsampling operation is performed to obtain a second-layer video image of the Gaussian pyramid with a shape of n×36×36×3; the second-layer video image of the Gaussian pyramid is Gaussian blurred using a Gaussian convolution kernel, and a downsampling operation is performed to obtain a third-layer video image of the Gaussian pyramid with a shape of n×18×18×3; Step 3.5: The first-layer video image of the Gaussian pyramid obtained in step 3.4 is subjected to two TSM convolution module operations and a maximum pooling operation, and its shape becomes n×36×36×32, and then stacked with the second-layer video image of the Gaussian pyramid that has undergone a TSM convolution module operation to obtain a feature map with a shape of n×36×36×64; Step 3.6: The feature map obtained in step 3.5 is subjected to two TSM convolution module operations and a maximum pooling operation, and its shape becomes n×18×18×64, and then stacked with the third layer video image of the Gaussian pyramid that has undergone two TSM convolution module operations to obtain a feature map with a shape of n×18×18×128; Step 3.7: Perform two TSM convolution module operations and one channel attention module operation on the feature map obtained in step 3.6 to obtain a feature map with a shape of n×18×18×128; Step 3.8: Flatten the feature map obtained in step 3.7 after performing a maximum pooling operation, use a fully connected layer for feature output, and obtain a feature map rPPG signal with a length of n as the network output.

4. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 11 is specifically as follows: Step 11.1: In the rPPG signal analysis, the time interval between adjacent peaks represents the duration of a single cardiac cycle. The index AVNN reflecting the heart rate characteristics can be obtained by extracting the time parameters of the cardiac cycle in one minute from the filtered rPPG signal obtained in step 7 and performing mean calculation. Its mathematical expression is: In the formula, RR i is the duration of the ith heartbeat cycle, and n is the total number of heartbeat cycles.

5. The method for real-time physiological parameter non-sensing monitoring based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 12 is specifically as follows: Step 12.1: The filtered rPPG signal obtained in step 7 is used as the input signal of the network, and its shape is a vector of length n. The input signal is normalized to eliminate baseline drift and high-frequency noise; Step 12.2: Use a 1D convolutional layer to extract the features of the input signal in the time dimension, and use the ReLU function for activation to obtain a time domain feature map; Step 12.3: Perform fast Fourier transform on the input signal obtained in step 12.1, convert the time domain signal into a frequency domain signal, and extract harmonic features; Step 12.4: Use a 1D convolutional layer to perform a convolution operation on the frequency domain signal and use the ReLU function for activation to obtain a frequency domain feature map; Step 12.5: Concatenate the feature map obtained in step 12.2 and the feature map obtained in step 12.4 to obtain a time-frequency mixed feature map; Step 12.6: Use the long short-term memory module LSTM to capture the time dependency of the input signal on the time-frequency mixed feature map obtained in step 12.5 to obtain a new feature map; Step 12.7: Use a fully connected layer to output the feature map obtained in step 12.6, use the Sigmoid function as the activation function, limit the output range to between 0 and 1, and obtain the percentage value of blood oxygen saturation.

6. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 15 is specifically as follows: Step 15.1: The obtained rPPG signal is used as the input signal of the network. It first passes through a 1D convolutional layer to extract spatial features, and then uses the ReLU function as an activation function for nonlinear transformation to obtain a preliminary feature map representation; Step 15.2: Perform residual processing on the feature map obtained in step 15.1, using a 4-level residual module network architecture, where each level of residual module contains two convolutional layers and skip connections, and gradually extracts deep physiological features through hierarchical feature transformation to obtain a new feature map; Step 15.3: Use two fully connected layers to output the feature maps obtained in step 15.2, where the output of one fully connected layer points to the systolic pressure, and the output of the other fully connected layer points to the diastolic pressure. The systolic pressure and diastolic pressure are used as the two output values ​​of the network.

7. A real-time physiological parameter non-sensing monitoring method based on visual perception and intelligent computing as claimed in claim 1, characterized in that: The step 19 is specifically as follows: Step 19.1: Take the heart rate value, heart rate variability, blood oxygen saturation, systolic blood pressure and diastolic blood pressure of the subjects as the input layer of the network, and standardize the input parameters to eliminate the dimension difference; Step 19.2: Use two fully connected layers as hidden layers, use Leaky ReLU as the activation function, and add a Dropout layer to prevent overfitting; Step 19.3: Use a fully connected layer as the output layer and the Sigmoid function as the activation function to limit the output to the range of 0 to 1. The larger the value, the higher the degree of fatigue.

Citation Information

Cited By

  • Non-contact health monitoring system based on RPPG technology

    CN120770769A