Non-contact body physiological signal acquisition system and method

By using a non-contact physiological signal acquisition system and deep learning and video signal processing technologies, the problems of poor portability, reliance on professional operation, and high cost of existing equipment have been solved, enabling convenient and accurate physiological signal monitoring and health analysis.

WO2026091503A1PCT designated stage Publication Date: 2026-05-07XIAMEN NACHITOZ BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
XIAMEN NACHITOZ BIOTECHNOLOGY CO LTD
Filing Date
2025-05-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing physiological signal acquisition devices suffer from problems such as poor portability, the need for professional operation, high cost, limited accuracy, reliance on user behavior, and inability to achieve remote real-time monitoring.

Method used

A non-contact body physiological signal acquisition system is adopted, including a facial age and gender recognition subsystem, a non-contact remote photoplethysmography (rPPG) pulse wave subsystem, and a heart rate variability and respiratory rate monitoring subsystem based on rPPG. The system uses a high-definition camera and a deep learning model to perform facial image analysis and video signal processing, extract rPPG signals, and perform filtering and frequency analysis.

Benefits of technology

It enables convenient and efficient physiological signal monitoring without skin contact, and can accurately detect health indicators in real time. It is suitable for individuals and various scenarios, reduces hardware and operating costs, reduces energy consumption, and provides comprehensive health analysis and risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025097695_07052026_PF_FP_ABST
    Figure CN2025097695_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention is a non-contact body physiological signal acquisition system. The system comprises: a, a facial age group and gender recognition function subsystem; b, a non-contact remote photoplethysmography function subsystem; and c, an rPPG-based heart rate variability and respiratory rate monitoring function subsystem. Further provided in the present invention is a non-contact body physiological signal acquisition method. The method comprises the following steps: step A, executing a facial age group and gender recognition method; step B, performing non-contact photoplethysmography measurement; and step C, implementing an rPPG-based heart rate variability and respiratory rate monitoring function. The present invention overcomes the defects in the prior art; and compared with a traditional health monitoring method, accurate detection can be realized by means of non-contact body physiological signal monitoring, without the need for a wearable device or skin contact. Such technology can effectively avoid problems such as cross infection, and can also improve the convenience and comfort of health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Non-contact body physiological signal acquisition system and method Technical Field

[0001] This invention relates to the field of vital sign monitoring and data fusion, specifically to a non-contact system and method for acquiring bodily physiological signals. Background Technology

[0002] Currently, physiological signal acquisition devices on the market come in several forms:

[0003] The first type is desktop monitoring equipment, used in hospital departments by professional physicians and nurses, typically for monitoring the routine physiological characteristics of critically ill patients. For example, the C100 cardiovascular monitor manufactured by Shenzhen Coman Medical Equipment Co., Ltd.

[0004] The shortcomings of desktop monitoring equipment are also obvious:

[0005] 1. Poor portability: It is bulky and can usually only be used in fixed locations.

[0006] 2. Complex to use: Requires professional operation and is suitable for medical institutions, not home use.

[0007] 3. High cost: The equipment is expensive and the maintenance cost is also high.

[0008] The second type is wearable monitoring devices, which can enable early detection and prediction of common diseases and sudden illnesses in people with sub-health conditions or latent diseases. However, the data collected by these devices cannot be transmitted in real time, so they cannot achieve remote real-time monitoring and can only be used as designated medical devices in hospitals. For example, the Life Shirt produced by Vivo Metrics in the United States.

[0009] The shortcomings of wearable monitoring devices are also obvious:

[0010] 1. Limited accuracy: Due to the size and design limitations of the sensor, it may be slightly inferior to professional equipment in terms of data accuracy.

[0011] 2. Battery life: Frequent monitoring and data transmission may result in limited device battery life.

[0012] 3. Dependence on user behavior: Deviations in wearing method and position may affect the accuracy of the data.

[0013] The third type is portable monitoring devices, such as the Bioharness portable physiological parameter measurement system from Zephyr in the United States, which embeds sensors into a chest strap. Its use has expanded from being operated by professional physicians in hospital departments to being used by individuals at home.

[0014] The shortcomings of portable monitoring devices are also obvious:

[0015] 1. Inconvenient to use: Although portable, it still requires manual operation and may not be as convenient to use as wearable devices.

[0016] 2. Requires skin contact: Some people with sensitive skin may experience severe discomfort.

[0017] 3. Professional operation is required: Some equipment may require certain professional knowledge or training to use correctly. Summary of the Invention

[0018] To address and resolve the above problems, this invention provides a non-contact body physiological signal acquisition system and method.

[0019] To achieve the above objectives, the technical solution provided by the present invention is as follows:

[0020] A non-contact body physiological signal acquisition system, comprising:

[0021] a. Facial age and gender recognition subsystem, used to enhance the intelligence and personalized service capabilities of the non-contact physiological parameter acquisition system. It uses a high-definition camera to capture the user's facial image and automatically identify the individual's gender, male or female, and age group from the facial image.

[0022] The facial age and gender recognition subsystem includes: a data preprocessing module: during image preprocessing, the input image OriginalImage is converted into a 4D tensor that can be processed by a deep learning model;

[0023] b. Non-contact remote photoplethysmography (rPPG) subsystem, used for rPPG signal extraction based on facial video, optimizing rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies; the non-contact remote photoplethysmography subsystem includes: video input and preprocessing module, time segmentation module, data augmentation module, signal extraction module, model optimization module and total loss function calculation module;

[0024] The video input and preprocessing module is used to detect, align, and crop the face regions in each frame given a facial video sequence.

[0025] Time segmentation module: used to segment the aligned video V into multiple segments, each segment containing T frames; Data augmentation module: used for spatial augmentation: for x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. The data augmentation process randomly selects a video segment from the facial video as input, while retaining the neighboring temporal videos of other segments as input; spatial augmentation and learnable frequency augmentation (LFA) are applied to the input respectively to generate positive and negative video samples;

[0026] Signal extraction module: used to extract positive and negative samples X p and X n The inputs are respectively fed into the local rPPG expert aggregation (REA) module to estimate the corresponding rPPG signals; a local rPPG expert aggregation (REA) module was designed to extract rPPG signals from positive and negative video samples.

[0027] Model optimization module: used for frequency contrast loss: in Y p Between signals in or in Y p and Y n Frequency contrast loss is applied between the input positive and negative sample videos and the rPPG signals generated from temporal neighborhood videos;

[0028] Total Loss Function Calculation Module: Used to calculate the total loss L total The model framework can learn rPPG signals from unlabeled facial videos;

[0029] c. Heart rate variability and respiratory rate monitoring subsystem based on rPPG: This subsystem is used to filter and denoise the obtained rPPG signal. A bandpass filter can be applied to remove noise and retain the heart rate-related frequency components. The maximum peak frequency, i.e., the heart rate, is obtained by analyzing the power spectral density of the rPPG signal through fast Fourier transform.

[0030] The rPPG-based subsystem for monitoring heart rate variability and respiratory rate includes:

[0031] The rPPG signal acquisition and filtering module is used to extract the rPPG signal of the face region from the video frame and predict it through the model. The obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise.

[0032] Heart rate detection module: used to find the maximum peak frequency in the power spectral density, and the corresponding frequency is the heart rate;

[0033] Frequency feature calculation module: used to calculate frequency features using Lomb-Scargle periodograms, and extract low-frequency LF and high-frequency HF power from IBI sequences;

[0034] Respiratory rate (RR) monitoring module: Represented by the high-frequency peak frequency of heart rate variability.

[0035] A non-contact method for acquiring bodily physiological signals includes the following steps:

[0036] Step A: Facial Age and Gender Recognition Method. A high-definition camera is used to capture the user's facial image; the individual's gender (male or female) and age group are automatically identified from the facial image.

[0037] Includes the following steps:

[0038] Step A-1, Data Preprocessing: In the image preprocessing process, the input image OriginalImage is converted into a 4D tensor that can be processed by the deep learning model through a series of steps; these steps include adjusting the image size, performing mean subtraction, RGB channel swapping, and normalization.

[0039] Step A-2: Adjust the size: Adjust the size of the input image to the fixed size of 256x256 pixels that the model accepts, ensuring that the input image size is consistent to meet the structural requirements of the model;

[0040] Step A-3, Mean Subtraction: Subtract a predefined mean from each pixel value to reduce the impact of lighting variations on the model. The predefined mean values ​​for the three channels of BGR are [104, 117, 123].

[0041] NewImage[c,h,w]=OrigianlImage[c,h,w]-mean[c,h,w]

[0042] Where c is the channel index, and h and w are the height and width coordinates of the image, respectively;

[0043] RGB channel swapping: The RGB channels of an image need to be swapped to BGR channels because some models are trained based on the BGR format;

[0044] Step A-4, Normalization: Normalize the range of pixel values ​​in the image to [0,1], that is, divide each pixel value by 255; normalization makes the model easier to train and infer, and reduces the fluctuation of the numerical range.

[0045] Step A-5, Face Detection: By inputting the preprocessed image into the FaceNet face detection model, the face detection results are obtained, and face boxes are selected based on a confidence threshold of 0.7.

[0046] Step A-6, Model forward propagation:

[0047] detections=FaceNet.forward()

[0048] To train the FaceNet face detection model, a triplet loss function was used. A triplet consists of an anchor image, a positive sample image, and a negative sample image; the goal is to ensure that the anchor image is closer to the positive sample image than to the negative sample image, and the difference is at least greater than a predefined margin α.

[0049] f(x a The embedding vector of the anchor image, f(x) p f(x) is the embedding vector of the positive sample image. p1 ) represents the embedding vector of the negative sample image; when an image is input into the FaceNet face detection model, the model processes the image and outputs the detection results. detections is the detection result output by the model; it is a multi-dimensional array containing the confidence, location, and size information of each detected candidate box.

[0050] Step A-7, Confidence Calculation:

[0051] confidence=detections[0,0,i,2];

[0052] In the detection results, the confidence value represents the degree of certainty that the model believes a certain region contains a face; the higher the confidence, the greater the probability that the region contains a face; the confidence value is a value in the detection results, located in the i-th detection result of the detections array;

[0053] Calculation of bounding box coordinates: x1 = int(detection[0,0,i,3]×frameWidth); y1 = int(detection[0,0,i,4]×frameHeight); x2 = int(detection[0,0,i,5]×frameWidth); y2 = int(detection[0,0,i,6]×frameHeight);

[0054] The coordinates of the detection box are proportional values ​​output by the model and need to be multiplied by the width and height of the original image to convert them into actual pixel coordinates. Here, x1, y1 are the coordinates of the top-left corner of the detection box, and x2, y2 are the coordinates of the bottom-right corner. frameWidth and frameHeight are the width and height of the input image, used to convert the relative coordinates of the detection box into absolute pixel coordinates.

[0055] Step A-8, Result Filtering: The detection result is considered valid only if the confidence level exceeds the set threshold of 0.7;

[0056] Step A-9, Gender Detection: The GenderNet model is a deep learning model for gender classification. It takes a face image as input and predicts whether the person in the image is male or female.

[0057] Step A-10, Model Architecture: The GenderNet gender detection model is based on a deep convolutional neural network architecture, containing multiple convolutional layers, pooling layers, and fully connected layers to progressively extract high-level features from the image. The last layer is a Softmax layer, which outputs two nodes, corresponding to male and female respectively. The Softmax function converts the network's output into probability values ​​for the two categories; by selecting the category corresponding to the node with the higher probability value, the model determines the gender of the image as male or female; that is:

[0058] Where z j Category j represents the score for male or female; z k This refers to the scores of all possible categories, namely male and female, which is the sum of the scores of all categories.

[0059] Step A-11, Age Detection: The AgeNet age detection model is also based on a deep convolutional neural network. The last layer of the model, the Softmax layer, outputs a vector containing 8 nodes, each corresponding to an age range: 0-3 years, 4-77 years, 8-14 years, 15-24 years, 25-37 years, 38-47 years, 48-53 years, and 54-100 years. Each range represents a classification output.

[0060] Where z a It is the score for the j-th age group output by the model; z k This refers to the score of all possible categories (8 categories), that is, the sum of the scores of all categories.

[0061] It also includes the following steps:

[0062] Step B, Non-contact Remote Photoplethysmography (rPPG) Measurement: A frequency-dependent deep convolutional neural network (DNN) model is used for remote rPPG signal extraction based on facial videos. This model learns to optimize rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies; specifically including:

[0063] Step B-1, Video Input and Preprocessing: Given a sequence of facial videos, the MTCNN face detector is first used to detect, align, and crop the facial regions in each frame. The aligned video is represented as: V = {v1, v2, ..., v...} T};

[0064] Where vT This represents the T-th frame in the sequence;

[0065] Step B-2, Time Segmentation: Divide the aligned video V into multiple segments, each segment containing T frames. The set of segments is represented as {V1, V2, ..., V...} K}, where V i Let x represent the i-th segment in the sequence. A segment is randomly selected from this sequence as the primary input x. a The remaining segments are considered as time neighbors, denoted as {x} n1 ,x n2 ,...,x nk};

[0066] Step B-3, Data Augmentation: Spatial Augmentation: For x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. Spatial augmentation does not affect the intrinsic rPPG signal in the sample, i.e.:

[0067] Step B-4, Learnable Frequency Enhancement LFA: The LFA module modulates x a The rPPG signal generates a negative sample set. The frequency of the negative sample rPPG signal and x a different:

[0068] Where r i ∈R={r1,r2,...,r M} represents the frequency ratio of LFA module applications;

[0069] Step B-5, Signal Extraction: Extract positive and negative samples X p and X n The inputs are respectively fed into the Local rPPG Expert Aggregation (REA) module to estimate the corresponding rPPG signals:

[0070] in, and

[0071] Step B-6, Model Optimization: Frequency Contrast Loss: In Y p Between signals in or in Y p and Y n Loss due to frequency of application:

[0072] Step B-7, Frequency Ratio Consistency Loss: This loss is used to constrain x. aThe rPPG signals of its time neighbors are similar in frequency, that is:

[0073] Cross-video frequency consistency loss:

[0074] Step B-8, Total Loss Calculation: L total =λ contrast L contrast +λ ratio L ratio +λ cross L cross ;

[0075] Where, λ contrast , λ ratio , λ cross The weighting factor is used; through the above stages, the model framework can learn the rPPG signal from unlabeled facial videos.

[0076] It also includes the following steps:

[0077] Step C, Implementation of heart rate variability and respiratory rate monitoring based on rPPG: This monitoring function is mainly achieved by filtering and denoising the obtained remote photoplethysmography (rPPG) signal. A bandpass filter can be applied to remove noise and retain heart rate-related frequency components. The maximum peak frequency, i.e., heart rate, is obtained by analyzing the power spectral density of the rPPG signal through fast Fourier transform. Heart rate variability includes three attributes: low-frequency (LF), high-frequency (HF), and the LF / HF ratio. These three attributes can be calculated by analyzing the intercardia-interval (IBI) sequence. Respiratory rate is related to LF.

[0078] Specifically, it includes:

[0079] Step C-1: Acquisition and Filtering of rPPG Signal: Extract the rPPG signal of the face region from the video frame and predict it using a model. The obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise. The frequency range of the bandpass filter is set to 0.6Hz to 4Hz. The filtered rPPG signal is represented as: rPPG(t) filtered =Butter_Bandpass(rPPG(t),;owcut=0.6Hz,highcut=4Hz);

[0080] Step C-2, heart rate calculation, is based on the analysis of the power spectral density (PSD) of the rPPG signal using Fast Fourier Transform (FFT); the specific calculation is as follows:

[0081] Fourier Transform: Calculate the FFT of the rPPG signal to obtain the frequency domain signal.

[0082] The Fourier transform operation is a mathematical method that converts a time-domain signal into a frequency-domain signal. It transforms a signal from a time-varying representation (time domain) to a representation of its frequency components (frequency domain).

[0083] Power spectral density (PSD): Calculates the power spectral density of a frequency domain signal: PSD(f) = |FFT(f) rPPG | 2 ;

[0084] Step C-3, Heart Rate Detection: Find the maximum peak frequency in the power spectral density; the corresponding frequency is the heart rate. The formula for calculating heart rate is: HR=argmaxf∈[0.67,3.33]HzPSD(f)×60;

[0085] HR stands for heart rate, measured in bpm (beats per minute). 0.67Hz and 3.33Hz correspond to 40 bpm and 200 bpm, respectively, and are commonly used heart rate frequency ranges.

[0086] Heart rate variability (HRV) frequency domain characteristics mainly include low-frequency LF and high-frequency HF power and their ratio;

[0087] The interbeat interval (IBI) refers to the time interval between consecutive heartbeat peaks given a sampling rate of sig_fps.

[0088] The formula for calculating IBI is:

[0089] Among them, t i It is the position of the i-th heart rate peak, IBI i It is the interval between the i-th and (i+1)-th peaks;

[0090] Step C-4, Calculation of frequency characteristics: Frequency characteristics are calculated using the Lomb-Scargle periodogram to extract low-frequency (LF) and high-frequency (HF) power from the IBI sequence; the specific steps are as follows:

[0091] Step C-4i, Spectrum Estimation: The IBI sequence is subjected to spectral analysis using the Lomb-Scargle method to obtain the frequency (freq) and power spectrum (power).

[0092] Lomb-Scargle(IBI)→(freq,power);

[0093] Step C-4ii, Low-frequency and high-frequency power: Select the power in the low-frequency range of 0.04Hz-0.15Hz and the high-frequency range of 0.15Hz-0.4Hz within the frequency range. power_LF=∑freq∈[0.04,0.15]power(freq); power_HF=∑freq∈(0.15,0.40]power(freq);

[0094] Step C-4iii, Peak frequency of the spectrum: Find the peak frequency corresponding to the highest power in the low frequency and high frequency regions respectively: freq_LF_peak=argmaxfreq∈[0.04,0.15]power(freq); freq_HF_peak=argmaxfreq∈(0.15,0.40]power(freq);

[0095] Step C-4iiiii, Frequency Domain Characteristics:

[0096] The low-frequency peak frequency is freq_LF_peak, which reflects the dominant frequency of the heart rate variability signal in the low-frequency range;

[0097] The high-frequency peak frequency is freq_HF_peak, which reflects the main frequency of the heart rate variability signal in the high-frequency range;

[0098] Low-frequency power is denoted as power_LF. Low-frequency power is usually used to measure the overall activity level of the sympathetic and parasympathetic nervous systems.

[0099] High-frequency power is denoted as power_HF and is typically used to measure the activity of the parasympathetic nervous system.

[0100] Normalize the low-frequency power and high-frequency power:

[0101] Step C-5, respiratory rate (RR) monitoring: respiratory rate can be represented by the high-frequency peak frequency of heart rate variability.

[0102] The present invention has the following beneficial effects at the application level:

[0103] This invention overcomes the shortcomings of existing technologies. Compared with traditional health monitoring methods, non-contact physiological signal monitoring achieves accurate detection without the need for wearable devices or skin contact. This technology effectively avoids cross-infection and other problems, while improving the convenience and comfort of health monitoring. It enables a high degree of automation, making the collection, processing, and analysis of monitoring data more convenient and efficient. By using big data analytics and artificial intelligence technologies, it can accurately identify various health indicators and provide more comprehensive health analysis and risk assessment services. It also has the advantage of real-time monitoring, allowing for better monitoring of disease progression and treatment effectiveness, while also helping people predict and prevent health problems in a timely manner. Furthermore, this technology is not only suitable for personal use but also has broad application prospects, such as in hospitals, highways, and sports events.

[0104] In addition to its scalability and compatibility, this invention can seamlessly integrate with other health management systems or smart devices to form a complete health monitoring ecosystem. Furthermore, it can be flexibly adjusted according to specific application scenarios, making it suitable for different groups, such as the elderly, patients with chronic diseases, or high-intensity athletes. Regarding energy saving and environmental protection, its non-contact design reduces energy consumption, minimizing battery usage and replacement needs, and features low power consumption, aligning with green environmental protection principles. In terms of efficiency and multitasking, this invention can not only monitor multiple physiological parameters in real time but also perform data analysis and risk assessment simultaneously, saving user time and improving diagnostic and treatment efficiency. Moreover, in terms of cost, compared to traditional vital sign monitoring devices, this invention reduces hardware costs, and the application of big data and artificial intelligence technologies further lowers long-term operating and maintenance costs, resulting in greater economic viability.

[0105] The present invention has the following advantages at the technical level:

[0106] 1) Enhanced Robustness Through Multi-ROI Extraction: This invention, based on rPPG signal detection, effectively reduces errors caused by environmental factors (such as uneven illumination or occlusion) affecting a single ROI by selecting multiple regions of interest (ROIs) on the skin for signal extraction. Through multi-region signal fusion processing, artifacts and noise interference are reduced, significantly improving the robustness of heart rate and other physiological parameter detection.

[0107] 2) Signal Enhancement and Noise Suppression: After extracting rPPG signals from multiple ROIs, an averaging process was used to further enhance the signal reliability while eliminating noise caused by local skin movement, micro-expression changes, and light reflection. This algorithm design makes the final extracted physiological signal more stable and has smaller errors, ensuring high-precision measurements in different application scenarios.

[0108] 3) Combination of time-domain and frequency-domain analysis: In the signal processing stage, this algorithm not only relies on time-domain information but also analyzes frequency-domain information through Fourier transform to accurately extract the frequency features of the target physiological signal, such as heart rate. This dual analysis method improves the accuracy of signal extraction, especially maintaining high detection performance in moving and non-stationary scenarios.

[0109] 4) Multi-frame fusion enhances signal quality: By fusing signals from multiple time frames, the stability and accuracy of the rPPG signal are further improved. Even if the signal may fluctuate slightly in a short period of time, the fusion of multiple frames can effectively smooth these fluctuations, ensuring the consistency and continuity of the output signal.

[0110] 5) Real-time performance and parallel computing optimization: This invention employs parallel processing technology, which reduces computation time while maintaining high computational accuracy when extracting and fusing multiple ROI signals, thereby enabling real-time monitoring of physiological signals. This is of great significance for applications requiring continuous monitoring, such as hospitals or sports settings.

[0111] 6) Algorithm Adaptability and Scene Optimization: This algorithm features scene adaptability, automatically adjusting signal extraction parameters based on different lighting conditions and skin characteristics (such as color and texture) to ensure reliable rPPG signals in various environments. This adaptive algorithm design allows the system to be widely applied across different populations and scenarios. Attached Figure Description

[0112] Figure 1 is a schematic diagram of the system framework of the present invention.

[0113] Figure 2 is a schematic diagram of the model framework of the present invention.

[0114] Figure 3 is a schematic diagram of the heart rate (HR), heart rate variability (high frequency peak frequency, low frequency peak frequency, high frequency power, high frequency power) and respiratory rate based on rPPG wave extraction according to the present invention. Detailed Implementation

[0115] The specific embodiments of the present invention will be described with reference to Figures 1-3 and the examples:

[0116] Example 1

[0117] A non-contact body physiological signal acquisition system, characterized by comprising:

[0118] a. The facial age and gender recognition subsystem is used to enhance the intelligence and personalized service capabilities of the non-contact physiological parameter acquisition system. It uses a high-definition camera to capture the user's facial image and automatically identifies the individual's gender, male or female, and age group from the facial image.

[0119] The facial age and gender recognition subsystem includes: a data preprocessing module: during image preprocessing, the input image OriginalImage is converted into a 4D tensor that can be processed by a deep learning model;

[0120] b. Non-contact remote photoplethysmography (rPPG) subsystem, used for rPPG signal extraction based on facial video, optimizing rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies; the non-contact remote photoplethysmography subsystem includes: video input and preprocessing module, time segmentation module, data augmentation module, signal extraction module, model optimization module and total loss function calculation module;

[0121] The video input and preprocessing module is used to detect, align, and crop the face regions in each frame given a facial video sequence.

[0122] Time segmentation module: used to segment the aligned video V into multiple segments, each segment containing T frames; Data augmentation module: used for spatial augmentation: for x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. The data augmentation process randomly selects a video segment from the facial video as input, while retaining the neighboring temporal videos of other segments as input; spatial augmentation and learnable frequency augmentation (LFA) are applied to the input respectively to generate positive and negative video samples;

[0123] Signal extraction module: used to extract positive and negative samples X p and X n The inputs are respectively fed into the local rPPG expert aggregation (REA) module to estimate the corresponding rPPG signals; a local rPPG expert aggregation (REA) module was designed to extract rPPG signals from positive and negative video samples.

[0124] Model optimization module: used for frequency contrast loss: in Y p Between signals in or in Y p and Y n Frequency contrast loss is applied between the input positive and negative sample videos and the rPPG signals generated from temporal neighborhood videos;

[0125] Total Loss Function Calculation Module: Used to calculate the total loss L total The model framework can learn rPPG signals from unlabeled facial videos;

[0126] c. Heart rate variability and respiratory rate monitoring subsystem based on rPPG: This subsystem is used to filter and denoise the obtained rPPG signal. A bandpass filter can be applied to remove noise and retain the heart rate-related frequency components. The maximum peak frequency, i.e., the heart rate, is obtained by analyzing the power spectral density of the rPPG signal through fast Fourier transform.

[0127] The rPPG-based subsystem for monitoring heart rate variability and respiratory rate includes:

[0128] The rPPG signal acquisition and filtering module is used to extract the rPPG signal of the face region from the video frame and predict it through the model. The obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise.

[0129] Heart rate detection module: used to find the maximum peak frequency in the power spectral density, and the corresponding frequency is the heart rate;

[0130] Frequency feature calculation module: used to calculate frequency features using Lomb-Scargle periodograms, and extract low-frequency LF and high-frequency HF power from IBI sequences;

[0131] Respiratory rate (RR) monitoring module: Represented by the high-frequency peak frequency of heart rate variability.

[0132] Example 2

[0133] 1. Facial age and gender recognition functionality implemented:

[0134] 1.1 Brief Description

[0135] Facial recognition for age and gender can enhance the intelligence and personalized service capabilities of contactless physiological parameter acquisition systems. Compared to traditional physiological monitoring equipment, contactless devices incorporating facial recognition can provide more accurate health monitoring by combining personal information without the user's awareness.

[0136] The main implementation methods are: first, to capture the user's facial image through a high-definition camera; and second, to automatically identify the individual's gender ("male" or "female") and age group (such as 0-2 years old, 4-6 years old, etc.) from the facial image.

[0137] 1.2 Technical Implementation Scheme:

[0138] Data preprocessing: In image preprocessing, the input image (OriginalImage) needs to go through a series of steps to be converted into a 4D tensor that can be processed by deep learning models. These steps include resizing the image, performing mean subtraction, RGB channel swapping, and normalization.

[0139] Resize: Adjust the size of the input image to a fixed size of 256x256 pixels that the model accepts, ensuring that the input image size is consistent to fit the structural requirements of the model.

[0140] Mean subtraction: Subtract a predefined mean from each pixel value to reduce the impact of lighting changes on the model. The predefined mean of the three channels of BGR is [104, 117, 123]: NewImage[c, h, w] = OriginalImage[c, h, w] - mean[c, h, w];

[0141] Where c is the channel index, and h and w are the height and width coordinates of the image, respectively.

[0142] RGB channel swapping: The RGB channels of an image need to be swapped to BGR channels because some models are trained based on the BGR format.

[0143] Normalization: Normalize the range of pixel values ​​in the image to [0,1], that is, divide each pixel value by 255.

[0144] Normalization makes models easier to train and infer, reducing fluctuations in the numerical range.

[0145] Face detection: By inputting the preprocessed image into the FaceNet face detection model, the face detection results are obtained, and face boxes are selected based on a confidence threshold of 0.7.

[0146] Model forward propagation:

[0147] detections=FaceNet.forward();

[0148] To train the FaceNet face detection model, a triplet loss function was used. A triplet consists of an anchor image, a positive sample image, and a negative sample image. The goal is to ensure that the anchor image is closer to the positive sample image than to the negative sample image, and the distance is at least greater than a predefined margin α. f(x) a The embedding vector of the anchor image, f(x) p f(x) is the embedding vector of the positive sample image. p1 ) represents the embedding vector of the negative sample image. When an image is input into the FaceNet face detection model, the model processes the image and outputs the detection results. `detections` is the detection result output by the model; it is a multi-dimensional array containing the confidence, location, and size information of each detected candidate box.

[0149] Confidence calculation:

[0150] confidence=detections[0,0,i,2];

[0151] In the detection results, the confidence value represents the model's degree of certainty that a certain region contains a face. The higher the confidence value, the greater the probability that the region contains a face. The confidence value is a value in the detection results, located in the i-th detection result of the `detections` array.

[0152] Calculation of bounding box coordinates: x1 = int(detection[0,0,i,3]×frameWidth); y1 = int(detection[0,0,i,4]×frameHeight); x2 = int(detection[0,0,i,5]×frameWidth); y2 = int(detection[0,0,i,6]×frameHeight);

[0153] The coordinates of the detection box are proportional values ​​output by the model and need to be multiplied by the width and height of the original image to convert them into actual pixel coordinates. Here, x1, y1 are the coordinates of the top-left corner of the detection box, and x2, y2 are the coordinates of the bottom-right corner. frameWidth and frameHeight are the width and height of the input image, used to convert the relative coordinates of the detection box into absolute pixel coordinates.

[0154] Result filtering: The detection result is considered valid only if the confidence level exceeds the set threshold of 0.7.

[0155] Gender Detection: GenderNet is a deep learning model for gender classification. It takes a face image as input and predicts whether the person in the image is male or female.

[0156] Model Architecture: The GenderNet gender detection model is based on a deep convolutional neural network architecture, containing multiple convolutional layers, pooling layers, and fully connected layers to progressively extract high-level features from the image. The final layer is a Softmax layer, which outputs two nodes, corresponding to male and female respectively. The Softmax function converts the network's output into probability values ​​for the two categories. By selecting the category corresponding to the node with the higher probability value, the model determines the gender of the image as either "male" or "female."

[0157] Where z j It is the score for category j (male or female); z kThis refers to the score of all possible categories (male and female), that is, the sum of the scores of all categories.

[0158] Age Detection: The AgeNet age detection model is also based on a deep convolutional neural network. The model's last layer, the Softmax layer, outputs a vector containing 8 nodes, each node corresponding to an age range (e.g., 0-3 years, 4-7 years, 8-14 years, 15-24 years, 25-37 years, 38-47 years, 48-53 years, 54-100 years). Each range represents a classification output.

[0159] Where z a It is the score for the a-th age group output by the model; z k This refers to the score of all possible categories (8 categories), that is, the sum of the scores of all categories.

[0160] The final input and output patterns of the face age and gender recognition function are shown in Figure 1:

[0161] Figure 1 shows the system framework of the present invention, which consists of three parts: FaceNet, a face detection model for recognizing face regions; GenderNet, a gender detection model for recognizing gender; and AgeNet, an age detection model for recognizing age.

[0162] 2. Non-contact measurement method - implementation of remote photoplethysmography (PPG) function:

[0163] 2.1 Brief Description

[0164] Remote photoplethysmography (rPPG) is a technique that uses optical sensors to detect changes in blood flow on the skin surface, commonly used for non-invasive measurement of physiological parameters such as heart rate and blood pressure. rPPG extracts pulse wave information related to cardiac activity by analyzing the light signals reflected from facial skin. We designed a novel frequency-dependent deep convolutional neural network (DNN) model for rPPG signal extraction based on facial videos. This model learns to optimize rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies. It has six main stages: video input and preprocessing, time segmentation, data augmentation, signal extraction, model optimization, and total loss function calculation.

[0165] 2.2 Technical Implementation Scheme

[0166] Video Input and Preprocessing: Given a sequence of facial videos, the MTCNN face detector is first used to detect, align, and crop the face regions in each frame. The aligned video is represented as: V = {v1, v2, ..., v...} T};

[0167] Where vT This represents the T-th frame in the sequence.

[0168] Time segmentation: The aligned video V is divided into multiple segments, each containing T frames. The set of segments is represented as {V1, V2, ..., V...} K}, where V i Let x represent the i-th segment in the sequence. A segment is randomly selected from this sequence as the primary input x. a The remaining segments are considered as time neighbors, denoted as {x} n1 ,x n2 ,...,x nk};

[0169] Data Augmentation: Spatial Augmentation: for x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. Spatial augmentation does not affect the intrinsic rPPG signal in the sample, i.e.:

[0170] Learnable Frequency Enhancement LFA: The LFA module modulates x a The rPPG signal generates a negative sample set. The frequency of the negative sample rPPG signal and x a different:

[0171] Where r i ∈R={r1,r2,...,r M} represents the frequency ratio of LFA module applications.

[0172] Signal extraction: This involves extracting positive and negative samples X. p and X n The inputs are respectively fed into the local rPPG expert aggregation (RE) module to estimate the corresponding rPPG signals:

[0173] in, and

[0174] Model optimization: Frequency contrast loss: in Y p Between signals in or in Y p and Y n Loss due to frequency of application:

[0175] Frequency ratio consistency loss: This loss is used to constrain x a The rPPG signals of its time neighbors are similar in frequency, that is:

[0176] Cross-video frequency consistency loss:

[0177] Total loss calculation: L total =λ contrast L contrast +λ ratio L ratio +λ cross L cross ;

[0178] Where, λ contrast , λ ratio , λ cross This is the weighting factor.

[0179] Through the above stages, the model framework can learn rPPG signals from unlabeled facial videos. The model overview is shown in Figure 2.

[0180] Figure 2 shows the model framework of this invention, which includes three main stages: 1) Data Augmentation: A video segment is randomly selected from the facial video as input, while neighboring temporal videos of other segments are retained. Spatial augmentation and learnable frequency augmentation (LFA) are applied to the input to generate positive and negative video samples. 2) Signal Extraction: A local rPPG expert aggregation (REA) module is designed to extract rPPG signals from the positive and negative video samples. 3) Model Optimization: The model is optimized for the rPPG signals generated from the input positive and negative sample videos and the temporal neighboring videos.

[0181] 3. Implementation of heart rate, heart rate variability, and respiratory rate monitoring functions based on rPPG

[0182] 3.1 Brief Description

[0183] This monitoring function is primarily achieved by filtering and denoising the obtained rPPG signal. A bandpass filter (typically between 0.5Hz and 4Hz) can be applied to remove noise and retain heart rate-related frequency components. The maximum peak frequency, i.e., the heart rate, is obtained by analyzing the power spectral density of the rPPG signal using a Fast Fourier Transform (FFT). Heart rate variability includes three attributes: low frequency (LF), high frequency (HF), and the LF / HF ratio. These three attributes can be calculated by analyzing inter-cardiac interval (IBI) sequences, while respiratory rate is correlated with LF.

[0184] 3.2 Technical Implementation Scheme

[0185] rPPG signal acquisition and filtering: The rPPG signal of the face region is extracted from the video frame and predicted using a model. The obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise. The frequency range of the bandpass filter is set to 0.6Hz to 4Hz. The filtered rPPG signal is represented as: rPPG(t) filtered =Butter_Bandpass(rPPG(t);owcut=0.6Hz, highcut=4Hz);

[0186] Heart rate is calculated based on the analysis of the power spectral density (PSD) of the rPPG signal using Fast Fourier Transform (FFT). The specific calculation is as follows:

[0187] Fourier Transform: Calculate the FFT of the rPPG signal to obtain the frequency domain signal.

[0188] The Fourier transform operation is a mathematical method that converts a time-domain signal into a frequency-domain signal. It transforms a signal from a time-varying representation (time domain) to a representation of its frequency components (frequency domain).

[0189] Power spectral density (PSD): Calculate the power spectral density of a frequency domain signal as PSD(f) = |FFT(f)|. rPPG | 2 ;

[0190] Heart rate detection: Find the maximum peak frequency in the power spectral density; the corresponding frequency is the heart rate. The formula for calculating heart rate is: HR=argmaxf∈[0.67,3.33]HzPSD(f)×60;

[0191] HR stands for heart rate, measured in bpm (heart beats per minute). 0.67Hz and 3.33Hz correspond to 40 bpm and 200 bpm, respectively, and are commonly used heart rate ranges.

[0192] Heart rate variability (HRV) frequency domain characteristics mainly include low-frequency (LF) and high-frequency (HF) power and their ratio.

[0193] The interbeat interval (IBI) is the time interval between consecutive heartbeat peaks. Given a sampling rate of sig_fps...

[0194] The formula for calculating IBI is:

[0195] Among them, t i It is the position of the i-th heart rate peak, IBI i It is the interval between the i-th and i+1-th peaks.

[0196] Frequency characteristics were calculated using a Lomb-Scargle periodogram to extract low-frequency (LF) and high-frequency (HF) power from the IBI sequence. The specific steps are as follows:

[0197] i. Spectrum estimation: The IBI sequence is subjected to spectrum analysis using the Lomb-Scargle method to obtain the frequency freq and power spectrum power; Lomb-Scargle(IBI)→(freq,power);

[0198] ii. Low-frequency and high-frequency power: Select low-frequency (0.04Hz-0.15Hz) and high-frequency power within the frequency range.

[0199] Power in the (0.15Hz-0.4Hz) region.

[0200] power_LF=∑freq∈[0.04,0.15]power(freq);

[0201] power_HF=∑freq∈[0.15,0.40]power(freq);

[0202] iii. Peak frequency of the spectrum: Find the peak frequency corresponding to the highest power in the low-frequency and high-frequency regions respectively.

[0203] freq_LF_peak=argmaxfreq∈[0.04,0.15]power(freq);

[0204] freq_HF_peak=argmaxfreq∈[0.15,0.40]power(freq);

[0205] iv. Frequency domain characteristics:

[0206] The low-frequency peak frequency, freq_LF_peak, reflects the dominant frequency of heart rate variability signals in the low-frequency range. It is typically associated with the activity of the sympathetic and parasympathetic nervous systems and is an indicator of heart rate regulation.

[0207] The high-frequency peak frequency, freq_HF_peak, reflects the dominant frequency of the heart rate variability signal in the high-frequency range. It is usually associated with the activity of the parasympathetic nervous system (especially the vagus nerve). The frequency corresponding to the high-frequency peak frequency usually matches the respiratory rate.

[0208] Low-frequency power, denoted as power_LF, is commonly used to measure the overall activity level of the sympathetic and parasympathetic nervous systems. Higher low-frequency power indicates increased activity of the sympathetic nervous system and is typically associated with blood pressure fluctuations and stress responses.

[0209] High-frequency power (power_HF) is commonly used to measure the activity of the parasympathetic nervous system. Higher high-frequency power indicates a stronger influence of the parasympathetic nervous system on cardiac activity and is usually associated with resting or relaxed states.

[0210] To reduce the impact of individual differences and signal strength, highlight the relative activity of the sympathetic and parasympathetic nervous systems, and eliminate the influence of environmental and measurement conditions, low-frequency and high-frequency power are usually normalized.

[0211] The low-frequency to high-frequency ratio (LF / HF) is an important indicator in the frequency domain analysis of heart rate variability. This ratio is often used to assess the relative balance between sympathetic and parasympathetic activity in the autonomic nervous system. A low LF / HF ratio (usually <1) generally indicates parasympathetic dominance, typically associated with relaxation, rest, or recovery. A moderate LF / HF ratio (usually between 1 and 3) suggests a relatively balanced sympathetic and parasympathetic nervous system. This is generally considered a "normal" or healthy state of autonomic nervous system balance. A high LF / HF ratio (usually >3) may indicate sympathetic dominance, potentially associated with stress, anxiety, physical activity, or sympathetic hyperactivity.

[0212] Respiratory rate (RR) monitoring: Because respiratory activity directly affects the heart rhythm, this phenomenon is called respiratory sinus arrhythmia (RSA). During different phases of the respiratory cycle, heart rate naturally fluctuates, with these fluctuations concentrated in the high-frequency (HF) range of the HRV spectrum. Therefore, respiratory rate can be represented by the high-frequency peak frequency of heart rate variability.

[0213] Based on this, the flowchart of the technical implementation is shown in Figure 3:

[0214] Figure 3 shows the heart rate (HR), heart rate variability (high-frequency peak frequency, low-frequency peak frequency, high-frequency power, high-frequency power), and respiratory rate extracted based on rPPG waves according to the present invention.

[0215] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the present invention without departing from its novel spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A non-contact body physiological signal acquisition system, characterized in that... include: a. Facial age and gender recognition subsystem, used to enhance the intelligence and personalized service capabilities of the non-contact physiological parameter acquisition system. It uses a high-definition camera to capture the user's facial image and automatically identify the individual's gender, male or female, and age group from the facial image. The facial age and gender recognition subsystem includes: a data preprocessing module: during image preprocessing, the input image OriginalImage is converted into a 4D tensor that can be processed by a deep learning model; b. Non-contact remote photoplethysmography (rPPG) subsystem, used for rPPG signal extraction based on facial video, optimizing rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies; the non-contact remote photoplethysmography subsystem includes: video input and preprocessing module, time segmentation module, data augmentation module, signal extraction module, model optimization module and total loss function calculation module; The video input and preprocessing module is used to detect, align, and crop the face regions in each frame given a facial video sequence. Time segmentation module: used to segment the aligned video V into multiple segments, each segment containing T frames; Data augmentation module: used for spatial augmentation: for x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. The data augmentation process randomly selects a video segment from the facial video as input, while retaining the neighboring temporal videos of other segments as input; spatial augmentation and learnable frequency augmentation (LFA) are applied to the input respectively to generate positive and negative video samples; Signal extraction module: used to extract positive and negative samples X p and X n The inputs are respectively fed into the local rPPG expert aggregation (REA) module to estimate the corresponding rPPG signals; a local rPPG expert aggregation (REA) module was designed to extract rPPG signals from positive and negative video samples. Model optimization module: used for frequency contrast loss: in Y p Between signals in or in Y p and Y n Frequency contrast loss is applied between the input positive and negative sample videos and the rPPG signals generated from temporal neighborhood videos; Total Loss Function Calculation Module: Used to calculate the total loss L total The model framework can learn rPPG signals from unlabeled facial videos; c. Heart rate variability and respiratory rate monitoring subsystem based on rPPG: This subsystem is used to filter and denoise the obtained rPPG signal. A bandpass filter can be applied to remove noise and retain the heart rate-related frequency components. The maximum peak frequency, i.e., the heart rate, is obtained by analyzing the power spectral density of the rPPG signal through fast Fourier transform. The rPPG-based subsystem for monitoring heart rate variability and respiratory rate includes: The rPPG signal acquisition and filtering module is used to extract the rPPG signal of the face region from the video frame and predict it through the model. The obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise. Heart rate detection module: used to find the maximum peak frequency in the power spectral density, and the corresponding frequency is the heart rate; Frequency feature calculation module: used to calculate frequency features using Lomb-Scargle periodograms, and extract low-frequency LF and high-frequency HF power from IBI sequences; Respiratory rate (RR) monitoring module: Represented by the high-frequency peak frequency of heart rate variability.

2. A non-contact method for acquiring bodily physiological signals, characterized in that... Includes the following steps: Step A: Facial Age and Gender Recognition Method. A high-definition camera is used to capture the user's facial image; the individual's gender (male or female) and age group are automatically identified from the facial image. Includes the following steps: Step A-1, Data Preprocessing: In the image preprocessing process, the input image OriginalImage is converted into a 4D tensor that can be processed by the deep learning model through a series of steps; these steps include adjusting the image size, performing mean subtraction, RGB channel swapping, and normalization. Step A-2: Adjust the size: Adjust the size of the input image to the fixed size of 256x256 pixels that the model accepts, ensuring that the input image size is consistent to meet the structural requirements of the model; Step A-3, Mean Subtraction: Subtract a predefined mean from each pixel value to reduce the impact of lighting changes on the model; the predefined mean values ​​for the three channels of BGR are [104, 117, 123]: NewImage[c,h,w]=OrigianImage[c,h,w]-mean[c,h,w]; Where c is the channel index, and h and w are the height and width coordinates of the image, respectively; RGB channel swapping: The RGB channels of an image need to be swapped to BGR channels because some models are trained based on the BGR format; Step A-4, Normalization: Normalize the range of pixel values ​​in the image to [0,1], that is, divide each pixel value by 255; normalization makes the model easier to train and infer, and reduces the fluctuation of the numerical range. Step A-5, Face Detection: By inputting the preprocessed image into the FaceNet face detection model, the face detection results are obtained, and face boxes are selected based on a confidence threshold of 0.

7. Step A-6, Model forward propagation: detections=FaceNet.forward(); To train the FaceNet face detection model, a triplet loss function was used; a triplet consists of an anchor image, a positive sample image, and a negative sample image; the goal is to ensure that the anchor image is closer to the positive sample image than to the negative sample image, and the difference is at least greater than a predefined margin α. f(x a The embedding vector of the anchor image, f(x) p f(x) is the embedding vector of the positive sample image. p1 ) represents the embedding vector of the negative sample image; when the image is input into the FaceNet face detection model, the model processes the image and outputs the detection results; detections is the detection results output by the model, which is a multi-dimensional array containing the confidence, location, and size information of each detected candidate box; Step A-7, Confidence Calculation: confidence=detections[0,0,i,2]; In the detection results, the confidence value represents the degree of certainty that the model believes a certain region contains a face; the higher the confidence, the greater the probability that the region contains a face; the confidence value is a value in the detection results, located in the i-th detection result of the detections array; Bounding Box Coordinate Calculation: x1=int(detection[0,0,i,3]×frameWidth); y1=int(detection[0,0,i,4]×frameHeight); x2=int(detection[0,0,i,5]×frameWidth); y2=int(detection[0,0,i,6]×frameHeight); The coordinates of the detection box are proportional values ​​output by the model and need to be multiplied by the width and height of the original image to convert them into actual pixel coordinates. Here, x1, y1 are the coordinates of the top left corner of the detection box, and x2, y2 are the coordinates of the bottom right corner of the detection box. frameWidth and frameHeight are the width and height of the input image, used to convert the relative coordinates of the detection box into absolute pixel coordinates. Step A-8, Result Filtering: The detection result is considered valid only if the confidence level exceeds the set threshold of 0.7; Step A-9, Gender Detection: The GenderNet model is a deep learning model for gender classification; it takes a face image as input and predicts whether the person in the image is male or female. Step A-10, Model Architecture: The GenderNet gender detection model is based on a deep convolutional neural network architecture, containing multiple convolutional layers, pooling layers, and fully connected layers to progressively extract high-level features from the image. The last layer is a Softmax layer, which outputs two nodes, corresponding to male and female respectively. The Softmax function converts the network's output into probability values ​​for the two categories. By selecting the category corresponding to the node with the higher probability value, the model determines whether the gender in the image is male or female; that is: Where z j Category j represents the score for male or female; z k This refers to the scores of all possible categories, namely male and female, which is the sum of the scores of all categories. Step A-11, Age Detection: The AgeNet age detection model is also based on a deep convolutional neural network. The last layer of the model, the Softmax layer, outputs a vector containing 8 nodes, each corresponding to an age range: 0-3 years, 4-7 years, 8-14 years, 15-24 years, 25-37 years, 38-47 years, 48-53 years, and 54-100 years. Each range represents a classification output. Where z a It is the score for the a-th age group output by the model; z k This refers to the scores of all possible categories, that is, the scores of all 8 categories, which is the sum of the scores of all categories.

3. The non-contact body physiological signal acquisition method according to claim 2, characterized in that... It also includes the following steps: Step B, Non-contact Remote Photoplethysmography (rPPG) Measurement: A frequency-dependent deep convolutional neural network (DNN) model is used for remote rPPG signal extraction based on facial videos. This model learns to optimize rPPG estimation from multiple enhanced videos with different signal frequencies and temporally adjacent videos with similar signal frequencies; specifically including: Step B-1, Video Input and Preprocessing: Given a facial video sequence, the MTCNN face detector is first used to detect, align, and crop the facial regions in each frame; the aligned video is represented as follows: V={v1,v2,...,v T }; Where v T This represents the T-th frame in the sequence; Step B-2, Time Segmentation: Divide the aligned video V into multiple segments, each segment containing T frames; the set of segments is represented as {V1, V2, ..., V...} K }, where V i Let x represent the i-th segment in the sequence; randomly select a segment from it as the primary input x. a The remaining segments are considered as time neighbors, denoted as {x} n1 ,x n2 ,...,x nk }; Step B-3, Data Augmentation: Spatial Augmentation: For x a Spatial augmentation is applied, which involves six image rotation methods: 0°, 90°, 180°, and 270°, as well as horizontal and vertical flipping, to obtain multiple sets of positive samples. Spatial augmentation does not affect the intrinsic rPPG signal in the sample, i.e.: Step B-4, Learnable Frequency Enhancement (LF): The LFA module modulates x a The rPPG signal generates a negative sample set. The frequency of the negative sample rPPG signal and x a different: Where r i ∈R={r1,r2,...,r M } represents the frequency ratio of LFA module applications; Step B-5, Signal Extraction: Extract positive and negative samples X p and X n The inputs are respectively fed into the Local rPPG Expert Aggregation (REA) module to estimate the corresponding rPPG signals: in, and Step B-6, Model Optimization: Frequency Contrast Loss: In Y p Between signals in or in Y p and Y n Loss due to frequency of application: Step B-7, Frequency Ratio Consistency Loss: This loss is used to constrain x. a The rPPG signals of its time neighbors are similar in frequency, that is: Cross-video frequency consistency loss: Step B-8, Total Loss Calculation: L total =λ contrast L contrast +λ ratio L ratio +λ cross L cross ; Where, λ contrast , λ ratio , λ cross The weighting factor is used; through the above stages, the model framework can learn the rPPG signal from unlabeled facial videos.

4. The non-contact body physiological signal acquisition method according to claim 2 or 3, characterized in that... It also includes the following steps: Step C, Implementation of heart rate variability and respiratory rate monitoring based on rPPG: This monitoring function is mainly achieved by filtering and denoising the obtained remote photoplethysmography (rPPG) signal. A bandpass filter can be applied to remove noise and retain heart rate-related frequency components. The maximum peak frequency, i.e., heart rate, is obtained by analyzing the power spectral density of the rPPG signal through fast Fourier transform. Heart rate variability includes three attributes: low-frequency (LF), high-frequency (HF), and the LF / HF ratio. These three attributes can be calculated by analyzing the intercardia-interval (IBI) sequence. Respiratory rate is related to LF. Specifically, it includes: Step C-1, rPPG signal acquisition and filtering: Extract the rPPG signal of the face region from the video frame and predict it using a model; the obtained rPPG signal is then processed by a bandpass filter to remove low-frequency and high-frequency noise; the frequency range of the bandpass filter is set to 0.6Hz to 4Hz; the filtered rPPG signal is represented as: rPPG(t) filtered =Butter_Bandpass(rPPG(t),owcut=0.6Hz,highcut=4Hz); Butter_Bandpass uses a Butterworth filter to implement bandpass filtering. The Butterworth filter is characterized by a smooth passband response without ripples, and it can better preserve the signal shape compared to other filters. Step C-2, heart rate calculation, is based on the analysis of the power spectral density (PSD) of the rPPG signal using Fast Fourier Transform (FFT); the specific calculation is as follows: Fourier Transform: Calculate the FFT of the rPPG signal to obtain the frequency domain signal. The Fourier transform operation is a mathematical method that converts a time-domain signal into a frequency-domain signal. It transforms a signal from a time-varying representation (time domain) to a representation of its frequency components (frequency domain). Power Spectral Density (PSD): Calculates the power spectral density of a frequency domain signal. PSD(f)=|FFT(f) rPPG | 2 ; Step C-3, Heart Rate Detection: Find the maximum peak frequency in the power spectral density; the corresponding frequency is the heart rate. The formula for calculating heart rate is: HR=argmaxf∈[0.67,3.33]HzPSD(f)×60; HR stands for heart rate, measured in bpm (beats per minute). 0.67Hz and 3.33Hz correspond to 40 bpm and 200 bpm, respectively, and are commonly used heart rate frequency ranges. Heart rate variability (HRV) frequency domain characteristics mainly include low-frequency LF and high-frequency HF power and their ratio; The interbeat interval (IBI) refers to the time interval between consecutive heartbeat peaks, and is usually calculated when the sampling rate sig_fps is known. The formula for calculating IBI is: Among them, t i It is the position of the i-th heart rate peak, IBI i It is the interval between the i-th and (i+1)-th peaks; Step C-4, Calculation of frequency characteristics: Frequency characteristics are calculated using the Lomb-Scargle periodogram to extract low-frequency (LF) and high-frequency (HF) power from the IBI sequence; the specific steps are as follows: Step C-4i, Spectrum Estimation: The IBI sequence is subjected to spectral analysis using the Lomb-Scargle method to obtain the frequency (freq) and power spectrum (power). Lomb-Scargle(IBI)→(freq,power); Step C-4ii, Low-frequency and high-frequency power: Select the power in the low-frequency range of 0.04Hz-0.15Hz and the high-frequency range of 0.15Hz-0.4Hz within the frequency range; power_LF=∑freq∈[0.04,0.15]power(freq); power_HF=∑freq∈(0.15,0.40]power(freq); Step C-4iii, Peak Frequency of the Spectrum: Find the peak frequency corresponding to the highest power in both the low-frequency and high-frequency regions: freq_LF_peak=argmaxfreq∈[0.04,0.15]power(freq); freq_HF_peak=argmaxfreq∈(0.15,0.40]power(freq); Step C-4iiiii, Frequency Domain Characteristics: The low-frequency peak frequency is freq_LF_peak, which reflects the dominant frequency of the heart rate variability signal in the low-frequency range; The high-frequency peak frequency is freq_HF_peak, which reflects the main frequency of the heart rate variability signal in the high-frequency range; Low-frequency power is denoted as power_LF. Low-frequency power is usually used to measure the overall activity level of the sympathetic and parasympathetic nervous systems. High-frequency power is denoted as power_HF and is typically used to measure the activity of the parasympathetic nervous system. Normalize the low-frequency power and high-frequency power: Step C-5, respiratory rate (RR) monitoring: respiratory rate can be represented by the high-frequency peak frequency of heart rate variability.