SYSTEM AND METHOD FOR DETERMINING PHYSIOLOGICAL VITAL PARAMETERS OF A USER IN A VEHICLE

A semi-supervised learning model decouples spatial and temporal feature extraction to address the challenges of remote vital sign monitoring, achieving efficient and accurate physiological parameter detection in vehicles.

DE102024137784A1Pending Publication Date: 2026-06-11MERCEDES BENZ GROUP AG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
MERCEDES BENZ GROUP AG
Filing Date
2024-12-14
Publication Date
2026-06-11

AI Technical Summary

Technical Problem

Existing methods for remote physiological vital sign monitoring in vehicles face challenges such as the need for large labeled datasets, computational intensity, and performance degradation under varying lighting and skin tones, leading to inaccurate and energy-intensive solutions.

Method used

A semi-supervised learning model (SSMLM) decouples spatial and temporal feature extraction using two-dimensional and one-dimensional convolutional layers, trained with a combination of supervised and unsupervised techniques, to generate remote photoplethysmography (rPPG) signals from facial video data, minimizing contrast loss and adapting to diverse conditions.

Benefits of technology

The SSMLM reduces computational and memory load, enhances generalization, and improves interpretability, enabling accurate vital parameter detection in resource-constrained environments, suitable for real-time monitoring in vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides a system (100) and a method (300) for determining the physiological vital parameters of a user (104) in a vehicle (102). The system (100) determines the physiological vital parameters of the user (104) based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning (SSMLM) model. The SSMLM generates the rPPG signals based on spatial and temporal features extracted from facial video data acquired by a sensor array (110). The SSMLM includes separate layers for spatial and temporal feature extraction. Furthermore, the SSMLM is trained using semi-supervised learning, which reduces the dependence on large amounts of labeled training data and the challenges associated with synchronizing labels with video data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates to the field of physiological vital sign monitoring of a user in a vehicle. In particular, the present disclosure provides a system and a method for physiological vital sign monitoring using a semi-supervised deep learning model for remote photoplethysmography (rPPG) in a vehicle. BACKGROUND

[0002] "Vital sensing" in vehicles is an emerging field that aims to improve driver safety and comfort by monitoring physiological parameters such as heart rate, respiratory rate, and stress levels. Such systems are particularly important for preventing road accidents caused by fatigue, drowsiness, or medical emergencies. Integrating vital sensors into vehicle systems can provide early warnings, trigger autonomous safety functions, or alert emergency services, thus creating a more responsive and intelligent in-car environment. This innovation aligns with the broader goals of intelligent mobility and personalized health monitoring.One of the ongoing challenges in this field, however, is the accurate detection and prediction of a vehicle user's vital signs, particularly for devices that are remotely controlled or where the user / owner / driver does not need to wear the device. Some solutions utilize machine learning based on area / video data, but such solutions are inaccurate and require large amounts of labeled data, which are only available to a limited extent.

[0003] Conventional methods using machine learning models created with supervised learning or training techniques require large amounts of labeled data, posing a challenge in data acquisition and labeling. Furthermore, inconsistencies between recorded baseline truth (GT), obtained with pulse oximeters, electrocardiograms, etc., and various regions of interest (ROIs) can lead to poor performance, as noise in GT does not correlate with ROI. Labeled data is obtained by capturing images / videos and using devices (e.g., photoplethysmography (PPG) devices) that measure baseline physiological data. However, synchronizing the baseline truth data and the images / videos presents a significant challenge.Furthermore, the data can also be noisy, rendering it unusable for training machine learning models. Dynamic conditions within the vehicle, such as motion artifacts and changing lighting conditions, further complicate the accurate monitoring of vital parameters.

[0004] Furthermore, the use of red-green-blue (RGB) sources to generate remote photoplethysmography (rPPG) signals is complicated because the intensity of the user's facial image is significantly affected by ambient light. This also leads to noisy data acquisition.

[0005] The non-patented publication "Deep Learning Methods for Remote Heart Rate Measurement: A Review and Future Research Agenda" by Cheng et al. describes deep learning methods for remote heart rate (HR) measurement using rPPG. The cited document describes end-to-end and hybrid deep learning methods. The end-to-end deep learning methods process the input video images directly to estimate heart rate or rPPG signals without intermediate steps. These methods are highly efficient and simplify the optimization process by integrating feature extraction, signal processing, and HR estimation into a single pipeline. Models such as two-dimensional (2D) foldable neural networks (CNNs), 3D CNNs, or attentional mechanisms enhance the model's ability to capture spatial, temporal, and physiological patterns.However, these methods rely heavily on large, labeled datasets, are computationally intensive, and pose a challenge in interpreting the models for clinical validation. Furthermore, their "black box" nature can limit their practical application in healthcare and complicate interpretability.

[0006] Furthermore, hybrid deep learning methods combine conventional algorithms with deep learning techniques to optimize specific phases of the HR measurement pipeline, such as signal extraction, noise filtering, or HR estimation. For example, deep learning-based approaches can refine skin detection, improve ROI quality, or filter rPPG signals for noise reduction. This offers flexibility, as these methods can target specific challenges (e.g., motion artifacts or skin tone variations) and are compatible with existing non-DL techniques. However, disadvantages include increased system complexity due to the integration of multiple methods, reliance on the limitations of conventional algorithms, and potentially lower performance compared to complete end-to-end methods in complex scenarios.

[0007] The unpatented publication "Remote Heart Rate Monitoring in Smart Environments from Videos with Self-supervised Pre-training" by Gupta et al. addresses remote heart rate (HR) monitoring using videos in smart environments, employing self-supervised contrastive learning to reduce reliance on labeled data. The cited document describes the use of a two-step approach: self-supervised pre-training followed by supervised fine-tuning for rPPG signal estimation. However, supervised learning methods are highly dependent on large, annotated datasets, making data collection and labeling a time-consuming and costly process. The models exhibit significant performance degradation when only a few labeled data are available and often struggle to generalize to unseen conditions, such as varying lighting or skin tones.

[0008] Patent document CN117835900A describes an imaging photoplethysmography (iPPG) system for remote monitoring of vital parameters using skin images captured by a camera. The system employs a supervised learning model based on a neural network with recurrent layers. The neural network is designed to process multidimensional time-series signals generated from the image sequence acquired by the system. The recurrent layers aid in the accurate estimation of PPG waveforms and vital signs by capturing both spatial and temporal information.

[0009] The models in the documents cited above require a much larger dataset and longer training times. This results in higher memory consumption and suboptimal performance. Furthermore, the supervised models in the cited documents rely heavily on labeled data and struggle with inconsistencies between the basic truth and the areas of interest, leading to suboptimal performance. Some records capture highly noisy basic truth data, rendering them unsuitable for supervised learning. Additionally, the existing solutions are computationally intensive and consume significant energy.

[0010] To overcome at least the limitations mentioned above, there is a need for systems and methods to determine physiological vital parameters using rPPG signals generated by computationally efficient models that are easy to train (i.e., with minimal labeled data). SUBJECT OF THE PRESENT DISCLOSURE

[0011] A general objective of the present disclosure is to provide a system and method for physiological vital sign measurement, such as heart rate measurement, using a semi-supervised learning model.

[0012] Another objective of this disclosure is to reduce the computational complexity and memory consumption of the machine learning implementation of the method and the system.

[0013] Another objective of the present disclosure is to overcome the influence of ambient light in the acquisition of remote photoplethysmography (rPPG) signals from the area of ​​a user in a vehicle. SUMMARY

[0014] Aspects of the present disclosure relate generally to the physiological vital sign monitoring of a user in a vehicle. In particular, the present disclosure provides a system and a method for determining the physiological vital signs of a user in a vehicle using a semi-supervised learning model.

[0015] In one aspect, a system for determining a user's physiological vital parameters in a vehicle comprises a processor and a memory that is communicatively coupled to the processor. The memory contains processor-executable instructions which, when executed, cause the processor to determine a user's physiological vital parameters based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning (SSMLM) model. The SSMLM is configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data captured by a sensor array.

[0016] In one embodiment, the SSMLM includes one or more two-dimensional convolutional layers configured to extract spatial features from facial video data. The SSMLM also includes one or more one-dimensional convolutional layers configured to extract temporal features from the spatial features extracted over a predefined time interval. Furthermore, the SSMLM includes a decoder layer configured to generate rPPG signals based on the extracted temporal features.

[0017] In one embodiment, the SSMLM is trained with a labeled dataset containing the facial video data that is associated with a corresponding basic truth vital signal determined by a PPG sensor.

[0018] In one embodiment, the SSMLM is trained based on an objective function to minimize contrast loss. The contrast loss is determined based on the similarity result between each pair of rPPG signal samples extracted from the rPPG signals generated for various facial video data using a spatiotemporal sampling method.

[0019] In one embodiment, the generated rPPG signals can be sampled based on a uniform spatiotemporal sampling.

[0020] In one embodiment, the facial video data includes near-infrared video images captured by an NIR image sensor.

[0021] In one embodiment, the physiological vital parameters include at least one of the parameters heart rate, respiratory rate or heart rate variability.

[0022] In another aspect, a method for determining a user's physiological vital parameters in a vehicle involves determining the user's physiological vital parameters by a processor based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning (SSMLM) model, the SSMLM being configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data captured by a sensor array.

[0023] Various objects, features, aspects and advantages of the subject matter according to the invention will become clearer from the following detailed description of preferred embodiments together with the accompanying drawing figures, in which the same numbers represent the same components. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings serve to further understand the present disclosure and are an integral part of this description. The drawings illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. Fig. Figure 1 shows a block diagram of a system for determining physiological vital parameters of a user in a vehicle according to the embodiments of the present disclosure. Fig. Figure 2 shows a simplified block diagram of a semi-supervised learning model (SSMLM). Fig. Figure 3 shows a flowchart for the implementation of an example procedure for determining physiological vital parameters of a user in a vehicle according to an embodiment of the present disclosure. Fig. Figure 4 shows an exemplary computer system in which or with which embodiments of the system according to the embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0025] A detailed description of the embodiments of the disclosure illustrated in the accompanying drawings follows. The embodiments are described in sufficient detail to clearly convey the disclosure. However, the necessary level of detail is not intended to limit foreseeable variations of embodiments; on the contrary, it is intended to cover all modifications, equivalents, and alternatives that fall within the scope of this disclosure as defined by the accompanying claims.

[0026] The embodiments described herein relate to the determination of a user's physiological vital parameters in a vehicle. In particular, the present disclosure provides a system and a method for determining a user's physiological vital parameters in a vehicle using a semi-supervised learning model.

[0027] In one aspect, the present disclosure provides a system and a method for determining the physiological vital parameters of a user in a vehicle. The system determines the physiological vital parameters of a user based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning (SSMLM) model. The SSMLM is configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data acquired by a sensor array. Furthermore, the SSMLM decouples the spatial and temporal feature extraction into separate layers, thereby optimizing real-time edge unfolding performance.

[0028] The various embodiments of the disclosure are described with reference to the Fig. 1-3.

[0029] Referring to the block diagram in Fig. Figure 1 shows a system (100) for determining the physiological vital parameters of a user (104) in a vehicle (102). The system (100) comprises a processor (106), a memory (108), and a sensor array (110). The system (100) can be implemented in the vehicle (102), for example, in an electronic control unit (ECU) or a vehicle control unit (VCU) of the vehicle (102). In some embodiments, the sensor (110) can include a visible light camera. In some arrangements, the sensor array (110) includes one or more near-infrared (NIR) image sensors. The NIR image sensor can include a sensor array for capturing images of the user's (104) area. In some embodiments, the sensor array (110) can be configured to capture (facial) video images or video data of the user (104), which may consist of several temporally associated images.

[0030] The system (100) can be configured to determine the physiological vital parameters of multiple users simultaneously in the vehicle (102). For example, facial video data can be acquired by a wide-angle NIR sensor that captures images / videos of multiple users, which can then be segmented to determine the physiological vital parameters of each user. In some embodiments, the system (100) can include a semi-supervised machine learning (SSMLM) model configured to determine the physiological vital parameters from the facial video images or data acquired by the sensor array (110).In some embodiments, the SSMLM can be adapted to determine the vital metrics of each user (104) regardless of, but not limited to, their sex, ethnicity, and skin color, and in a variety of situations representing, but not limited to, different vehicle speeds, user body temperatures, ambient temperatures, ambient pressures, and ambient light conditions. In some embodiments, the SSMLM includes separate layers to extract spatial and temporal features from the input facial data. In such embodiments, the SSMLM performs spatial feature extraction (via two-dimensional convolutions) and temporal feature extraction (via one-dimensional convolutions), enabling efficient computation and reduced memory usage. Further details of the SSMLM are provided in the present disclosure with reference to [reference to relevant document]. Fig. 2 described in detail. In some embodiments, the specified physiological vital parameters may be any one or a combination of heart rate, respiratory rate, and heart rate variability parameters, but are not limited to these. The SSMLM can be adapted to determine all physiological vital parameters that can be extracted from the facial video data.

[0031] In some embodiments, the SSMLM can be configured to extract spatial features from the facial video data. It is important to note that the spatial features are sections or areas of interest (ROIs) within the image. Examples of spatial features include sections of the video frames corresponding to the nose, forehead, cheeks, T-zones, ears, chin, and the like of the user (104). Because the SSMLM separates spatial and temporal feature extraction into distinct layers, the spatial feature extraction layers can be optimized independently to focus on each of the spatial features or different ROIs within the facial video data. The processor (106) prepares the ROI by applying preprocessing techniques, such as object detection, segmentation, or temporal filtering, to isolate relevant sections of the data.In a facial video dataset used for remote photoplethysmography (rPPG), the processor can, for example, use facial recognition algorithms to identify the area in each image. It can then crop and align these regions to focus on areas with prominent blood flow signals.

[0032] Furthermore, the SSMLM can be configured to determine temporal features from spatial features extracted over a predefined time interval. For example, spatial features extracted from 10 seconds of video data can be used to extract temporal features. The temporal features correspond to the changes in the spatial features over the predefined time interval. This time interval can be defined by a predefined number of video frames. In one example, the temporal features can show the changes in individual spatial features, such as color, shape, size, orientation, and the like, over a specific period. In some embodiments, the spatial and temporal features can be represented by vectors or embeddings. Based on the temporal features, the SSMLM can be configured to generate the rPPG signal.The processor (106) can receive the rPPG signal and use it to determine physiological vital parameters. For example, if the spatial features indicative of the eyes are identified as drooping or closing over a specific period of time, the user (104) can be identified as drowsy. Similarly, the number of changes in the user's (104) skin color can indicate the user's (104) heart rate.

[0033] The system (100) may include the processor(s) (106), which may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logic circuits, and / or any devices that process data based on operating instructions. Among other capabilities, the processor(s) (106) may be configured to retrieve and execute computer-readable / processor-executable instructions stored in the memory (108) of the system (100). The memory (108) may be configured to store one or more computer-readable instructions or routines on a non-volatile, computer-readable storage medium, which can be retrieved and executed to create or exchange data packets with components of the system (100).The memory (108) can contain any non-volatile storage device, such as volatile memory like random-access memory (RAM) or non-volatile memory like erasable programmable read-only memory (EPROM), flash memory, and the like. In some embodiments, the instructions stored in the memory (108) can cause the processor (106) to determine the physiological vital parameters using the SSMLM. In some embodiments, the SSMLM can contain one or more weights assigned to each of its layers. The weights of the SSMLM can be stored in a database (not shown). In some embodiments, the SSMLM can be trained using semi-supervised learning techniques, wherein the SSMLM can be trained partly using supervised learning techniques and partly using unsupervised learning techniques.In supervised learning, the SSMLM can be trained on a labeled dataset.

[0034] In some embodiments, the system (100) may also include a PPG sensor (112). The PPG sensor (112) may be configured to detect physiological vital parameters that correspond to the basic truth. The processor (106) may also be communicatively coupled to the PPG sensor (112) to detect the physiological vital parameters and generate a labeled data set. Once sufficient data has been collected to generate the labeled data set, the PPG sensors (112) may be removed.

[0035] In some embodiments, the PPG sensor (112) may include invasive or non-invasive PPG sensors. Examples of PPG sensors include wearable PPG sensors such as fitness trackers, smartwatches, medical wearables, pulse oximeters, cameras, and the like.

[0036] In some embodiments, the processor (106) can be configured to generate the rPPG signals based on the facial video data, while the PPG sensors (112) acquire the physiological vital parameters. In some arrangements, the processor (106) can be configured to acquire NIR images / pictures / video data using the sensor arrangement (110) and PPG signals using the PPG sensors (112) from a multitude of users (104) under a multitude of situations and store them in memory (108). The multitude of users represents at least different genders, ethnicities, and skin colors. The multitude of situations represents at least different vehicle speeds, user body temperatures, road textures, ambient temperatures, ambient pressure, and ambient light conditions.In some embodiments, the ground-truth vital data are vital data measured by the PPG sensors (112). Ground-truth heart rate labels are the actual, accurate heart rate values ​​used as reference data for training and validating a model. These values ​​are typically measured using devices or methods that meet the gold standard and are known for their high precision. They serve as a benchmark against which the system's predictions (PPG or rPPG) are compared. Therefore, the labeled dataset can be a diverse dataset, enabling the SSMLM to generate the rPPG signal in a variety of situations for a variety of demographic groups.

[0037] Once the marked data set is generated, the system (100) can be configured to generate the rPPG signals and thus determine the physiological vital parameters. The determined physiological vital parameter can be compared with the corresponding baseline physiological vital parameters detected by the PPG sensors (112) at the same time as the facial video data was acquired by the sensor array (110). The difference between the determined and baseline physiological vital parameters can be used to determine a loss / error value, which can be used to update the SSMLM weights. The weights can be updated to minimize the loss value in subsequent iterations.

[0038] In some embodiments, the SSMLM can also be trained using unsupervised learning techniques when physiological vital parameters are unavailable. In such embodiments, the processor (106) is also configured to train the SSMLM using the generated rPPG signals. In some embodiments, the rPPG signals can be generated based on the spatial and temporal features produced in the corresponding layers of the SSMLM (such as two-dimensional convolutional layers and one-dimensional convolutional layers). The one or more two-dimensional convolutional layers are configured to extract the spatial features from the facial video data. Furthermore, the one or more one-dimensional convolutional layers are configured to extract the temporal features from the spatial features extracted over a predefined time interval.Furthermore, the SSMLM includes a decoder layer configured to generate rPPG signals based on the extracted temporal features. For unsupervised learning, rPPG signals can be generated for a variety of users (104) and / or in a variety of situations.

[0039] In some embodiments, the unsupervised learning method used to train the SSMLM can be a contrastive learning method. Contrastive learning techniques are a class of self-supervised learning methods that aim to learn meaningful representations by juxtaposing positive and negative pairs of data samples. The primary goal is to bring similar (positive) pairs closer together in the learned representation space and to push dissimilar (negative) pairs apart. Positive pairs typically consist of extended versions of the same data pattern, while negative pairs are drawn from different patterns. Contrastive learning often employs techniques such as data extension (e.g., cropping, mirroring, or color shifting) to generate different views of the same data, enabling the model to learn invariant and robust features.Key algorithms include SimCLR (Simple Contrastive Learning of Representations), which uses contrastive losses with large stack sizes to generate different negatives, and MoCo (Momentum Contrast), which uses a memory bank to maintain a pool of negative samples. These techniques have proven highly successful in areas such as computer vision, where they excel in downstream tasks by learning rich, transferable features without relying on labeled data.

[0040] In such embodiments, rPPG signal samples can be extracted from the rPPG signals generated for different sets of facial video data (e.g., for different users (104) or the same users (104) in different situations or environmental conditions). In some embodiments, the rPPG signals generated for the different users (104) / situations can be sampled using a spatiotemporal sampling method. In one embodiment, the SSMLM is trained based on an objective function to minimize contrast loss. The contrast loss is determined based on a similarity result between each pair of extracted rPPG signal samples using the spatiotemporal sampling method. In some embodiments, the generated rPPG signals are sampled based on a uniform spatiotemporal sampling approach.Furthermore, the similarity result can be determined by comparing at least one of the rPPG signals with the power spectral densities (PSDs) associated with the sampled extracted spatial features. Since the spatial and temporal features can be represented as vectors, the resulting rPPG signals can also be represented as vectors in some embodiments. In such embodiments, any method / technique for determining the similarity result between any pair of rPPG signal samples can be used, such as, but is not limited to, cosine similarity, Euclidean distance, Manhattan distance, and the like. In some embodiments, the contrastive loss function aims to minimize the distance between positive pairs of rPPG signal samples while maximizing the distance between negative pairs.Positive pairs are generated from spatiotemporal regions of the same video or user in similar contexts, while negative pairs are drawn from different videos, users, or environmental conditions. This ensures robust learning of physiological features across various scenarios. Furthermore, the rPPG signals of the baseline truth (where available) and the generated rPPG signals are continuously compared to maintain signal integrity and accuracy. PSD features are used for precise similarity assessment of both labeled and unlabeled data.

[0041] The SSMLM can therefore be a machine learning model trained on a dataset containing both labeled data (with basic truth labels) and unlabeled data (without labels). It uses the labeled data to guide the learning process and the unlabeled data to improve the model's performance by capturing the underlying data structure or patterns. The training process combines supervised learning techniques for labeled data and unsupervised methods (such as clustering or self-training) for unlabeled data. To improve model performance, at least one of the following techniques can be employed: pseudo-labeling, consistency regulation, and contrastive learning.

[0042] In some embodiments, the SSMLM can be implemented within the vehicle (102). In some embodiments, the SSMLM can be trained by a training module (not shown). In some embodiments, the training module can be implemented as a processing resource or a processing machine / module of the processor (106). In some embodiments, the training module can be implemented in the system (100) within the vehicle (102). In such embodiments, the training module can be configured to train the SSMLM in real time or at regular intervals based on the rPPG signals generated during inference. In other embodiments, the training module can be located outside the vehicle (102). In such embodiments, the SSMLM can be trained by the training module in an external computing device and deployed to the system (100) for inference.The weights of the SSMLM can be updated and downloaded via the system (100) in the vehicle (102).

[0043] Further details of the SSMLM architecture will be provided with reference to Fig. 2 described.

[0044] With reference to the in Fig. Figure 2 (200) shows a simplified block diagram of the semi-supervised learning model (SSMLM).

[0045] In some embodiments, one or more inputs (202) can be provided to the SSMLM. The inputs (202) can be the face video data / frames. In some embodiments, the face video data can comprise a plurality of NIR images acquired by the sensor array (110). In some embodiments, the face video data can be represented in the form of matrices / vectors of the form (B, T, H, W). Here, B refers to the stack size, T is the number of images (in the temporal dimension), and H and W are the height and width (spatial dimension) of the video images.

[0046] Furthermore, the SSMLM can include an initial convolution layer (204) that processes the input video frames individually using one or more two-dimensional convolution layers with a predefined kernel (filter) size. In one example, the kernel size can be 5x5. This is followed by stack normalization and an exponential layer activation function (ELU) to introduce nonlinearity. The initial convolution layer (204) improves the local spatial features for each frame before proceeding to deeper layers.

[0047] Additionally, the SSMLM (200) includes a spatial feature extraction block / layer (206) configured to extract spatial features from the initial convolution (204). These features are then further processed by one or more two-dimensional convolution layers, such as Loop-1 (206A), Encoder-1 (206BA), Encoder-2 (206BB), and Loop-4 (206C). Loop-1 (206A) performs 2D convolutions and extracts features with higher spatial resolution. Encoder-1 (206BA) and Encoder-2 (206BB) reduce the spatial dimensions (height and width) while maintaining the same number of channels through convolution and averaging. This allows the SSMLM (200) to retain essential information while reducing computational complexity. Loop 4 (206C) is another layer of 2D folding for further refinement of the spatial features.

[0048] Furthermore, the SSMLM (200) includes a temporal convolution block (208) with one or more one-dimensional convolution layers, such as temporal convolution layer (TC)-1 (208A) and TC-2 (208B). After spatial feature extraction (206), the data are restructured to prepare them for temporal processing. For example, spatial features associated with a predetermined time interval can be combined for temporal feature extraction. TC-1 (208A) applies a 1D convolution across the temporal dimension to capture temporal changes in the spatial features. TC-2 (208B) follows with another convolution to further refine the temporal features.

[0049] Furthermore, the SSMLM contains a decoding block (210) with two decoders, Decoder-1 and Decoder-2. These decoders decode the temporal features. The temporal convolutions are passed through these layers, which include one-dimensional convolution with ELU activation to reconstruct the signal. These layers help prepare the signal for final extraction by refining the temporal features derived from the temporal convolution layers.

[0050] Furthermore, a final layer (212) blocks max pooling over the time dimension, followed by 1D convolution to reduce the output to a single channel. This output represents the extracted rPPG signal over time, which is typically used to estimate heart rate.

[0051] Finally, one output (214) represents the generated rPPG signal, which is a temporal waveform indicating the change in physiological vital signs, e.g., blood volume over time or heart rate, extracted from the video images of the face.

[0052] It should be noted that the in Fig. The two blockages shown are examples and should not be understood as limiting the scope of this disclosure. The arrangement and components shown can be modified or supplemented with additional elements to improve or optimize the model's performance, accommodate specific hardware requirements, or meet other implementation needs without deviating from the scope of this disclosure.

[0053] As described, the SSMLM can be configured to use different layers for extracting spatial and temporal features. Separate layers for spatial and temporal feature extraction reduce the number of parameters and thus the computational and memory load compared to other (three-dimensional) neural networks due to the smaller number of operations achieved through the smaller form of the kernels used for the convolution operations. Furthermore, the SSMLM completely decouples the layers for spatial and temporal feature extraction (by using two-dimensional convolutions for spatial feature extraction and one-dimensional convolutions for temporal feature extraction) to minimize the number of parameters, reduce the computational and memory load, and simplify the optimization process, making the SSMLM lightweight (i.e.,SSMLM can be trained independently and effectively. Due to its lower computational and memory requirements, it is also suitable for edge applications, especially in resource-constrained environments. Furthermore, SSMLM offers flexibility in handling the time window, as temporal convolutions can be set independently of spatial convolutions. This allows for greater adaptability to capture different temporal dynamics without impacting spatial feature extraction. The separation of layers also optimizes real-time performance for edge unfolding.

[0054] The SSMLM also exhibits better generalization with smaller datasets due to its reduced number of parameters. The SSMLM architecture reduces the risk of overfitting by enabling better generalization through its modular design. This makes it particularly suitable for fitness, medical, and physiological applications, where labeled data is often scarce. Furthermore, the clear separation between the convolutional layers that extract spatial and temporal features allows for specialized optimization of each layer type, improving performance in tasks requiring distinct processing of spatial and temporal features. Additionally, the SSMLM offers enhanced interpretability through the separate processing of spatial and temporal components, aiding in the analysis of which parts of the SSMLM are responsible for static versus dynamic variations, for example.For example, when tracking impulse changes across multiple frames.

[0055] In some embodiments, physiological vital parameters can be used to determine the condition of the user (104), who may be either an occupant or a driver. The design / architecture of SSMLM enables the acquisition of vital data in real time by utilizing computationally efficient spatiotemporal processing, making it suitable for deployment in vehicles with limited computing power. Detecting physiological vital parameters can reduce road accidents caused by fatigue, drowsiness, or medical emergencies, as such conditions of the user (104) can be detected early. Integrating vital data into vehicle systems allows for early warnings, the triggering of autonomous safety functions, or the alerting of emergency services, resulting in a more responsive and intelligent in-vehicle environment.

[0056] In Fig. Figure 3 shows a flowchart for implementing an example method (300) for determining the physiological vital parameters of a user (104) in a vehicle (102). The method highlights the two-stage SSMLM processing, in which spatial features are first extracted independently of each video frame, and temporal features are processed across frames to efficiently generate rPPG signals. In some embodiments, the method (300) can be implemented by the system (100) or any general-purpose processor known to those skilled in the art.

[0057] In step (302), the method (300) comprises the determination of the user's (104) physiological vital parameters by a processor based on rPPG signals generated by a SSMLM. The determined physiological vital parameters include one or a combination of, but not limited to, heart rate, respiratory rate, heart rate variability parameters, and the like. The rPPG signals can be extended by suitable adaptable SSMLMs to detect other metrics, such as blood oxygen content. In some embodiments, the SSMLM is configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data acquired by a sensor array, such as the sensor array (110) of Fig. 1. The facial video data may be near-infrared video images captured by an NIR image sensor.

[0058] In some embodiments, the SSMLM can include one or more two-dimensional convolution layers configured to extract spatial features from the face video data, one or more one-dimensional convolution layers configured to extract temporal features from the spatial features extracted over a predefined time interval, and a decoder layer configured to generate the rPPG signals based on the extracted temporal features. The SSMLM can be similar to the one described in Fig. The model shown and described in section 2 can be implemented. In some embodiments, the SSMLM can be trained using semi-supervised learning techniques that include a combination of supervised and unsupervised learning techniques / methods. In some embodiments, the SSMLM can be trained using a labeled dataset containing the facial video data mapped to a corresponding basic truth vital signal determined by a PPG sensor (e.g., PPG sensor (112)), if available. Alternatively or additionally, the SSMLM can also be trained based on an objective function to minimize a contrastive loss value. The contrastive loss value can be determined based on a similarity result between each pair of rPPG signal samples extracted from the rPPG signals generated for different facial video data using a spatiotemporal sampling method.In some embodiments, the generated rPPG signals can be sampled based on a uniform spatiotemporal sampling.

[0059] Back to Fig. 3: To determine the physiological vital parameters using the SSMLM in step (302), the method (300) may comprise substeps (302A to 302D). In step (302A), the method (300) comprises extracting the spatial features from the facial video data using the SSMLM. In some embodiments, the facial video data may be preprocessed to extract different segments of the video images using segmentation techniques / methods known to those skilled in the art. In step (302B), the method (300) comprises extracting the temporal features from the spatial features extracted in step (302A) over a predefined time interval using the SSMLM. The predefined time interval (time window) may be appropriately adjusted based on the requirements of the application.In step (302C), procedure (300) includes the generation of rPPG signals by the SSMLM, based on the temporal features extracted in step (302B). Furthermore, in step (302D), procedure (300) includes the determination of physiological vital parameters by the processor (106) based on the rPPG signals generated in step (302C). For example, the rPPG signals can represent a waveform of the user's heart rate (104). The waveform can be processed (e.g., by counting the number of peaks in the waveform) to derive the heart rate.

[0060] In some embodiments, the system (100) can be implemented in a computer system. Fig.Figure 4 shows a block diagram of a computer system (400) comprising an external storage device (410), a bus (420), main memory (430), read-only memory (440), mass storage device (450), a communication port (460), and a processor (470). A person skilled in the art will understand that the computer system (400) may include more than one processor (470) and communication ports (460). The processor (470) may contain various modules, which are associated with the embodiments of this disclosure. The communication port (460) may be a recommended standard 232 port for use with a modem-based dial-up connection, a 10 / 100 Ethernet port, a Gigabit or 10 Gigabit port over copper or fiber optic cable, a serial port, a parallel port, or other existing or future ports. The port (460) may be chosen depending on a network, e.g.,a Local Area Network (LAN), a Wide Area Network (WAN) or any other network to which the system (400) is connected.

[0061] In one embodiment, the memory (430) can be RAM or other dynamic storage device generally known in the art. The read-only memory (ROM) (440) can be any static storage device, e.g., a programmable read-only memory (PROM) for storing static information, without being limited to that. The mass storage device (450) can be any current or future mass storage solution that can be used to store information and / or instructions. Exemplary mass storage solutions include, but are not limited to, parallel Advanced Technology Attachment (PATA) or serial Advanced Technology Attachment (SATA) hard disk drives or solid-state drives (internal or external, e.g., with Universal Serial Bus (USB) and / or FireWire interfaces), one or more optical disks, redundant array of independent disks (RAID) storage, e.g.,an array of hard drives (e.g. SATA arrays).

[0062] In one embodiment, the bus (420) communicatively couples the processor(s) (470) with the other storage, repository, and communication blocks. The bus (420) can be, for example, a Peripheral Component Interconnect (PCI) / PCI-Extended (PCI-X) bus, a Small Computer System Interface (SCSI), USB, or similar, for connecting expansion cards, drives, and other subsystems, as well as other buses, such as the front-end bus (FSB), which connects the processor (470) to the computer system (400).

[0063] In another embodiment, operator and management interfaces, such as a display device, a keyboard, and a cursor control unit, can also be connected to the bus (420) to support direct operator interaction with the computer system (400). Other operator and management interfaces can be provided via network connections connected through the communication port (460). In some embodiments, the external storage device (410) can be any type of external hard disk drive, floppy disk drive, Compact Disc - Read Only Memory (CD-ROM), Compact Disc - Re-Writable (CD-RW), or Digital Video Disc - Read Only Memory (DVD-ROM). The components described above are given only as examples of various possibilities. The exemplary computer system (400) described above is not intended to limit the scope of this disclosure in any way.

[0064] While the foregoing describes various embodiments of the present disclosure, other and further embodiments of the present disclosure may be developed without departing from the basic scope. The scope of the present disclosure is determined by the following claims. The present disclosure is not limited to the described embodiments, versions, or examples, which are included to enable a person with ordinary technical knowledge to manufacture and use the present disclosure when combined with the information and knowledge available to such a person. BENEFITS OF THE PRESENT DISCLOSURE

[0065] The present disclosure enables a significant reduction in the number of parameters, thereby reducing the computational and memory load compared to three-dimensional convolutional neural networks.

[0066] The present disclosure offers smaller kernel forms for convolution operations, resulting in fewer operations and lower resource consumption.

[0067] The present disclosure offers a simplified optimization process by decoupling spatial and temporal feature extraction, which enables simpler and more effective training of the SSMLM.

[0068] The present disclosure offers flexibility in handling the time window, so that the temporal convolutions can be set independently of the spatial convolutions in order to achieve better adaptability to different temporal dynamics.

[0069] The present disclosure offers improved generalization to smaller datasets due to the reduced number of parameters, making it particularly suitable for scenarios with limited data availability, such as medical and physiological signal extraction.

[0070] The present disclosure offers a specialized optimization of spatial and temporal layers, thereby improving performance for tasks that require different treatment of spatial and temporal features.

[0071] The lower complexity of the model reduces the risk of overfitting, leading to better generalization across different datasets.

[0072] The present disclosure offers improved interpretability through the separate processing of spatial and temporal components, enabling the analysis of static and dynamic variations, such as pulse tracking across frames.

[0073] The present disclosure offers maximum data utilization; even rejected data can be used to develop a semi-supervised model that helps with cost optimization.

[0074] The present disclosure uses NIR-based rPPG facial imaging, which is not affected by ambient light.

[0075] The present disclosure provides a model that is optimized for edge unfolding or unfolding in resource-poor environments because it requires less computational effort.

[0076] This disclosure enables the simultaneous monitoring of multiple users / occupants of the vehicle.

[0077] The present disclosure ensures robustness against environmental fluctuations such as changes in light, temperature and pressure.

[0078] The current disclosure offers the possibility of including various demographic groups, generally for all genders, ethnicities and skin colors.

[0079] The present disclosure ensures the robustness of contrastive learning by minimizing the distance between positive pairs and maximizing the distance between negative pairs.

[0080] The present disclosure offers scalability for various applications, including fitness, health, and automotive systems.

[0081] The present disclosure offers energy efficiency suitable for battery-powered and energy-limited devices.

[0082] This disclosure enables integration with security systems, e.g., the detection of fatigue and drowsiness.

[0083] The present disclosure offers possibilities for extending the measurement to include additional parameters, such as the oxygen content in the blood. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] CN 117835900A

[0008]

Claims

[1] System (100) for determining physiological vital parameters of a user (104) in a vehicle (102), wherein the system (100) comprises: a processor (106); and a memory (108) that is communicatively coupled to the processor (106), wherein the memory (108) contains processor-executable instructions which, when executed, cause the processor (106) to: to determine the physiological vital parameters of the user (104) based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning (SSMLM) model, wherein the SSMLM is configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data captured by a sensor array (110). [2] System (100) according to claim 1, wherein the SSMLM comprises: one or more two-dimensional convolution layers configured to extract spatial features from facial video data; one or more one-dimensional convolution layers configured to extract temporal features from spatial features extracted over a predefined time interval; and a decoder layer configured to generate the rPPG signals based on the extracted temporal features. [3] System (100) according to claim 1, wherein the SSMLM is trained using a labeled data set comprising the facial video data mapped to a corresponding basic truth vital signal determined by a PPG sensor. [4] System (100) according to claim 1, wherein the SSMLM is trained on the basis of an objective function to minimize a contrastive loss value, and wherein the contrastive loss value is determined on the basis of a similarity result between each pair of rPPG signal samples extracted from the rPPG signals generated for different face video data using a spatiotemporal sampling method. [5] System (100) according to claim 4, wherein the generated rPPG signals are sampled on the basis of a uniform spatiotemporal sampling. [6] System (100) according to claim 1, wherein the face video data are near-infrared video images captured by a NIR image sensor. [7] System (100) according to claim 1, wherein the physiological vital parameters comprise one or a combination of heart rate, respiratory rate and heart rate variability parameters. [8] Method (300) for determining physiological vital parameters of a user (104) in a vehicle (102), the method comprising: Determination of the user's physiological vital parameters (104) by a processor based on remote photoplethysmography (rPPG) signals generated by a semi-supervised machine learning model (SSMLM), wherein the SSMLM is configured to generate the rPPG signals based on spatial and temporal features extracted from facial video data captured by a sensor array (110). [9] The method of claim 8, wherein the SSMLM comprises: one or more two-dimensional convolution layers configured to extract spatial features from facial video data; one or more one-dimensional convolution layers configured to extract temporal features from spatial features extracted over a predefined time interval; and a decoder layer configured to generate the rPPG signals based on the extracted temporal features.